slime joins Megatron training to SGLang rollouts
slime coordinates the loop that turns generated responses into model updates. Megatron trains the policy, SGLang produces rollouts, a data buffer carries prompts and results, and Ray places workers across the cluster. Custom generation and reward functions can add search, tools, sandboxes, verifiers, or multi-agent interactions. The repository also covers checkpoint conversion, weight synchronization, evaluation, tracing, and recovery, which are the unglamorous parts that determine whether a long training job can continue.
The framework is intentionally narrow about engines. Every installed SGLang argument can pass through with an --sglang- prefix, and Megatron parameters stay visible on the training side. That gives specialists access to routing, caching, parallelism, checkpointing, and memory controls without waiting for a generic adapter. Teams standardized on vLLM or another rollout server should consider a different framework, because backend interchangeability is not the design goal stated in the README.
The quick start assumes a prepared GPU cluster
The recommended setup pulls the project's Docker image and starts it with all GPUs, host IPC, and 16 GB of shared memory. Users then download model weights and datasets, convert Hugging Face checkpoints into Megatron's distributed format, source a model-specific parameter script, and launch through Ray. The guide's standard example assigns 4 GPUs to training and another 4 to rollouts; its colocated example shares 8 GPUs and requires memory tuning.
Model scripts reduce typing but do not remove validation work. The guide tells users to compare rotary settings with the exact model and notes that checkpoint conversion may need a manually supplied vocabulary size. Batch sizes must keep generated and consumed sample counts equal. The FAQ covers stuck placement groups, out-of-memory failures, port conflicts, bad stop tokens, garbled output, and NaN gradients because those are credible operating conditions for this stack.
What happened when we ran it
Our sandbox installed 191 Python packages in 121 seconds, consuming 5,866 MB on disk. The build then passed in 6 seconds. The checkout at commit 624b824 had 608 files, roughly 85,633 lines of source, and a 10.9 MB working tree before dependencies. Pip-audit found 0 known vulnerabilities in the installed environment. These checks used Python 3.12 on Debian with 3 CPUs and 8 GB of RAM, without secrets.
Tests failed with exit code 1 after 392 seconds. Pytest reported 526 passed, 26 failed, 16 skipped, and 29 collection or setup errors of 581. The log tail named a Qwen test with megatron.core.__spec__ is None. Several rollout-data tests raised AttributeError because the loaded ray module had no get attribute. The log does not establish why those conditions occurred, so our finding stops at the failed suite in the stated sandbox.
SGLang focus buys control and creates dependency coupling
Release v0.3.1 added memory work for colocated training, faster checkpoint saving, external rollout engines, delta weight synchronization, sampling alignment, and fused PPO calculations. It also removed Megatron Bridge and points users who need broader model support toward Miles. The same release includes coding-agent work that can select Claude Code or Codex harnesses. That integration is useful for training agents on test-based rewards, but it adds sandboxes and long message histories to an already complex system.
Direct access to upstream switches reduces wrapper friction, while upgrades remain a compatibility problem shared across slime, SGLang, Megatron, PyTorch, CUDA, Ray, and model kernels. Open issue 2332 gives a small example: W&B 0.28.2 removed a function used during initialization, so training can fail before it starts. A pull request was already open on August 26, 2026, which shows fast response and the cost of moving dependencies.
Current reports include OOM and nonfinite model states
Issue 2253 reports full-vocabulary FP32 logits consuming about 10.8 GB for one roughly 11,600-token microbatch in a stated Qwen3.5 configuration. Issue 2209 reports NaN or infinite rollout weights after delta synchronization on a Qwen3.5 122B mixture-of-experts model, reproduced in 4 of 4 runs by the reporter. These are specific configurations and reporter measurements, not results from our sandbox, but both concern expensive failures during the exact workloads slime targets.
Issue 2201 isolates nonfinite gradients to a GLM-5.2 SparseMLA backward kernel on a 16-node H200 setup. The reporter produced finite reference gradients and a slow fallback, then asked for a supported safer path. A team adopting slime should keep numerical checks enabled, save enough state for reproduction, and stage new kernels or weight-transfer modes. Silently replacing bad values or disabling checks may let a job continue while invalidating the training result.
Active development does not make this beginner infrastructure
GitHub recorded 8,260 stars, 450 combined open issues and pull requests, and a last push on August 26, 2026. Release v0.3.1 shipped on August 6. The large open queue includes detailed questions, fixes, performance work, and reports across many models, so it signals both use and a wide support burden. Apache 2.0 permits commercial adoption, and the project maintains English documentation alongside Chinese material.
slime is worth evaluating for teams already committed to Megatron and SGLang, especially when custom agent rollouts need tight control over training data and weight updates. Our 526 passing tests show meaningful coverage, while the 26 failures and 29 errors rule out calling this checkout clean in a general CPU sandbox. Budget for GPUs, conversion, observability, reproducibility work, and engineers who can read the failure below the framework layer.

