mrkeyoor.com_
Wed 30 Sept 20:34 UTC
LLM Toolsevaluationupdated 26 Aug 2026

slime review

slime is a framework for improving large language models with reinforcement learning after initial training. It connects Megatron training, SGLang rollouts, rewards, checkpoints, and custom agent environments in one distributed loop for teams that would otherwise assemble those systems themselves.

+35stars / 7d
Verdict

Our slime run installed 191 packages and built successfully, but pytest finished with 26 failures and 29 errors after 392 seconds. Adopt it only when Megatron and SGLang are deliberate choices and your team can investigate distributed training failures at model, kernel, and cluster level. For a first post-training experiment or a backend-portable stack, TRL, veRL, or OpenRLHF asks less of the operator.

We ran it

Lab card: what happened when we ran slimeScreenshot of slime (thudm.github.io/slime)
Install✓ · 121s191 packages · 5866 MB
Build✓ · 6s
Tests✗ · 392s526 passed · 26 failed · 16 skipped · 29 errors of 581 (pytest)
Known vulns0(pip-audit)
Repo608 files~85,633 lines of source · 10.9 MB · 5 CI workflows · tests dir

Answers from our run

Does slime build from source?

Dependencies installed in 121 seconds (191 packages), and the build succeeded in 6 seconds. We cloned commit 624b824 into a clean Debian container with 3 CPUs and no project-specific setup.

Do slime's tests pass?

Not all of them: 526 of 581 passed and 26 failed when we ran the project's own test command (pytest), with 29 collection errors. Some failures need services or credentials a bare container does not have.

Does slime have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use slime?

Developers seeking a laptop RL tutorial: the quick start uses GPU containers, model and dataset downloads, Megatron conversion, Ray, and multi-GPU layouts.

What are the alternatives to slime?

veRL, OpenRLHF, TRL. Our slime run installed 191 packages and built successfully, but pytest finished with 26 failures and 29 errors after 392 seconds.

Setup2/55,866 MB install plus model conversion and multi-GPU configuration
Docs4/5Detailed setup, debugging, topology, and customization guides
Community5/58,260 stars with same-day pushes and active technical discussion
Maturity3/5Major training use, but our suite had 26 failures and 29 errors

Who it’s for

Model-training teams already operating multi-GPU infrastructure, Ray, Megatron, and SGLang.
Researchers running GRPO, PPO, GSPO, or custom reward workflows on supported GLM, Qwen, DeepSeek, and Llama families.
Teams training coding, search, tool-using, or multi-agent behavior through custom rollout functions.
Engineers who want direct access to SGLang and Megatron controls instead of a backend-neutral wrapper.

Who it’s NOT for

Developers seeking a laptop RL tutorial: the quick start uses GPU containers, model and dataset downloads, Megatron conversion, Ray, and multi-GPU layouts.
Teams requiring interchangeable inference backends: the README says slime intentionally chooses SGLang as its single rollout backend.
Operators who require all accelerators to receive the same validation: H100 and H200 paths have stated CI coverage, while B-series lacks that protection and AMD follows separate instructions.
Buyers who require a clean general-purpose test run on fresh Debian: our suite had 26 failures and 29 collection or setup errors.
Long-context teams unable to debug low-level memory and numerical faults: open issues 2253, 2209, and 2201 concern OOM, NaN weights, and nonfinite gradients in specific large-model configurations.

Setup reality

Our sandbox installed 191 Python packages in 121 seconds and used 5,866 MB. The build passed in 6 seconds. Tests failed after 392 seconds: pytest reported 526 passed, 26 failed, 16 skipped, and 29 collection or setup errors of 581. Pip-audit found 0 known vulnerabilities.

A training run still needs model weights, datasets, Hugging Face or ModelScope access where required, Megatron-format conversion, reward logic, and optional W&B credentials. External rollout engines, sandboxes, and multi-node storage add more addresses and secrets.

The recommended route is the project's GPU Docker image. Documented examples use NVIDIA GPUs, 16 GB shared memory, Ray placement, and manual model parameters; native installation owns the PyTorch, CUDA, SGLang, Megatron, Ray, and kernel compatibility mix.

slime joins Megatron training to SGLang rollouts

slime coordinates the loop that turns generated responses into model updates. Megatron trains the policy, SGLang produces rollouts, a data buffer carries prompts and results, and Ray places workers across the cluster. Custom generation and reward functions can add search, tools, sandboxes, verifiers, or multi-agent interactions. The repository also covers checkpoint conversion, weight synchronization, evaluation, tracing, and recovery, which are the unglamorous parts that determine whether a long training job can continue.

The framework is intentionally narrow about engines. Every installed SGLang argument can pass through with an --sglang- prefix, and Megatron parameters stay visible on the training side. That gives specialists access to routing, caching, parallelism, checkpointing, and memory controls without waiting for a generic adapter. Teams standardized on vLLM or another rollout server should consider a different framework, because backend interchangeability is not the design goal stated in the README.

The quick start assumes a prepared GPU cluster

The recommended setup pulls the project's Docker image and starts it with all GPUs, host IPC, and 16 GB of shared memory. Users then download model weights and datasets, convert Hugging Face checkpoints into Megatron's distributed format, source a model-specific parameter script, and launch through Ray. The guide's standard example assigns 4 GPUs to training and another 4 to rollouts; its colocated example shares 8 GPUs and requires memory tuning.

Model scripts reduce typing but do not remove validation work. The guide tells users to compare rotary settings with the exact model and notes that checkpoint conversion may need a manually supplied vocabulary size. Batch sizes must keep generated and consumed sample counts equal. The FAQ covers stuck placement groups, out-of-memory failures, port conflicts, bad stop tokens, garbled output, and NaN gradients because those are credible operating conditions for this stack.

What happened when we ran it

Our sandbox installed 191 Python packages in 121 seconds, consuming 5,866 MB on disk. The build then passed in 6 seconds. The checkout at commit 624b824 had 608 files, roughly 85,633 lines of source, and a 10.9 MB working tree before dependencies. Pip-audit found 0 known vulnerabilities in the installed environment. These checks used Python 3.12 on Debian with 3 CPUs and 8 GB of RAM, without secrets.

Tests failed with exit code 1 after 392 seconds. Pytest reported 526 passed, 26 failed, 16 skipped, and 29 collection or setup errors of 581. The log tail named a Qwen test with megatron.core.__spec__ is None. Several rollout-data tests raised AttributeError because the loaded ray module had no get attribute. The log does not establish why those conditions occurred, so our finding stops at the failed suite in the stated sandbox.

SGLang focus buys control and creates dependency coupling

Release v0.3.1 added memory work for colocated training, faster checkpoint saving, external rollout engines, delta weight synchronization, sampling alignment, and fused PPO calculations. It also removed Megatron Bridge and points users who need broader model support toward Miles. The same release includes coding-agent work that can select Claude Code or Codex harnesses. That integration is useful for training agents on test-based rewards, but it adds sandboxes and long message histories to an already complex system.

Direct access to upstream switches reduces wrapper friction, while upgrades remain a compatibility problem shared across slime, SGLang, Megatron, PyTorch, CUDA, Ray, and model kernels. Open issue 2332 gives a small example: W&B 0.28.2 removed a function used during initialization, so training can fail before it starts. A pull request was already open on August 26, 2026, which shows fast response and the cost of moving dependencies.

Current reports include OOM and nonfinite model states

Issue 2253 reports full-vocabulary FP32 logits consuming about 10.8 GB for one roughly 11,600-token microbatch in a stated Qwen3.5 configuration. Issue 2209 reports NaN or infinite rollout weights after delta synchronization on a Qwen3.5 122B mixture-of-experts model, reproduced in 4 of 4 runs by the reporter. These are specific configurations and reporter measurements, not results from our sandbox, but both concern expensive failures during the exact workloads slime targets.

Issue 2201 isolates nonfinite gradients to a GLM-5.2 SparseMLA backward kernel on a 16-node H200 setup. The reporter produced finite reference gradients and a slow fallback, then asked for a supported safer path. A team adopting slime should keep numerical checks enabled, save enough state for reproduction, and stage new kernels or weight-transfer modes. Silently replacing bad values or disabling checks may let a job continue while invalidating the training result.

Active development does not make this beginner infrastructure

GitHub recorded 8,260 stars, 450 combined open issues and pull requests, and a last push on August 26, 2026. Release v0.3.1 shipped on August 6. The large open queue includes detailed questions, fixes, performance work, and reports across many models, so it signals both use and a wide support burden. Apache 2.0 permits commercial adoption, and the project maintains English documentation alongside Chinese material.

slime is worth evaluating for teams already committed to Megatron and SGLang, especially when custom agent rollouts need tight control over training data and weight updates. Our 526 passing tests show meaningful coverage, while the 26 failures and 29 errors rule out calling this checkout clean in a general CPU sandbox. Budget for GPUs, conversion, observability, reproducibility work, and engineers who can read the failure below the framework layer.

Alternatives

ProjectWhat it isPick it when
veRL gh↗A distributed language-model RL framework with several worker and rollout arrangements.pick this instead when backend choice or an existing veRL worker setup matters more than slime's SGLang focus.
OpenRLHFA Ray-based RLHF framework covering supervised tuning, reward models, PPO, and newer policy methods.pick this instead when a broader RLHF workflow matters more than close Megatron and SGLang coupling.
TRLHugging Face trainers for post-training models through the Transformers ecosystem.pick this instead when approachability and Hugging Face conventions matter more than distributed engine control.

What people are saying

  1. [github-trending] THUDM/slime

Sources

  1. slime README
  2. slime quick start
  3. slime v0.3.1 release
  4. Full-vocabulary logits OOM report
  5. Delta weight sync NaN report
  6. GLM-5.2 nonfinite gradients report

More llm tools reviews

agent-toolkit-for-aws · agent-memory · codex-astra-luna-orchestrator · okf-agent-memory · mlc-llm · awesome-openclaw-skills · the whole board →