mrkeyoor.com_
Tue 29 Sept 18:30 UTC
AI Toolsevaluationupdated 26 Aug 2026

Soup review

Soup is a Python command-line toolkit for fine-tuning, evaluating, and packaging language models from one YAML recipe. It tries to replace a pile of training scripts and infrastructure chores with a guided workflow that can run locally, including an experimental layer-streaming mode for GPUs with very little memory.

+421stars / 7d
Verdict

Our Soup install used 5,370 MB and its test suite reached only 31% before a 900-second timeout, so contributors should budget for a large environment and a long verification loop. Try it for a guided local path from dataset to evaluated adapter, starting with an ordinary LoRA job. Prove layer streaming, MLX, distributed recipes, and MCP execution separately on the hardware that will run them.

We ran it

Lab card: what happened when we ran SoupScreenshot of Soup (trysoup.dev)
Install✓ · 86s146 packages · 5370 MB
Build✓ · 7s
Tests✗ timed out · 900sran, no count parsed
Known vulns0(pip-audit)
Repo1029 files~373,342 lines of source · 20 MB · 5 CI workflows · Dockerfile · tests dir

Answers from our run

Does Soup build from source?

Dependencies installed in 86 seconds (146 packages), and the build succeeded in 7 seconds. We cloned commit e7fe9b2 into a clean Debian container with 3 CPUs and no project-specific setup.

Do Soup's tests pass?

We could not finish them: the suite was still running after 15 minutes in our container.

Does Soup have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use Soup?

Python 3.13 or newer environments: the README supports only Python 3.10 through 3.12 because that PyTorch stack has not been validated.

What are the alternatives to Soup?

Axolotl, LLaMA-Factory, Unsloth. Our Soup install used 5,370 MB and its test suite reached only 31% before a 900-second timeout, so contributors should budget for a large environment and a long verification loop.

Setup3/586-second install used 5,370 MB; build passed in 7 seconds
Docs5/5Detailed guides, recipes, caveats, measurements, and troubleshooting
Community4/5Pushed August 26 with 70 combined issues and pull requests
Maturity3/5Core workflow is useful; several advanced paths still carry sharp edges

Discussed on

  1. hnShow HN: Fine-tune an 8B model on a 4 GB laptop GPU139 points
  2. hnShow HN: A Soup-inspired runtime for a real fruit-fly connectome =)4 points

Who it’s for

Developers who want a guided path from training data to a LoRA adapter without assembling the Hugging Face stack by hand.
Researchers who value reproducible configs, built-in evaluation gates, and published measurement records.
Laptop GPU owners willing to test the beta layer-streaming path on their own model and sequence length.
Small teams that need training, chat, export, serving, data tools, and MCP access in one Python project.

Who it’s NOT for

Python 3.13 or newer environments: the README supports only Python 3.10 through 3.12 because that PyTorch stack has not been validated.
Teams that need a proven multi-GPU recipe without modification: issue 350 reports that the shipped 70B FSDP2 recipe cannot run because two required dtype settings cannot both be expressed in config.
Apple Silicon users expecting independent production proof for the MLX path: issue 23 says the real training loop lacks an end-to-end Apple Silicon test.
Anyone treating the 4 GB claim as a guarantee for long contexts: layer streaming is labeled beta, and issue 395 says its memory estimate under-predicts beyond roughly 4,700 tokens.
Contributors who need a quick full-suite result on modest CPU hardware: our 3-CPU test run reached only 31% before the 900-second cap.

Setup reality

Our install succeeded in 86 seconds, adding 146 packages and occupying 5,370 MB. The build passed in 7 seconds. Tests reached 31% and then timed out at 900 seconds; the log tail showed continuing pytest progress, not a named failure. Pip-audit found 0 known vulnerabilities.

Training needs Python 3.10 through 3.12, the train extra, model access, a valid dataset, substantial storage and system RAM, and usually a CUDA GPU. CPU training is described as experimental and very slow.

The 1,029-file checkout includes Docker, Compose, 5 CI workflows, and tests. Layer streaming, MLX, FSDP, DeepSpeed, serving, and export add separate extras and hardware assumptions. Pin versions and run a representative smoke job before committing expensive training time.

One YAML file follows a model from data to release

Soup packages language-model fine-tuning into a command-line workflow. A YAML file names the base model, dataset, task, LoRA settings, quantization, and output. Commands then cover data inspection, training, chatting, evaluation, merging, export, serving, and publishing artifacts. The useful idea is one reviewable recipe rather than another notebook with hidden state.

The audience is a developer who understands fine-tuning but does not want to wire Transformers, PEFT, TRL, dataset loaders, quantization, logging, and export for each experiment. Soup includes templates for supervised training, preference methods, reasoning, vision, audio, and embeddings. soup ship compares a tuned model with its base on a task evaluation and bundled regression suites, then records a ship or do-not-ship result.

Layer streaming fits an 8B model into a documented 4 GB trial

Soup's distinctive feature is layer streaming. Rather than keep a frozen base on the GPU, it feeds decoder layers to the device while training the adapter. The README reports 3.32 GB peak GPU memory for an 8B model on an RTX 3050 Laptop and labels the mode beta. It also says the published laptop speed predates a correctness repair and has not been rerun on that 4 GB card.

The repository includes a memory-capped notebook, measurement records, and withdrawn readings rather than presenting one chart as universal proof. Still, issue 395 says the preflight formula under-predicts memory beyond roughly 4,700 tokens. Sequence length, batch shape, quantization, architecture, host RAM, disk speed, and preference loss all affect fit. Run a short job with the real recipe before paying for a long one.

What happened when we ran it

Our sandbox installed commit e7fe9b2 in 86 seconds, adding 146 packages and consuming 5,370 MB. The build completed in 7 seconds. Pip-audit found 0 known vulnerabilities. The machine was an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets or GPU.

Pytest did not finish within the 900-second cap. Its output kept printing passing progress markers and had reached 31% when the harness stopped it. No failed test name or traceback appeared in the supplied tail, so the result supports only one conclusion: the complete suite did not finish on our CPU allocation within 15 minutes.

The checkout itself was 20 MB, with 1,029 files and about 373,342 source lines. We found 5 CI workflow files, a Dockerfile, Compose, and a tests directory. Buildability was good; verification time was the expensive part. A contributor using similar CPU resources needs a longer test window or a documented narrower test target.

Training still needs models, data, storage, and a supported backend

The bare soup-cli install handles configuration and data tools without PyTorch. soup-cli[train] adds the training stack, while separate extras cover serving, evaluation, MLX, DeepSpeed, and other paths. The initializer, templates, automatic batch sizing, soup doctor, and a published container remove setup chores.

Python support stops at 3.12. Model weights may require credentials and a large disk allocation. Evaluation needs held-out data and enough compute to run the base and adapter. CPU training is described as experimental and very slow, which agrees with our 900-second partial test run but does not measure training speed. Pin Soup and dependency versions because releases frequently change recipe validation and scoring behavior.

Advanced backends need separate acceptance tests. Issue 23 records the absence of an end-to-end MLX training-loop test on Apple Silicon. Issue 350 says a bundled 70B FSDP2 recipe cannot express both dtype settings needed by its 4-bit base and LoRA adapter. Those reports do not indict the ordinary CUDA LoRA path; they mean a recipe listing is not hardware certification.

MCP execution uses confirmation tokens for expensive actions

Soup can serve an OpenAI-compatible API, export several formats, track experiments, prepare data, and expose MCP tools. The server begins with read or planning access. Training and export execution require an explicit flag and a single-use server-generated confirmation token. Config snapshots and protected-path digests reduce the chance that a planned model is swapped before execution.

The MCP tag matters because this is not a passive documentation server. A successful call can start a costly, long-running job. Put it behind authentication, limit filesystem and model-hub credentials, preserve run state outside the client session, and keep an independent scheduler or process supervisor. Confirmation is a useful control, but it does not replace resource quotas and job recovery.

An August 26 push follows release v0.73.3

GitHub recorded the last push on August 26, 2026. Release v0.73.3 arrived on August 18, and GitHub showed 3,102 stars with 70 open issues and pull requests combined when fetched. The release notes say all 24 included pull requests came from contributors other than the maintainer and detail fixes for silent masking, configuration fields that did nothing, MLX detection, and MCP run state.

The documentation is unusually direct about supported Python versions, beta status, corrections, optional extras, and hardware assumptions. That earns Soup a trial for a small team that wants a path from dataset to evaluated adapter. Our 5,370 MB install and unfinished suite also say this is a large ML project. Adopt one backend at a time, and keep the exact evidence produced by the recipe you ship.

Alternatives

ProjectWhat it isPick it when
AxolotlA configuration-driven toolkit for fine-tuning many language-model families and training methods.pick this instead when you want a widely used training-focused project and are comfortable owning more of the configuration surface.
LLaMA-Factory gh↗A broad fine-tuning framework with a web interface, many model families, and numerous training recipes.pick this instead when model coverage and a browser-based workflow matter more than Soup's opinionated train-to-ship path.
Unsloth gh↗A training library focused on reducing memory use and increasing fine-tuning speed.pick this instead when training efficiency is the main goal and you are willing to build evaluation and release checks separately.

What people are saying

  1. [github-trending] MakazhanAlpamys/Soup
  2. [producthunt] Soup CLI

Sources

  1. Soup repository and README
  2. Soup v0.73.3 release
  3. FSDP2 four-bit recipe failure, issue 350
  4. MLX end-to-end test gap, issue 23
  5. Layer-streaming memory under-prediction, issue 395

More ai tools reviews

voltagent · InferenceX · Bonsai-demo · qwen-audio-agent · wechat-intelligence-hub · dlss5-visual-enhancer · the whole board →