One YAML file follows a model from data to release
Soup packages language-model fine-tuning into a command-line workflow. A YAML file names the base model, dataset, task, LoRA settings, quantization, and output. Commands then cover data inspection, training, chatting, evaluation, merging, export, serving, and publishing artifacts. The useful idea is one reviewable recipe rather than another notebook with hidden state.
The audience is a developer who understands fine-tuning but does not want to wire Transformers, PEFT, TRL, dataset loaders, quantization, logging, and export for each experiment. Soup includes templates for supervised training, preference methods, reasoning, vision, audio, and embeddings. soup ship compares a tuned model with its base on a task evaluation and bundled regression suites, then records a ship or do-not-ship result.
Layer streaming fits an 8B model into a documented 4 GB trial
Soup's distinctive feature is layer streaming. Rather than keep a frozen base on the GPU, it feeds decoder layers to the device while training the adapter. The README reports 3.32 GB peak GPU memory for an 8B model on an RTX 3050 Laptop and labels the mode beta. It also says the published laptop speed predates a correctness repair and has not been rerun on that 4 GB card.
The repository includes a memory-capped notebook, measurement records, and withdrawn readings rather than presenting one chart as universal proof. Still, issue 395 says the preflight formula under-predicts memory beyond roughly 4,700 tokens. Sequence length, batch shape, quantization, architecture, host RAM, disk speed, and preference loss all affect fit. Run a short job with the real recipe before paying for a long one.
What happened when we ran it
Our sandbox installed commit e7fe9b2 in 86 seconds, adding 146 packages and consuming 5,370 MB. The build completed in 7 seconds. Pip-audit found 0 known vulnerabilities. The machine was an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets or GPU.
Pytest did not finish within the 900-second cap. Its output kept printing passing progress markers and had reached 31% when the harness stopped it. No failed test name or traceback appeared in the supplied tail, so the result supports only one conclusion: the complete suite did not finish on our CPU allocation within 15 minutes.
The checkout itself was 20 MB, with 1,029 files and about 373,342 source lines. We found 5 CI workflow files, a Dockerfile, Compose, and a tests directory. Buildability was good; verification time was the expensive part. A contributor using similar CPU resources needs a longer test window or a documented narrower test target.
Training still needs models, data, storage, and a supported backend
The bare soup-cli install handles configuration and data tools without PyTorch. soup-cli[train] adds the training stack, while separate extras cover serving, evaluation, MLX, DeepSpeed, and other paths. The initializer, templates, automatic batch sizing, soup doctor, and a published container remove setup chores.
Python support stops at 3.12. Model weights may require credentials and a large disk allocation. Evaluation needs held-out data and enough compute to run the base and adapter. CPU training is described as experimental and very slow, which agrees with our 900-second partial test run but does not measure training speed. Pin Soup and dependency versions because releases frequently change recipe validation and scoring behavior.
Advanced backends need separate acceptance tests. Issue 23 records the absence of an end-to-end MLX training-loop test on Apple Silicon. Issue 350 says a bundled 70B FSDP2 recipe cannot express both dtype settings needed by its 4-bit base and LoRA adapter. Those reports do not indict the ordinary CUDA LoRA path; they mean a recipe listing is not hardware certification.
MCP execution uses confirmation tokens for expensive actions
Soup can serve an OpenAI-compatible API, export several formats, track experiments, prepare data, and expose MCP tools. The server begins with read or planning access. Training and export execution require an explicit flag and a single-use server-generated confirmation token. Config snapshots and protected-path digests reduce the chance that a planned model is swapped before execution.
The MCP tag matters because this is not a passive documentation server. A successful call can start a costly, long-running job. Put it behind authentication, limit filesystem and model-hub credentials, preserve run state outside the client session, and keep an independent scheduler or process supervisor. Confirmation is a useful control, but it does not replace resource quotas and job recovery.
An August 26 push follows release v0.73.3
GitHub recorded the last push on August 26, 2026. Release v0.73.3 arrived on August 18, and GitHub showed 3,102 stars with 70 open issues and pull requests combined when fetched. The release notes say all 24 included pull requests came from contributors other than the maintainer and detail fixes for silent masking, configuration fields that did nothing, MLX detection, and MCP run state.
The documentation is unusually direct about supported Python versions, beta status, corrections, optional extras, and hardware assumptions. That earns Soup a trial for a small team that wants a path from dataset to evaluated adapter. Our 5,370 MB install and unfinished suite also say this is a large ML project. Adopt one backend at a time, and keep the exact evidence produced by the recipe you ship.

