A 9B model returns fixed choices instead of prose
Nimble takes text plus a flat schema, then picks one allowed answer for each field. A field can be boolean or an enum with as many as 255 choices in the latest checkpoint. The scorer reads the model's logits for one-token answer codes and turns them into probabilities, so callers do not parse generated JSON. That is a useful shape for request routing, policy checks, and ordered ratings where the permitted outcomes are known before inference.
The narrow contract is the appeal and the limit. Nimble cannot write an explanation, copy text from the source, inspect images, return nested objects, or let one field depend on another. Its latest checkpoint accepts up to 8,192 tokens, including schema text. A team that wants an open-ended agent will fight those boundaries. A team that needs HIGH, LOW, or NO_MATCH can build simpler downstream code around them.
What happened when we ran it
Our run at commit dcfdbd9 installed 35 packages in 12 seconds and occupied 37 MB. The build completed in 7 seconds. Our measurement setup was an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Pip-audit reported 0 known vulnerabilities. Those figures cover repository setup and checks, not the model download, inference speed, or decision quality.
Pytest exited with failure after 11 seconds. Of 143 tests, 125 passed, 2 failed, 4 were skipped, and 16 hit collection or setup errors. The two failures were in curation recovery and schema-tuning tests. The log names missing pydantic for the first and missing torch for the second. The 16 errors shown in the tail also include schema-training cases that could not import torch.
The log does not prove why those modules were absent, so blaming packaging, optional dependency selection, or the tests would go beyond the evidence. It does prove that the full command did not pass in our fresh container. The repository has a tests directory, yet it has 0 CI workflow files and no Dockerfile. Anyone adopting Nimble should reproduce the intended environment and decide which test groups belong in their own gate.
Local inference needs 9B-class hardware and separate environments
The README supports two local paths. Apple users need Apple Silicon, native macOS Python, and the MLX scorer. Linux users need an NVIDIA GPU with BF16 support and the CUDA scorer. Python 3.12 is specified. The unquantized 9B weights alone take about 18 GB, according to the project, with more memory needed at runtime and during adapter merging. A CPU-only Linux server is outside the documented path.
Dependency separation adds work that the 12-second base install does not reveal. Nimble documents distinct MLX, PyTorch, and curator environments because their package versions differ. Adapter releases may need the pinned base model merged on CPU, while the resolved revision and prompt contract must travel with the result. Local inference needs no TypeSafe or generation API credential, though model files still come from Hugging Face.
The 3,000 published examples make the recipe inspectable
Bespoke Labs publishes 2,676 training examples and a frozen 324-example holdout. The method creates contrastive pairs: a small factual change flips the correct label while the policy and surrounding text stay fixed. The repository also includes verification metadata, training code, and evaluation guides. That gives researchers something more useful than a model card alone: they can inspect how the classifier was shaped and replay parts of the process.
All 3,000 labels are synthetic, and the README says no person reviewed them. The holdout covers 6 source families from a broader set of 10 subject categories. Bespoke Labs reports that Nimble matched 292 of 324 reference labels in its own held-out comparison, but those reference labels came from the same synthetic pipeline. Treat that result as project evidence, then run a human-labeled set from the domain where mistakes carry a cost.
A probability threshold still needs your labeled data
The latest checkpoint uses temperature 1.0 and has no separate temperature fit. Nimble's probabilities add up across the choices supplied, but the README warns that 0.9 does not mean 90 percent accuracy on a new workload. If none of the choices may fit, the schema needs an explicit no-match option. Any threshold for automatic action should come from labeled examples that resemble real traffic.
Earlier revisions used a fitted temperature of 2.179, and changing temperature alters probabilities without changing the selected answer. The README also notes that merged weights can produce slightly different logits and were not separately checked for that older fit. Revision pinning therefore matters twice: it fixes model behavior and tells the scorer which probability treatment belongs to the checkpoint.
The October 5 push is recent, but licensing is unresolved
GitHub showed 2,062 stars on October 6, 2026, less than 3 weeks after the repository was created. The last push was October 5. Its 4 open items consisted of 3 issues and 1 pull request, including an AMD GPU request and a license question. That is active early interest rather than evidence of a settled support record.
There is no tagged GitHub release, and GitHub's license field is empty. Open issue 14 points out that the Hugging Face model card says Apache-2.0 while the source repository has no LICENSE file. Until the maintainer answers, a company should not assume that the model's license automatically covers every file in the training and serving recipe.
Nimble is worth a controlled evaluation because its fixed-choice interface removes a lot of parsing and prompting machinery. Our 143-test result keeps it out of the drop-in category. The best buyer has supported hardware, a narrow decision schema, and enough labeled cases to challenge both the answer and its probability before either can trigger an automated action.

