TensorFold serves seven named model configurations, not any checkpoint
TensorFold's README lists 7 model rows across Nemotron, Qwen, GLM, and Gemma families. Each row names a checkpoint, backend, quantization assumptions, and drafting method. That specificity is the product. The server is not a general loader that happens to run faster on some machines. It carries custom kernels and verification rules for the exact layouts it understands, and it refuses unsupported Gemma layouts before downloading them. Pick the checkpoint first, then decide whether TensorFold fits.
The interface is familiar enough to test without rewriting a client. TensorFold listens on 127.0.0.1:8080 by default and exposes models, health, chat completions, and raw completions routes. Python 3.11 or newer is required. Apple machines use MLX 0.32.2 or newer, while the NVIDIA route starts inside NVIDIA's PyTorch container and compiles extensions for the GPU present. Model downloads and their licenses remain separate from TensorFold's MIT license.
Exact drafting is the reason to choose this narrower server
Speculative decoding normally asks a cheaper draft path to propose tokens and lets the main model verify them. TensorFold makes a stricter promise: a draft is accepted only when it matches what the same engine would have produced serially. The comparison holds only when the weights, runtime, sampling settings, prompt, and hardware path stay fixed. The README explicitly says it does not promise identical output between MLX and CUDA or between different quantizations.
That boundary is honest and useful. A developer can send the same request with drafting enabled and with draft: false, then compare the results under one engine. TensorFold also checks concurrent MLX streams against their solo output and limits shared execution where a family cannot reproduce serial arithmetic. We did not run a supported model or measure generation speed in our 3-CPU sandbox, so this review makes no claim about tokens per second, prompt speed, or memory use.
What happened when we ran it
Our sandbox installed TensorFold in 22 seconds, adding 51 packages and using 129 MB. The package build succeeded in 3 seconds. The checkout contained 476 files, about 73,794 lines of source, and 11 MB before installation. Pip-audit reported 0 known vulnerabilities. Those results show that its Python package can install and build on a fresh Debian host even when that host cannot provide either documented acceleration backend.
The test step failed with exit code 1 after 192 seconds. Pytest reported 777 passed, 75 failed, 157 skipped, and 2 collection or setup errors in the supplied summary of 854. The log tail repeatedly showed ModuleNotFoundError: No module named 'mlx' in resume-memory, snapshot, HTTP tool, sampling, and unsupported-checkpoint cases. Two resume-chunk tests also expected populated results and received empty lists. The tail does not identify the cause of all 75 failures or either error.
The repository has a tests directory but no CI workflow file and no Dockerfile. That makes our mixed result harder to interpret as a release gate. The pyproject.toml installs MLX only on Darwin, while the runbook says Linux serving needs a supported NVIDIA CUDA environment. A plain Debian runner therefore exercises plenty of Python logic but is not one of the documented serving targets. The right follow-up is a Mac or GPU run, not a claim that the server itself cannot work.
Memory admission protects the host but narrows usable context
On MLX, the default process budget is capped at 70% of RAM and the GPU's recommended working set, with model-family exceptions documented for GLM. TensorFold reserves room for weights, cache growth, reply tokens, and prefill work before accepting a request. It may evict a retained prefix or queue another stream when memory is tight. A model fitting in storage or memory does not establish that your intended prompt and reply will fit together.
The public memory table still marks its context and peak-footprint cells as TBD. That is better than filling a matrix with estimates, but it pushes qualification onto the buyer. The CUDA defaults also differ by family, and some 2-rank modes need identical checkpoints, context settings, networking, and drafting choices on both machines. Before adoption, test your actual prompt lengths, reply limits, concurrency, restart behavior, and cache persistence on the hardware you plan to keep.
OpenAI-compatible routes do not cover every agent behavior
The API accepts text chat, streaming, tools, reasoning fields, and common sampling controls. It rejects image, audio, and video requests with HTTP 400. Several fields are backend-specific: reasoning effort and thinking budgets are MLX features, while CUDA does not enforce all MLX-only controls. Multiple choices through n and log probabilities are unsupported. Compatibility here means a useful request shape, not full behavioral equivalence with OpenAI's service.
Two open reports matter for agent workloads. Issue 52 says tool_choice: required is accepted while the model may still answer with plain text. Issue 60 describes a complete tool call inside an unclosed think block being placed in reasoning_content, leaving content and tool_calls empty. Both were filed on September 28, 2026. If tools can change files or infrastructure, replay your own traces and reject empty or noncompliant turns at the client boundary.
Same-day release activity is strong, while the package remains alpha
GitHub showed 538 stars, 54 forks, and 23 open issues and pull requests on September 28, 2026. The repository was pushed that day, and v0.3.6.1 was also released that day with a CUDA container build fix. Several open pull requests touched CUDA packaging, scheduling, caching, and model support. That is active maintenance, but the package classifier still says Alpha and the pace means behavior can move quickly between tags.
TensorFold earns a trial when its short supported list matches your hardware and model choice. The 22-second install and 3-second build make that trial cheap at the package level, but our failed suite and the open tool-call reports argue for a workload replay before deployment. Pin the release, checkpoint revision, container, context, and sampling settings together. Without those details, even an exact decoder comparison is measuring a different system.

