mrkeyoor.com_
Mon 28 Sept 15:19 UTC
LLM Toolsevaluationupdated 28 Sept 2026

TensorFold review

TensorFold is a local language-model server for a small set of large models on Apple Silicon and NVIDIA GPUs. It exposes OpenAI-style chat and completion routes while using model-specific kernels and speculative decoding to produce the same tokens as its own serial path.

Verdict

Our TensorFold run installed 51 packages and built in 25 seconds combined, but its test step finished with 75 failures and 2 errors in a Debian container without MLX. Try it only when you own the specified Apple or NVIDIA hardware and one of its named checkpoints is exactly what you want to serve. Do not treat its OpenAI-shaped routes as full OpenAI behavior until your tool-calling and reasoning cases pass.

We ran it

Lab card: what happened when we ran TensorFoldScreenshot of TensorFold (github.com/ashhart/TensorFold)
Install✓ · 22s51 packages · 129 MB
Build✓ · 3s
Tests✗ · 192s777 passed · 75 failed · 157 skipped · 2 errors of 854 (pytest)
Known vulns0(pip-audit)
Repo476 files~73,794 lines of source · 11 MB · 0 CI workflows · tests dir

Answers from our run

Does TensorFold build from source?

Dependencies installed in 22 seconds (51 packages), and the build succeeded in 3 seconds. We cloned commit 34bae79 into a clean Debian container with 3 CPUs and no project-specific setup.

Do TensorFold's tests pass?

Not all of them: 777 of 854 passed and 75 failed when we ran the project's own test command (pytest), with 2 collection errors. Some failures need services or credentials a bare container does not have.

Does TensorFold have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use TensorFold?

CPU-only Linux users: the runbook requires Apple Silicon for MLX or a supported NVIDIA CUDA environment.

What are the alternatives to TensorFold?

vLLM, llama.cpp, MLX LM. Our TensorFold run installed 51 packages and built in 25 seconds combined, but its test step finished with 75 failures and 2 errors in a Debian container without MLX.

Setup2/5Build passed, but hardware-specific serving needs careful setup
Docs4/5Runbook names model, memory, backend, and API limits
Community3/5538 stars with active issues and pull requests on September 28
Maturity2/5Alpha package with 75 failures in our platform-neutral run

Who it’s for

Apple Silicon owners who want an MLX server tuned for one of TensorFold's named model families.
NVIDIA operators willing to use the prescribed PyTorch container and compile CUDA extensions for the installed GPU.
Developers who need an OpenAI-compatible local endpoint and care about drafted output matching the same engine's serial output.
Inference engineers prepared to test checkpoint layout, memory admission, context limits, and tool behavior on their own workload.

Who it’s NOT for

CPU-only Linux users: the runbook requires Apple Silicon for MLX or a supported NVIDIA CUDA environment.
Teams that need arbitrary Hugging Face models to load by name: TensorFold documents a fixed family table and refuses unsupported formats.
Agent systems that depend on tool_choice: required: open issue 52 says the server accepts that value without enforcing a tool call.
Workflows that cannot tolerate a missed tool call inside reasoning: issue 60 reports complete tool markup being returned as reasoning and an empty turn when the think block stays open.
Buyers who require a platform-neutral green test suite before evaluation: our Debian run ended with 75 failures and 2 errors, including repeated missing-MLX errors.

Setup reality

Our sandbox installed 51 packages in 22 seconds and used 129 MB. The build succeeded in 3 seconds. Tests exited 1 after 192 seconds: pytest reported 777 passed, 75 failed, 157 skipped, and 2 collection or setup errors in the harness summary of 854. Pip-audit found 0 known vulnerabilities.

Serving a model requires Python 3.11 or newer plus either Apple Silicon with MLX 0.32.2 or newer, or a supported NVIDIA environment. The CUDA route uses NVIDIA's PyTorch container rather than a package extra. You must choose and download a supported checkpoint; two-rank jobs need matching files and settings on both hosts.

The failing log repeatedly showed ModuleNotFoundError: No module named 'mlx' in resume, snapshot, HTTP tool, sampling, and checkpoint tests. Two chunk-resume assertions also returned empty lists. The log tail does not explain every failure or the 2 errors. No CI workflow or Dockerfile was present, although the 476-file repository has a tests directory.

TensorFold serves seven named model configurations, not any checkpoint

TensorFold's README lists 7 model rows across Nemotron, Qwen, GLM, and Gemma families. Each row names a checkpoint, backend, quantization assumptions, and drafting method. That specificity is the product. The server is not a general loader that happens to run faster on some machines. It carries custom kernels and verification rules for the exact layouts it understands, and it refuses unsupported Gemma layouts before downloading them. Pick the checkpoint first, then decide whether TensorFold fits.

The interface is familiar enough to test without rewriting a client. TensorFold listens on 127.0.0.1:8080 by default and exposes models, health, chat completions, and raw completions routes. Python 3.11 or newer is required. Apple machines use MLX 0.32.2 or newer, while the NVIDIA route starts inside NVIDIA's PyTorch container and compiles extensions for the GPU present. Model downloads and their licenses remain separate from TensorFold's MIT license.

Exact drafting is the reason to choose this narrower server

Speculative decoding normally asks a cheaper draft path to propose tokens and lets the main model verify them. TensorFold makes a stricter promise: a draft is accepted only when it matches what the same engine would have produced serially. The comparison holds only when the weights, runtime, sampling settings, prompt, and hardware path stay fixed. The README explicitly says it does not promise identical output between MLX and CUDA or between different quantizations.

That boundary is honest and useful. A developer can send the same request with drafting enabled and with draft: false, then compare the results under one engine. TensorFold also checks concurrent MLX streams against their solo output and limits shared execution where a family cannot reproduce serial arithmetic. We did not run a supported model or measure generation speed in our 3-CPU sandbox, so this review makes no claim about tokens per second, prompt speed, or memory use.

What happened when we ran it

Our sandbox installed TensorFold in 22 seconds, adding 51 packages and using 129 MB. The package build succeeded in 3 seconds. The checkout contained 476 files, about 73,794 lines of source, and 11 MB before installation. Pip-audit reported 0 known vulnerabilities. Those results show that its Python package can install and build on a fresh Debian host even when that host cannot provide either documented acceleration backend.

The test step failed with exit code 1 after 192 seconds. Pytest reported 777 passed, 75 failed, 157 skipped, and 2 collection or setup errors in the supplied summary of 854. The log tail repeatedly showed ModuleNotFoundError: No module named 'mlx' in resume-memory, snapshot, HTTP tool, sampling, and unsupported-checkpoint cases. Two resume-chunk tests also expected populated results and received empty lists. The tail does not identify the cause of all 75 failures or either error.

The repository has a tests directory but no CI workflow file and no Dockerfile. That makes our mixed result harder to interpret as a release gate. The pyproject.toml installs MLX only on Darwin, while the runbook says Linux serving needs a supported NVIDIA CUDA environment. A plain Debian runner therefore exercises plenty of Python logic but is not one of the documented serving targets. The right follow-up is a Mac or GPU run, not a claim that the server itself cannot work.

Memory admission protects the host but narrows usable context

On MLX, the default process budget is capped at 70% of RAM and the GPU's recommended working set, with model-family exceptions documented for GLM. TensorFold reserves room for weights, cache growth, reply tokens, and prefill work before accepting a request. It may evict a retained prefix or queue another stream when memory is tight. A model fitting in storage or memory does not establish that your intended prompt and reply will fit together.

The public memory table still marks its context and peak-footprint cells as TBD. That is better than filling a matrix with estimates, but it pushes qualification onto the buyer. The CUDA defaults also differ by family, and some 2-rank modes need identical checkpoints, context settings, networking, and drafting choices on both machines. Before adoption, test your actual prompt lengths, reply limits, concurrency, restart behavior, and cache persistence on the hardware you plan to keep.

OpenAI-compatible routes do not cover every agent behavior

The API accepts text chat, streaming, tools, reasoning fields, and common sampling controls. It rejects image, audio, and video requests with HTTP 400. Several fields are backend-specific: reasoning effort and thinking budgets are MLX features, while CUDA does not enforce all MLX-only controls. Multiple choices through n and log probabilities are unsupported. Compatibility here means a useful request shape, not full behavioral equivalence with OpenAI's service.

Two open reports matter for agent workloads. Issue 52 says tool_choice: required is accepted while the model may still answer with plain text. Issue 60 describes a complete tool call inside an unclosed think block being placed in reasoning_content, leaving content and tool_calls empty. Both were filed on September 28, 2026. If tools can change files or infrastructure, replay your own traces and reject empty or noncompliant turns at the client boundary.

Same-day release activity is strong, while the package remains alpha

GitHub showed 538 stars, 54 forks, and 23 open issues and pull requests on September 28, 2026. The repository was pushed that day, and v0.3.6.1 was also released that day with a CUDA container build fix. Several open pull requests touched CUDA packaging, scheduling, caching, and model support. That is active maintenance, but the package classifier still says Alpha and the pace means behavior can move quickly between tags.

TensorFold earns a trial when its short supported list matches your hardware and model choice. The 22-second install and 3-second build make that trial cheap at the package level, but our failed suite and the open tool-call reports argue for a workload replay before deployment. Pin the release, checkpoint revision, container, context, and sampling settings together. Without those details, even an exact decoder comparison is measuring a different system.

Alternatives

ProjectWhat it isPick it when
vLLM gh↗A widely used GPU inference server focused on throughput and broad model support.pick this instead when ecosystem breadth, production integrations, and multi-user serving matter more than TensorFold's narrow exact-drafting design.
llama.cpp gh↗A portable C and C++ runtime for running many quantized models locally.pick this instead when CPU support, hardware portability, or the GGUF ecosystem matters more than these specific MLX and CUDA kernels.
MLX LMApple's MLX toolkit for generating, fine-tuning, and serving language models on Apple Silicon.pick this instead when you want the upstream Apple-focused toolkit and broader experimentation rather than TensorFold's selected server families.

What people are saying

  1. [github-trending] ashhart/TensorFold

Sources

  1. TensorFold README
  2. TensorFold installation runbook
  3. TensorFold API fields
  4. TensorFold v0.3.6.1 release
  5. Issue 52: required tool choice is not enforced
  6. Issue 60: tool call dropped inside thinking

More llm tools reviews

ai-evaluation-framework · llm-wiki-compiler · claude-skills · Humanizer-zh · agent-beacon · MiMo-Code · the whole board →