mrkeyoor.com_
Fri 02 Oct 14:57 UTC
LLM Toolsevaluationupdated 02 Oct 2026

Edge0 review

Edge0 is a local inference project for sparse mixture-of-experts language models, using SSD-streamed expert weights, trained routing prediction, and LoRA adapters to reduce active memory. The repository now contains separate Python, macOS, iOS, Android, and Windows engines, but they do not yet share the unified interface described in its roadmap.

Verdict

Our Edge0 run installed 61 packages and built in 30 seconds combined, but pytest stopped with 7 collection/setup errors because libmlx.so was unavailable in the Debian sandbox. Try v0.1.0 for research or local use on its documented Apple Silicon Python path, where SSD-streamed MoE models are the specific attraction. Choose a more established runtime for Linux, tool-calling agents, or one consistent API across platforms.

We ran it

Lab card: what happened when we ran Edge0Screenshot of Edge0 (github.com/Edge0-AI/Edge0)
Install✓ · 26s61 packages · 197 MB
Build✓ · 4s
Tests✗ · 6s4 passed · 0 failed · 7 errors of 11 (pytest)
Known vulns0(pip-audit)
Repo1125 files~163,870 lines of source · 9.4 MB · 1 CI workflows · tests dir

Answers from our run

Does Edge0 build from source?

Dependencies installed in 26 seconds (61 packages), and the build succeeded in 4 seconds. We cloned commit 88f121d into a clean Debian container with 3 CPUs and no project-specific setup.

Do Edge0's tests pass?

Yes: 4 of 11 passed when we ran the project's own test command (pytest), with 7 collection errors. Some failures need services or credentials a bare container does not have.

Does Edge0 have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use Edge0?

Linux or NVIDIA users expecting the Python package to work today: its documented backend supports macOS on Apple Silicon, while CUDA remains in the roadmap and an open pull request.

What are the alternatives to Edge0?

MLX LM, llama.cpp, Ollama. Our Edge0 run installed 61 packages and built in 30 seconds combined, but pytest stopped with 7 collection/setup errors because `libmlx.

Setup2/5Install and build pass, but 7 test modules cannot load on Debian
Docs4/5Detailed model, architecture, platform, and reproduction notes
Community3/52,280 stars with 13 issues and 13 PRs open
Maturity2/5Alpha v0.1.0 with no tagged release and active interface gaps

Who it’s for

Apple Silicon developers willing to use the Python MLX stack and download an Edge0 checkpoint.
Researchers studying SSD expert offload, routing prediction, and quantized mixture-of-experts inference.
App teams prepared to evaluate a separate native engine for iOS, macOS, Android, or Windows.
Local-model users who want an OpenAI-style chat endpoint without hosted inference fees.

Who it’s NOT for

Linux or NVIDIA users expecting the Python package to work today: its documented backend supports macOS on Apple Silicon, while CUDA remains in the roadmap and an open pull request.
Agent servers that require OpenAI tool calling: issue 115 says tools is ignored and tool_calls is never returned.
Small disks: the documented 4-bit checkpoints are about 4.2 GB and 23 GB before caches or app files.
Capacity planners relying on the README's active-memory figure as total process memory: issue 106 says the metric excludes several allocations.
Conservative production teams: v0.1.0 is marked alpha, has no tagged GitHub release, and had 26 open issues and pull requests.

Setup reality

Our sandbox installed commit 88f121d in 26 seconds, adding 61 packages and using 197 MB. The build passed in 4 seconds. Tests exited 1 after 6 seconds: 4 passed, 4 were deselected, and 7 hit collection/setup errors because libmlx.so could not be opened. Pip-audit found 0 known vulnerabilities.

The Python route needs macOS on Apple Silicon, Python 3.10 or newer, and exact MLX-related pins. Running a model also requires a public checkpoint download of about 4.2 GB or 23 GB and a local path; no hosted inference credential is required after the files are present.

There is no Dockerfile. Native macOS, iOS, Android, and Windows engines live in separate directories with their own stacks. The single cross-platform access layer is a roadmap item rather than current behavior.

Python v0.1.0 supports Apple Silicon, not general Linux

Edge0's Python package is an MLX inference engine for macOS on Apple Silicon. Version 0.1.0 loads one of two released sparse mixture-of-experts checkpoints, streams expert weights from SSD, applies a trained routing predictor, and keeps LoRA recovery adapters separate from the quantized base. The command-line interface can run a prompt or expose /v1/chat/completions. Python 3.10 or newer is required, with MLX 0.30.6 and mlx-lm 0.31.0 pinned exactly.

The narrow platform support is easy to miss because the top-level repository also has native iOS, macOS, Android, and Windows directories. For the Python framework, the README requirements name M1 through M4 hardware and say no other platform is currently supported. CUDA has a reserved backend directory and an open implementation pull request, while the published package still resolves to MLX. Linux users should read that as unavailable, not as a routine pip portability issue.

SSD streaming moves the constraint to storage and cache behavior

The 4-bit Edge0 checkpoints are about 4.2 GB for the smaller tier and 23 GB for the larger one. Expert tensors are memory-mapped and read when routing selects them, so the whole model need not sit in active accelerator memory. A trained prerouter predicts which experts will be needed on the next step, letting reads overlap computation. LoRA adapters seek to recover quality lost during quantization without merging changes into the base weights.

That design makes the machine around the model part of the result. SSD latency, page cache, physical memory, prompt length, and routing pattern affect what a user experiences. Issue 106 argues that the README's peak active-memory figure is the MLX allocator high-water mark rather than total process memory, and names allocations it excludes. Capacity-plan from process measurements on your own device, not the headline active-memory number.

What happened when we ran it

Our sandbox installed commit 88f121d in 26 seconds, pulling 61 packages and using 197 MB on disk. The build succeeded in 4 seconds. Pytest exited with status 1 after 6 seconds: 4 tests passed, 4 were deselected, and 7 modules stopped during collection or setup. Pip-audit reported 0 known vulnerabilities. The checkout contained 1,125 files and roughly 163,870 lines of source.

Every reported error ended with ImportError: libmlx.so: cannot open shared object file: No such file or directory. The affected files covered CLI flags, model specifications, prefill hooks, the registry, sampling, the server, and streaming math. The log does not show failed assertions in those 7 modules, and it does not identify a package fix. Our Debian container was outside the documented Apple Silicon Python platform, so the defensible result is simply that the full suite did not collect there.

The OpenAI-style server ignores tool definitions

Edge0 can serve chat completions on port 8000, including streaming responses and reasoning content. That endpoint is useful for plain local chat, but OpenAI compatibility has a material gap. Issue 115 demonstrates that a request containing tools and tool_choice is accepted while the definition is ignored, and the response never contains tool_calls. An open pull request proposes the missing behavior.

That gap rules out drop-in use for many agent loops. A client may receive an ordinary text answer where it expected a structured function request, without a clear unsupported-feature error. Another open pull request addresses malformed streamed text when a character spans multiple byte tokens. Neither report came from our sandbox, but both concern the HTTP boundary buyers are likely to treat as familiar because it uses an OpenAI-shaped route. Test your exact client features before migration.

Four native engines do not yet provide one interface

commit 88f121d reorganized Edge0 as a multi-platform repository with separate engines for iOS, macOS, Android, and Windows. Their stacks differ: Swift and MLX Swift on iOS, Rust on macOS, Kotlin plus native code on Android, and C++ with Vulkan on Windows. That is more platform work than the original Python package suggests, and each directory needs its own build and runtime review. Our 4-second build covered the measured Python project, not all four native applications.

The README places a unified access layer in its Q4 2026 roadmap. Treat that as intended work, since the current repository still directs readers to platform-specific instructions. The same caution applies to CUDA and later model architectures. A directory placeholder or roadmap entry does not give today's Python client another backend. Choose the engine that exists for your device, then verify its model format, API surface, and app permissions independently.

October 1 activity is current, while 26 items remain open

The repository was pushed on October 1, 2026, one day before this review. GitHub showed 2,280 stars and 26 open issues and pull requests, split evenly between 13 issues and 13 pull requests. There is one CI workflow, a Python tests directory, and active work on CUDA, tool calling, documentation corrections, streaming output, and model-conversion scripts. The latest-release API returned no tagged release.

Edge0 is worth examining for one specific reason: SSD-streamed MoE inference paired with trained routing prediction. The code, models, paper, and platform experiments give researchers plenty to inspect. A production buyer should require a green suite on the intended device, full process-memory measurements, client-level server tests, and a pinned model directory before adoption. Our 7 collection errors make Linux Python an especially clear no today.

Alternatives

ProjectWhat it isPick it when
MLX LMApple's MLX toolkit for running and fine-tuning language models on Apple Silicon.pick this instead when broad MLX model support matters more than Edge0's SSD-streamed MoE recipe.
llama.cpp gh↗A widely used local inference engine for GGUF models across CPUs, GPUs, and operating systems.pick this instead when hardware coverage and model choice matter more than trained expert prediction.
Ollama gh↗A local model manager and server with a simple pull-and-run workflow.pick this instead when easy model management and a stable local API are the priority.
vLLM gh↗A production server focused on high-throughput accelerator inference.pick this instead when NVIDIA deployment and concurrent server throughput define the job.

What people are saying

  1. [hf-trending] Edge0/Audio8-ASR-Infinite (trending model on Hugging Face)
  2. [hf-trending] Edge0/Edge0-35B-A3B-preview (trending model on Hugging Face)
  3. [velocity-scout] Edge0-AI/Edge0

Sources

  1. Edge0 README
  2. Python package configuration
  3. Issue 106: active memory metric
  4. Issue 115: tool calling is ignored
  5. Commit 88f121d: multi-platform layout

More llm tools reviews

whatsapp-mcp · deepseek-recipe · ag-ui · awesome-codex-plugins · claude-style-patch · agent-toolkit-for-aws · the whole board →