Python v0.1.0 supports Apple Silicon, not general Linux
Edge0's Python package is an MLX inference engine for macOS on Apple Silicon. Version 0.1.0 loads one of two released sparse mixture-of-experts checkpoints, streams expert weights from SSD, applies a trained routing predictor, and keeps LoRA recovery adapters separate from the quantized base. The command-line interface can run a prompt or expose /v1/chat/completions. Python 3.10 or newer is required, with MLX 0.30.6 and mlx-lm 0.31.0 pinned exactly.
The narrow platform support is easy to miss because the top-level repository also has native iOS, macOS, Android, and Windows directories. For the Python framework, the README requirements name M1 through M4 hardware and say no other platform is currently supported. CUDA has a reserved backend directory and an open implementation pull request, while the published package still resolves to MLX. Linux users should read that as unavailable, not as a routine pip portability issue.
SSD streaming moves the constraint to storage and cache behavior
The 4-bit Edge0 checkpoints are about 4.2 GB for the smaller tier and 23 GB for the larger one. Expert tensors are memory-mapped and read when routing selects them, so the whole model need not sit in active accelerator memory. A trained prerouter predicts which experts will be needed on the next step, letting reads overlap computation. LoRA adapters seek to recover quality lost during quantization without merging changes into the base weights.
That design makes the machine around the model part of the result. SSD latency, page cache, physical memory, prompt length, and routing pattern affect what a user experiences. Issue 106 argues that the README's peak active-memory figure is the MLX allocator high-water mark rather than total process memory, and names allocations it excludes. Capacity-plan from process measurements on your own device, not the headline active-memory number.
What happened when we ran it
Our sandbox installed commit 88f121d in 26 seconds, pulling 61 packages and using 197 MB on disk. The build succeeded in 4 seconds. Pytest exited with status 1 after 6 seconds: 4 tests passed, 4 were deselected, and 7 modules stopped during collection or setup. Pip-audit reported 0 known vulnerabilities. The checkout contained 1,125 files and roughly 163,870 lines of source.
Every reported error ended with ImportError: libmlx.so: cannot open shared object file: No such file or directory. The affected files covered CLI flags, model specifications, prefill hooks, the registry, sampling, the server, and streaming math. The log does not show failed assertions in those 7 modules, and it does not identify a package fix. Our Debian container was outside the documented Apple Silicon Python platform, so the defensible result is simply that the full suite did not collect there.
The OpenAI-style server ignores tool definitions
Edge0 can serve chat completions on port 8000, including streaming responses and reasoning content. That endpoint is useful for plain local chat, but OpenAI compatibility has a material gap. Issue 115 demonstrates that a request containing tools and tool_choice is accepted while the definition is ignored, and the response never contains tool_calls. An open pull request proposes the missing behavior.
That gap rules out drop-in use for many agent loops. A client may receive an ordinary text answer where it expected a structured function request, without a clear unsupported-feature error. Another open pull request addresses malformed streamed text when a character spans multiple byte tokens. Neither report came from our sandbox, but both concern the HTTP boundary buyers are likely to treat as familiar because it uses an OpenAI-shaped route. Test your exact client features before migration.
Four native engines do not yet provide one interface
commit 88f121d reorganized Edge0 as a multi-platform repository with separate engines for iOS, macOS, Android, and Windows. Their stacks differ: Swift and MLX Swift on iOS, Rust on macOS, Kotlin plus native code on Android, and C++ with Vulkan on Windows. That is more platform work than the original Python package suggests, and each directory needs its own build and runtime review. Our 4-second build covered the measured Python project, not all four native applications.
The README places a unified access layer in its Q4 2026 roadmap. Treat that as intended work, since the current repository still directs readers to platform-specific instructions. The same caution applies to CUDA and later model architectures. A directory placeholder or roadmap entry does not give today's Python client another backend. Choose the engine that exists for your device, then verify its model format, API surface, and app permissions independently.
October 1 activity is current, while 26 items remain open
The repository was pushed on October 1, 2026, one day before this review. GitHub showed 2,280 stars and 26 open issues and pull requests, split evenly between 13 issues and 13 pull requests. There is one CI workflow, a Python tests directory, and active work on CUDA, tool calling, documentation corrections, streaming output, and model-conversion scripts. The latest-release API returned no tagged release.
Edge0 is worth examining for one specific reason: SSD-streamed MoE inference paired with trained routing prediction. The code, models, paper, and platform experiments give researchers plenty to inspect. A production buyer should require a green suite on the intended device, full process-memory measurements, client-level server tests, and a pinned model directory before adoption. Our 7 collection errors make Linux Python an especially clear no today.

