mrkeyoor.com_
Wed 30 Sept 20:36 UTC
LLM Toolsevaluationupdated 26 Aug 2026

llama-swap review

llama-swap is a local reverse proxy that starts the right AI model server when a request names a model and stops conflicting servers when memory is needed elsewhere. It gives OpenAI- and Anthropic-compatible clients one address for text, embeddings, audio, images, reranking, and several local backends without keeping every model loaded at once.

+38stars / 7d
Verdict

Our llama-swap build took 44 seconds and all 17 tested Go packages passed in 25 seconds, giving the proxy itself a clean baseline. Use it when local model variety exceeds available memory and your clients already speak OpenAI or Anthropic APIs. Choose a simpler runner when you want model downloads handled for you, or a serving engine when sustained throughput matters more than swapping.

We ran it

Lab card: what happened when we ran llama-swapScreenshot of llama-swap (github.com/mostlygeek/llama-swap)
Install✓ · 37s171 packages
Build✓ · 44s
Tests✓ · 25s17 passed · 0 failed of 17 (go test)
Repo481 files~56,985 lines of source · 5.6 MB · 8 CI workflows

Answers from our run

Does llama-swap build from source?

Dependencies installed in 37 seconds (171 packages), and the build succeeded in 44 seconds. We cloned commit 0bdb372 into a clean Debian container with 3 CPUs and no project-specific setup.

Do llama-swap's tests pass?

Yes: 17 of 17 passed when we ran the project's own test command (go test). Some failures need services or credentials a bare container does not have.

Who should not use llama-swap?

Anyone expecting models or inference quality from the proxy itself: llama-swap starts other servers and does not supply model weights.

What are the alternatives to llama-swap?

Ollama, LocalAI, vLLM. Our llama-swap build took 44 seconds and all 17 tested Go packages passed in 25 seconds, giving the proxy itself a clean baseline.

Setup4/5One binary and YAML, but every upstream still needs configuration
Docs5/5Endpoints, installs, routing, containers, and proxy caveats are clear
Community5/55,473 stars with current releases and detailed issue activity
Maturity4/5Green 17-package run and broad APIs, with scheduling gaps

Who it’s for

Local AI users with more models than their RAM or VRAM can hold simultaneously.
Developers who want one OpenAI- or Anthropic-shaped endpoint in front of llama.cpp, vLLM, ComfyUI, or similar servers.
Homelab operators who need model TTLs, manual unload controls, logs, metrics, and API-key protection.
Advanced users willing to describe process commands and allowed model combinations in YAML.

Who it’s NOT for

Anyone expecting models or inference quality from the proxy itself: llama-swap starts other servers and does not supply model weights.
Multi-user services that need per-key permissions or file-based key rotation today: keys live in YAML or environment variables, and issue 1009 requests more flexible storage.
Operators wanting automatic RAM or VRAM scheduling: the current matrix requires configured compatibility rules, while memory-budget scheduling remains issue 1004.
Clients that cannot tolerate model-load delays or tune timeouts: issue 726 reports a coding client canceling after 120 seconds while a large prompt was still processing.
Teams seeking measured inference speed from this review: our lab measured install, build, and tests only, not tokens per second or swap latency.

Setup reality

Our sandbox installed 171 Go packages in 37 seconds. The build succeeded in 44 seconds, and all 17 tested Go packages passed in 25 seconds. The source checkout was 5.6 MB. These measurements cover llama-swap itself, not model downloads or inference servers.

The minimum configuration names a model and supplies the command that starts its upstream server with an assigned port. You must install that server or use a container image, obtain model files, choose backend flags, and ensure the process fits the available CPU, RAM, GPU, and VRAM.

A public deployment needs API keys, TLS, and careful log exposure. Nginx buffering must be disabled for streamed responses. Concurrent model sets require a matrix, containers may need explicit stop commands, and slow loads need client and proxy timeouts that match the hardware.

llama-swap starts an upstream from the requested model ID

llama-swap sits between an AI client and one or more local inference servers. A request arrives at an OpenAI- or Anthropic-compatible route, the proxy reads its model identifier, and the configured process for that model is started. If a conflicting process is already running, llama-swap replaces it. Clients keep one base URL while the machine loads only the model or compatible set needed for the current request.

The proxy covers more than chat. Its documented routes include completions, Responses, embeddings, speech, transcription, voices, image generation and editing, Anthropic messages, reranking, code infill, stable-diffusion.cpp, audio.cpp, and a ComfyUI path. A web interface exposes a playground, request captures, metrics, logs, and manual load controls. Version v251 also added vLLM speculative-decoding metrics and changed disconnected-client recording to HTTP 499.

One binary still depends on every model server you configure

A minimal YAML file needs a model ID and a cmd such as a llama-server invocation containing the supplied ${PORT} placeholder. llama-swap can then allocate the port and proxy the request. It does not download the chosen weights, decide a quantization, or tune context and batch settings. Operators must install the upstream executable, store model files, and test the exact flags for their hardware. The proxy simplifies lifecycle and routing, not inference engineering.

Distribution choices are broad. Release binaries cover Linux, macOS, Windows, and FreeBSD, while Homebrew, MacPorts, and WinGet paths are documented. Container users can choose unified nightly images that bundle llama-server and several media backends, or legacy images closer to llama.cpp. The recommended unified images target CUDA or Vulkan. Even there, you mount the model directory and configuration, and GPU runtime compatibility remains your responsibility.

What happened when we ran it

Our Go 1.24 sandbox installed 171 packages in 37 seconds. The source checkout held 481 files, roughly 56,985 lines of source, and occupied 5.6 MB. The build completed in 44 seconds on 3 CPUs with 8 GB of memory. This is a small development footprint for a proxy with a bundled web interface and many API routes, especially compared with the model weights and GPU software that sit outside it.

The test command succeeded in 25 seconds. Go reported 17 passing packages and 0 failures out of 17. That green result applies to commit 0bdb372 and is the strongest reason to trust the proxy code enough for a local trial. It says nothing about a particular llama.cpp build, vLLM environment, model file, CUDA driver, or tokens-per-second result because those components and measurements were outside the lab run.

The repository had 8 CI workflow files, with no top-level Dockerfile or conventional tests directory detected by the harness. Those structural signals do not conflict with the documented container images or Go tests; they describe the checked-out layout. The useful finding is the completed path: install, build, and all 17 tested packages passed without secrets in an unprivileged Debian container.

Swapping saves memory and adds cold-start delay

In the basic mode, only one model runs at a time. A ttl can unload an idle model, and an unload timeout gives the upstream a chance to stop cleanly. Profiles change routing names at runtime. A matrix DSL describes which model combinations may coexist, allowing an embedding model and chat model to remain loaded together when hardware permits. Hooks can preload models at startup, trading memory use for a shorter first request.

The proxy cannot remove load time. Issue 726 describes Kilo Code canceling near 120 seconds while a large-context model was still processing, after which llama-swap recorded a 502 for the canceled request. The report concerns one Windows configuration. It illustrates the client boundary rather than establishing behavior for every deployment. Server, upstream, and client timeouts must cover loading plus prompt processing, and user interfaces need a useful loading state.

Memory scheduling is also manual. Issue 1004 asks for per-model RAM declarations and automatic eviction under a global budget because the matrix becomes harder to maintain across 8 or more models. Today, operators describe valid combinations themselves. That is precise when you know the hardware and model footprints; it is tedious when models, context sizes, or offload settings change often.

API compatibility is broad, while access control stays basic

llama-swap can require API keys and rewrite selected request parameters before forwarding. Aliases let a local model answer to a familiar client-facing name. Logs can stream globally, by proxy or upstream, or for one model. Prometheus metrics and a health endpoint help with operations. These features are enough for a trusted LAN or single-user workstation when the listening address and keys are configured intentionally.

Key management is less suited to a large shared service. Issue 1009 says keys currently come from YAML or environment variables and requests file-backed rotation that would avoid reloading configuration and model state. The request also refers to future per-key permissions. Until such controls exist, place authorization in a gateway if users need different rights. Logs and captured requests may contain sensitive prompts or outputs, so do not expose the UI or log endpoints merely because the inference API has a key.

Streaming needs a proxy exception, and v251 is active

The README warns that Nginx response buffering breaks Server-Sent Events and streamed chat output. It recommends disabling buffering and caching on the relevant routes, even though llama-swap sends X-Accel-Buffering: no as a safeguard. Verify streaming through the actual public URL, since a working direct request does not prove the reverse proxy passes chunks promptly. Container stop behavior also benefits from cmdStop so Docker or Podman processes shut down cleanly.

GitHub recorded the last push on August 25, 2026, 2 days after v251 shipped. The repository had 5,473 stars and 79 open issues and pull requests when fetched. Current activity covers metrics, error bodies, audio builds, logs, scheduling, video routes, and client behavior. llama-swap is active and technically focused. Its 17-package green run makes it easy to recommend for a local proof, as long as model performance and hardware fit are tested separately.

Alternatives

ProjectWhat it isPick it when
Ollama gh↗A local model runner with model packaging, downloads, and a simple API.pick this instead when you want one supported model runtime and easier model management rather than supervising several server types.
LocalAI gh↗A local OpenAI-compatible server that brings many inference backends into one project.pick this instead when you want a bundled multi-backend inference service rather than a process-switching proxy.
vLLM gh↗A GPU-focused server for high-throughput language-model inference.pick this instead when throughput for a smaller set of resident language models matters more than local hot swapping.

What people are saying

  1. [github-trending] mostlygeek/llama-swap

Sources

  1. llama-swap repository
  2. llama-swap v251 release
  3. Memory-constrained scheduling request
  4. API key file request
  5. Kilo Code timeout report
  6. Nginx streaming issue

More llm tools reviews

agent-toolkit-for-aws · agent-memory · codex-astra-luna-orchestrator · okf-agent-memory · mlc-llm · awesome-openclaw-skills · the whole board →