mrkeyoor.com_
Tue 06 Oct 15:55 UTC
LLM Toolsevaluationupdated 06 Oct 2026

llm-d-router review

llm-d Router decides which Kubernetes model-serving pod should receive each inference request. It can route by model, current load, reusable prompt cache, and priority, while also coordinating split prefill and decode workers through Envoy or the Kubernetes Gateway API.

Verdict

Our llm-d Router run installed 383 packages and built in 239 seconds, then passed 210 of 221 tests before 11 failures kept the suite red. It is worth evaluating when cache locality, priority queues, or split prefill and decode workers already cost your Kubernetes fleet real money or latency. For one model service, use a simpler gateway; for a serious multi-pod inference platform, test llm-d Router with your own request shapes and failure drills before adopting its scheduling policy.

We ran it

Lab card: what happened when we ran llm-d-routerScreenshot of llm-d-router (llm-d.ai/docs/architecture/core/router)
Install✓ · 47s383 packages
Build✓ · 239s
Tests✗ · 263s210 passed · 11 failed of 221 (go test)
Repo1520 files~258,009 lines of source · 13.2 MB · 23 CI workflows · tests dir

Answers from our run

Does llm-d-router build from source?

Dependencies installed in 47 seconds (383 packages), and the build succeeded in 239 seconds. We cloned commit 2539e00 into a clean Debian container with 3 CPUs and no project-specific setup.

Do llm-d-router's tests pass?

Not all of them: 210 of 221 passed and 11 failed when we ran the project's own test command (go test). Some failures need services or credentials a bare container does not have.

Who should not use llm-d-router?

A developer serving one model on one machine: the Endpoint Picker, Envoy integration, CRDs, and Helm charts add work without a pool to schedule.

What are the alternatives to llm-d-router?

AIBrix, vLLM Production Stack, Gateway API Inference Extension. Our llm-d Router run installed 383 packages and built in 239 seconds, then passed 210 of 221 tests before 11 failures kept the suite red.

Setup2/5Needs Kubernetes, CRDs, Helm, Envoy or a gateway, and model pods
Docs5/5Architecture, sizing, HA, testing, TLS, and artifact guides are specific
Community4/5October 6 push with 214 issues and 122 pull requests open
Maturity3/5v0.11.0 has signed images, but our complete suite failed

Who it’s for

Platform teams already running several vLLM or compatible model-server pods on Kubernetes.
Operators whose request volume justifies cache-aware placement, admission control, priorities, and separate prefill or decode workers.
Organizations already comfortable with Helm, Gateway API resources, Envoy, metrics, and Kubernetes failure modes.
Teams that need a pluggable endpoint picker behind shared cluster gateways.

Who it’s NOT for

A developer serving one model on one machine: the Endpoint Picker, Envoy integration, CRDs, and Helm charts add work without a pool to schedule.
Teams without Kubernetes operations experience: the development path requires Make, Go, Docker or Podman, Kind, and kubectl, while production Gateway mode adds cluster gateway infrastructure.
Clusters that need approximate prefix routing with fully active-active replicas: the operations guide warns that unsynchronized prefix state can reduce cache hits.
Release gates that require the full supplied suite to pass in a plain Go container: our run ended with 11 failures out of 221 tests.
Custom Envoy deployments that cannot use FULL_DUPLEX_STREAMED request and response body modes, the only supported setting documented by the project.
Buyers expecting signed Helm artifacts: the project documents provenance and SBOMs for container images, but not for its charts.

Setup reality

Our sandbox installed commit 2539e00 in 47 seconds, adding 383 Go packages. The build succeeded in 239 seconds. Tests failed in 263 seconds: 210 passed and 11 failed out of 221. The checkout held 1,520 files, about 258,009 source lines, and occupied 13.2 MB.

A realistic setup needs Kubernetes, Gateway API Inference Extension CRDs, Helm, model-serving pods, and either the included Envoy path or a supported L7 gateway. The current development guide also calls for Make 4+, Go 1.25+, Docker or Podman, Kind, and kubectl 1.25 or newer.

The failing log tail names the standalone Helm chart test, EPP integration tests, and a 20-second sidecar end-to-end test. It does not show the underlying assertions, so we cannot assign a cause. The repo had 23 CI workflow files and a tests directory, but no Dockerfile in our scan.

KV-cache locality is the reason to add this router

llm-d Router earns its place when ordinary round-robin routing wastes expensive inference work. Its Endpoint Picker filters candidate pods, scores the survivors, and chooses a destination using signals such as model compatibility, current load, and reusable prefix-cache state. Request priorities and flow control can hold or reject work under saturation. For split inference, a sidecar coordinates encode, prefill, and decode stages. Those decisions matter across a fleet; one model on one server gives the scheduler nothing useful to choose.

The routing logic is pluggable. Administrators define filters, scorers, data producers, and scheduling profiles in an EndpointPickerConfig resource. A simple deployment can use one profile, while disaggregated prefill and decode need separate profiles. The current configuration API is llm-d.ai/v1; v1alpha1 is deprecated and will lose support in a later release. That is a manageable migration, but it signals that this v0.11.0 project is still tightening its public surface.

Standalone mode includes Envoy; Gateway mode joins shared infrastructure

Standalone mode packages the Endpoint Picker with a self-managed Envoy proxy, either as a sidecar or a separately scaled service. It is the sensible entry point for a local evaluation or one model service. Gateway mode connects an InferencePool to an HTTPRoute on a shared Kubernetes Gateway and is the project's production recommendation. Both paths require Gateway API Inference Extension CRDs before the Helm charts can install.

A custom proxy must also meet a precise protocol condition: the README says only FULL_DUPLEX_STREAMED is supported for both request and response body handling in Envoy's external processor. The separate service topology can fail open and send traffic directly to model servers while the Endpoint Picker recovers. That keeps requests moving, but it also means operators must decide whether unscheduled inference is safer than an outage for each workload.

What happened when we ran it

Our sandbox installed commit 2539e00 in 47 seconds and added 383 Go packages. The build completed successfully in 239 seconds. We used a fresh Debian container with 3 CPUs, 8 GB of RAM, Go 1.24, no secrets, and no elevated privileges. The repository checkout contained 1,520 files, about 258,009 source lines, and used 13.2 MB. These results show that dependency resolution and compilation worked in the environment we provided.

The test step failed after 263 seconds. Go test reported 210 passed and 11 failed out of 221. The last lines name TestStandaloneChart, failures under test/integration/epp, and TestE2E in the sidecar package, which ran for 20 seconds. The excerpt does not include the assertion messages or setup errors behind those failures. We therefore cannot say that Kind, Helm, networking, or the code itself caused them.

Active-active replicas can weaken approximate prefix routing

The container sizing guide warns that active-active replicas do not share approximate prefix state. Each replica sees only the requests it handled, which can reduce cache hits. Active-passive keeps one scheduling view but adds standby capacity without raising routing throughput. Teams need to choose between a unified state view and horizontal request processing instead of treating replica count as a generic availability setting.

Flow-control queues also live in Endpoint Picker memory. The documented default allows 5,000 requests and 1G of buffered bodies per priority band, while the global byte limit defaults to unlimited. The project tells operators to set a global cap below the container memory limit. Its own sizing guide recommends 4 to 6 GiB for 50 to 100 requests per second with 1,000 output tokens, and more than 20 GiB for longer outputs. Those are project measurements, not results from our sandbox.

Container images carry attestations; Helm charts do not

The project publishes Endpoint Picker, sidecar, and coordinator images with signed SLSA provenance and SPDX SBOMs. Its artifact verification guide gives digest-based commands for GitHub attestation verification and Cosign. The same table says Helm charts have neither provenance attestations nor SBOMs, and release images from v0.10.0 or earlier lack both. That split is clear enough for a supply-chain policy to enforce.

Our repository scan found 23 CI workflow files and a tests directory, though no Dockerfile. Current development docs require Make 4+, Go 1.25+, Docker or Podman, Kind, and kubectl 1.25+. The Kind path creates a cluster with the Endpoint Picker, a vLLM simulator, and a gateway implementation. Integration tests expect that environment, while the end-to-end command creates and later deletes its own temporary cluster. Local development is a cluster exercise, not a single Go binary loop.

Version 0.11.0 shipped while 336 items remained open

Release v0.11.0 arrived on September 27, 2026, and the repository was pushed on October 6. GitHub showed 336 open issues and pull requests: 214 issues and 122 pull requests. The release added stage-aware flow control, latency-based endpoint scoring, queue time limits, and other scheduler work. Combined with 23 workflows, that activity shows a project moving quickly, with a large review and issue queue to match.

Choose llm-d Router after the fleet has made cache-aware placement or multi-stage inference a concrete operating problem. AIBrix covers a wider set of inference infrastructure concerns, while vLLM Production Stack starts from deploying vLLM across Kubernetes. llm-d Router is the focused choice when the endpoint decision itself needs policy and live model signals. Our 11 failed tests mean the next evaluation should reproduce the full suite inside the documented Kind setup before the router touches production traffic.

Alternatives

ProjectWhat it isPick it when
AIBrixA set of pluggable Kubernetes infrastructure components for generative AI inference.pick this instead when you want a broader inference platform around routing, autoscaling, and model management.
vLLM Production StackA Kubernetes reference stack for cluster-wide vLLM deployment.pick this instead when your main goal is deploying and operating vLLM rather than adopting a specialized endpoint picker.
Gateway API Inference ExtensionThe Kubernetes inference-routing API and protocol that llm-d Router implements.pick this instead when you are evaluating the shared API contract or building around another conforming implementation.

What people are saying

  1. [github-trending] llm-d/llm-d-router

Sources

  1. llm-d Router README
  2. llm-d Router architecture
  3. llm-d Router development guide
  4. llm-d Router container sizing guide
  5. llm-d Router v0.11.0 release
  6. llm-d Router artifact verification guide

More llm tools reviews

Rapid-MLX · simple-jev · kev · openjev-sglang · OptMem · SemIf-OpenJev · the whole board →