KV-cache locality is the reason to add this router
llm-d Router earns its place when ordinary round-robin routing wastes expensive inference work. Its Endpoint Picker filters candidate pods, scores the survivors, and chooses a destination using signals such as model compatibility, current load, and reusable prefix-cache state. Request priorities and flow control can hold or reject work under saturation. For split inference, a sidecar coordinates encode, prefill, and decode stages. Those decisions matter across a fleet; one model on one server gives the scheduler nothing useful to choose.
The routing logic is pluggable. Administrators define filters, scorers, data producers, and scheduling profiles in an EndpointPickerConfig resource. A simple deployment can use one profile, while disaggregated prefill and decode need separate profiles. The current configuration API is llm-d.ai/v1; v1alpha1 is deprecated and will lose support in a later release. That is a manageable migration, but it signals that this v0.11.0 project is still tightening its public surface.
Standalone mode includes Envoy; Gateway mode joins shared infrastructure
Standalone mode packages the Endpoint Picker with a self-managed Envoy proxy, either as a sidecar or a separately scaled service. It is the sensible entry point for a local evaluation or one model service. Gateway mode connects an InferencePool to an HTTPRoute on a shared Kubernetes Gateway and is the project's production recommendation. Both paths require Gateway API Inference Extension CRDs before the Helm charts can install.
A custom proxy must also meet a precise protocol condition: the README says only FULL_DUPLEX_STREAMED is supported for both request and response body handling in Envoy's external processor. The separate service topology can fail open and send traffic directly to model servers while the Endpoint Picker recovers. That keeps requests moving, but it also means operators must decide whether unscheduled inference is safer than an outage for each workload.
What happened when we ran it
Our sandbox installed commit 2539e00 in 47 seconds and added 383 Go packages. The build completed successfully in 239 seconds. We used a fresh Debian container with 3 CPUs, 8 GB of RAM, Go 1.24, no secrets, and no elevated privileges. The repository checkout contained 1,520 files, about 258,009 source lines, and used 13.2 MB. These results show that dependency resolution and compilation worked in the environment we provided.
The test step failed after 263 seconds. Go test reported 210 passed and 11 failed out of 221. The last lines name TestStandaloneChart, failures under test/integration/epp, and TestE2E in the sidecar package, which ran for 20 seconds. The excerpt does not include the assertion messages or setup errors behind those failures. We therefore cannot say that Kind, Helm, networking, or the code itself caused them.
Active-active replicas can weaken approximate prefix routing
The container sizing guide warns that active-active replicas do not share approximate prefix state. Each replica sees only the requests it handled, which can reduce cache hits. Active-passive keeps one scheduling view but adds standby capacity without raising routing throughput. Teams need to choose between a unified state view and horizontal request processing instead of treating replica count as a generic availability setting.
Flow-control queues also live in Endpoint Picker memory. The documented default allows 5,000 requests and 1G of buffered bodies per priority band, while the global byte limit defaults to unlimited. The project tells operators to set a global cap below the container memory limit. Its own sizing guide recommends 4 to 6 GiB for 50 to 100 requests per second with 1,000 output tokens, and more than 20 GiB for longer outputs. Those are project measurements, not results from our sandbox.
Container images carry attestations; Helm charts do not
The project publishes Endpoint Picker, sidecar, and coordinator images with signed SLSA provenance and SPDX SBOMs. Its artifact verification guide gives digest-based commands for GitHub attestation verification and Cosign. The same table says Helm charts have neither provenance attestations nor SBOMs, and release images from v0.10.0 or earlier lack both. That split is clear enough for a supply-chain policy to enforce.
Our repository scan found 23 CI workflow files and a tests directory, though no Dockerfile. Current development docs require Make 4+, Go 1.25+, Docker or Podman, Kind, and kubectl 1.25+. The Kind path creates a cluster with the Endpoint Picker, a vLLM simulator, and a gateway implementation. Integration tests expect that environment, while the end-to-end command creates and later deletes its own temporary cluster. Local development is a cluster exercise, not a single Go binary loop.
Version 0.11.0 shipped while 336 items remained open
Release v0.11.0 arrived on September 27, 2026, and the repository was pushed on October 6. GitHub showed 336 open issues and pull requests: 214 issues and 122 pull requests. The release added stage-aware flow control, latency-based endpoint scoring, queue time limits, and other scheduler work. Combined with 23 workflows, that activity shows a project moving quickly, with a large review and issue queue to match.
Choose llm-d Router after the fleet has made cache-aware placement or multi-stage inference a concrete operating problem. AIBrix covers a wider set of inference infrastructure concerns, while vLLM Production Stack starts from deploying vLLM across Kubernetes. llm-d Router is the focused choice when the endpoint decision itself needs policy and live model signals. Our 11 failed tests mean the next evaluation should reproduce the full suite inside the documented Kind setup before the router touches production traffic.

