mrkeyoor.com_
Sun 04 Oct 15:18 UTC
Open Source6 min read

Experiential Adds 570 Stars for an Agent-Trace Router

Experiential turns agent traces into routing policy, but its most ambitious features require explicit evaluation, paid calls, and careful handling of captured data.

A GitHub Trending snapshot collected October 4 showed Experiential adding 570 stars in a day. The interesting part sits beyond the familiar promise of one endpoint for many models: this gateway wants yesterday's agent runs to decide which model handles tomorrow's task.

That idea gives teams a possible route out of static model selection. A support agent may need an expensive model for a messy escalation and a cheaper one for a routine lookup. Experiential records traffic, builds test scenarios from those traces, evaluates candidate models, and fits a router against quality, speed, and cost. The repository is Apache-2.0 licensed, and the Python package includes a Rust gateway data plane.

The star burst deserves attention because the project joins two jobs that are usually bought or built separately. It handles live model traffic, then uses evidence from that traffic to prepare a routing policy. The second job is also where the operational cost and privacy questions begin.

One endpoint, two operating modes

The quickest local setup is unusually short. The README gives two commands:

pip install experiential
exp

The first-run wizard asks for a provider, model, public alias, caller identity, and command budget. It then starts an authenticated OpenAI-compatible endpoint on 127.0.0.1:8000 and prints a one-time gateway credential. Existing code can point an OpenAI client at that address and keep using familiar chat-completions calls. The gateway also implements Anthropic's Messages surface.

Local operation matters here. Provider credentials are stored in a user-data file rather than the project catalog, according to the provider documentation. Teams can connect OpenAI, Anthropic, Gemini, OpenRouter, Azure, Bedrock, Vertex AI, or another OpenAI-compatible server. That last route is how a local inference server can sit behind the same gateway as a hosted model.

Experiential also runs a managed gateway. Its setup flow can connect bring-your-own provider credentials to the company's hosted endpoint, while a local install keeps the gateway and its SQLite accounting on the user's machine. Those are materially different trust choices. The project's "zero markup" description is a vendor claim about pass-through model use, not evidence that every Experiential workflow is free. Evaluations still call models, and managed fine-tuning uses an outside service.

For live traffic, the gateway has more substance than a URL switcher. The documented release scope covers virtual credentials, per-identity grants, monthly spending limits, exact-model pools, bounded provider fallback, and content-free usage accounting. Guardrails are available but stay off until an operator assigns a policy. That default is sensible for compatibility, though it means installation alone does not create a safety layer.

The router does not learn in the background

The repository description says Experiential "learns from your traffic." In practice, the process has visible stages and operator decisions. The usage guide says a build imports gateway or file traces, mines scenarios, and prepares retrieval data plus a simulated environment for tool-using tasks. Router optimization then runs candidate models and a judge across a bounded evaluation before fitting a frozen policy and checking held-out cases.

That distinction is useful. A model gateway silently changing routes after every request would be difficult to audit. Experiential instead produces versioned artifacts and requires an explicit command:

exp optimize router support-agent

The command can spend money. Experiential calculates an estimate first, uses a $50 warning budget by default, and asks for confirmation when an estimate crosses its thresholds. Dry runs and exact replays avoid new provider calls, while a fresh optimization can pay for candidate responses, simulated observations, and judging. The final report separates operating cost per successful task from the money spent running the experiment.

Quality still depends on the evaluation design. The built-in judge starts with provisional provenance, and the documentation recommends human calibration without requiring it. Traces from a narrow week can produce a narrow scenario set. A weak simulated environment can also reward a model for succeeding in a test that does not behave like production. The software records those boundaries, but it cannot choose representative work on a team's behalf.

There is a second optimization path. With the optional sft dependencies, exp optimize model prepares a project-bound dataset and hands managed fine-tuning to Thinking Machines' Tinker service. That path can yield a specialized model alias, according to the release scope. It should be read as an integration, not a fully local training stack. Teams that only want routing can leave that dependency out.

Capture changes the privacy calculation

Experiential can ingest existing OpenTelemetry traces, which is the cleanest option for teams that already instrument agent calls. It also offers exp capture for supported macOS applications. The capture route intercepts selected OpenAI and Anthropic traffic, creates a private certificate authority restricted to chosen provider hosts, and asks the user to approve a network extension and certificate trust.

This feature is experimental. The usage documentation says captured traces include prompts, responses, and tool content. Credential headers are excluded. Supported trace bodies can be uploaded to Experiential's platform, with bounded local retry files used when delivery fails. A client that rejects the capture certificate falls back to encrypted pass-through for the rest of that run and is no longer observed.

That is a clear boundary, and teams should treat it as one. Source code visibility does not make uploaded prompts harmless. Anyone testing capture on work repositories should first check whether source fragments, customer records, or tool results may appear in agent conversations. Starting with exported, scrubbed OpenTelemetry traces gives an operator more control over the first evaluation corpus.

Anonymous aggregate product telemetry is a separate channel and is enabled by default. The README says it excludes prompts, traces, credentials, paths, model names, and raw customer content. It can be switched off with exp config telemetry disable. The distinction between aggregate telemetry and trace upload is easy to miss, so deployment notes should record both settings.

A fast release train still carries caution labels

The public package reached version 0.7.151 on PyPI on October 4 and requires Python 3.12 or newer. Its native dependency publishes wheels for mainstream macOS, Linux, and Windows targets. Capture has the tighter requirement: macOS and Python 3.13 or newer.

The same day's 0.7.151 release changed admission under saturated lanes. Paying callers may overflow an authored capacity bound, while free callers receive a 429 response. It also added calling-agent attribution headers and adjusted how split Anthropic thinking blocks map into the chat reasoning carrier. These are details from a project dealing with real gateway problems, including fairness under load and protocol translation.

Yet the Rust crate still describes its gateway data plane as a proof of concept, and the release-scope document carefully limits claims to behavior exercised on the exact checkout. Capture's authors say broader application and network compatibility needs more testing. Those labels do not cancel the code behind them. They tell an operator where to run a trial before placing the gateway in front of production agents.

Experiential is most compelling for teams with repeated agent tasks, several viable models, and enough trace volume to test routing choices. A developer who only needs one model endpoint would inherit a large policy and evaluation system without much benefit. For a team already comparing models by hand, the frozen router and held-out report offer a more repeatable method.

The next evidence to watch is a reproducible workload showing how often the fitted router beats one fixed model after experiment spend is counted. Also watch whether macOS capture loses its experimental label and whether independent contributors begin owning meaningful parts of the code. The 570-star day shows strong curiosity. Those results will show whether trace-trained routing earns a permanent place in the agent stack.

We reviewed this

  1. experiential — our honest review

Sources

  1. Experiential GitHub repository
  2. Experiential README
  3. Experiential usage guide
  4. Experiential release scope
  5. Experiential model provider documentation
  6. Experiential 0.7.151 on PyPI
  7. Experiential native gateway on PyPI
  8. Experiential v0.7.151 release