mrkeyoor.com_
Mon 05 Oct 07:14 UTC
AI Toolsevaluationupdated 05 Oct 2026

jev-experiments review

Jev Experiments is a collection of 21 small applications showing how TypeSafe's Jev model can make fast, probability-backed decisions without writing prose. Each demo applies that pattern to a concrete job such as support triage, command safety, search, moderation, or routing, but the repository is a gallery rather than one reusable product.

Verdict

Our Agent Assist run installed 141 packages, built in 13 seconds, and passed all 14 tests in 7 seconds, making it a credible reference implementation rather than a screenshot collection. Use Jev Experiments to borrow narrow design patterns, especially the split between model judgment and code-enforced policy. Do not adopt the repository wholesale: choose one demo, verify it with a real key and your own cases, and settle licensing before carrying code into a product.

We ran it

Lab card: what happened when we ran jev-experimentsScreenshot of jev-experiments (github.com/dabit3/jev-experiments)
Install✓ · 22s141 packages · 180 MB
Build✓ · 13s
Tests✓ · 7s14 passed · 0 failed of 14 (vitest)
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo728 files~59,458 lines of source · 136.4 MB · 1 CI workflows

Answers from our run

Does jev-experiments build from source?

Dependencies installed in 22 seconds (141 packages), and the build succeeded in 13 seconds. We cloned commit c469e5b into a clean Debian container with 3 CPUs and no project-specific setup.

Do jev-experiments's tests pass?

Yes: 14 of 14 passed when we ran the project's own test command (vitest). Some failures need services or credentials a bare container does not have.

Does jev-experiments have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use jev-experiments?

Teams seeking one supported package to install across a product: the repository contains 21 separate demos with their own dependencies and commands.

What are the alternatives to jev-experiments?

NanoJev, AI SDK, BAML. Our Agent Assist run installed 141 packages, built in 13 seconds, and passed all 14 tests in 7 seconds, making it a credible reference implementation rather than a screenshot collection.

Setup4/522-second install and passing build; live decisions need a hosted key
Docs4/5Agent Assist explains requests, gates, fixtures, and run commands clearly
Community3/5398 stars and 7 open pull requests, but no open issue discussion
Maturity2/5Working demos, no tagged release, no license, and narrow CI coverage

Who it’s for

TypeScript developers evaluating Jev's Choice, Score, and Noul primitives in working interfaces.
Product teams looking for concrete patterns that keep business rules in code and use a model only for judgment.
Frontend engineers who learn faster from runnable applications than from SDK reference pages.
Support-tool builders who want to inspect a confidence-gated macro picker before designing their own.

Who it’s NOT for

Teams seeking one supported package to install across a product: the repository contains 21 separate demos with their own dependencies and commands.
Contact centers that need independently verified classification quality: Agent Assist uses 8 scripted chats and 40 messages, and the README says its seeded judgments were checked by hand.
Privacy programs that cannot send conversation text and customer attributes to a hosted service: the live Agent Assist request includes the last 6 messages, plan, tenure, prior-ticket count, and macro summaries.
Benchmark buyers who need a neutral Jev-versus-LLM comparison: the 4-second LLM wait is explicitly simulated, and our secret-free run did not reproduce the README's live latency figures.
Companies that require a declared open-source license: GitHub reports no license, and the repository root has no LICENSE, LICENSE.md, or COPYING file.

Setup reality

Our sandbox entered agent-assist/ at commit c469e5b and installed 141 npm packages in 22 seconds, using 180 MB. The build succeeded in 13 seconds. Vitest finished in 7 seconds with 14 passed and 0 failed, and npm audit reported 0 known vulnerabilities.

Live use requires Node 22 or newer and a TypeSafe API key. The development command starts a Node proxy on port 8787 and Vite on 5173. MOCK=1 substitutes labelled keyword answers for interface work, but those answers do not test Jev.

The complete checkout had 728 files, about 59,458 source lines, and occupied 136.4 MB. It has no Dockerfile and no tests directory, though Agent Assist keeps Vitest files beside its source. The sole CI workflow only checks the separate say/ macOS app, so Agent Assist has no visible GitHub Actions gate.

Twenty-one demo folders make this a pattern catalog

The root is deliberately thin: a short README points to 21 top-level demo directories, each built around a latency-sensitive judgment. The names reveal the range, including commit-sentry, jev-shell-guard, modstream, turbo-rerank, and agent-assist. There is no shared root package that turns them into one platform. You pick an example, enter its directory, and treat it as a separate application.

That shape is useful when your question is, "Where would a direct decision model fit?" It is less useful when you need a maintained SDK abstraction or a production service. The examples mix web apps, developer tools, and a macOS project. Dependencies and operating assumptions vary with each folder. The repository's value is concrete product patterns, especially cases where probabilities arrive quickly enough to alter an interface while a person is still acting.

Agent Assist asks 9 judgments in one hosted request

The measured subproject is a support console for 8 concurrent scripted chats. After each customer message, its server sends one Jev request containing 9 independent questions. Those cover a 13-way macro choice, a 10-option intent, 2 four-level scores, and 5 yes-or-no judgments. The response drives churn and escalation badges, queue ordering, and a suggested reply without asking the model to compose customer-facing text.

That boundary is the best decision in the demo. Twelve macro bodies remain client-side, and ordinary TypeScript determines refund eligibility from plan and tenure. Jev selects a macro from summaries, then code auto-fills it only when confidence reaches 0.6. A none choice sends the agent back to freehand writing. The model judges ambiguous language, while policy, display order, thresholds, and final wording stay visible in source.

What happened when we ran it

Our sandbox ran commit c469e5b from agent-assist/ because the repository has no root application. The checkout contained 728 files, about 59,458 lines of source, and used 136.4 MB. npm installed 141 packages in 22 seconds and occupied another 180 MB. The build succeeded in 13 seconds on 3 CPUs and 8 GB of RAM inside an unprivileged Node 22 container with no secrets.

Vitest completed in 7 seconds: 14 tests passed and 0 failed. npm audit found 0 known vulnerabilities at every reported severity. Our scan found 1 CI workflow, no Dockerfile, and no tests directory. The test files live beside the Agent Assist library source, so the missing directory does not mean tests are absent. The workflow is scoped to say/, however, and does not run the 14 Agent Assist tests.

Fourteen tests check policy wiring, not Jev's judgment quality

The passing suite covers useful seams. It checks the 0.6 macro gate, none handling, probability ordering, latency percentiles, refund rules, queue priority, all 9 question IDs, 12 macro choices plus none, the 6-message window, and the 8-chat, 40-message fixture. These tests guard application logic around the model and would catch several easy integration mistakes.

They do not call Jev or score its answers against labelled support data. The README describes hand-checked outcomes on the seeded scripts, including a duplicate charge, an account lockout, a cancellation threat, and a data-erasure request. That is enough for a demo walkthrough. A contact center still needs its own confusion cases, policy language, languages, and false-escalation costs before trusting the queue ordering or allowing the 1.5-second hands-free auto-send path.

Live mode sends 6 messages and customer attributes off the box

Starting the real path requires TYPESAFE_API_KEY. The key stays in the Node proxy and does not enter the browser bundle, which is the right local boundary. Each request still sends the last 6 conversation messages, the customer's plan, tenure in months, prior-ticket count, and 12 macro summaries to TypeSafe. Names are stripped from the customer object, but message bodies can contain personal or regulated data.

The server permits at most 8 Jev calls in flight and configures retries for rate limits and server errors. Without a key, /api/judge returns 503. MOCK=1 keeps interface work offline with a visible label and 5-millisecond canned delay. That mode is useful for CSS and state transitions; its keyword answers say nothing about model quality or live latency. A real evaluation needs approved sample data and a provider agreement that fits the data involved.

The claimed 93 ms median was outside our secret-free run

The Agent Assist README reports a 93 ms median panel refresh and 223 ms at the 95th percentile across 40 messages on a Linux VM. It also labels its 4-second language-model baseline as simulated. Those disclosures are better than presenting the animation as a benchmark, but neither number came from our sandbox. We measured installation, build, tests, audit, and repository signals only.

This distinction changes the buying decision. The demo supports the argument that one request can answer several typed questions and update an interface. It does not establish performance from your region, under your rate limits, with your conversation sizes. Run the same 8-chat control against the account and network you plan to use, then add real cases. Compare end-to-end paint time and answer quality, not the artificial 4-second countdown.

Seven open pull requests have not produced a release or license

GitHub showed 398 stars and 7 open items on October 5, 2026, all of them pull requests. The repository was last pushed on September 21, while three testing and documentation pull requests opened September 25. That combination shows outside work after the last main-branch update. It also leaves the proposed adversarial policy probes and cross-demo pattern guide unmerged.

There is no tagged GitHub release, and GitHub detects no license. The only workflow builds the say/ macOS app when its paths change. For evaluation code, those gaps are manageable. For copied production code, they are decision points. Agent Assist earned its place as a readable, passing example in our run; it still needs independent live results, an explicit license, and a CI path before it can serve as more than a well-made reference.

Alternatives

ProjectWhat it isPick it when
NanoJev gh↗A public 0.6B decision-model checkpoint with training data and local inference code.pick this instead when you want to inspect or self-host a decision model rather than call the hosted Jev service.
AI SDK gh↗A TypeScript toolkit for building applications on many generative model providers.pick this instead when text generation, tool calls, streaming, and provider choice matter more than Jev-specific decision patterns.
BAMLA typed language and runtime for extracting structured outputs from language models.pick this instead when you need typed structured generation from conventional language models across providers.

What people are saying

  1. [velocity-scout] dabit3/jev-experiments

Sources

  1. Jev Experiments repository
  2. Agent Assist README
  3. Agent Assist package manifest
  4. Say-only GitHub Actions workflow
  5. Open adversarial policy probes pull request

More ai tools reviews

OrcaBonsai-27B-Uncensored · NanoJev · uplifting-biomolecular-modeling · procedural-film · jev-review · Dream-RSI · the whole board →