Seven rules turn field accuracy into a reviewable decision
The framework scores each named output field with one of 7 rules. Identifiers and enum values can require exact matches. Dates are normalized before comparison. Numeric values accept a chosen tolerance, while names can use containment or token overlap. A field can also be expected to be missing, and a case may list alternate accepted renderings. This is a better definition of accuracy than asking a model to grade another model with an unexplained rubric.
The worked example is invoice extraction, but the runner only sees text input, expected JSON fields, and a prompt. An OCR dump, transcript, or PDF text can use the same structure. For each model and case, it records the returned fields, per-field correctness, token usage, request latency, and computed cost. The JSON output retains cell-level details, which makes a surprising aggregate score traceable to the exact answer that produced it.
The grid measures sequential calls rather than production traffic
The runner loops through each model and then each case, awaiting one request before starting the next. That gives a simple p50 and p95 across the observed calls, but it does not test concurrency, queues, rate limits, warm worker pools, or shared production contention. A 3-case invoice sample is enough to demonstrate the mechanics, not to establish a stable tail-latency claim.
Provider support is deliberately narrow. Model IDs beginning with claude route to Anthropic, and every other ID routes to OpenAI unless the code is changed. An OpenAI-compatible gateway can be supplied through a base URL. Both paths ask for JSON and parse it directly. There is no retry policy, judge model, tool-call loop, trace ingestion, or adapter system in the 772-line codebase. That keeps the behavior readable and limits the jobs it can represent.
What happened when we ran it
Our sandbox install completed in 5 seconds, installed 0 packages, and used 1 MB on disk. The repository has no build script or target, so the build step was skipped. That is appropriate for native Node.js modules that run directly, and it should not be reported as a successful build.
The available test command passed in 5 seconds. The measurement did not supply a test count, so the honest result is simply that the project-defined test step succeeded. The README says its scripted mock covers every scoring rule, percentile and cost calculations, plus the feedback store and HTTP endpoints. No live provider key is required for that path.
Npm audit found 0 known vulnerabilities: 0 critical, 0 high, 0 moderate, and 0 low. Our repository scan found 19 files, about 772 lines of source, a tests directory, 0 CI workflow files, and no Dockerfile. The small dependency-free surface is easy to inspect, while the absence of CI means adopters must decide where those passing tests will run on future changes.
Cost numbers are placeholders until you replace the table
The bundled price table identifies itself as placeholder-2026-09-01. It contains illustrative per-million-token rates, and any unlisted model falls back to a default row. Output marks fallback costs with a tilde, which is an honest warning, but even unmarked rows should be replaced with current provider prices before a purchasing decision. The framework computes cost from reported token usage; it does not measure the provider invoice.
Caching uses a hash of the model, task prompt and fields, and case input. Changing a scoring tolerance can reuse the old model answer, while changing the prompt invalidates the relevant cache entry. That is useful for cheap rubric work. It also means latency from a cached cell describes the earlier request, not a new call. Use the force option when comparing current provider behavior or timing.
The feedback endpoint is a local review inbox
The sample feedback server listens on 127.0.0.1:8790 and defaults to a wildcard CORS origin. It accepts votes and optional context, writes append-only JSONL, and exposes stats, raw export, and negative entries reshaped as candidate cases. A thumbs-down must include a comment. The reviewer then fills in the expected result before moving that failure into the curated evaluation set.
This separation is sensible. User complaints are sparse and biased, so raw votes should not silently become ground truth. The server is still a development component. It has no authentication or user authorization, and the export can contain whatever input and output context the host page submitted. Keep it on loopback or place it behind a secured application endpoint with data minimization and retention rules.
A September 2 snapshot is too young for a health verdict
The repository was created and last pushed on September 2, 2026. GitHub showed 119 stars and 0 combined open issues or pull requests on September 28. There is no published release, while package.json calls the code version 0.1.0. Those facts describe a clean first snapshot, not an established maintenance record. Zero open issues can mean quiet code just as easily as proven code when a project is 26 days old.
The Apache-2.0 license, passing local test step, explicit limitations, and compact source make a low-cost trial reasonable. Pin commit 93a7a23, add the test command to your CI, replace the price table, and start with cases whose correct fields a domain expert can defend. If the evaluation expands into agent traces, security tests, human grading, or production monitoring, DeepEval, Promptfoo, or OpenAI Evals offers a broader starting point.

