mrkeyoor.com_
Mon 28 Sept 07:44 UTC
LLM Toolsevaluationupdated 28 Sept 2026

ai-evaluation-framework review

AI Evaluation Framework is a small Node.js harness for comparing model outputs against field-level ground truth. It reports accuracy, request latency percentiles, and estimated token cost, then collects user complaints as candidates for future test cases.

Verdict

Our run installed 0 packages and passed the available tests in 5 seconds, which fits a 772-line harness that an engineer can read before trusting. Use it for structured, field-level comparisons when you want to own the cases, scoring rules, and result files. Choose a larger system for agent traces, qualitative judging, production observability, concurrent load, or a secured feedback service.

We ran it

Lab card: what happened when we ran ai-evaluation-frameworkScreenshot of ai-evaluation-framework (github.com/dreamers-laboratory/ai-evaluation-framework)
Install✓ · 5s0 packages · 1 MB
Buildn/ano build script
Tests✓ · 5sran, no count parsed
Known vulns00 critical · 0 high · 0 moderate · 0 low (npm audit)
Repo19 files~772 lines of source · 0.1 MB · 0 CI workflows · tests dir

Answers from our run

Does ai-evaluation-framework build from source?

Dependencies installed in 5 seconds (0 packages), and the project has no separate build step. We cloned commit 93a7a23 into a clean Debian container with 3 CPUs and no project-specific setup.

Do ai-evaluation-framework's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does ai-evaluation-framework have known vulnerabilities in its dependencies?

npm audit found none in the dependency tree at the time of our run.

Who should not use ai-evaluation-framework?

Teams evaluating free-form prose, agent trajectories, retrieval quality, or tool use: the runner expects a JSON object and scores named fields.

What are the alternatives to ai-evaluation-framework?

DeepEval, Promptfoo, OpenAI Evals. Our run installed 0 packages and passed the available tests in 5 seconds, which fits a 772-line harness that an engineer can read before trusting.

Setup5/50 packages installed and tests passed in 5 seconds
Docs4/5The example explains scoring, cost caveats, cache, and feedback
Community2/5119 stars, with no issue activity or contributor history yet
Maturity2/5Version 0.1.0 has tests but no releases, CI, or Dockerfile

Who it’s for

Developers evaluating structured extraction or classification where each output field has a known correct value.
Small teams that want an auditable JavaScript harness instead of an evaluation service.
Engineers comparing OpenAI, Anthropic, or an OpenAI-compatible gateway on the same cases.
Product teams willing to review negative user feedback by hand before adding it to a benchmark.

Who it’s NOT for

Teams evaluating free-form prose, agent trajectories, retrieval quality, or tool use: the runner expects a JSON object and scores named fields.
Load-testing or production-latency work: the grid calls every model and case sequentially, so it does not reproduce concurrent traffic.
Buyers who need current pricing out of the box: the shipped table is explicitly placeholder data and unknown models use a marked fallback.
Public feedback endpoints without another security layer: the server has no authentication and exposes stats, review candidates, and raw JSONL exports.
Organizations that need releases, CI, containers, dashboards, or a long maintenance record: the repository has no release, CI workflow, or Dockerfile.

Setup reality

Our sandbox install succeeded in 5 seconds, installed 0 packages, and occupied 1 MB. There was no build script or target, so the build step was skipped. The available tests passed in 5 seconds. Npm audit reported 0 known vulnerabilities across all severity levels.

Live model runs need an OpenAI or Anthropic API key, or an OpenAI-compatible base URL. You must create task cases, field rules, and expected values. Cost figures require replacing the bundled placeholder price table with current provider rates.

The optional feedback server writes append-only JSONL and serves a browser widget from 127.0.0.1:8790. It has no built-in authentication. Exposing it outside the local machine requires your own proxy, access control, retention policy, and review process for captured input and output context.

Seven rules turn field accuracy into a reviewable decision

The framework scores each named output field with one of 7 rules. Identifiers and enum values can require exact matches. Dates are normalized before comparison. Numeric values accept a chosen tolerance, while names can use containment or token overlap. A field can also be expected to be missing, and a case may list alternate accepted renderings. This is a better definition of accuracy than asking a model to grade another model with an unexplained rubric.

The worked example is invoice extraction, but the runner only sees text input, expected JSON fields, and a prompt. An OCR dump, transcript, or PDF text can use the same structure. For each model and case, it records the returned fields, per-field correctness, token usage, request latency, and computed cost. The JSON output retains cell-level details, which makes a surprising aggregate score traceable to the exact answer that produced it.

The grid measures sequential calls rather than production traffic

The runner loops through each model and then each case, awaiting one request before starting the next. That gives a simple p50 and p95 across the observed calls, but it does not test concurrency, queues, rate limits, warm worker pools, or shared production contention. A 3-case invoice sample is enough to demonstrate the mechanics, not to establish a stable tail-latency claim.

Provider support is deliberately narrow. Model IDs beginning with claude route to Anthropic, and every other ID routes to OpenAI unless the code is changed. An OpenAI-compatible gateway can be supplied through a base URL. Both paths ask for JSON and parse it directly. There is no retry policy, judge model, tool-call loop, trace ingestion, or adapter system in the 772-line codebase. That keeps the behavior readable and limits the jobs it can represent.

What happened when we ran it

Our sandbox install completed in 5 seconds, installed 0 packages, and used 1 MB on disk. The repository has no build script or target, so the build step was skipped. That is appropriate for native Node.js modules that run directly, and it should not be reported as a successful build.

The available test command passed in 5 seconds. The measurement did not supply a test count, so the honest result is simply that the project-defined test step succeeded. The README says its scripted mock covers every scoring rule, percentile and cost calculations, plus the feedback store and HTTP endpoints. No live provider key is required for that path.

Npm audit found 0 known vulnerabilities: 0 critical, 0 high, 0 moderate, and 0 low. Our repository scan found 19 files, about 772 lines of source, a tests directory, 0 CI workflow files, and no Dockerfile. The small dependency-free surface is easy to inspect, while the absence of CI means adopters must decide where those passing tests will run on future changes.

Cost numbers are placeholders until you replace the table

The bundled price table identifies itself as placeholder-2026-09-01. It contains illustrative per-million-token rates, and any unlisted model falls back to a default row. Output marks fallback costs with a tilde, which is an honest warning, but even unmarked rows should be replaced with current provider prices before a purchasing decision. The framework computes cost from reported token usage; it does not measure the provider invoice.

Caching uses a hash of the model, task prompt and fields, and case input. Changing a scoring tolerance can reuse the old model answer, while changing the prompt invalidates the relevant cache entry. That is useful for cheap rubric work. It also means latency from a cached cell describes the earlier request, not a new call. Use the force option when comparing current provider behavior or timing.

The feedback endpoint is a local review inbox

The sample feedback server listens on 127.0.0.1:8790 and defaults to a wildcard CORS origin. It accepts votes and optional context, writes append-only JSONL, and exposes stats, raw export, and negative entries reshaped as candidate cases. A thumbs-down must include a comment. The reviewer then fills in the expected result before moving that failure into the curated evaluation set.

This separation is sensible. User complaints are sparse and biased, so raw votes should not silently become ground truth. The server is still a development component. It has no authentication or user authorization, and the export can contain whatever input and output context the host page submitted. Keep it on loopback or place it behind a secured application endpoint with data minimization and retention rules.

A September 2 snapshot is too young for a health verdict

The repository was created and last pushed on September 2, 2026. GitHub showed 119 stars and 0 combined open issues or pull requests on September 28. There is no published release, while package.json calls the code version 0.1.0. Those facts describe a clean first snapshot, not an established maintenance record. Zero open issues can mean quiet code just as easily as proven code when a project is 26 days old.

The Apache-2.0 license, passing local test step, explicit limitations, and compact source make a low-cost trial reasonable. Pin commit 93a7a23, add the test command to your CI, replace the price table, and start with cases whose correct fields a domain expert can defend. If the evaluation expands into agent traces, security tests, human grading, or production monitoring, DeepEval, Promptfoo, or OpenAI Evals offers a broader starting point.

Alternatives

ProjectWhat it isPick it when
DeepEvalA larger LLM evaluation framework with a broad set of metrics and testing workflows.pick this instead when you need more evaluation methods than field-by-field ground truth scoring.
Promptfoo gh↗A prompt and agent testing toolkit with CI integration and security testing.pick this instead when red teaming, declarative matrices, and CI checks matter more than a tiny readable runner.
OpenAI EvalsA framework and benchmark registry for evaluating language models and systems.pick this instead when shared benchmark definitions and a larger evaluation ecosystem are the goal.

What people are saying

  1. [velocity-scout] dreamers-laboratory/ai-evaluation-framework

Sources

  1. AI Evaluation Framework README
  2. Evaluation runner source
  3. Feedback server source
  4. Placeholder pricing table

More llm tools reviews

llm-wiki-compiler · claude-skills · Humanizer-zh · agent-beacon · MiMo-Code · pi-claude-bridge · the whole board →