InferenceX compares serving stacks on 12 supported accelerator classes
InferenceX exists because a fixed inference result ages quickly. A new server image, kernel, quantization path, scheduler, or driver can move throughput and latency without any hardware change. The repository stores recipes and automation for vLLM, SGLang, TensorRT-LLM, CUDA, ROCm, NVIDIA systems, AMD Instinct systems, and TPU hardware, then sends validated results to a public dashboard.
The README listed 12 supported hardware entries on September 29, 2026, from H100 and MI300X through GB300 NVL72 and TPUv7x. This is a data-center research platform. Its 64.9 MB checkout contains launchers, workflow generation, evaluation, power collection, network and kernel experiments, result processing, and bilingual operating guides. An ordinary application benchmark uses only a thin slice of that machinery.
A standalone AgentX profile still runs for one hour
The easiest route avoids InferenceX's CI and Slurm orchestration. Its standalone AgentX guide pins a separate harness, points it at an existing OpenAI-compatible server, and runs a trace workload. The sample command uses concurrency 8, 393 dataset entries, a 3,600-second measurement, and separate preparation, warmup, and drain time. You still provide the model server and tokenizer access.
Full participation is much heavier. The contribution flow asks for a green full sweep with evaluations, CODEOWNER evidence, an append-only performance changelog entry, and explicit reuse of expensive artifacts at merge. Recipes record server images, model and hardware topology, concurrency, sequence lengths, power measurements, and provenance. Those rules make public comparisons easier to audit, while also making this a poor fit for casual timing tests.
What happened when we ran it
Our sandbox installed the Python project under inferencex-e2e in 16 seconds. It added 41 packages and used 47 MB on disk, then built in 10 seconds. Pip-audit reported 0 known vulnerabilities. The fresh Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, no secrets, and no elevated privileges.
The test step failed with exit code 1 after 179 seconds. Pytest's own summary reported 1,605 passed, 8 failed, 1 skipped, and 128 collection/setup errors out of 1,741. The pytest summary clock was 143.20 seconds. A large passing majority is useful evidence of covered local behavior, but it does not make the complete command green.
commit f437f7b contained 4,011 files, about 814,444 lines counted as source, and 23 CI workflow files. The scan found no Dockerfile and no repository-root tests directory, although the Python project keeps suites under paths such as infx/tests and utils. The source moved again later on September 29, so our measurement is a commit-specific snapshot.
Five failures stopped at the missing tabulate module
Five of the 8 named failures ended with ModuleNotFoundError: No module named 'tabulate'. They came from reusable-sweep artifact validation cases, including metric selection, filename ordering, and decoder-failure handling. The project manifest places tabulate in its optional results dependencies. The log proves the import was absent in our environment, but we did not rerun with another dependency selection.
Three launch-layout tests failed assertions for historical, nested, and stale-submodule layouts. Their assertion text is truncated in the supplied tail. The tail also names utils/srt-slurm/tests/test_vllm_router_frontend.py among the collection errors without showing its traceback. Assigning all 128 errors to tabulate, Slurm, missing GPUs, or system packages would go beyond the evidence.
Python tests cannot prove a GPU sweep
InferenceX's testing guide draws a useful boundary. Local checks can validate YAML, schemas, matrix generation, result transforms, and focused Python behavior. A smoke run can prove one allocation, server start, workload, and artifact path. Only a full sweep plus evaluation covers the selected concurrency space on the target hardware, and reviewers still have to inspect the jobs and artifacts.
That distinction is why the red sandbox result matters without settling the product verdict. Our 1,605 passing tests say a lot about local control logic. They say nothing about H200 allocation, a Blackwell kernel, ROCm communication, power-sample coverage, or whether a model server produced a fair curve. Anyone publishing numbers needs the exact image, model, topology, workload, and validation path attached.
Same-day commits coexist with 291 open issues and pull requests
GitHub showed 1,785 stars, 105 open issues, and 186 open pull requests on September 29, 2026. The repository was pushed that day, after our measured commit, with fixes for a GB300 recipe and an MI300X Slurm partition. Issue 3425 reports a broken B300 multi-node path, and issue 3322 documents a GB200 curve that lost 18 of 31 points.
That queue is evidence of active hardware work and a large maintenance burden. GitHub returned no latest release, so consumers follow commits, pinned images, and workflow evidence rather than a stable tagged package. InferenceX is worth that pace when cross-vendor performance research is your job. If you only need to load-test one OpenAI-compatible endpoint, its governance and cluster surface area will cost more than the comparison is worth.

