mrkeyoor.com_
Tue 06 Oct 04:13 UTC
Dev Toolsevaluationupdated 06 Oct 2026

hyperfine review

Hyperfine is a command-line tool for timing shell commands across repeated runs and comparing the results. It handles run counts, warmups, setup commands, outlier warnings, parameter sweeps, and machine-readable exports so you do not have to build a benchmark loop from shell scripts.

Verdict

Our Hyperfine run built in 35 seconds and passed all 238 tests in another 35 seconds, making commit 79aea88 an easy recommendation for local command benchmarking. Use it to replace improvised timing loops and export evidence you can inspect. Choose a harness or continuous benchmarking system when you need function-level measurements, concurrent load, or historical regression policy.

We ran it

Lab card: what happened when we ran hyperfineScreenshot of hyperfine (github.com/sharkdp/hyperfine)
Install✓ · 6s59 packages
Build✓ · 35s
Tests✓ · 35s238 passed · 0 failed of 238 (cargo test)
Repo66 files~8,101 lines of source · 0.5 MB · 1 CI workflows · tests dir

Answers from our run

Does hyperfine build from source?

Dependencies installed in 6 seconds (59 packages), and the build succeeded in 35 seconds. We cloned commit 79aea88 into a clean Debian container with 3 CPUs and no project-specific setup.

Do hyperfine's tests pass?

Yes: 238 of 238 passed when we ran the project's own test command (cargo test). Some failures need services or credentials a bare container does not have.

Who should not use hyperfine?

Load testers: hyperfine runs benchmark commands sequentially, and concurrent execution remains an open feature request.

What are the alternatives to hyperfine?

bench, Criterion.rs, Bencher. Our Hyperfine run built in 35 seconds and passed all 238 tests in another 35 seconds, making commit 79aea88 an easy recommendation for local command benchmarking.

Setup5/56-second install, broad packages, and no service configuration
Docs5/5Clear examples for caches, shells, parameters, exports, and limits
Community5/528,950 stars with an October 5 release and active pull requests
Maturity5/5v1.21.0 and 238 passing tests in our measured commit

Discussed on

  1. hnHyperfine: A command-line benchmarking tool241 points
  2. hnShow HN: Hyperfine – a command-line benchmarking tool126 points
  3. hnShow HN: Hyperfine – a command-line benchmarking tool96 points
  4. hnHyperfine: A command-line benchmarking tool4 points
  5. hnHyperfine: A command-line benchmarking tool3 points

Who it’s for

Developers comparing two CLI implementations or checking whether an optimization changed wall-clock time.
Maintainers who want repeatable local measurements that can be exported to JSON, CSV, Markdown, or AsciiDoc.
Engineers tuning thread counts, compiler choices, or other parameters across the same command.
Teams that will control caches, background load, inputs, and machine state instead of treating one result as universal.

Who it’s NOT for

Load testers: hyperfine runs benchmark commands sequentially, and concurrent execution remains an open feature request.
Teams that need built-in historical baselines and regression gates across CI runs: issue 607 still asks for saved-run comparison, while Bencher is designed for continuous benchmarking.
Windows users who need peak-memory numbers: the README says memory_peak_resident is currently unsupported there.
Anyone who needs the simultaneous peak memory of a whole process tree: on Unix, the exported figure is the largest per-process peak among the command and waited-for children.
Microbenchmarks that belong inside a Rust test harness: Criterion.rs can measure functions without shell startup, process launch, and CLI I/O in the path.
Readers who expect a timing tool to remove experimental bias automatically: warm cache, cold cache, shell choice, preparation, CPU contention, and input data remain your decisions.

Setup reality

Our sandbox installed commit 79aea88 in 6 seconds, adding 59 packages. The Rust build succeeded in 35 seconds. Cargo test also finished in 35 seconds with 238 passed and 0 failed of 238.

Most users can install a binary or use an OS package manager. Building this measured commit from source needs Rust 1.97 or newer. No service, account, database, port, or API credential is required.

The checkout was 0.5 MB, with 66 files and about 8,101 source lines. It had a tests directory and 1 CI workflow, but no Dockerfile. That is a sensible shape for a portable CLI whose benchmarked commands need direct access to the host.

At least 10 runs make a better default than a shell loop

Hyperfine turns an arbitrary shell command into a repeated benchmark and reports its timing distribution. By default it performs at least 10 measured runs and estimates a count that targets roughly 3 seconds. You can set an exact run count when a comparison needs a fixed protocol. Multiple commands receive a relative comparison, while CSV, JSON, Markdown, and AsciiDoc exports keep the result available after the terminal scrolls away.

That makes Hyperfine a good replacement for loops around time. It also handles warmup runs, a preparation command before each measurement, and parameter sweeps for values such as thread count. These controls save work, but they do not decide the experiment. A warm cache answers a different question from a cold one. Changing the input file, machine load, power mode, or compiler flags can matter more than a neat relative-speed line.

The 238-test result supports a small, mature codebase

Our sandbox installed commit 79aea88 in 6 seconds and added 59 packages. The Rust build succeeded in 35 seconds. Cargo test took another 35 seconds and reported 238 passed with 0 failed of 238. We used an unprivileged container with 3 CPUs, 12 GB of RAM, and no secrets. These are repository checks, not a benchmark of Hyperfine's timing overhead.

The checkout occupied 0.5 MB and contained 66 files with roughly 8,101 lines of source. It had a tests directory and 1 CI workflow. There was no Dockerfile, which suits the job: benchmarking inside a container would answer questions about that container's scheduler, filesystem, and limits. Hyperfine is generally more useful installed on the same host and environment where the target command actually runs.

What happened when we ran it

Our run completed every supplied stage at commit 79aea88. Installation took 6 seconds for 59 packages, the build took 35 seconds, and all 238 Cargo tests passed in 35 seconds. No failure appeared in the provided results. We did not compare commands or claim a minimum timing error, because the lab block measured project setup and tests rather than benchmark accuracy.

The successful suite matters because Hyperfine performs calibration, statistics, formatting, shell handling, and cross-platform process measurement. Version 1.21.0 also changed timing-related code, including per-command peak-memory collection on Unix and Windows CPU-time precision. A green 238-test run cannot prove every operating system path, but it is relevant evidence for the exact commit we reviewed.

Commands under 5 ms need the no-shell option

On Unix, Hyperfine normally starts commands through sh; Windows uses cmd.exe. It measures empty-shell startup and subtracts that calibration from results. For very fast commands under 5 ms, the README recommends --shell=none, also available as -N, because shell correction can become a meaningful part of the number. Direct execution then loses shell syntax such as globs and home-directory expansion.

This is the sort of limitation good benchmark tooling should expose. The default is convenient for pipelines and redirection. Direct execution gives a cleaner path for tiny programs. Neither mode removes process startup, operating-system scheduling, CPU frequency changes, or background tasks. If you need to measure a Rust function rather than a whole executable, Criterion.rs puts the benchmark closer to the code and avoids pretending a process-level tool sees inside it.

Warmups and preparation commands answer different questions

A warmup count runs the command before recorded measurements, which is useful when caches and just-in-time work should already be settled. A preparation command runs before each timed execution and can restore files or clear state. Hyperfine also separates one-time setup for each parameterized command. Mixing these phases into the measured command can turn cleanup work into apparent program time.

The parameter tools are especially practical. A numeric scan can vary a value from 1 through 12, while a list can compare named compilers or Git branches. $HYPERFINE_ITERATION supplies a unique zero-based value to each run when outputs must not collide. These features make a benchmark repeatable enough to share, provided the command, input, warmup policy, machine details, and Hyperfine version travel with the result.

JSON exports do not provide a benchmark history

Hyperfine can write every measured duration plus wall-clock, user, and system time to JSON. On Unix and macOS, it can include peak resident memory. That memory field is the largest per-process peak among the command and waited-for children, rather than their simultaneous total; Windows memory is unsupported. Those definitions matter when a child-heavy program is being compared with a single process.

The tool does not manage a long-lived baseline by itself. Open issue 607 asks for saving a run and loading it into a later comparison. Hyperfine's README points to Bencher for continuous benchmarking and Chronologer for plotting measurements through Git history. Exporting JSON into one of those systems is the right move when a team needs regression thresholds and results that survive across machines or releases.

Version 1.21.0 arrived with active maintenance

GitHub showed 28,950 stars and 60 open issues and pull requests on October 6, 2026. Release v1.21.0 shipped on October 5, and the repository was pushed later that day. The release added time-unit aliases and iteration access in preparation commands, while fixing export behavior, invalid run counts, memory reporting, Windows precision, and the included Welch's t-test script. This is current maintenance, not a dormant utility coasting on old popularity.

Some boundaries remain deliberate or unresolved. Concurrent load testing is still an open request, CPU affinity is requested, and historical saved-run comparison is not built in. Those are reasons to choose another tool for a different job, not defects in Hyperfine's main use. For local whole-command comparisons, the 238 passing tests and careful documentation make Hyperfine the first tool I would install.

Alternatives

ProjectWhat it isPick it when
benchA compact command-line benchmark runner that inspired hyperfine.pick this instead when you want a smaller command runner and do not need hyperfine's wider export and parameter features.
Criterion.rsA statistics-driven Rust library for benchmarks that live inside a codebase.pick this instead when you need function-level Rust microbenchmarks rather than whole-process command timings.
BencherA continuous benchmarking system for storing results and detecting regressions over time.pick this instead when CI history, thresholds, and team-visible performance tracking are the main job.

What people are saying

  1. [velocity-scout] sharkdp/hyperfine

Sources

  1. Hyperfine repository and README
  2. Hyperfine v1.21.0 release
  3. Issue 58: concurrent execution and load testing
  4. Issue 607: saved benchmark comparisons
  5. Measured commit 79aea88

More dev tools reviews

vscodium · stripe-cli · build123d · OpenCore-Legacy-Patcher · cli · learning-python · the whole board →