mrkeyoor.com_
Mon 05 Oct 06:27 UTC
Automationevaluationupdated 05 Oct 2026

foreman review

Foreman runs beside a coding agent and asks TypeSafe's Jev model whether the work is complete, stuck, off track, ready for verification, or waiting on a person. Python policy turns those scores into actions such as continuing, steering, stopping, retrying, verifying, or finishing one worker's job.

Verdict

Our Foreman run passed 259 of 260 tests, with the lone failure finding 24 coalesced output events where the test expected 50, so the project is substantial but did not clear a clean release gate in our sandbox. Try it when you already run Codex, OpenCode, or Hermes and want explicit supervisory policy around one worker. Wait if you need isolated execution, proven assessment accuracy, concurrent workers, or a settled human-escalation recovery flow.

We ran it

Lab card: what happened when we ran foremanScreenshot of foreman (thruwire.ai)
Install✓ · 26s63 packages · 88 MB
Build✓ · 4s
Tests✗ · 47s259 passed · 1 failed of 260 (pytest)
Known vulns0(pip-audit)
Repo95 files~12,609 lines of source · 0.6 MB · 1 CI workflows · tests dir

Answers from our run

Does foreman build from source?

Dependencies installed in 26 seconds (63 packages), and the build succeeded in 4 seconds. We cloned commit e5d1aa4 into a clean Debian container with 3 CPUs and no project-specific setup.

Do foreman's tests pass?

Not all of them: 259 of 260 passed and 1 failed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does foreman have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use foreman?

Teams expecting proven autonomous correctness: the README says Jev's accuracy for this job is unproven and its scores need calibration.

What are the alternatives to foreman?

OpenHands, Aider, LangGraph. Our Foreman run passed 259 of 260 tests, with the lone failure finding 24 coalesced output events where the test expected 50, so the project is substantial but did not clear a clean release gate in our sandbox.

Setup3/526-second install; real runs need two authenticated model tools
Docs5/5Detailed runtime, routing, evidence, safety, and backend docs
Community3/5671 stars with active September releases and October issue work
Maturity3/5v0.4.1 has broad tests, but one failed and limits remain explicit

Who it’s for

Developers already using Codex, OpenCode, or Hermes for repository work who want a separate supervisor.
Teams willing to tune evidence, thresholds, and time limits for their own definition of done.
Codex App Server users who want a supervisor to steer an active turn and preserve an event trail.
Researchers testing whether a fast decision model can supervise a slower coding model.

Who it’s NOT for

Teams expecting proven autonomous correctness: the README says Jev's accuracy for this job is unproven and its scores need calibration.
Operators who need several coding workers at once: Foreman currently runs one worker at a time.
Security-sensitive users expecting Foreman to isolate untrusted jobs: workers and configured evidence commands inherit local permissions, and the OpenCode backend runs with --auto.
OpenCode or Hermes users who require live in-turn steering: only the Codex App Server backend supports it; the others fall back to stop and retry.
Release gates that require a clean fresh-container suite: our run ended with 1 failed integration test out of 260.
Teams that need a clear resume path after human escalation: issue 40 reports that version 0.4.1 can stop without a concrete question or managed-run resume command.

Setup reality

Our fresh Debian sandbox installed commit e5d1aa4 in 26 seconds, adding 63 packages and using 88 MB. The build passed in 4 seconds. Pytest exited 1 after 47 seconds: 259 passed and 1 failed out of 260. Pip-audit reported 0 known vulnerabilities.

Real supervision needs Python 3.11 or newer, a TypeSafe API key, and an authenticated Codex, OpenCode, or Hermes CLI. The deterministic demo and tests need no credential, network, coding-agent process, or external repository.

Foreman runs workers with local permissions rather than adding its own isolation. Live steering needs Codex App Server; the exec, OpenCode, and Hermes paths cannot accept an in-flight steering message. State and events are stored under .foreman or the configured global data directory.

Foreman supervises one worker instead of replacing it

Foreman sits beside a coding agent and watches its output, repository changes, test evidence, recent events, and elapsed time. The worker still chooses tools and edits files. A separate Jev request judges narrower questions such as whether progress is meaningful, requirements are satisfied, verification is needed, or a person must intervene. The design is aimed at a familiar problem: an agent can keep producing plausible work after the job has gone wrong or should have stopped.

The runtime currently manages one coding worker at a time. Codex is the default, with OpenCode and Hermes available through an environment setting. A verifier is another independent worker pass, not a formal proof. Foreman keeps observations bounded by default: the diff is capped at 20,000 characters, each output tail at 12,000 characters, the event window at 30 entries, and worker history at 10.

Ten global checks feed one deterministic action

Jev does not directly stop a process or declare success. It returns probabilities for ten global checks, plus a documentation check when that responsibility is routed. Python responsibilities propose directives, then a fixed arbiter chooses one permitted action. Human need and iteration limits come before drift, stuck-work handling, completion, verification, and continued work. That separation is good engineering: model output informs policy, while ordinary code owns the side effect.

Defaults still encode judgment calls. A needs_human, worker_stuck, work_off_track, or AGENTS.md-drift score reaches its action threshold at 0.80. Verification begins at 0.65, while readiness, requirement satisfaction, and test sufficiency use 0.75. Teams should treat those values as starting hypotheses. The README explicitly says assessment accuracy is unproven for this use and needs calibration.

What happened when we ran it

Our sandbox installed commit e5d1aa4 in 26 seconds, pulling 63 packages and using 88 MB on disk. The build succeeded in 4 seconds. We used a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Pip-audit reported 0 known vulnerabilities.

The test step failed with exit code 1 after 47 seconds. Pytest reported 259 passed and 1 failed out of 260, with 40.71 seconds inside pytest itself. The failing case was test_noisy_events_are_coalesced in tests/test_integration.py. It expected 50 output events and received 24. The log does not establish whether timing, implementation, or the test expectation caused that mismatch, so we will not choose a cause for it.

The repository is much more than a wrapper script: 95 files, roughly 12,609 source lines, a tests directory, and one CI workflow. It has no Dockerfile. The deterministic demo exercises the same runtime, policy, persistence, event stream, and interface without a TypeSafe key, network connection, coding-agent CLI, or external repository. That is the right first run before allowing Foreman to launch a real worker.

Live steering belongs to the Codex App Server path

With the default Codex App Server backend, Foreman can steer or interrupt an active turn because the thread remains addressable. The older codex exec transport stays available, but it cannot receive live steering. OpenCode and Hermes also degrade to stop and retry. Backend choice therefore changes the supervisor's abilities, not merely the command used to start a worker.

A real run needs Python 3.11 or newer, TYPESAFE_API_KEY, and a working authenticated agent CLI. The Jev timeout defaults to 10 seconds, the minimum assessment interval to 5 seconds, and the whole factory timeout to 7,200 seconds. Transient supervisor failures can occur 3 times before escalation. Foreman's retry logic keeps the worker moving during those misses, within the configured boundaries.

Local permissions remain the security boundary

Foreman asks Codex for its workspace-write sandbox, but Foreman itself does not isolate the job. OpenCode runs with --auto, which approves requests unless its own permission configuration denies them. Hermes likewise runs as a local CLI. Evidence commands inherit Foreman's environment and local permissions, though the project invokes them without a shell and applies timeout and output limits.

That makes untrusted repositories and untrusted job text a poor fit. The README says so directly. Persistent state helps inspection and recovery, but is not production-grade durable execution. Each managed repository receives state.json and events.jsonl under .foreman/runs/<run-id>, while attached sessions live in the global Foreman data directory. Review both the worker configuration and evidence commands before turning on autonomous runs.

Version 0.4.1 is active, while escalation recovery is unsettled

Foreman 0.4.1 was released September 27, 2026, and the repository was pushed the following day. GitHub showed 671 stars, 52 forks, and 9 open issues and pull requests. Issue activity continued on October 4 with a report about managed runs that escalate for human input but provide no concrete question and expose no answer, approve, or resume command.

That open case lands on Foreman's hardest product problem. Stopping for a person is safer than inventing permission, but the operator needs to know what Foreman is asking and how to continue without erasing history. Until that workflow is clear, use Foreman on bounded jobs where a stopped run can be inspected and restarted deliberately. The 259 passing tests show serious work; the failed coalescing test and explicit operational limits say this is still a system to supervise while it supervises yours.

Alternatives

ProjectWhat it isPick it when
OpenHands gh↗A coding-agent platform that handles software tasks directly rather than supervising another CLI.pick this instead when you want the worker and its execution platform in one project.
Aider gh↗A terminal pair programmer built around direct, human-guided repository edits.pick this instead when you want to stay in the loop and do not need a second model judging the worker.
LangGraph gh↗A library for building stateful agent workflows with explicit nodes and transitions.pick this instead when you want to design the whole control graph rather than adopt Foreman's policy.
AgentOpsAgent monitoring and evaluation focused on traces, costs, and recorded sessions.pick this instead when observation matters more than automatic steering and stopping.

What people are saying

  1. [velocity-scout] thruwire/foreman

Sources

  1. Foreman README
  2. Foreman 0.4.1 release
  3. Foreman runtime and event flow
  4. Issue 40: recovery after human escalation

More automation reviews

jianying-headless · ok-wuthering-waves · jev-ultrafast · typesafe-computer-use · jev-trader · omniget · the whole board →