Foreman supervises one worker instead of replacing it
Foreman sits beside a coding agent and watches its output, repository changes, test evidence, recent events, and elapsed time. The worker still chooses tools and edits files. A separate Jev request judges narrower questions such as whether progress is meaningful, requirements are satisfied, verification is needed, or a person must intervene. The design is aimed at a familiar problem: an agent can keep producing plausible work after the job has gone wrong or should have stopped.
The runtime currently manages one coding worker at a time. Codex is the default, with OpenCode and Hermes available through an environment setting. A verifier is another independent worker pass, not a formal proof. Foreman keeps observations bounded by default: the diff is capped at 20,000 characters, each output tail at 12,000 characters, the event window at 30 entries, and worker history at 10.
Ten global checks feed one deterministic action
Jev does not directly stop a process or declare success. It returns probabilities for ten global checks, plus a documentation check when that responsibility is routed. Python responsibilities propose directives, then a fixed arbiter chooses one permitted action. Human need and iteration limits come before drift, stuck-work handling, completion, verification, and continued work. That separation is good engineering: model output informs policy, while ordinary code owns the side effect.
Defaults still encode judgment calls. A needs_human, worker_stuck, work_off_track, or AGENTS.md-drift score reaches its action threshold at 0.80. Verification begins at 0.65, while readiness, requirement satisfaction, and test sufficiency use 0.75. Teams should treat those values as starting hypotheses. The README explicitly says assessment accuracy is unproven for this use and needs calibration.
What happened when we ran it
Our sandbox installed commit e5d1aa4 in 26 seconds, pulling 63 packages and using 88 MB on disk. The build succeeded in 4 seconds. We used a fresh unprivileged Debian container with 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets. Pip-audit reported 0 known vulnerabilities.
The test step failed with exit code 1 after 47 seconds. Pytest reported 259 passed and 1 failed out of 260, with 40.71 seconds inside pytest itself. The failing case was test_noisy_events_are_coalesced in tests/test_integration.py. It expected 50 output events and received 24. The log does not establish whether timing, implementation, or the test expectation caused that mismatch, so we will not choose a cause for it.
The repository is much more than a wrapper script: 95 files, roughly 12,609 source lines, a tests directory, and one CI workflow. It has no Dockerfile. The deterministic demo exercises the same runtime, policy, persistence, event stream, and interface without a TypeSafe key, network connection, coding-agent CLI, or external repository. That is the right first run before allowing Foreman to launch a real worker.
Live steering belongs to the Codex App Server path
With the default Codex App Server backend, Foreman can steer or interrupt an active turn because the thread remains addressable. The older codex exec transport stays available, but it cannot receive live steering. OpenCode and Hermes also degrade to stop and retry. Backend choice therefore changes the supervisor's abilities, not merely the command used to start a worker.
A real run needs Python 3.11 or newer, TYPESAFE_API_KEY, and a working authenticated agent CLI. The Jev timeout defaults to 10 seconds, the minimum assessment interval to 5 seconds, and the whole factory timeout to 7,200 seconds. Transient supervisor failures can occur 3 times before escalation. Foreman's retry logic keeps the worker moving during those misses, within the configured boundaries.
Local permissions remain the security boundary
Foreman asks Codex for its workspace-write sandbox, but Foreman itself does not isolate the job. OpenCode runs with --auto, which approves requests unless its own permission configuration denies them. Hermes likewise runs as a local CLI. Evidence commands inherit Foreman's environment and local permissions, though the project invokes them without a shell and applies timeout and output limits.
That makes untrusted repositories and untrusted job text a poor fit. The README says so directly. Persistent state helps inspection and recovery, but is not production-grade durable execution. Each managed repository receives state.json and events.jsonl under .foreman/runs/<run-id>, while attached sessions live in the global Foreman data directory. Review both the worker configuration and evidence commands before turning on autonomous runs.
Version 0.4.1 is active, while escalation recovery is unsettled
Foreman 0.4.1 was released September 27, 2026, and the repository was pushed the following day. GitHub showed 671 stars, 52 forks, and 9 open issues and pull requests. Issue activity continued on October 4 with a report about managed runs that escalate for human input but provide no concrete question and expose no answer, approve, or resume command.
That open case lands on Foreman's hardest product problem. Stopping for a person is safer than inventing permission, but the operator needs to know what Foreman is asking and how to continue without erasing history. Until that workflow is clear, use Foreman on bounded jobs where a stopped run can be inspected and restarted deliberately. The 259 passing tests show serious work; the failed coalescing test and explicit operational limits say this is still a system to supervise while it supervises yours.

