The prompt pack is tied to named Codex model lines
The main README is in Chinese, with a synchronized English version and matching diagrams. gpt-instruct does not present one timeless system prompt. It maintains a stable gpt-5.6-sol-v45 package alongside gpt-6-astra-v2-rc1 and gpt-6.1-sol-v1-rc2 prereleases. Each file is bound to a model identity and evaluation history, which is the right way to document instructions whose behavior can change when the underlying model changes.
The intended change is direct: make Codex begin tool or file work earlier, carry confirmed state into later turns, verify produced artifacts, and leave a tested rollback. A Python installer previews or applies a chosen package, accepts a custom ZIP or Markdown file, and can reset the one configuration field it owns. It uses Codex's normal model_instructions_file setting rather than patching a binary or intercepting traffic.
The release gates expose incomplete results
The project separates its tests into A, B, and C. A contains 4 cases and requires two fresh runs plus artifact checks. B calls for 66 cases and 74 turns. C contains 120 medium-reasoning cases and runs only after A and B pass. The two current prerelease prompts passed their fresh A gates but fell short of the complete B requirement, so C was not run.
That disclosure matters more than a chart trending upward. The README reports 42 of 50 non-cloud B cases for the Astra candidate and 34 of 50 for the 6.1 candidate, with cloud attempts listed separately. It refuses to merge unlike runs into a prettier fraction. Still, these are project-authored cases with manual decisions. A team adopting the prompt should repeat representative tasks against stock Codex and score them without knowing which configuration produced each response.
An optional JailbreakBench track is kept apart from A, B, and C. Its official judge requires a specific Together-hosted model and leaves rows unjudged when that backend is missing. Human fields separately record refusal, cheating, and protocol violations. That method is careful, but it also shows the labor behind the headline scores: a ZIP is easy to install, while a credible comparison needs frozen inputs, credentials, repeated runs, and review.
What happened when we ran it
Our sandbox installed commit 3ab84df in 10 seconds. The Python environment added 35 packages and occupied 37 MB on disk. A build step completed in 1 second, and pip-audit found 0 known vulnerabilities. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.
Tests did not fail because none were started. The measured checkout exposed no test script or target, so the lab skipped that step. It contained 93 files, about 2,816 lines of source, and 4.9 MB of checked-out data, with 2 CI workflow files but no Dockerfile or tests directory. Those facts describe commit 3ab84df; they do not validate the README's current prompt scores or newer evaluation folders.
A 10-second install therefore answers only the packaging question. It does not tell us whether the instructions improve first-turn completion, preserve legitimate safety boundaries, or behave consistently across accounts and model revisions. The repository's own method treats those as model runs followed by human review, which our dependency and build check did not perform.
Applying it changes one sensitive configuration field
The deployment script can preview its work before writing, and it saves the previous state. Reset is scoped to model_instructions_file, leaving provider, model, authentication, and unrelated settings alone. Full snapshot restore is a separate explicit action. These choices lower the recovery cost if a candidate prompt makes Codex less useful.
Recovery is still worth rehearsing. A custom instruction file affects every request handled under that Codex home, not just the experiment that motivated the install. Use a separate CODEX_HOME, compare the generated diff, and keep experimental accounts and workspaces away from ordinary production access. The README itself says custom instructions can create account risk and advises use only in environments you are entitled to operate.
October activity is high, but model drift is the product risk
GitHub showed 9,298 stars and 30 open issues and pull requests on October 7, 2026. The last push was October 4, and issue 89 received a response and closed on October 5. That is healthy recent activity for a repository created in July 2026. The latest GitHub release, however, is the September 6 gpt-6-astra-v1 tag; newer release candidates are distributed from the repository rather than a newer release page.
Several issue titles ask whether older prompt revisions have stopped working or whether a request was refused. Issue 88 concerns gpt-5.6-sol-v42 behavior on a reverse-engineering request, while issue 83 reports a refusal involving self-hiding behavior. Those are individual reports, not controlled measurements, but they point to the core maintenance burden. The service and model can change without this repository changing, so yesterday's prompt result can expire outside Git.
Use the evidence format before the prompt
gpt-instruct is most valuable as an example of how to keep model identity, prompt bytes, run conditions, artifacts, provider blocks, and human judgments separate. Its authors publish missed gates and unrun stages instead of calling every candidate stable. That record is more useful to a research team than the promise that one instruction file will make an agent comply more often.
For daily coding, stock Codex is the safer baseline. If you still test gpt-instruct, isolate the Codex home, start with dry-run, define allowed tasks, and compare against the unchanged agent. A candidate that cannot beat the baseline on your own legitimate work is not rescued by 9,298 stars or a detailed internal scorecard.

