mrkeyoor.com_
Wed 07 Oct 14:39 UTC
LLM Toolsevaluationupdated 07 Oct 2026

gpt-instruct review

gpt-instruct is a Chinese-first project, with a full English README, that packages custom Codex instructions and the evaluation machinery used to compare them. It tries to make Codex act on complex requests earlier, preserve work across turns, verify artifacts, and keep a runnable rollback.

Verdict

Our run installed 35 packages in 10 seconds and built in 1 second, but commit 3ab84df exposed no test target, so the easy setup does not prove the instruction pack works. Treat gpt-instruct as a detailed prompt-research archive, not a dependable upgrade to Codex. Its rollback tooling and candid gates are useful, yet the current prereleases miss full B and never run C, while the authors themselves warn about account risk.

We ran it

Lab card: what happened when we ran gpt-instructScreenshot of gpt-instruct (mdx-tom.github.io/gpt-instruct)
Install✓ · 10s35 packages · 37 MB
Build✓ · 1s
Testsn/ano test script
Known vulns0(pip-audit)
Repo93 files~2,816 lines of source · 4.9 MB · 2 CI workflows

Answers from our run

Does gpt-instruct build from source?

Dependencies installed in 10 seconds (35 packages), and the build succeeded in 1 seconds. We cloned commit 3ab84df into a clean Debian container with 3 CPUs and no project-specific setup.

Does gpt-instruct have tests you can run?

Not through a standard command: the project exposes no test script or target that our harness could run.

Does gpt-instruct have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use gpt-instruct?

Anyone who depends on a stable, supported Codex configuration: the README labels the gpt-6-astra and gpt-6.1 lines as prereleases, and neither passed the full B gate or ran C.

What are the alternatives to gpt-instruct?

Codex, promptfoo, JailbreakBench. Our run installed 35 packages in 10 seconds and built in 1 second, but commit 3ab84df exposed no test target, so the easy setup does not prove the instruction pack works.

Setup4/510-second install, dry-run, backup, and reset; config still changes
Docs5/5Paired Chinese and English guides explain gates and rollback
Community4/59,298 stars, an October 4 push, and recent issue replies
Maturity2/5Measured commit had no test target; current lines remain prerelease

Who it’s for

AI-safety researchers studying how custom instructions change coding-agent behavior.
Codex power users who will inspect the prompt, use dry-run mode, and keep the generated backup.
Prompt evaluators who want the project's A, B, C, and JailbreakBench evidence format as a reference.
Bilingual teams comfortable checking the Chinese source README against its English counterpart.

Who it’s NOT for

Anyone who depends on a stable, supported Codex configuration: the README labels the gpt-6-astra and gpt-6.1 lines as prereleases, and neither passed the full B gate or ran C.
Users unwilling to risk account action: the README explicitly warns that custom instructions carry account risk, and issue 10 remains an open first-person ban report.
Teams seeking an independent security benchmark: the project authors its prompts, test gates, and manual judgments, so outside validation is still needed.
People who want a clean lab test signal: our measured commit had no test script or target, so its test step was skipped.
Users who only want stock Codex behavior and official defaults: this project changes model_instructions_file in Codex configuration.

Setup reality

Our fresh Debian run installed commit 3ab84df in 10 seconds, adding 35 packages and using 37 MB. The build passed in 1 second. The lab found no test script or target, so tests were skipped; pip-audit reported 0 known vulnerabilities.

Applying a package writes a custom model_instructions_file into Codex configuration. The installer has dry-run, backup, reset, and explicit snapshot-restore paths. Full model evaluation needs Codex access, while the optional official JailbreakBench judge needs a TOGETHER_API_KEY.

The measured checkout had 93 files, about 2,816 source lines, 2 CI workflows, no Dockerfile, and no tests directory. Current model packages move quickly, and the README warns that using custom instructions can put an account at risk.

The prompt pack is tied to named Codex model lines

The main README is in Chinese, with a synchronized English version and matching diagrams. gpt-instruct does not present one timeless system prompt. It maintains a stable gpt-5.6-sol-v45 package alongside gpt-6-astra-v2-rc1 and gpt-6.1-sol-v1-rc2 prereleases. Each file is bound to a model identity and evaluation history, which is the right way to document instructions whose behavior can change when the underlying model changes.

The intended change is direct: make Codex begin tool or file work earlier, carry confirmed state into later turns, verify produced artifacts, and leave a tested rollback. A Python installer previews or applies a chosen package, accepts a custom ZIP or Markdown file, and can reset the one configuration field it owns. It uses Codex's normal model_instructions_file setting rather than patching a binary or intercepting traffic.

The release gates expose incomplete results

The project separates its tests into A, B, and C. A contains 4 cases and requires two fresh runs plus artifact checks. B calls for 66 cases and 74 turns. C contains 120 medium-reasoning cases and runs only after A and B pass. The two current prerelease prompts passed their fresh A gates but fell short of the complete B requirement, so C was not run.

That disclosure matters more than a chart trending upward. The README reports 42 of 50 non-cloud B cases for the Astra candidate and 34 of 50 for the 6.1 candidate, with cloud attempts listed separately. It refuses to merge unlike runs into a prettier fraction. Still, these are project-authored cases with manual decisions. A team adopting the prompt should repeat representative tasks against stock Codex and score them without knowing which configuration produced each response.

An optional JailbreakBench track is kept apart from A, B, and C. Its official judge requires a specific Together-hosted model and leaves rows unjudged when that backend is missing. Human fields separately record refusal, cheating, and protocol violations. That method is careful, but it also shows the labor behind the headline scores: a ZIP is easy to install, while a credible comparison needs frozen inputs, credentials, repeated runs, and review.

What happened when we ran it

Our sandbox installed commit 3ab84df in 10 seconds. The Python environment added 35 packages and occupied 37 MB on disk. A build step completed in 1 second, and pip-audit found 0 known vulnerabilities. The unprivileged Debian container had 3 CPUs, 8 GB of RAM, Python 3.12, and no secrets.

Tests did not fail because none were started. The measured checkout exposed no test script or target, so the lab skipped that step. It contained 93 files, about 2,816 lines of source, and 4.9 MB of checked-out data, with 2 CI workflow files but no Dockerfile or tests directory. Those facts describe commit 3ab84df; they do not validate the README's current prompt scores or newer evaluation folders.

A 10-second install therefore answers only the packaging question. It does not tell us whether the instructions improve first-turn completion, preserve legitimate safety boundaries, or behave consistently across accounts and model revisions. The repository's own method treats those as model runs followed by human review, which our dependency and build check did not perform.

Applying it changes one sensitive configuration field

The deployment script can preview its work before writing, and it saves the previous state. Reset is scoped to model_instructions_file, leaving provider, model, authentication, and unrelated settings alone. Full snapshot restore is a separate explicit action. These choices lower the recovery cost if a candidate prompt makes Codex less useful.

Recovery is still worth rehearsing. A custom instruction file affects every request handled under that Codex home, not just the experiment that motivated the install. Use a separate CODEX_HOME, compare the generated diff, and keep experimental accounts and workspaces away from ordinary production access. The README itself says custom instructions can create account risk and advises use only in environments you are entitled to operate.

October activity is high, but model drift is the product risk

GitHub showed 9,298 stars and 30 open issues and pull requests on October 7, 2026. The last push was October 4, and issue 89 received a response and closed on October 5. That is healthy recent activity for a repository created in July 2026. The latest GitHub release, however, is the September 6 gpt-6-astra-v1 tag; newer release candidates are distributed from the repository rather than a newer release page.

Several issue titles ask whether older prompt revisions have stopped working or whether a request was refused. Issue 88 concerns gpt-5.6-sol-v42 behavior on a reverse-engineering request, while issue 83 reports a refusal involving self-hiding behavior. Those are individual reports, not controlled measurements, but they point to the core maintenance burden. The service and model can change without this repository changing, so yesterday's prompt result can expire outside Git.

Use the evidence format before the prompt

gpt-instruct is most valuable as an example of how to keep model identity, prompt bytes, run conditions, artifacts, provider blocks, and human judgments separate. Its authors publish missed gates and unrun stages instead of calling every candidate stable. That record is more useful to a research team than the promise that one instruction file will make an agent comply more often.

For daily coding, stock Codex is the safer baseline. If you still test gpt-instruct, isolate the Codex home, start with dry-run, define allowed tasks, and compare against the unchanged agent. A candidate that cannot beat the baseline on your own legitimate work is not rescued by 9,298 stars or a detailed internal scorecard.

Alternatives

ProjectWhat it isPick it when
Codex gh↗OpenAI's terminal coding agent with its supported configuration and release path.pick this instead when you want the official agent without a third-party behavior override.
promptfoo gh↗A test and red-team framework for comparing prompts, agents, and model providers.pick this instead when repeatable cross-model evaluation matters more than installing one prompt pack.
JailbreakBenchAn open benchmark and dataset for evaluating language-model jailbreaks.pick this instead when you need a standard research benchmark without changing your daily Codex instructions.

What people are saying

  1. [github-trending] MDX-Tom/gpt-instruct

Sources

  1. gpt-instruct repository and Chinese README
  2. gpt-instruct English README
  3. gpt-6-astra v1 release
  4. Open account-risk report
  5. gpt-5.6-sol v42 behavior report

More llm tools reviews

minorun-marp-skill · jev-skill · jev-pruner · llm-d-router · Rapid-MLX · simple-jev · the whole board →