The Forge · 21 min mission

Recursive Self-Improvement in Codex

Turn repeated delivery failures into one measured Codex skill change.

codexrecursive improvementskillsevalsFact-checked 2026-09-22
On this page

A coding agent can finish a feature and still leave the system that produced it unchanged. Here you will take a repeated failure in software delivery, encode one repair in a Codex skill, and test whether that repair helps on later tasks.

The worked case is a navigation product. A team asks for temporary road closures. The agent adds closed-road markers to the map; the route search still chooses the closed edge and voice guidance still tells the driver to take it. That is a product failure. If the same integration miss appears across tasks, it is also a candidate harness failure. We will test a route-constraint verification skill, then try it on a different restriction such as vehicle height. The same protocol works for permissions, payment flows, migrations, refactors, and CI repair.

You need: a Git repository, the Codex CLI, one executable acceptance check or a named human reviewer, and permission to use two disposable worktrees. Commands that say npm run ... are examples: replace them with your repository's real commands. The evaluation script and fixtures live outside the agent's writable checkout. This guide is the implementation manual; the shorter story explains why the two loops matter.

The experiment file boundary

One frozen navigation task feeds separate H0 and H1 worktrees. Only H1 adds a skill. A controller outside the write boundary runs the same route evaluator on both results, then holds an audit task and a human promotion gate.
Pan the architecture plate on a narrow screen or enlarge it. The task fixtures and evaluator sit outside both agent worktrees.
ObjectOwnerMay H1 change it?
Application codeCodex during each task runYes, inside its run worktree
Harness: AGENTS.md, skills, tool setupEngineer proposing H1Only the single candidate change
Task briefs, fixtures, acceptance code, scoring ruleController or reviewerNo
Promotion decisionHuman ownerNo

1. Turn the user need into a frozen task

Start with behavior a user can observe. “Show closure pins” is too narrow for a navigation feature. A route planner consumes the same road restriction in search, displayed route geometry, ETA, and spoken maneuvers. Name those surfaces in the brief. Freeze it before either H0 or H1 run. H0 is the current instructions and skills; H1 is exactly H0 plus one candidate skill.

Put the task pool and evaluator in a controller directory outside both worktrees. Select a few development tasks to diagnose the pattern, separate validation tasks to choose the candidate, and untouched audit tasks to check transfer. Include a control task where the skill should stay inactive. For a first pilot, a small pool teaches you the mechanics; it does not establish a general performance claim.

Controller-owned tasks/NAV-41.md
markdown
# NAV-41 — temporary road closure
User: When a road closes, do not send me through it or tell me to turn onto it.
Given: a closure for edge E17 active from 09:00 to 11:00.
Deliver: ingest the restriction, exclude E17 during route search, update route geometry/ETA, and generate directions from the chosen route.
Acceptance: the 09:30 route and maneuvers contain no E17; a 12:00 control route may use E17; missing closure data is reported, not silently treated as safe.
Out of scope: redesigning the map UI.
Evidence: exact commit, command output, evaluator report, screenshot or route trace.
Stop: acceptance, a real blocker, or the agreed run budget.

The next fixture format is deliberately small. Your adapter must drive the real product (API, service, simulator, or browser), save its returned routeEdgeIds and maneuvers, and keep its invocation identical for H0 and H1. A hand-written result JSON does not test the product. Keep the fixture and evaluator unavailable for editing from the run worktrees; a read-only mount or a separate evaluator job works.

Controller-owned fixtures/NAV-41.json
json
{
  "caseId": "NAV-41",
  "departure": "09:30",
  "mustAvoidEdgeIds": ["E17"],
  "mustIncludeEdgeIds": [],
  "mustHaveRoute": true
}
Controller-owned fixtures/NAV-41-control.json
json
{
  "caseId": "NAV-41-control",
  "departure": "12:00",
  "mustAvoidEdgeIds": [],
  "mustIncludeEdgeIds": ["E17"],
  "mustHaveRoute": true
}
Controller-owned evaluate-route.mjs
javascript
import assert from 'node:assert/strict';
import { readFileSync } from 'node:fs';
 
const [fixturePath, resultPath] = process.argv.slice(2);
if (!fixturePath || !resultPath) throw new Error('usage: node evaluate-route.mjs fixture.json result.json');
const fixture = JSON.parse(readFileSync(fixturePath, 'utf8'));
const result = JSON.parse(readFileSync(resultPath, 'utf8'));
assert.equal(result.caseId, fixture.caseId);
assert.ok(Array.isArray(result.routeEdgeIds), 'adapter must return routeEdgeIds[]');
assert.ok(Array.isArray(result.maneuvers), 'adapter must return maneuvers[]');
if (fixture.mustHaveRoute) assert.ok(result.routeEdgeIds.length > 0, 'no route returned');
const spokenEdges = result.maneuvers.map((m) => m.edgeId);
for (const edge of fixture.mustAvoidEdgeIds) {
  assert.ok(!result.routeEdgeIds.includes(edge), `route uses restricted edge ${edge}`);
  assert.ok(!spokenEdges.includes(edge), `guidance uses restricted edge ${edge}`);
}
for (const edge of fixture.mustIncludeEdgeIds) {
  assert.ok(result.routeEdgeIds.includes(edge), `control route omitted ${edge}`);
}
console.log(`PASS ${fixture.caseId}`);

That evaluator catches an unsafe route and an unsafe voice instruction separately. It does not prove the adapter used real closure data or that ETA is correct: add integration checks for those claims before using them as acceptance criteria. A failed adapter or unavailable simulator is blocked, never a pass. Adapt the fields to your domain while keeping the principle: the oracle observes the product, not Codex's prose.

2. Capture a failure trace before writing a rule

Run current H0 on actual work. Save codex exec --json events, the final report, the diff, and the external evaluator result. The JSONL contains tool and usage events; -o saves the final message. That final message is a claim, not acceptance. In the navigation case, inspect which hop failed: restriction ingestion, graph-edge eligibility, route construction, guidance generation, or test coverage. A single bad implementation does not justify a permanent skill. Look for the same mechanism across multiple tasks, including a case where the agent could have checked it with available tools.

Write a one-line hypothesis that can be falsified: “When route behavior changes, a skill that traces restriction data through both route geometry and guidance reduces accepted misses without hurting unrelated tasks.” This is more useful than “be thorough.”

Trace evidenceLikely repairDo not confuse it with
Agent can inspect the graph and run route tests, but repeatedly stops at UI markersA narrow verification skill or task templateA one-off application bug
Route test cannot start because map data or a service is missingPinned service fixture or setup commandA reasoning failure
Agent reports green because existing tests only assert marker renderingImprove the controller-owned acceptance oracle, then freeze it before H0/H1Evidence that H1 helped
Agent lacks the tool or permission needed to inspect routing stateOne separately tested tool or permission changeA skill that merely says “check harder”
Capture one H0 trace (run from the controller shell)
bash
LAB="$PWD/rsi-lab"
APP="$PWD/my-app"
mkdir -p "$LAB/runs"
git -C "$APP" status --short
git -C "$APP" rev-parse HEAD > "$LAB/app-start.commit"
codex --version > "$LAB/codex.version"
codex exec -C "$APP" --sandbox workspace-write --json \
  -o "$LAB/runs/NAV-41-H0.final.txt" \
  "$(cat "$LAB/tasks/NAV-41.md")" \
  > "$LAB/runs/NAV-41-H0.events.jsonl"

Set LAB to a directory outside APP in your real run; the sample paths are placeholders. Put the frozen brief at $LAB/tasks/NAV-41.md and do this diagnostic run in a disposable checkout. Save the baseline test output first. If git status --short prints changes, preserve them and create a clean worktree at a recorded commit. The --sandbox workspace-write flag permits edits in the worktree; it does not protect a task fixture placed inside that same worktree.

3. Make one versioned harness change

Put stable repository facts and commands in AGENTS.md. Codex reads a chain of instruction files from the Git root toward the working directory; nearer instructions can override broader ones. Put this narrow procedure in .agents/skills/, with required name and description frontmatter. Its trigger matters: a skill that fires on every task can inflate cost and distract the agent. Keep H0 unchanged and store the H1 file separately until testing.

H1 only: .agents/skills/verify-route-constraints/SKILL.md
markdown
---
name: verify-route-constraints
description: Use when a task changes route selection, road restrictions, ETA, or turn-by-turn guidance.
---
 
For a route-affecting task:
1. Trace each restriction from ingestion to graph-edge eligibility. Name the source and validity window.
2. Verify that the chosen route geometry and every spoken maneuver refer to the same eligible edge sequence.
3. Run one negative case in which a restricted edge would otherwise be shortest, plus an unrestricted control.
4. Record the route edge IDs, maneuver edge IDs, departure time, evaluator command, and exact candidate commit.
5. Mark unavailable map data or simulator runs BLOCKED. Do not claim a pass from a rendered closure marker.
 
Skip this procedure for changes that cannot affect routing behavior. Do not edit task briefs, fixtures, or the external evaluator.

Do not add a new model, permission, hook, MCP server, and skill together. If H1 wins, you need to know which change caused it. If the failure is missing tool access rather than a verification habit, test a tool or environment repair as a separate candidate. A skill cannot enforce a security boundary; keep evaluator files and promotion rights under controller ownership.

Build a real experiment record

Experiment workbench

Produce the files. Compare paired runs.

The navigation case and its result rows are illustrative; the listed report files do not exist here. Replace them with your own task, then copy a contract, a candidate skill, a frozen experiment manifest, and a decision record. Enter results from independent checks; this page runs no code.

1. State the repeat failure and one change

2. Freeze the comparison before either run

Complete the fixed fields before treating the manifest or decision record as ready. Missing values are marked in the output.

3. Copy the repository artifacts

tasks/<task-id>.md
# Task contract — replace this example with one real task

User outcome: A route must never traverse a road edge closed at departure time.

Independent acceptance (owned outside the agent worktree):
- Run: npm run test:routes && npm run test:e2e -- route-closures
- Inspect the user journey and a control case.
- Record pass, fail, blocked, or unreviewed and the evidence path.

Frozen before the run: task brief, start commit, evaluator, permissions, model, environment, and budget.
Protected from the candidate: task briefs, fixtures, acceptance script, and scoring rules.
Stop at acceptance, a real blocker, or the budget. An unrun check is never a pass.

4. Record independent, paired outcomes

Each row is one task run from the same application commit under H0 and H1. Acceptance comes from the protected evaluator or a reviewer, not the agent's final message. “Audit” means a task unseen during candidate design and validation.

Candidate for human review: inspect the exact diffs and independent evidence before promotion.

Validation: 1 H1 wins / 0 losses / 1 ties. Fresh audit: 1 wins / 0 losses / 0 ties. Human rework: H0 75 min; H1 15 min.

The frozen setup is incomplete; no promotion decision can be made.

Replace the navigation example with your own failure. Copy the contract, candidate skill, manifest, and decision JSON. The paired table rejects unmatched starts, blocked checks, missing reviews, and critical regressions; it does not run your evaluator or approve H1.

4. Replay matched tasks in isolated Codex runs

For each validation task, create two fresh worktrees from that task's same application commit. H0 receives the current harness; H1 receives only the candidate skill. Keep the exact task brief, model, Codex version, dependencies, sandbox, allowed services, timeout, and evaluator commit fixed. Run each in a fresh codex exec process. Do not resume a prior chat: its context may carry the answer or the failed attempt.

Below is the worktree skeleton for one pair. It assumes APP points to a clean Git repo, BASE is a commit, and LAB is outside the repo. Save the H1 skill shown above at $LAB/candidate/verify-route-constraints/SKILL.md; keep it absent from H0. Run the two Codex commands serially if they share external services or state. The task brief is injected as text; the agent does not need write access to the controller directory.

Pair one task from equal starting code
bash
APP="/absolute/path/to/my-app"
LAB="/absolute/path/to/rsi-lab"
BASE="$(git -C "$APP" rev-parse HEAD)"
mkdir -p "$LAB/runs"
git -C "$APP" worktree add --detach "$LAB/NAV-41-H0" "$BASE"
git -C "$APP" worktree add --detach "$LAB/NAV-41-H1" "$BASE"
mkdir -p "$LAB/NAV-41-H1/.agents/skills/verify-route-constraints"
cp "$LAB/candidate/verify-route-constraints/SKILL.md" \
  "$LAB/NAV-41-H1/.agents/skills/verify-route-constraints/SKILL.md"
 
for ARM in H0 H1; do
  RUN="$LAB/NAV-41-$ARM"
  codex exec -C "$RUN" --sandbox workspace-write --json \
    -o "$LAB/runs/NAV-41-$ARM.final.txt" \
    "$(cat "$LAB/tasks/NAV-41.md")" \
    > "$LAB/runs/NAV-41-$ARM.events.jsonl"
  git -C "$RUN" diff > "$LAB/runs/NAV-41-$ARM.diff"
done

The example tests a newly added skill absent from the base commit. If your repository already has that skill, define H0/H1 at explicit harness commits instead of copying over it. Pin the model through your Codex configuration or CLI option supported by your installed version, and record the resolved value. Do not use --ephemeral if you rely on Codex's saved rollout history; the explicit event files above remain your portable trace. Never put secrets in shared trace files.

After each run, the controller drives the app with its adapter at 09:30 and 12:00, saves two product response files, and calls node "$LAB/evaluate-route.mjs" "$LAB/fixtures/NAV-41.json" "$LAB/runs/NAV-41-H0.route.json" plus the matching control fixture and result (then repeats for H1). The 12:00 fixture assumes E17 is the shortest open edge, so its required inclusion is meaningful. Run broader repository tests and inspect the diff as well. Do not let the worktree write $LAB/fixtures or $LAB/evaluate-route.mjs. Repeat on validation and control tasks, then on untouched audit tasks after candidate selection.

5. Compare per-task evidence and decide

Log one pair per task: split, start commit, harness fingerprint, resolved model and CLI version, tool permissions, environment/lockfile, evaluator version, raw trace, independent acceptance, usage or cost, human rework minutes, and any critical regression. Compare paired wins, losses, and ties. An aggregate pass rate can conceal that H1 fixed one task and broke another. Report blocked and unreviewed runs separately. A candidate with one dramatic win on a task used to design it has not demonstrated transfer.

Predeclare a rule before seeing H1: for example, validation wins must exceed losses; audit must have no net loss; critical safety or privacy regressions reject immediately; extra rework must remain under a stated limit. Small sets are directional. For a consequential change, use more tasks, repeat runs to expose variance, and get a fresh domain review. A human promotes the exact H1 commit only after reading failed cases. Keep the H0 commit for rollback and evaluate the next harness change against the promoted version, not against a moving mixture.

One paired result: controller-owned runs/NAV-41.pair.json
json
{
  "taskId": "NAV-41",
  "split": "validation",
  "applicationStartCommit": "<same SHA for both>",
  "evaluatorCommit": "<fixed SHA>",
  "model": "<resolved model ID>",
  "codexVersion": "<codex --version>",
  "environment": "<lockfile + service image hashes>",
  "budget": "<same time and token limit>",
  "H0": { "acceptance": "fail", "evidence": "runs/NAV-41-H0.report.json", "usageTokens": null, "humanReworkMinutes": 45 },
  "H1": { "acceptance": "pass", "evidence": "runs/NAV-41-H1.report.json", "usageTokens": null, "humanReworkMinutes": 10 },
  "criticalRegression": false,
  "reviewer": "<independent owner>"
}

Replace placeholders and null usage with observed values from the run artifacts; keep fail, blocked, and unreviewed distinct. This example pair is illustrative, not a measured claim about Codex.

Observed resultAction
H1 passes route-closure cases but breaks vehicle-height routingReject or narrow the trigger; rerun all pairs.
A map-data service was unavailableMark both affected runs blocked; restore the environment and replay.
Validation improves; untouched audit has no gainKeep H0 or gather more representative tasks.
Validation and audit improve, no critical regression, acceptable reworkSubmit the exact H1 skill and evidence for human promotion.

Adapt the method to any software task

Change the oracle, not the experimental controls. For a payment feature, check ledger entries, retries, and duplicate charges. For an authorization bug, test owner and unauthorized identities. For a database migration, verify old and new reads, rollback, and production-shaped data. For a refactor, assert invariant behavior and measure performance. For CI repair, evaluate a cold checkout in the target runner. In each case: user outcome → independent acceptance → repeated trace diagnosis → one harness change → matched replay → fresh audit → human decision.

If you code by directing agents rather than writing tests: ask a developer or a separate reviewer to own the acceptance checklist and run the same user journey on both worktrees. Save screenshots, API responses, and failure descriptions with task IDs. Do not ask the same agent that built H1 to judge its own result. A human checklist is slower than an automated evaluator, but it is still usable when the observations are concrete and recorded.

Troubleshooting: If Codex ignores the skill, check its path and frontmatter, the working directory and instruction chain, and test an explicit $verify-route-constraints invocation. If results vary, pin versions and repeat the pair. If tests are green but behavior is wrong, fix the external acceptance oracle before testing another candidate. A schema-constrained final report can improve log parsing, but it cannot certify the delivered code.

Reach the end and this star joins your charted sky.