The Navigator · 21 min mission
Recursive Self-Improvement in Claude Code
Build a candidate skill, replay matched software tasks, and judge it against product behavior.
On this page
This is a field procedure for making Claude Code better at later software tasks. You will diagnose a repeat failure from traces, change one part of the project harness, replay matched tasks, and decide from independent product evidence whether the change stays.
The example is a navigation feature. Users need temporary road closures respected by route search and spoken directions. The first implementation shows the closure on a map but still drives across it. The proposed improvement is a narrow route-constraint skill that forces the agent to trace a restriction through ingestion, graph eligibility, route geometry, and guidance. The audit task switches to a vehicle-height restriction, so the candidate must transfer beyond the exact closure it was written for.
You need a Git repo, Claude Code CLI, an executable route or browser check (or a named reviewer with a concrete checklist), and room for two disposable worktrees. Replace the navigation adapter and commands with your application's real interfaces. For the product story, read the blog; for the equivalent Codex mechanics, see the Codex guide.
The experiment file boundary
1. Draw the control boundary
There are two changes. Delivery edits application code for a user task. Improvement edits Claude Code's future instructions, skills, or checks. The proposer may diagnose and draft H1, but it cannot rewrite the task pool, fixtures, evaluator, score rule, or promotion decision. Put those controller files outside the agent's writable worktrees. CLAUDE.md is useful context, not an access-control mechanism; file permissions, tool policy, and workspace isolation enforce boundaries.
H0 is today's project harness. H1 is H0 plus exactly one candidate change. A richer prompt, new model, extra tools, and a skill added together make the result uninterpretable. Start with a bounded skill; test other changes separately.
| Pool | When used | Navigation case |
|---|---|---|
| Development | Find and classify repeated misses; H1 may be designed from these | Closure marker without route exclusion |
| Validation | Choose or reject the candidate; select before inspecting H1 results | Other closure windows, detours, and an unrelated map-copy control |
| Fresh audit | Check transfer after the candidate is frozen | Truck height restriction on a different edge and route |
Write an outcome-based contract before the agent starts. The brief belongs to the controller and is identical in both arms. Example:
# NAV-41 — temporary road closure
User need: never route me through a road closed at departure time, including voice instructions.
Input: edge E17 closes 09:00–11:00; the shortest unrestricted route uses E17.
Acceptance: at 09:30, routeEdgeIds and maneuver edgeIds exclude E17; at 12:00 a control route may use it. The ETA is calculated from the selected route. Missing restriction data is reported.
Evidence: application commit, real product route response, evaluator result, test output, and unresolved findings.
Stop: accepted result, real blocker, or the agreed turn/cost budget.
Out of scope: a map UI redesign.Build the independent oracle around the application, not Claude's summary. For this case an adapter calls the real route API or simulator and writes { "caseId": "NAV-41", "routeEdgeIds": ["..."], "maneuvers": [{"edgeId":"..."}], "etaSeconds": 420 }. An evaluator outside the worktree checks that neither route nor spoken maneuvers use restricted E17, and checks the control route and ETA separately. The Codex guide contains a complete Node evaluator for the edge checks. Copy it unchanged for both Claude arms, then add an ETA assertion for your application's expected tolerance. A hand-written response file only tests the file, not the product.
2. Collect a real H0 trace
Record a clean starting commit, claude --version, resolved model, dependency lockfile hash, service image versions, permissions, and the task brief. Run the current harness on several tasks and preserve the per-run stream. claude -p is the scriptable entry point; --output-format stream-json --verbose exposes messages and tool events. The agent's final answer and tool transcript help diagnose the failure, but the external evaluator determines acceptance.
The command below is an example for a repository whose tests are npm test and npm run test:routes. --tools limits built-in tool availability; --allowedTools preapproves listed uses; it does not restrict tools on its own. --disallowedTools "mcp__*" removes MCP tools for this controlled run. dontAsk denies any unresolved permission request, so a missing command grant becomes a visible blocked run. Keep the permission set equal for H0 and H1.
RUN="/absolute/path/to/NAV-41-H0"
LAB="/absolute/path/to/rsi-lab"
mkdir -p "$LAB/runs"
(
cd "$RUN"
claude -p "$(cat "$LAB/tasks/NAV-41.md")" \
--output-format stream-json --verbose \
--max-turns 20 --max-budget-usd 5 \
--permission-mode dontAsk \
--tools "Read,Glob,Grep,Edit,Write,Bash" \
--allowedTools "Read,Glob,Grep,Edit,Write,Bash(npm test *),Bash(npm run test:routes *)" \
--disallowedTools "mcp__*" \
> "$LAB/runs/NAV-41-H0.events.jsonl"
)Do not add --bare here: it skips project instructions, skills, hooks, and MCP, so it would erase the harness difference under test. Store event streams away from the worktrees and redact credentials before sharing them. If a run ends at the turn cap, hits the cost cap, or cannot invoke a needed tool, record that exact result; do not convert it into “pass.” Claude's cost field is an estimate, not a business-cost measure.
Read the trace like a developer, not like a prompt editor. Identify the first unsupported assumption. Did Claude miss the closure ingestion code? Did it connect the map marker to the routing graph? Did it run a test that only checked UI state? Could it have reached the route API? Record the evidence path and the smallest recurrent cause. If only one task failed, first repair the application or oracle. A durable harness change needs repeated evidence or a high-impact failure with a clear, testable mechanism.
| What the trace shows | Candidate repair | What to verify first |
|---|---|---|
| Available route code and tests were ignored across tasks | A task-triggered skill | The same miss appears in more than one trace |
| A Bash command was denied or simulator unavailable | A narrow permission or service setup change | The command works outside the agent and the grant is acceptable |
| Claude ran tests, but they only covered the map marker | A stronger controller-owned route oracle | Freeze the new oracle before comparing H0/H1 |
| One implementation broke despite adequate checks | Fix that application task | Whether a recurring harness mechanism exists at all |
3. Encode one candidate in a project skill
Keep stable repository commands and constraints in CLAUDE.md or .claude/CLAUDE.md. Put the narrow procedure in .claude/skills/<name>/SKILL.md. The skill description controls when Claude considers it. Use a positive trigger and an exclusion so a route check does not consume time on unrelated map styling. Do not hardcode NAV-41's answer or fixture edge E17 into the skill; the point is transfer.
The project skill below may be invoked when its description matches. In a diagnostic session you can call it explicitly with /verify-route-constraints to check discovery; the trial itself should use ordinary user requests so it measures natural selection.
---
name: verify-route-constraints
description: Use when a task changes route selection, road restrictions, ETA, or turn-by-turn guidance.
---
For a route-affecting task:
1. Trace restriction data from ingestion through validity time and graph-edge eligibility.
2. Check that displayed route geometry and spoken maneuvers describe the same eligible edge sequence.
3. Test a restricted edge that would otherwise be chosen; also test an unrestricted control.
4. Save the observed route edge IDs, maneuver edge IDs, departure time, verifier command, and candidate commit.
5. Report absent map data, simulator access, or unrun checks as BLOCKED. A closure icon is not route evidence.
Skip tasks that cannot change routing behavior. Never edit the task brief, fixtures, or external evaluator.If the root cause is a permission or environment failure, a skill may be the wrong repair. Test a tool grant, service fixture, or build command as its own H1. Hooks can log or veto specific tool calls: PreToolUse may block before a call, while PostToolUse can only observe after success. A command hook must exit 2 to block. Hooks still do not make a writable evaluator trustworthy; keep evaluation outside the worktree.
Prepare the H0/H1 files and evidence
Experiment workbench
Produce the files. Compare paired runs.
The navigation case and its result rows are illustrative; the listed report files do not exist here. Replace them with your own task, then copy a contract, a candidate skill, a frozen experiment manifest, and a decision record. Enter results from independent checks; this page runs no code.
1. State the repeat failure and one change
2. Freeze the comparison before either run
Complete the fixed fields before treating the manifest or decision record as ready. Missing values are marked in the output.
3. Copy the repository artifacts
# Task contract — replace this example with one real task User outcome: A route must never traverse a road edge closed at departure time. Independent acceptance (owned outside the agent worktree): - Run: npm run test:routes && npm run test:e2e -- route-closures - Inspect the user journey and a control case. - Record pass, fail, blocked, or unreviewed and the evidence path. Frozen before the run: task brief, start commit, evaluator, permissions, model, environment, and budget. Protected from the candidate: task briefs, fixtures, acceptance script, and scoring rules. Stop at acceptance, a real blocker, or the budget. An unrun check is never a pass.
4. Record independent, paired outcomes
Each row is one task run from the same application commit under H0 and H1. Acceptance comes from the protected evaluator or a reviewer, not the agent's final message. “Audit” means a task unseen during candidate design and validation.
Candidate for human review: inspect the exact diffs and independent evidence before promotion.
Validation: 1 H1 wins / 0 losses / 1 ties. Fresh audit: 1 wins / 0 losses / 0 ties. Human rework: H0 75 min; H1 15 min.
The frozen setup is incomplete; no promotion decision can be made.
4. Replay equal tasks from equal starts
Each paired task gets two clean worktrees at the same application commit. H1 gets only the candidate .claude/skills/verify-route-constraints/SKILL.md; H0 does not. Launch a fresh claude -p process for each arm. Pin the model, CLI version, max turns, budget, permissions, environment and services. Reset databases and caches between arms. Run serially when a shared simulator or service could leak state. Do not reuse conversation history or let H1 see the H0 diff.
This shell skeleton prepares one pair. It assumes the baseline commit has no skill at that path. Save the H1 skill shown above at $LAB/candidate/verify-route-constraints/SKILL.md and the brief at $LAB/tasks/NAV-41.md, outside the app repo.
APP="/absolute/path/to/my-app"
LAB="/absolute/path/to/rsi-lab"
BASE="$(git -C "$APP" rev-parse HEAD)"
mkdir -p "$LAB/runs"
git -C "$APP" worktree add --detach "$LAB/NAV-41-H0" "$BASE"
git -C "$APP" worktree add --detach "$LAB/NAV-41-H1" "$BASE"
mkdir -p "$LAB/NAV-41-H1/.claude/skills/verify-route-constraints"
cp "$LAB/candidate/verify-route-constraints/SKILL.md" \
"$LAB/NAV-41-H1/.claude/skills/verify-route-constraints/SKILL.md"
git -C "$LAB/NAV-41-H0" rev-parse HEAD
git -C "$LAB/NAV-41-H1" rev-parse HEADfor ARM in H0 H1; do
RUN="$LAB/NAV-41-$ARM"
(
cd "$RUN"
claude -p "$(cat "$LAB/tasks/NAV-41.md")" \
--output-format stream-json --verbose \
--max-turns 20 --max-budget-usd 5 \
--permission-mode dontAsk \
--tools "Read,Glob,Grep,Edit,Write,Bash" \
--allowedTools "Read,Glob,Grep,Edit,Write,Bash(npm test *),Bash(npm run test:routes *)" \
--disallowedTools "mcp__*" \
> "$LAB/runs/NAV-41-$ARM.events.jsonl"
)
git -C "$RUN" diff > "$LAB/runs/NAV-41-$ARM.diff"
doneAfter each run, the controller calls the real application adapter, then the protected evaluator. Save its JSON report, raw route response, worktree diff, stream trace, and a reviewer verdict. For each pair record pass/fail/blocked/unreviewed, human rework minutes, and estimated model cost. A pass requires the same frozen acceptance checks on both arms. A closure case that improved while a truck-height audit case failed is a regression, not a clean win.
Use the workbench's paired rows to record the outcomes. Pick a decision rule before reading H1: validation wins must exceed losses, the untouched audit must not regress, critical safety failures reject immediately, and extra rework or cost must stay below your limit. Small samples are a screening signal. Repeat runs and add domain coverage before making a broad productivity claim. A human owner reviews the exact H1 commit and keeps H0 available for rollback.
{
"taskId": "NAV-41",
"split": "validation",
"applicationStartCommit": "<same SHA for both>",
"evaluatorCommit": "<fixed SHA>",
"model": "<resolved model ID>",
"claudeVersion": "<claude --version>",
"permissions": "<same --tools and allow/deny rules>",
"budget": { "maxTurns": 20, "maxUsd": 5 },
"H0": { "acceptance": "fail", "evidence": "runs/NAV-41-H0.report.json", "estimatedUsd": null, "humanReworkMinutes": 45 },
"H1": { "acceptance": "pass", "evidence": "runs/NAV-41-H1.report.json", "estimatedUsd": null, "humanReworkMinutes": 10 },
"criticalRegression": false,
"reviewer": "<independent owner>"
}Fill the cost fields from the run metadata and keep raw traces alongside the report. The values above illustrate the record shape; they are not a measured Claude Code result. Preserve blocked and unreviewed as distinct outcomes.
5. Use Claude's plugin eval for a skill probe
If your candidate is a reusable skill, Claude Code v2.1.269+ has claude plugin eval. Package a copy of the skill as a test plugin, add realistic cases, and compare runs with and without the plugin. It starts fresh isolated sessions and normally repeats each case three times per arm. This can expose a skill that never activates or harms unrelated tasks. It is a supplement to the product-level worktree test: plugin graders over text or created files do not prove your navigation service routes safely.
Create a small plugin directory with .claude-plugin/plugin.json, skills/verify-route-constraints/SKILL.md (a copy of the project skill), and evals/<case>/prompt.md. The case prompt should describe a real request without naming the skill. A skill-invocation grader is useful for diagnosis but cannot be the outcome score: without the plugin it is impossible to invoke that skill, which would inflate the measured delta. Grade the result, include an unrelated control case, and use the separate product oracle for release. For code tasks, seed the same Git fixture in each run with a trusted context.scaffold_script and --scaffold; that script runs as you, so review it first.
{
"name": "route-constraints-lab",
"version": "0.1.0"
}schema_version: "1.1"
name: road-closure
context:
scaffold_script: fixture.sh#!/usr/bin/env bash
set -euo pipefail
git clone --no-hardlinks /absolute/path/to/my-app .
git checkout YOUR_FROZEN_APPLICATION_COMMIT
npm ci---
max_turns: 20
timeout_seconds: 600
allowed_tools: [Read, Glob, Grep, Skill]
---
Implement the temporary road-closure feature in this fixture application. A closure active at departure time must affect the selected route, displayed geometry, ETA, and spoken maneuvers. Run the supplied tests and report remaining blockers.---
type: llm
---
PASS if the final response identifies an implemented route-search eligibility check, guidance derived from the selected route, and the exact tests it ran.
FAIL if it reports only a map marker, omits route or guidance behavior, or calls an unrun check passing.This rubric detects an obvious workflow miss, but it judges a response. It can be fooled by an incorrect claim; keep the controller-run route evaluator as the acceptance source. Add a second control case for a map-copy edit where the route skill should be inert. If you create a tool_used: Skill grader for debugging invocation, read its indicator separately from the with/without outcome score.
claude --version
cd /absolute/path/to/route-skill-plugin
claude plugin eval . --scaffold --allow-tools Write Edit "Bash(npm test *)" --no-publish
# For cheap grader iteration only:
claude plugin eval . --case road-closure --runs 1 --ablation none --scaffold --allow-tools Write Edit "Bash(npm test *)" --no-publishThe first command's with/without delta reflects the plugin eval suite and its graders; the second is a single-arm diagnostic, not a promotion result. If your Claude build does not expose plugin eval or says it is unavailable, use the manual worktree experiment above. For a skill-only plugin test, remember that Claude's isolated eval workspace does not automatically contain your full project's CLAUDE.md or dependencies; scaffold those when they are part of the behavior under test.
Transfer this to your own work
For an auth change, the oracle uses two identities and checks denied operations leave no data behind. For a data migration, it checks old/new reads, backfill counts, and rollback on production-shaped fixtures. For a refactor, it freezes behavior tests and performance limits. For a new feature, it tests the complete user journey and a neighboring journey the feature could break. The harness candidate changes the repeated working method (for example, test forbidden identities or trace derived state), not the answer to one task.
If you direct agents without writing code, keep the same structure: record a real user outcome, ask an engineer or separate reviewer to own a concrete checklist, collect a screenshot or API response for each arm, and keep the same model, task, budget, and starting commit. You can judge a feature manually; you cannot infer improvement from Claude saying it followed the skill.
Troubleshooting: A missing skill may be at the wrong path or have a weak description; test discovery in a fresh session. A permission denial is a blocked run, not a model failure. A large plugin-eval delta with unchanged product acceptance means the grader is measuring the wrong thing. If H1 changes a protected test or task brief, discard that pair and repair the boundary before replaying it.
Reach the end and this star joins your charted sky.