Recursive Self-Improvement: Build a Development System That Learns From Its Own Work
AI SystemsSep 22, 20263 min read

Recursive Self-Improvement: Build a Development System That Learns From Its Own Work

A road closure appears on the map, but the active route still crosses it. That gap can change both the next navigation feature and the way the development system designs features.

On this page
Back to blogHussam Ahmed

A road-closure marker that leaves turn-by-turn guidance untouched shows how a release can teach a development system to design better navigation features.

The ticket says, “Show temporary road closures.” An agent adds a red closure marker to the map. The map test passes; the marker appears exactly where the feed says it should.

A driver following an active route still hears, “Turn right onto Bridge Street.” Bridge Street is closed. The driver pulls over and finds another way. The map knew about the closure. The route and spoken instruction did not.

Follow the route all the way to arrival

The driver needs a feasible path to the destination, not a warning icon. Closure data has to exclude the affected edge from routing for the right direction and time window. The planner must find another path, the instruction engine must announce the new maneuver in time, and the arrival estimate must change with the route. If the closure feed is stale, the system needs an explicit way to handle that uncertainty.

A driver relying on spoken guidance may miss the marker. Guidance has to change before the blocked turn. The map, graph, guidance, and timing have to agree.

The release leaves a trail the team can inspect: a route replay whose selected path crosses a closed edge, a spoken instruction that contradicts the map, and a driver report about the detour. An agent can use those signals with the approved product brief to propose missing behavior. The product owner checks the driver need, data quality, and acceptable trade-offs before choosing what to ship.

Change how the next navigation feature begins

Fixing the closure flow improves this release. The development system can also revise the procedure that let “show a closure” stop at the map layer. A candidate feature-brief skill might require the agent to trace each new road condition through the graph, active route, spoken guidance, arrival estimate, and failure cases. Keep that procedure versioned so later tasks can use it.

Now give the agent a different request: “Avoid low-clearance bridges for trucks.” Does its brief ask for vehicle height, restriction data, a feasible alternative, warning distance, and what happens when the map lacks a reliable clearance value? Those are product questions to resolve with drivers and domain experts before code is written.

Let another route test the new method

Illustrative navigation feature experiment. Route replay evidence informs a candidate feature-brief procedure that traces road conditions through routing, guidance, and arrival estimates. A later truck-routing request tests the procedure under independent acceptance.Click to inspect full size

Call the current procedure H0 and the candidate H1. Give both the same product context and later feature requests. Keep the model, starting code, budget, and acceptance criteria fixed. Review the resulting plans and implementations with route scenarios: feasible arrivals, missed restrictions, late instructions, unwanted scope, rework, and cost. Include fresh requests that played no part in drafting H1.

Promote H1 only if those later tasks justify it. H1 then shapes future features, including the next proposal to change H1. The version number matters only if later routes improve.

Research such as SICA and the Darwin Gödel Machine examines agents changing their operating machinery under controlled tests. It does not establish that this navigation procedure improves a real product. That claim needs route-level acceptance and evidence from drivers.

The Codex guide and Claude Code guide show how to version a candidate procedure, compare it with the current one, and record a promotion decision. Their worked examples use a verification failure; the same experiment structure can test feature ideation for navigation systems.

Hussam Ahmed

Building large-scale systems by day, exploring the universe by night.

Keep reading

AI systemsJun 18, 2026

Make Your Coding Agent Work Like Fable 5: A Step-by-Step Guide

Most of Fable 5's quality came from the order it worked in, not its weights: it read before editing, checked after editing, and changed course when a tool result broke the plan. That order shows up in session logs, so you can measure it and move it onto the model you already use — with seven copy-paste prompts, a CLAUDE.md playbook, and a test hook.

Read article
AI SystemsJun 21, 2026

Loop Engineering for AI Coding Agents

Loop engineering is the practice of giving an AI coding agent a trigger, goal, execution mode, verifier, review gate, state log, and stop rule.

Read article

Featured project

See the Map Knowledge Graph reason about a live driving scene.

An interactive simulator with scenario switching, graph traversal, and step-by-step decision playback.

Open simulator

Follow new posts

I share build logs on AI systems, execution, and astrophotography as they ship — no schedule, only substance.