The Forge · 11 min mission
Programmatic Codex: SDK & Subagents
Drive Codex from TypeScript or Python and fan tasks out to parallel sub-agents.
On this page
- When the SDK beats codex exec
- TypeScript: @openai/codex-sdk
- Resuming threads and streaming progress
- Python: openai-codex, sync and async
- Structured output: get JSON back, not prose
- Subagents: one orchestrator, many workers
- Set a deliberate concurrency ceiling
- Senior scenario: an async PR-review bot over 50 repos
- When to reach for the SDK over the CLI
codex exec is great until you need a program around the agent. The moment you want to loop over 50 repositories, parse the agent's answer as typed data, run ten reviews at once, or wire Codex into a webhook handler, shelling out to a CLI and scraping its stdout stops being clever and starts being a liability.
The Codex SDK is the same agent — same model, same sandbox, same config — exposed as a library you import. You get a thread object, you call run(), you get a result back in your own process. No subprocess parsing, no fragile string munging, and you can fan the work out across subagents that run in parallel. This guide is about that: when the SDK beats the CLI, the exact TypeScript and Python surface, how to get structured output back, and how to orchestrate a swarm of agents without melting your machine.
When the SDK beats codex exec
codex exec "..." is the right tool for one-shot automation: a CI step, a git hook, a shell pipeline. It streams progress to stderr and prints only the final agent message to stdout, so codex exec "summarize the diff" | pbcopy just works. Reach for the SDK the moment any of these is true:
- You need the result as typed data, not prose to re-parse — a JSON object with fields you can branch on.
- You need to run many agents concurrently and collect their results (the async/fan-out case).
- The agent is one step in a larger program — a server, a queue worker, a bot — where managing a child process per request is the wrong abstraction.
- You want to resume a long-lived thread across separate invocations and keep its context.
The dividing line is simple: if a single string in and a single string out is enough, use codex exec. If you need control flow, types, or parallelism around the agent, use the SDK.
codex exec vs. the SDK
codex exec (CLI)
Shape: one prompt in, final message on stdout.
codex exec "fix the failing test" \
--sandbox workspace-write \
--output-schema ./schema.json \
-o result.jsonPerfect for CI steps, git hooks, and shell pipelines. You get JSON Lines with --json and a final message you can pipe. But control flow lives in bash, and parallelism means juggling child processes.
@openai/codex-sdk / openai-codex
Shape: a thread object in your process; run() returns a typed result.
const codex = new Codex();
const thread = codex.startThread();
const turn = await thread.run("fix the failing test");
console.log(turn.finalResponse);Control flow lives in your language. Loop, await, Promise.all, branch on a parsed object. This is the only sane path once you have dozens of tasks or need the answer as data.
TypeScript: @openai/codex-sdk
Install @openai/codex-sdk and you get three calls that cover almost everything: construct a client, start a thread, run a turn.
import { Codex } from "@openai/codex-sdk";
const codex = new Codex(); // uses your existing Codex auth/config
const thread = codex.startThread(); // a fresh conversation
const turn = await thread.run("Make a plan to diagnose and fix the CI failures");
console.log(turn.finalResponse); // the agent's final message
console.log(turn.items); // every item it produced this turnA thread is a conversation with memory; a turn is one run() and the items it produced. run() resolves to a turn object whose finalResponse is the agent's last message and whose items array holds everything that happened — reasoning, command executions, file changes. Authentication and configuration are inherited from your normal Codex setup, so a script that runs locally needs no extra wiring; in CI you pass credentials through the environment the SDK reads.
startThread() takes options to pin the thread to a project and relax the git guardrail:
const thread = codex.startThread({
workingDirectory: "/path/to/project",
skipGitRepoCheck: true,
});Both workingDirectory and skipGitRepoCheck are SDK options — the second is the SDK equivalent of the CLI's --skip-git-repo-check, which you need when the agent runs somewhere that isn't a git repo.
Current TypeScript releases also accept a typed config object and raw TOML-compatible configOverrides on the Codex constructor. Raw overrides win when both set the same key. Use the typed object for ordinary settings; reserve raw overrides for a key the SDK types do not expose yet:
const codex = new Codex({
config: { model: "gpt-5.6-terra" },
configOverrides: ["model_reasoning_effort=\"max\""],
});Resuming threads and streaming progress
Threads are persisted to ~/.codex/sessions. If your process restarts — or a webhook fires a follow-up an hour later — you do not lose the conversation. Reconstruct it from its id and keep going:
// First invocation
const thread = codex.startThread();
await thread.run("Start refactoring the auth module");
const savedThreadId = /* persist this id somewhere durable */;
// Later, in a fresh process
const thread2 = codex.resumeThread(savedThreadId);
await thread2.run("Now add tests for what you changed");resumeThread(threadId) reconnects to an existing thread by id and returns a thread you can run() again, with all prior context intact.
For long turns where you want to react to intermediate progress — show a tool call, stream tokens to a UI, surface file diffs as they happen — use runStreamed() instead of run(). It hands back an async iterable of events:
const { events } = await thread.runStreamed("Audit the codebase for N+1 queries");
for await (const event of events) {
switch (event.type) {
case "item.completed":
// a reasoning step, command, or file edit finished
break;
case "turn.completed":
// the whole turn is done
break;
}
}The same event vocabulary (thread.started, item.completed, turn.completed, error) is what the CLI emits with --json — runStreamed is that stream, in your language, without parsing JSON Lines by hand.
Python: openai-codex, sync and async
The Python package is openai-codex, and its import namespace is openai_codex. It ships two clients used as context managers: Codex (synchronous) and AsyncCodex (asyncio). You start a thread with thread_start(...), which takes model and sandbox directly:
from openai_codex import Codex, Sandbox
with Codex() as codex:
thread = codex.thread_start(model="gpt-5.6-terra", sandbox=Sandbox.workspace_write)
result = thread.run("Make a plan to diagnose and fix the CI failures")
print(result.final_response)Note the snake_case: the Python result exposes final_response where TypeScript exposes finalResponse. The sandbox argument takes one of the presets below — Sandbox.workspace_write lets the agent edit files inside the workspace, which is what you want for "fix this" tasks.
The async client is the one that matters for scale. AsyncCodex gives you awaitable thread_start and run, which means you can launch dozens of independent agents and gather their results with asyncio.gather — no subprocess pool, no thread pool, just coroutines. That is the senior scenario later in this guide.
| SDK preset (Python) | CLI flag | What the agent can touch | Use it for |
|---|---|---|---|
Sandbox.read_only | --sandbox read-only | Read files only — no writes, no network | Review, audit, Q&A over a repo |
Sandbox.workspace_write | --sandbox workspace-write | Read and write inside the workspace | Fixes, refactors, codegen — the common case |
Sandbox.full_access | --sandbox danger-full-access | Unrestricted filesystem and network | Only in throwaway/controlled environments |
Structured output: get JSON back, not prose
The single biggest reason to drive Codex programmatically is to stop parsing English. Both the CLI and the SDK can enforce a JSON Schema on the final answer so run() hands you data your code can branch on.
On the CLI, --output-schema <file> points at a JSON Schema and the final message is guaranteed to match it:
{
"type": "object",
"properties": {
"project_name": { "type": "string" },
"languages": { "type": "array", "items": { "type": "string" } }
},
"required": ["project_name", "languages"]
}codex exec "Extract this repo's metadata" --output-schema ./schema.json -o output.jsonIn the TypeScript SDK the same idea is a per-turn option — pass outputSchema to run() and the agent's answer conforms to it:
const schema = {
type: "object",
properties: { severity: { type: "string" }, files: { type: "array", items: { type: "string" } } },
required: ["severity", "files"],
additionalProperties: false,
};
const turn = await thread.run("Triage this PR's risk", { outputSchema: schema });
const verdict = JSON.parse(turn.finalResponse); // typed, branchableYou do not have to hand-write the schema: generate it from a Zod schema with zod-to-json-schema (target "openAi") and keep one source of truth for both validation and the agent contract.
Two knobs shape how hard the agent thinks before it answers. Reasoning effort is set on the CLI with -c model_reasoning_effort=<level> and via config in the SDK. Current releases recognize low, medium, high, xhigh, max, and ultra, subject to model support; Ultra can use subagents and is not a cheap way to make a routine classification slightly better. The practical pattern is a two-phase pipeline: a low-effort pass to classify or filter (does this PR even need review?), then a higher-effort pass with a strict outputSchema only on the items that survived.
Subagents: one orchestrator, many workers
A single thread is one worker. Subagents let a primary agent spawn child agents that run in their own context and report back — so a big task splits into parallel pieces instead of one long serial slog. Codex ships three built-in agents:
default— the general-purpose fallback agent.worker— an execution-focused agent for implementation and fixes.explorer— a read-heavy agent tuned for codebase exploration.
The mental model: an explorer maps the territory (where does auth live? which files import this?), a worker changes it (apply the fix, write the test), and the orchestrator stitches their results together. You spawn explorers to investigate in parallel without polluting the main thread's context, then hand the findings to workers.
You define custom agents as standalone TOML files — ~/.codex/agents/ for personal agents, .codex/agents/ for project-scoped ones you commit with the repo. Each file needs name (how it's spawned), description (when to use it), and developer_instructions (the core behavior), with optional model, model_reasoning_effort, sandbox_mode, mcp_servers, and skills.config:
# .codex/agents/migration-checker.toml
name = "migration-checker"
description = "Audits a service for a specific framework migration and reports gaps."
model = "gpt-5.6-terra"
model_reasoning_effort = "high"
sandbox_mode = "read-only"
developer_instructions = """
You audit one repository for the v2 migration. Check imports, config keys, and
deprecated calls. Report only concrete, file-anchored findings — never speculate.
"""| Key | Default | What it controls |
|---|---|---|
enabled | true | Whether the primary agent can spawn subagents |
max_concurrent_threads_per_session | Codex chooses | Explicit per-session concurrency ceiling when you set one |
max_threads | Legacy alias | Backward-compatible name for the concurrency ceiling |
default_subagent_model | Current configured model | Model used when a spawned agent does not override one |
default_subagent_reasoning_effort | Model/config dependent | Reasoning effort for spawned agents without an override |
interrupt_message | true | Allows the parent to interrupt a child with a follow-up message |
Set a deliberate concurrency ceiling
Only add a numeric ceiling after considering the host and account limits. This example chooses six; six is an operator decision, not a documented Codex default:
[agents]
enabled = true
max_concurrent_threads_per_session = 6
default_subagent_model = "gpt-5.6-terra"
default_subagent_reasoning_effort = "high"
interrupt_message = trueThe current reader-facing reference does not document max_depth, job_max_runtime_seconds, spawn_agents_on_csv, or report_agent_job_result. Do not build a public integration around those names unless the Codex version you deploy documents them. For a tabular batch, read the rows in your own program, start one SDK thread per row, and enforce concurrency and timeouts in that program.
Senior scenario: an async PR-review bot over 50 repos
You run platform engineering for an org with 50 services and a shared @company/auth library that just shipped a breaking v2. You need a same-day report: for each repo, is it still on v1, and what exactly has to change? Doing this by hand is a day of grep. codex exec in a bash loop is serial and gives you 50 blobs of prose to read. The right tool is AsyncCodex + asyncio.gather — 50 read-only agents, each scoped to one repo, each returning a typed verdict.
import asyncio
import json
from pathlib import Path
from openai_codex import AsyncCodex, Sandbox
REPOS = [str(path) for path in Path("/srv/checkouts").iterdir() if path.is_dir()]
SCHEMA = {
"type": "object",
"properties": {
"on_v1": {"type": "boolean"},
"blocking_changes": {"type": "array", "items": {"type": "string"}},
"risk": {"type": "string", "enum": ["none", "low", "medium", "high"]},
},
"required": ["on_v1", "blocking_changes", "risk"],
"additionalProperties": False,
}
# Six is this program's chosen limit, not a Codex SDK default.
gate = asyncio.Semaphore(6)
async def review(codex: AsyncCodex, repo: str) -> dict:
async with gate:
thread = await codex.thread_start(
model="gpt-5.6-terra",
sandbox=Sandbox.read_only,
cwd=repo,
)
result = await asyncio.wait_for(
thread.run(
f"Audit {repo} for the @company/auth v2 migration. "
"Report whether it still uses v1 and the exact blocking changes.",
output_schema=SCHEMA,
),
timeout=1800,
)
return {"repo": repo, **json.loads(result.final_response)}
async def main() -> None:
async with AsyncCodex() as codex:
reports = await asyncio.gather(*(review(codex, r) for r in REPOS))
blocked = [r for r in reports if r["on_v1"] and r["risk"] in ("medium", "high")]
print(f"{len(blocked)}/{len(reports)} repos need urgent migration work")
asyncio.run(main())Three things make this safer than an unbounded batch. Sandbox.read_only prevents a reviewer from mutating the repo. asyncio.Semaphore(6) makes six an explicit application limit, so you do not open every connection at once; tune it to the account and host you actually operate. asyncio.wait_for(..., timeout=1800) stops one pathological repo from holding the batch indefinitely. Finally, output_schema turns every answer into a dictionary you can filter, sort, and gate a deploy on. Swap gather for as_completed when you want results to reach a dashboard as each repo finishes.
Designing your own fan-out
Pick the weakest sandbox that works
Read-only for review/audit/Q&A;
workspace_writeonly when agents must edit. Neverfull_accessin a fan-out — one bad prompt multiplies across every worker.Bound concurrency to your real limits
Choose an explicit
asyncio.Semaphoreor worker-pool size for SDK threads. Six is a reasonable example, not a promised Codex default. Raise it only after watching for rate-limit and host-pressure errors.Make every worker return a schema
Give each agent an
output_schema/outputSchemaso results are dicts, not prose. Compute the final verdict in code — the whole point of going programmatic is that the decision is deterministic.Cap runtime per worker
Wrap each task in
asyncio.wait_foror the equivalent in your language so one stuck agent cannot hold the whole batch hostage.Separate SDK concurrency from subagent concurrency
agents.max_concurrent_threads_per_sessionlimits subagents spawned inside a Codex session. It does not throttle independent SDK threads that your program starts; bound both layers if you use both.
Orchestrate the swarm
Watch delegation happen
The orchestrator hands a slice of work to each subagent. Every subagent runs in its own context window, does the noisy part — searching, reviewing, running tests — and returns only a short summary. Dispatch them and watch the work fan out, then the results pulse home.
The roster
Searches and maps the codebase without editing.
model · Haiku 4.5
Read-only pass for bugs, style, and risk.
model · Sonnet 5
Runs the suite and reports failures.
model · Haiku 4.5
Writes the focused change end to end.
model · Opus 5
Idle. Four subagents waiting for the orchestrator to dispatch work.
Knowledge check
You write a bot that reviews 50 repos with AsyncCodex and asyncio.gather. The Codex config sets agents.max_concurrent_threads_per_session = 6, and each independent SDK reviewer uses Sandbox.read_only. All 50 reviewers start at once. What is the likely problem, and the cleanest fix?
When to reach for the SDK over the CLI
Going programmatic is not about replacing codex exec — it is about earning types, parallelism, and control flow when a single string in and out is no longer enough. Drive a thread with startThread() / thread_start, get data back with outputSchema, persist and resumeThread across invocations, and when the work is wide, fan it out with subagents or AsyncCodex under an explicit concurrency bound. Pick the weakest sandbox that still works, set application-level timeouts, and make every worker answer with a schema.
Reach the end and this star joins your charted sky.