The Forge · 11 min mission

Programmatic Codex: SDK & Subagents

Drive Codex from TypeScript or Python and fan tasks out to parallel sub-agents.

sdksubagentscodexFact-checked 2026-06-13
On this page

codex exec is great until you need a program around the agent. The moment you want to loop over 50 repositories, parse the agent's answer as typed data, run ten reviews at once, or wire Codex into a webhook handler, shelling out to a CLI and scraping its stdout stops being clever and starts being a liability.

The Codex SDK is the same agent — same model, same sandbox, same config — exposed as a library you import. You get a thread object, you call run(), you get a result back in your own process. No subprocess parsing, no fragile string munging, and you can fan the work out across subagents that run in parallel. This guide is about that: when the SDK beats the CLI, the exact TypeScript and Python surface, how to get structured output back, and how to orchestrate a swarm of agents without melting your machine.

When the SDK beats codex exec

codex exec "..." is the right tool for one-shot automation: a CI step, a git hook, a shell pipeline. It streams progress to stderr and prints only the final agent message to stdout, so codex exec "summarize the diff" | pbcopy just works. Reach for the SDK the moment any of these is true:

  • You need the result as typed data, not prose to re-parse — a JSON object with fields you can branch on.
  • You need to run many agents concurrently and collect their results (the async/fan-out case).
  • The agent is one step in a larger program — a server, a queue worker, a bot — where managing a child process per request is the wrong abstraction.
  • You want to resume a long-lived thread across separate invocations and keep its context.

The dividing line is simple: if a single string in and a single string out is enough, use codex exec. If you need control flow, types, or parallelism around the agent, use the SDK.

codex exec vs. the SDK

codex exec (CLI)

Shape: one prompt in, final message on stdout.

bash
codex exec "fix the failing test" \
  --sandbox workspace-write \
  --output-schema ./schema.json \
  -o result.json

Perfect for CI steps, git hooks, and shell pipelines. You get JSON Lines with --json and a final message you can pipe. But control flow lives in bash, and parallelism means juggling child processes.

@openai/codex-sdk / openai-codex

Shape: a thread object in your process; run() returns a typed result.

ts
const codex = new Codex();
const thread = codex.startThread();
const turn = await thread.run("fix the failing test");
console.log(turn.finalResponse);

Control flow lives in your language. Loop, await, Promise.all, branch on a parsed object. This is the only sane path once you have dozens of tasks or need the answer as data.

TypeScript: @openai/codex-sdk

Install @openai/codex-sdk and you get three calls that cover almost everything: construct a client, start a thread, run a turn.

ts
import { Codex } from "@openai/codex-sdk";
 
const codex = new Codex();                  // uses your existing Codex auth/config
const thread = codex.startThread();         // a fresh conversation
const turn = await thread.run("Make a plan to diagnose and fix the CI failures");
 
console.log(turn.finalResponse);            // the agent's final message
console.log(turn.items);                    // every item it produced this turn

A thread is a conversation with memory; a turn is one run() and the items it produced. run() resolves to a turn object whose finalResponse is the agent's last message and whose items array holds everything that happened — reasoning, command executions, file changes. Authentication and configuration are inherited from your normal Codex setup, so a script that runs locally needs no extra wiring; in CI you pass credentials through the environment the SDK reads.

startThread() takes options to pin the thread to a project and relax the git guardrail:

ts
const thread = codex.startThread({
  workingDirectory: "/path/to/project",
  skipGitRepoCheck: true,
});

Both workingDirectory and skipGitRepoCheck are SDK options — the second is the SDK equivalent of the CLI's --skip-git-repo-check, which you need when the agent runs somewhere that isn't a git repo.

Current TypeScript releases also accept a typed config object and raw TOML-compatible configOverrides on the Codex constructor. Raw overrides win when both set the same key. Use the typed object for ordinary settings; reserve raw overrides for a key the SDK types do not expose yet:

ts
const codex = new Codex({
  config: { model: "gpt-5.6-terra" },
  configOverrides: ["model_reasoning_effort=\"max\""],
});

Resuming threads and streaming progress

Threads are persisted to ~/.codex/sessions. If your process restarts — or a webhook fires a follow-up an hour later — you do not lose the conversation. Reconstruct it from its id and keep going:

ts
// First invocation
const thread = codex.startThread();
await thread.run("Start refactoring the auth module");
const savedThreadId = /* persist this id somewhere durable */;
 
// Later, in a fresh process
const thread2 = codex.resumeThread(savedThreadId);
await thread2.run("Now add tests for what you changed");

resumeThread(threadId) reconnects to an existing thread by id and returns a thread you can run() again, with all prior context intact.

For long turns where you want to react to intermediate progress — show a tool call, stream tokens to a UI, surface file diffs as they happen — use runStreamed() instead of run(). It hands back an async iterable of events:

ts
const { events } = await thread.runStreamed("Audit the codebase for N+1 queries");
for await (const event of events) {
  switch (event.type) {
    case "item.completed":
      // a reasoning step, command, or file edit finished
      break;
    case "turn.completed":
      // the whole turn is done
      break;
  }
}

The same event vocabulary (thread.started, item.completed, turn.completed, error) is what the CLI emits with --json — runStreamed is that stream, in your language, without parsing JSON Lines by hand.

Python: openai-codex, sync and async

The Python package is openai-codex, and its import namespace is openai_codex. It ships two clients used as context managers: Codex (synchronous) and AsyncCodex (asyncio). You start a thread with thread_start(...), which takes model and sandbox directly:

python
from openai_codex import Codex, Sandbox
 
with Codex() as codex:
    thread = codex.thread_start(model="gpt-5.6-terra", sandbox=Sandbox.workspace_write)
    result = thread.run("Make a plan to diagnose and fix the CI failures")
    print(result.final_response)

Note the snake_case: the Python result exposes final_response where TypeScript exposes finalResponse. The sandbox argument takes one of the presets below — Sandbox.workspace_write lets the agent edit files inside the workspace, which is what you want for "fix this" tasks.

The async client is the one that matters for scale. AsyncCodex gives you awaitable thread_start and run, which means you can launch dozens of independent agents and gather their results with asyncio.gather — no subprocess pool, no thread pool, just coroutines. That is the senior scenario later in this guide.

SDK preset (Python)CLI flagWhat the agent can touchUse it for
Sandbox.read_only--sandbox read-onlyRead files only — no writes, no networkReview, audit, Q&A over a repo
Sandbox.workspace_write--sandbox workspace-writeRead and write inside the workspaceFixes, refactors, codegen — the common case
Sandbox.full_access--sandbox danger-full-accessUnrestricted filesystem and networkOnly in throwaway/controlled environments
Sandbox presets — the same three across the CLI and SDK, named workspace-write on the command line and Sandbox.workspace_write in Python. If you omit sandbox, the SDK uses the app-server default from your effective Codex configuration.

Structured output: get JSON back, not prose

The single biggest reason to drive Codex programmatically is to stop parsing English. Both the CLI and the SDK can enforce a JSON Schema on the final answer so run() hands you data your code can branch on.

On the CLI, --output-schema <file> points at a JSON Schema and the final message is guaranteed to match it:

json
{
  "type": "object",
  "properties": {
    "project_name": { "type": "string" },
    "languages":    { "type": "array", "items": { "type": "string" } }
  },
  "required": ["project_name", "languages"]
}
bash
codex exec "Extract this repo's metadata" --output-schema ./schema.json -o output.json

In the TypeScript SDK the same idea is a per-turn option — pass outputSchema to run() and the agent's answer conforms to it:

ts
const schema = {
  type: "object",
  properties: { severity: { type: "string" }, files: { type: "array", items: { type: "string" } } },
  required: ["severity", "files"],
  additionalProperties: false,
};
 
const turn = await thread.run("Triage this PR's risk", { outputSchema: schema });
const verdict = JSON.parse(turn.finalResponse); // typed, branchable

You do not have to hand-write the schema: generate it from a Zod schema with zod-to-json-schema (target "openAi") and keep one source of truth for both validation and the agent contract.

Two knobs shape how hard the agent thinks before it answers. Reasoning effort is set on the CLI with -c model_reasoning_effort=<level> and via config in the SDK. Current releases recognize low, medium, high, xhigh, max, and ultra, subject to model support; Ultra can use subagents and is not a cheap way to make a routine classification slightly better. The practical pattern is a two-phase pipeline: a low-effort pass to classify or filter (does this PR even need review?), then a higher-effort pass with a strict outputSchema only on the items that survived.

Subagents: one orchestrator, many workers

A single thread is one worker. Subagents let a primary agent spawn child agents that run in their own context and report back — so a big task splits into parallel pieces instead of one long serial slog. Codex ships three built-in agents:

  • default — the general-purpose fallback agent.
  • worker — an execution-focused agent for implementation and fixes.
  • explorer — a read-heavy agent tuned for codebase exploration.

The mental model: an explorer maps the territory (where does auth live? which files import this?), a worker changes it (apply the fix, write the test), and the orchestrator stitches their results together. You spawn explorers to investigate in parallel without polluting the main thread's context, then hand the findings to workers.

You define custom agents as standalone TOML files — ~/.codex/agents/ for personal agents, .codex/agents/ for project-scoped ones you commit with the repo. Each file needs name (how it's spawned), description (when to use it), and developer_instructions (the core behavior), with optional model, model_reasoning_effort, sandbox_mode, mcp_servers, and skills.config:

toml
# .codex/agents/migration-checker.toml
name = "migration-checker"
description = "Audits a service for a specific framework migration and reports gaps."
model = "gpt-5.6-terra"
model_reasoning_effort = "high"
sandbox_mode = "read-only"
developer_instructions = """
You audit one repository for the v2 migration. Check imports, config keys, and
deprecated calls. Report only concrete, file-anchored findings — never speculate.
"""
KeyDefaultWhat it controls
enabledtrueWhether the primary agent can spawn subagents
max_concurrent_threads_per_sessionCodex choosesExplicit per-session concurrency ceiling when you set one
max_threadsLegacy aliasBackward-compatible name for the concurrency ceiling
default_subagent_modelCurrent configured modelModel used when a spawned agent does not override one
default_subagent_reasoning_effortModel/config dependentReasoning effort for spawned agents without an override
interrupt_messagetrueAllows the parent to interrupt a child with a follow-up message
The documented [agents] controls. Values shown as “Codex chooses” have no stable numeric default in the current reference.

Set a deliberate concurrency ceiling

Only add a numeric ceiling after considering the host and account limits. This example chooses six; six is an operator decision, not a documented Codex default:

toml
[agents]
enabled = true
max_concurrent_threads_per_session = 6
default_subagent_model = "gpt-5.6-terra"
default_subagent_reasoning_effort = "high"
interrupt_message = true

The current reader-facing reference does not document max_depth, job_max_runtime_seconds, spawn_agents_on_csv, or report_agent_job_result. Do not build a public integration around those names unless the Codex version you deploy documents them. For a tabular batch, read the rows in your own program, start one SDK thread per row, and enforce concurrency and timeouts in that program.

Senior scenario: an async PR-review bot over 50 repos

You run platform engineering for an org with 50 services and a shared @company/auth library that just shipped a breaking v2. You need a same-day report: for each repo, is it still on v1, and what exactly has to change? Doing this by hand is a day of grep. codex exec in a bash loop is serial and gives you 50 blobs of prose to read. The right tool is AsyncCodex + asyncio.gather — 50 read-only agents, each scoped to one repo, each returning a typed verdict.

pr_review_bot.py — fan 50 read-only reviewers out with asyncio
python
import asyncio
import json
from pathlib import Path
from openai_codex import AsyncCodex, Sandbox
 
REPOS = [str(path) for path in Path("/srv/checkouts").iterdir() if path.is_dir()]
 
SCHEMA = {
    "type": "object",
    "properties": {
        "on_v1": {"type": "boolean"},
        "blocking_changes": {"type": "array", "items": {"type": "string"}},
        "risk": {"type": "string", "enum": ["none", "low", "medium", "high"]},
    },
    "required": ["on_v1", "blocking_changes", "risk"],
    "additionalProperties": False,
}
 
# Six is this program's chosen limit, not a Codex SDK default.
gate = asyncio.Semaphore(6)
 
async def review(codex: AsyncCodex, repo: str) -> dict:
    async with gate:
        thread = await codex.thread_start(
            model="gpt-5.6-terra",
            sandbox=Sandbox.read_only,
            cwd=repo,
        )
        result = await asyncio.wait_for(
            thread.run(
                f"Audit {repo} for the @company/auth v2 migration. "
                "Report whether it still uses v1 and the exact blocking changes.",
                output_schema=SCHEMA,
            ),
            timeout=1800,
        )
        return {"repo": repo, **json.loads(result.final_response)}
 
async def main() -> None:
    async with AsyncCodex() as codex:
        reports = await asyncio.gather(*(review(codex, r) for r in REPOS))
    blocked = [r for r in reports if r["on_v1"] and r["risk"] in ("medium", "high")]
    print(f"{len(blocked)}/{len(reports)} repos need urgent migration work")
 
asyncio.run(main())

Three things make this safer than an unbounded batch. Sandbox.read_only prevents a reviewer from mutating the repo. asyncio.Semaphore(6) makes six an explicit application limit, so you do not open every connection at once; tune it to the account and host you actually operate. asyncio.wait_for(..., timeout=1800) stops one pathological repo from holding the batch indefinitely. Finally, output_schema turns every answer into a dictionary you can filter, sort, and gate a deploy on. Swap gather for as_completed when you want results to reach a dashboard as each repo finishes.

running the fan-out
… scroll to run this session
Fifty read-only reviewers, six at a time, each returning schema-validated JSON. The bot prints a computed verdict, not 50 paragraphs to read.

Designing your own fan-out

  1. Pick the weakest sandbox that works

    Read-only for review/audit/Q&A; workspace_write only when agents must edit. Never full_access in a fan-out — one bad prompt multiplies across every worker.

  2. Bound concurrency to your real limits

    Choose an explicit asyncio.Semaphore or worker-pool size for SDK threads. Six is a reasonable example, not a promised Codex default. Raise it only after watching for rate-limit and host-pressure errors.

  3. Make every worker return a schema

    Give each agent an output_schema / outputSchema so results are dicts, not prose. Compute the final verdict in code — the whole point of going programmatic is that the decision is deterministic.

  4. Cap runtime per worker

    Wrap each task in asyncio.wait_for or the equivalent in your language so one stuck agent cannot hold the whole batch hostage.

  5. Separate SDK concurrency from subagent concurrency

    agents.max_concurrent_threads_per_session limits subagents spawned inside a Codex session. It does not throttle independent SDK threads that your program starts; bound both layers if you use both.

Orchestrate the swarm

Watch delegation happen

The orchestrator hands a slice of work to each subagent. Every subagent runs in its own context window, does the noisy part — searching, reviewing, running tests — and returns only a short summary. Dispatch them and watch the work fan out, then the results pulse home.

orchestrator · main thread
ready
exploreridlerevieweridletesteridleimplementeridleorchestrator

The roster

exploreridle

Searches and maps the codebase without editing.

model · Haiku 4.5

revieweridle

Read-only pass for bugs, style, and risk.

model · Sonnet 5

testeridle

Runs the suite and reports failures.

model · Haiku 4.5

implementeridle

Writes the focused change end to end.

model · Opus 5

Idle. Four subagents waiting for the orchestrator to dispatch work.

Choose an explicit concurrency ceiling, assign explorer/worker/default roles, and see which tasks run in parallel. Treat the numeric value as your configuration, not a product default.

Knowledge check

You write a bot that reviews 50 repos with AsyncCodex and asyncio.gather. The Codex config sets agents.max_concurrent_threads_per_session = 6, and each independent SDK reviewer uses Sandbox.read_only. All 50 reviewers start at once. What is the likely problem, and the cleanest fix?

When to reach for the SDK over the CLI

Going programmatic is not about replacing codex exec — it is about earning types, parallelism, and control flow when a single string in and out is no longer enough. Drive a thread with startThread() / thread_start, get data back with outputSchema, persist and resumeThread across invocations, and when the work is wide, fan it out with subagents or AsyncCodex under an explicit concurrency bound. Pick the weakest sandbox that still works, set application-level timeouts, and make every worker answer with a schema.

Reach the end and this star joins your charted sky.