Codex compatibility

Claude Code is this catalog’s native harness. Every plugin also ships a Codex-side manifest, and the repo carries a second marketplace — but ported and identical are different claims, so each plugin declares which one applies.

Two marketplaces, one catalog

Claude Code Codex

marketplace

.claude-plugin/marketplace.json

.agents/plugins/marketplace.json

per-plugin manifest

plugins/<name>/.claude-plugin/plugin.json

plugins/<name>/.codex-plugin/plugin.json

instructions file

CLAUDE.md

AGENTS.override.md, then AGENTS.md, then an existing CLAUDE.md. In this repo AGENTS.md is a symlink to CLAUDE.md, so both clients read identical bytes.

The Codex manifests are generated, never hand-written:

make codex          # regenerate from the Claude-side source of truth
make validate       # fails if they are stale

This is the one place we deliberately diverge from the reference implementation. Superpowers — which pioneered the cross-harness skills repo and supports six — hand-maintains a manifest per harness and keeps versions aligned with a nine-entry stamping list. That works until someone adds a seventh manifest and forgets the list. Deriving the Codex manifests from the Claude ones makes the drift impossible instead of merely discouraged, and the check is a diff.

Three tiers of portability

Tiers are computed from the tree, not hand-listed, so a plugin that gains hooks tomorrow is re-tiered by the next make codex. They appear in each Codex manifest under compatibility.

Tier Count What it means

1 — portable

9

Prose skills. Same behaviour on any harness; nothing to translate.

2 — needs tool translation

6

Uses agent dispatch or an external CLI runner. The skill ships references/codex-tools.md mapping Agent → spawn_agent, the result wait → wait_agent, corrections → followup_task, and client-specific CLI invocation. Multi-agent is needed only for workflows that dispatch agents. Role prompts transfer; runtime tool restrictions and model settings require explicit mapping.

3 — lifecycle hooks

7

Hook behavior is enabled only for plugins shipping a reviewed hooks/codex.json; other plugins retain advisory skill instructions.

Tier 3 is dev-crew, evolving-claude-md, learn-on-failure, memory-hygiene, progress-channel, prompt-coach and roles. Each ships an explicit hooks/codex.json alongside its Claude hooks/hooks.json.

Feature comparison

Platform surfaces

What changes between the two clients for every plugin that touches the surface. Each plugin page carries its own Claude Code vs Codex section with the specifics.

Surface Claude Code Codex

Slash commands

/plugin:skill and commands/*.md.

Ask in plain language; the command’s procedure runs unchanged.

Hook registration

hooks/hooks.json via ${CLAUDE_PLUGIN_ROOT}; active once the plugin is enabled.

hooks/codex.json via ${PLUGIN_ROOT}. Review and trust with /hooks after install, and again whenever a definition changes. A skill-only copy registers no hooks.

Hook events

SessionStart, UserPromptSubmit, PreToolUse, PostToolUse, Stop, PostCompact.

The same events. PostCompact findings arrive as a systemMessage, because that event doesn’t accept injected context.

Edits the hooks see

Write / Edit calls.

apply_patch, split per file by the adapter. One rejected file rejects the whole patch.

Shell

Bash.

exec_command / shell_command, mapped to Bash. Edits made from the shell are not a matched tool path on either client.

Instructions file

CLAUDE.md.

AGENTS.override.md → AGENTS.md → existing CLAUDE.md.

Project memory

Claude’s native memory, loaded automatically.

Plugin-managed .codex/memory/, loaded by a SessionStart hook. Not Codex’s internal memory.

Plugin workflow state

.claude/<plugin>/ in the consuming repo.

The same directory — shared on purpose, so learning carries across clients.

Subagents

agents/*.md register as agent types with runtime tool lists and model tiers.

spawn_agent with the full role body in the brief. Tool limits become instructions, not sandbox rules; enable multi-agent where supported.

Models

Anthropic model names and tiers.

Mapped to the client’s available models; Anthropic names are never sent.

Live progress

Status-line bar.

Browser dashboard or progress watch in a companion terminal.

Session identity

CLAUDE_CODE_SESSION_ID.

CODEX_THREAD_ID / CODEX_SESSION_ID (or PROGRESS_SESSION_ID).

Transcripts

Claude JSONL.

Codex rollout files — not a stable API; unknown formats degrade gracefully.

Plugin by plugin

Plugin Tier What differs on Codex

evolving-claude-md

3

Same five hooks through the adapter; instructions file resolves to AGENTS.*; patches are linted per file.

memory-hygiene

3

Guards plugin-managed .codex/memory/ instead of Claude’s native memory.

dev-crew

3

Roles travel in the brief rather than registering as agent types; the phase gate needs the exact dc-* name.

brainstorm-panel

2

Seats spawn via the client’s agent tool; without one the panel runs sequentially and says so.

learn-on-failure

3

Writes .codex/memory/ topic files; a Codex-only SessionStart hook loads the index.

implement-issue

1

Attribution names the client actually used.

maven-quality

1

Nothing.

security-audit

1

Nothing.

review-agents

2

No-write becomes an instruction, not a runtime guarantee.

research-sweep

2

Results are collected from child messages, not transcripts.

roles

3

Same SessionStart hook; commands become plain-language requests.

screenshot-tour

1

Nothing beyond tool names.

progress-channel

3

No status-line bar — use the dashboard or progress watch; tracker identical.

ticket-triage

2

Lane dispatch uses the client’s agent tool.

prompt-coach

3

Native UserPromptSubmit hook; state shared with Claude; register once per client.

mindmap-prompt

2

--ai-client codex runs a sealed read-only Codex with every MCP server disabled.

skill-linter

1

Automated rules are Claude-convention first; Codex packaging is checked by hand.

tune-repo-beta

1

Tunes AGENTS.md; never turns Claude allowlists into sandbox bypasses.

systemic-fix-beta

1

Nothing.

screenshot-sweep

1

Nothing (needs a client that can view images).

spring-batch

1

Nothing.

conductor

2

Waits and relays use the client’s agent status and follow-up tools.

How Codex support works

Hooks and the shared adapter

  • Every tier-3 plugin vendors the same codex-bridge.py. It sets SKILL_CLIENT=codex, turns Codex tool envelopes into the shapes the existing Claude handlers expect, and reshapes their output for Codex. The rules themselves live in one place for both clients.

  • Codex behavior is strictly opt-in: handlers branch only on SKILL_CLIENT == "codex", and only the bridge sets it. Claude’s hook files and full workflows are untouched.

  • Hooks cover matching exposed tool paths, not arbitrary shell edits.

  • Crew dispatch messages must include the exact role name so the handoff gate can identify the role.

Agents, runners and live progress

  • review-agents ships a discoverable skill; Codex dispatches the complete shipped specialist definition in the child brief. Markdown agent files are not registered Codex agent types, and prose tool limits are not runtime restrictions.

  • Mindmap’s --ai-client auto|claude|codex selects an implemented runner, preserving the original Claude CLI path. The Codex runner uses a read-only local sandbox and a final-message file; external MCP servers, plugins and apps are disabled, and a failed MCP configuration check stops it before any model request.

  • Progress keeps its full tracker, history, notifications, browser and terminal watch. Only the status-line bar is Claude-specific.

Tests

  • make test-codex-adapters — shipped adapters, role packaging and the unchanged Claude paths, with no model calls. Runs in CI.

  • make test-codex-runtime — the real Codex CLI blocking a malformed instruction patch through the shipped hook.

  • make test-coach-codex — the real Codex CLI running the prompt hook against a local mock model, spending no tokens.

The last two need an installed Codex and skip cleanly without one.

Verified behavior and remaining limits

Surface Evidence Limit

Instruction hooks

Real Codex rejects a malformed patch; adapter regression covers a 751-file patch and nested overrides.

Matches exposed tool paths, not arbitrary shell edits.

Mindmap expansion

Real Codex returns parsed ideas with both user and project MCP servers disabled; sentinel servers never start.

Model/provider availability remains local configuration; inspection errors stop expansion.

Agent workflows

Complete shipped role definitions remain the dispatch source.

Claude agent metadata is not registered as native Codex agent configuration; prompt restrictions are not runtime enforcement.

Memory

Codex index loading and format checks have adapter tests.

Plugin-managed files, not Codex’s internal memory database.

Progress

Existing 123-check suite and Codex identity/nudge regressions pass.

Live Codex bars use a companion terminal or browser, not Claude’s command status line.

Installation

All 22 generated manifests validate; each declares its skill directory.

End-to-end marketplace installation and every client’s tool surface still need verification.

What still needs verifying on the Codex side

The packaging mirrors a working example, but these are assumptions until a real Codex install confirms them:

  • Marketplace schema — .agents/plugins/marketplace.json field names and the policy block are copied from a shipped plugin, not from a published spec.

  • source.url shape — ours points at ./plugins/<name> (a multi-plugin repo); the reference example is a single-plugin repo pointing at ./.

  • Skill discovery — the manifests declare "skills": "./skills/"; confirm Codex reads SKILL.md frontmatter the same way, especially the description field that drives triggering.

  • references/ loading — the tier-2 translation only helps if Codex follows a relative link out of SKILL.md on demand.

  • Multi-agent tool names — these have shipped in more than one version; the translation files say so and tell the reader to trust their live tool list over any table.

  • Hook execution beyond the tested paths — the real Codex CLI is verified for the instruction-patch lint and the prompt hook. The remaining events (memory, progress, roles, crew gate) are covered by adapter contract tests only. Plugins without hooks/codex.json keep hooks: {} to suppress automatic discovery.