Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Use the local Codex CLI as an independent second agent. Two branches — (1) proactively run `codex review` for a second opinion after completing a substantive change, before presenting it as done or committing; (2) delegate a well-defined implementation task via `codex exec`, ONLY when the user explicitly asks for Codex to do it. Also covers how to prompt Codex.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Codex is an independent agent on PATH (codex — invoke as command codex if
a shell alias shadows it), sharing this working tree and already authenticated. It is a second opinion, not ground truth: verify what it
reports, own what it changes. It reads the same skills your repo carries.
codex is not installedWhen codex is missing from PATH, offer to install it — ask the user for
approval first, never install on your own initiative. On yes, follow the
current instructions at https://developers.openai.com/codex/cli. First-run
authentication is interactive — hand that step to the user. Verify with
codex --version before proceeding.
Prompt Codex like an operator, not a collaborator: compact, block-structured
with XML tags. State the task, what "done" looks like, and the few constraints
that matter. A tighter prompt beats a bigger run — improve the contract before
raising --effort.
<task> — the concrete job, the repo/failure context, the expected end
state. Nearly always present.<output_contract> — exact shape, highest-value first, compact.<default_follow_through> — take the low-risk interpretation and keep
going; stop only when a missing detail changes correctness, safety, or an
irreversible action.<verification_loop> — before finalizing, check the result against the
requirements and the changed files; revise rather than ship the first
draft. Any risky fix.<grounding> — ground every claim in code or tool output; label inferences
as inferences. Review and research.<action_safety> — keep the diff tightly scoped; no drive-by refactors.
Write tasks.Run a Codex review whenever you have a substantive diff you'd want a second set of eyes on — a refactor, a tricky algorithm, renderer work, a security-sensitive change — before declaring it done or committing. Skip it for trivial edits (typos, comments, doc-only).
Fix the diff scope before launching: --uncommitted for working-tree
changes, --base <branch> for a branch diff, --commit <sha> for a landed
commit. Resolve a landed commit to its full SHA. Use the non-interactive
review entry point with an explicit read-only sandbox and separate progress
and verdict files in a fresh temporary directory:
review_dir=$(mktemp -d "${TMPDIR:-/tmp}/codex-review.XXXXXX")
codex exec --sandbox read-only -C "$(git rev-parse --show-toplevel)" review \
--commit "$(git rev-parse HEAD)" --json -o "$review_dir/verdict.md" \
< /dev/null > "$review_dir/events.jsonl" 2> "$review_dir/stderr.log"
Background long runs and retain their process/session handle. Scope flags and custom instructions are mutually exclusive: a scope flag takes no prompt. For custom framing, specify the exact diff to inspect in the prompt; don't rely on default scope or traverse all refs/checkpoint history. Never state the answer you expect (unprimed, same discipline as screenshot-critique).
Monitor progress in events.jsonl and failures in stderr.log. An active
reviewer reading relevant code is not hung merely because five minutes
elapsed. If it stalls or drifts outside the diff and its dependencies,
inspect the last command/error before interrupting; correct the cause or
narrow the task before retrying. Completion requires exit status zero,
turn.completed, and a nonempty final verdict. A timeout, turn.failed, or
missing verdict is incomplete, never a clean review.
Triage every finding: confirm it against the code before acting. Preserve Codex's evidence boundaries — an inference it labelled is not a fact. Overlap with your own doubts is high-priority evidence; a finding you dismiss needs a stated reason, not silence.
Report the outcome to the user — what Codex flagged, what you fixed, what you dismissed and why. Done when every finding is either fixed or explicitly dismissed.
Delegate implementation to Codex only when the user names Codex for the task. Never hand it work on your own initiative, and never re-delegate follow-up work without a fresh ask.
codex exec --sandbox workspace-write "<task>".
Network is off; add -c sandbox_workspace_write.network_access=true only
when the task must fetch (e.g. new deps).codex exec --dangerously-bypass-approvals-and-sandbox "<task>". The
sandbox blocks localhost binds (listen EPERM on vite/playwright), so a
sandboxed codex ships code it never saw run. Bypass trades that blindness
for zero OS control: only in a dedicated git worktree, only with a prompt
you authored end-to-end (never relaying third-party text), and the diff
review you owe afterwards is the control.
Use -o <file> to capture the final message and background long calls;
suppress codex's stderr thinking-noise with 2>/dev/null so it doesn't
bloat your context (drop it only to debug a failing run), and add
--skip-git-repo-check to run outside a git repo. Non-interactive runs
never ask for approval either way.codex exec resume <session-id> "<follow-up>", taking the
id from the run header. resume --last means the most recent session
globally — a review or any other codex run in between will hijack it. If
two resume rounds don't converge, stop delegating and finish it yourself —
iterating a confused agent costs more than taking over.A backgrounded codex exec can wedge at startup: process alive at ~0% CPU, but
no session file under ~/.codex/sessions/<Y/M/D>/, no network socket, no tree
changes. "Process running" is NOT "working."
/dev/null, never an inherited terminal. Left
on an interactive/piped stdin with no prompt source, exec prints Reading additional input from stdin... and blocks forever — the most common hang.
Either pass the prompt as a positional arg with < /dev/null, or feed a
prompt file with codex exec [flags] - < prompt.txt (the - makes codex
read the task from stdin, and a missing file fails the redirect loudly —
unlike "$(cat prompt.txt)", which silently sends the fallback string as
the task). Add nohup/& as needed.grep -rl "<marker>" ~/.codex/sessions/<Y/M/D>/ finds no session within
~3 minutes. Relaunching after a kill reliably works.[projects."<worktree-path>"]\ntrust_level = "trusted" to
~/.codex/config.toml before exec'ing in one. This is the safe fix; the
bypass flags stay forbidden here.web/), every edit outside
it is rejected as "writing outside of the project" and a never approval
policy can't recover — the run burns with zero files changed. Pass
-C <root> to set the working root explicitly instead of cd-ing into it.--sandbox read-only for reviews, consultation and questions;
workspace-write only for delegated implementation.--dangerously-bypass-approvals-and-sandbox is reserved for tasks that must
run browsers/servers/full suites (above) — dedicated worktree, self-authored
prompt, mandatory diff review after. --full-auto has been removed (it
aliased workspace-write) — don't reach for it.-c model_reasoning_effort=high, -m <model> (the old --effort flag is gone
in current Codex). Leave both at their defaults unless the user asks —
tighten the prompt first.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Audit or rewrite AGENTS.md so it holds only lasting principles. Run it only when the user asks for it by name; never invoke it on your own.
日本語の概要は準備中です。原文の説明を表示しています。
Audit the choices an implementing agent made, not its diff — a pure decision audit that traces the session's history into a choices ledger, changes no code, and never blocks an unsupervised run. Working code still embeds architecture the user never chose; surface it because future work inherits it. Use when the user wants to review the decisions the AI made on their behalf, before merging or committing AI-implemented work, when integrating a delegated subagent's pass, or when a fix "works" but might be a point fix.
日本語の概要は準備中です。原文の説明を表示しています。
Audit and prioritize performance work through bounded-work and forward-progress checks. Use when asked to find performance issues, investigate freezes/500s/OOMs, review retries or polling for no-progress loops, rank a performance backlog, or implement the next simple performance fixes.
日本語の概要は準備中です。原文の説明を表示しています。
Audit whether tests earn their maintenance cost and which suite owns each contract. Use when pruning redundant tests, investigating implementation coupling or test-only production hooks, reviewing the value of proposed coverage, or auditing an entire subsystem. Use write-tests to implement the resulting test changes.
日本語の概要は準備中です。原文の説明を表示しています。
Run fast, progressive experiments to map tunable parameters and their effects. Use when optimizing code, prompts, configurations, or other artifacts against a goal or benchmark, or when a long research run needs a planned hypothesis queue, focused trials, and durable findings.
日本語の概要は準備中です。原文の説明を表示しています。
Use Claude Code as an independent `claude -p` subagent when the user explicitly asks for Claude, wants a second-agent opinion from Claude, or asks to delegate a well-scoped task to Claude. Supports selecting `--model` and thinking/effort level with defaults of `opus` and `high`.
日本語の概要は準備中です。原文の説明を表示しています。