本文へ移動
cccskills
無料GitHub で公開

ato-review

Before committing any non-trivial change, dispatch the diff to a reviewer runtime via ATO (`ato dispatch <reviewer> --session <id>`), parse the numbered/severity-tagged findings, apply or defer each one with a recorded justification, then commit. Fights the "build passes therefore ship it" failure mode — what Garry Tan calls the AI agent complexity ratchet. Place in the v2.16 stack: this skill is the LAST gate. `ato-warroom` decides the design; `ato-mission` runs the multi-step work and produces the diff; `ato-review` checks the diff before commit. When the review is part of a Mission, dispatch the review with `--require-tools read_file,grep,git_diff,git_log` so the reviewer can walk the source itself instead of reasoning from a paraphrase (PR-1.5 tool surface). Receipts land in `execution_logs` and the Mission narrative. Fires automatically before commits touching public surface (CLI subcommands, Tauri commands, MCP tools, schema migrations, security boundaries) or whenever a diff exceeds ~50 LOC of behavior change.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md11.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

When this skill fires

Before git commit, check whether any of these apply to the staged diff:

  • Adds or changes a public surface: CLI subcommand or flag, Tauri command, MCP tool, REST endpoint, exported function signature, schema migration (ALTER TABLE, CREATE TABLE), tauri.conf.json, package.json bin entries.
  • Touches a security boundary: shell-out / Command::new / IPC allowlist / authentication code / file-system writes outside the repo.
  • Is >50 lines of behavior change (not counting test fixtures, snapshots, or pure formatting).
  • Adds a new module / file that's larger than ~30 LOC.
  • Changes an existing schema or persistence shape.

If none apply (small typo fix, comment-only change, doc-only edit, dependency bump with no code change), skip this skill and commit normally. Note the reason in your turn message so the next reader sees you decided rather than forgot.

If any apply, run the procedure below.

Procedure

1. Capture the diff

# Capture both staged and unstaged so the review sees what the commit will look like.
git diff HEAD > /tmp/ato-review-$$.patch

The "When this skill fires" section above is authoritative. The wc -l heuristic is only a final tie-breaker for diffs that match none of the public-surface / security-boundary / new-file / schema triggers:

if [ $(wc -l < /tmp/ato-review-$$.patch) -lt 10 ] \
   && ! grep -qE '#\[tauri::command\]|server\.tool\(|CREATE TABLE|ALTER TABLE|Command::new|spawn\(' /tmp/ato-review-$$.patch; then
    rm /tmp/ato-review-$$.patch
    exit 0  # genuinely trivial
fi

A 9-line diff that adds a Tauri command or runs Command::new is not trivial. Trust the triggers over the line count.

2. Open or reuse a review session

The session keeps reviewer context across multiple turns of one feature, so the reviewer doesn't re-derive the codebase each commit.

# Try to find an existing review session for the active branch.
# Pass $BRANCH as a python argv to avoid nesting shell quotes inside a
# Python f-string (review-of-the-skill caught this — the earlier version
# had mismatched quotes that silently SyntaxError'd, so every commit
# opened a fresh session instead of reusing one).
BRANCH=$(git branch --show-current)
SID=$(ato sessions list --limit 20 2>/dev/null | python3 -c '
import sys, json
branch = sys.argv[1]
sessions = json.load(sys.stdin)
for s in sessions:
    if s.get("title", "").startswith("review/" + branch):
        print(s["id"]); break
' "$BRANCH" 2>/dev/null)

# If none, open one. Default reviewer is minimax; allow user to override
# via $ATO_REVIEWER env var.
REVIEWER="${ATO_REVIEWER:-minimax}"
if [ -z "$SID" ]; then
    SID=$(ato sessions new --runtime "$REVIEWER" --title "review/$BRANCH" 2>/dev/null \
          | python3 -c "import sys,json; print(json.load(sys.stdin)['id'])")
fi
echo "Review session: $SID  reviewer: $REVIEWER"

ATO's sessions new and sessions list commands emit pure JSON to stdout (diagnostics on stderr) — that's a stable contract, so direct piping into json.load is safe.

If ato isn't on PATH or the user has no $ATO_REVIEWER configured and no MiniMax / Grok / DeepSeek / Qwen / OpenRouter key, tell the user once and skip — don't block the commit on infrastructure they don't have.

3. Dispatch the diff for review

The prompt is the most load-bearing part. The reviewer needs to know:

  • The change is a real diff being committed today
  • What categories of issues to look for (calibrated to the diff)
  • The expected output format (numbered, severity-tagged, with concrete fixes)
DIFF=$(cat /tmp/ato-review-$$.patch)
PROMPT="You are a senior reviewer for a multi-runtime AI agent ops platform written in
Rust + TypeScript. This diff is about to be committed to main. Critique it.

Look specifically for:
1. **Security**: shell injection, IPC trust boundaries, SQL injection, path traversal,
   unbounded user input passed to Command::new / spawn / fs writes.
2. **Race conditions / cache invalidation**: any concurrent writers, stale reads, UI
   that re-renders before the backend confirms.
3. **Validation gaps**: input ranges, regex shapes, empty / null handling at edges.
4. **Contract drift**: a flag rename, a Tauri command shape change, a CLI argv shift —
   anything that would break a wrapper or earlier caller silently.
5. **Bugs you can spot from the diff alone**, especially around error handling and
   resource cleanup.

Reply with a numbered list. For each finding:
  N. **SEVERITY — short title** (HIGH / MEDIUM / LOW / INFO)
     Brief description.
     **Fix:** one concrete diff or sentence.

Be brief — 3–8 findings, not a wall of text. Skip the obvious. If a category has
nothing wrong, don't enumerate it. If a candidate finding is wrong on closer look,
say so explicitly. Do NOT invent findings to pad the list.

DIFF:
\`\`\`
$DIFF
\`\`\`"

# Bound the dispatch so a hung reviewer doesn't block the commit forever.
# Coreutils `timeout` is available on Linux; on macOS install via
# `brew install coreutils` (gtimeout) or fall back to running without.
TIMEOUT=${ATO_REVIEW_TIMEOUT:-180}
TIMEOUT_CMD=""
if command -v timeout >/dev/null 2>&1; then
    TIMEOUT_CMD="timeout $TIMEOUT"
elif command -v gtimeout >/dev/null 2>&1; then
    TIMEOUT_CMD="gtimeout $TIMEOUT"
fi
$TIMEOUT_CMD ato dispatch "$REVIEWER" "$PROMPT" --session "$SID" --human \
    | tee /tmp/ato-review-findings-$$.txt
DISPATCH_RC=${PIPESTATUS[0]}
if [ "$DISPATCH_RC" = "124" ]; then
    echo "Review timed out after ${TIMEOUT}s. Proceeding without review — note in commit."
fi

If the dispatch fails (network, quota, key missing), surface the error to the user but do NOT auto-retry — they may want to skip the review for this commit.

After the dispatch returns, verify the response actually contains review findings before triaging:

# Findings should be numbered + severity-tagged. If grep finds none, the
# reviewer probably returned prose or freeform text — surface that as a
# warning so we don't silently advance to commit thinking the review
# was clean.
if ! grep -qE '^\s*[0-9]+\.\s+\*\*[A-Z]+' /tmp/ato-review-findings-$$.txt; then
    echo "WARN: review output did not match the expected numbered+severity format."
    echo "      Eyeball /tmp/ato-review-findings-$$.txt before committing."
fi

4. Triage findings

For each numbered finding in the response:

  • HIGH → must apply before committing. Edit the code, rebuild, re-stage.
  • MEDIUM → apply unless there's a specific reason not to. Document the reason in the commit message under a Deferred from review: line.
  • LOW / INFO → judgment call. Apply if cheap (<5 LOC, no design implication). Defer otherwise.

Verify findings against the actual diff before applying. The reviewer can hallucinate — refer to the v2.3.38 dogfood pass which caught a real inline-code-fence bug AND surfaced multiple non-bugs that didn't apply to the actual code. Grep + read the cited file/line before committing the fix.

Anti-pattern: applying every finding mechanically. The signal-to-noise on these reviews is real but imperfect; you're the human-in-the-loop even when no human is in the loop.

5. Re-build + re-test

After applying findings:

# Match the project's QA §0:
cargo build --manifest-path apps/cli/Cargo.toml -p ato
cargo build --manifest-path apps/desktop/src-tauri/Cargo.toml
cargo test --manifest-path apps/cli/Cargo.toml -p ato
cd apps/desktop && npx vite build && cd -

If any step fails, fix the failure before committing. The pre-commit hook will catch this anyway, but catching it now means one fewer round-trip.

6. Commit with the review note

Include a ### Dogfood + review process section in the commit body listing:

  • Which reviewer ran (minimax / grok / etc.)
  • What the headline findings were
  • Which ones were applied vs deferred with justification

Example:

### Dogfood + review process

MiniMax-reviewed the diff before commit. Findings:
- MEDIUM: IPC validation on days/threshold → applied (bounds check in lock_ratchet)
- LOW: agent slug regex → applied (frontend regex guard)
- LOW: ato-binary fallback could be clearer → deferred (intentional graceful
  fallback so a post-startup ato install still works)
- INFO: targetKey computed in 3 sites → applied (extracted helper)

This commits the audit trail of the review, not just the code. A future reader can see why a finding wasn't applied without spelunking through chat history.

7. Cleanup

rm -f /tmp/ato-review-$$.patch /tmp/ato-review-findings-$$.txt

The session itself stays open — next commit on this branch reuses it via the review/<branch> title lookup in step 2.

Override / skip

There are legitimate reasons to skip review on a specific commit:

  • WIP commit you'll squash later — note [wip] in the subject.
  • Reviewer unreachable (no network, no key configured) — proceed with a Review skipped: <reason> line in the commit body.
  • Trivial / mechanical change the trigger heuristics caught as a false positive — note "trivial" in the commit body.

Don't skip silently. The point of the skill is to make "did I review?" a yes-or-no question with a recorded answer.

Why this skill exists

Tan's "AI Agent Complexity Ratchet" (May 2026) argues that AI agents make 90% test coverage free — agents don't experience effort writing the fourteenth edge-case test. The same principle applies to code review: dispatching a diff to a second runtime costs cents and seconds, but humans (and Claude in flow) routinely skip it because "build passes." Build passing isn't a review. Tests passing isn't a review. A second runtime reading the actual diff with a fresh prior is a review.

ATO's Phase 6 cluster (sessions, bridge, ratchet) was built to make this loop cheap and ergonomic. This skill makes "use it" the default rather than something to remember.

Pairs well with

  • ato ratchet check as a pre-deploy CI gate — quality floors complement per-commit review.
  • ato dispatch <runtime> --tag-bridge when a single reviewer's pass leaves you uncertain; bridge into a second runtime for a second opinion.
  • Activity feed (ato posts list --kind approval_request) to surface review spinning to a human when the bridge can't converge.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Turn any installed skill (gstack, custom, third-party) into an ATO agent the user can summon into war-rooms. Reads a SKILL.md, extracts the persona, strips runtime boilerplate, and writes an agent file at `.claude/agents/<slug>.md` (project-scoped) or `~/.claude/agents/<slug>.md` (global). Prompts for a model roster (primary + 1-2 alts) so cross-family dispatch in war-rooms produces real disagreement. Companion to `ato-warroom` — that skill summons agents this skill creates. Use when asked "turn this skill into an agent", "register X as a war-room agent", or when scoping a new persona before a war-room runs.

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

Before any multi-step work with a stated goal — a feature, a bugfix spanning multiple files, a QA sweep, a doc draft + iterations, a multi-day investigation — create an ATO Mission instead of doing it via bare `ato dispatch` calls. A Mission persists the goal + the verifiable success criteria, lets the coordinator tick drive the work across days, captures every event in a structured audit trail (SQLite + markdown narrative), and integrates parallel agents' work via merge strategies. Complement to `ato-warroom` (the cross-family decision before you start) and `ato-review` (the post-code-diff review). Missions is where multi-step work LIVES; war-rooms are where decisions ABOUT it get made; reviews are where the resulting commits get vetted. Fires when: the work has more than one decision point, a verifiable end state, or runs across more than one session. Use it for any ATO development that doesn't fit in a single dispatch.

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

Before any material decision — code chunk, plan, strategy, design, scope cut, push to GitHub — convene a war-room. The session driver takes the CEO seat: frame the tradeoff, summon specialist seats from whatever agent roster the user has built, dispatch a cross-family voice via `ato dispatch` so priors actually disagree, decide. A failure-mode filter (wrong assumptions / overcomplexity / orthogonal edits / imperative-over-declarative — Karpathy's four are one good default, swap in your own) runs on every dispatch. Place in the v2.16 stack: war-rooms DECIDE before code starts; `ato-mission` EXECUTES the work between decisions (multi-step, goal-driven, persisted across days); `ato-review` VERIFIES the resulting commits. Use a war-room for the design verdict, hand the verdict to a Mission, review the merged result. Fires before: sending a code draft to the user as final, opening a PR, pushing to a remote-tracking branch, committing >50 LOC of behavior change, or delivering a plan or strategic recommendation as the final answer.

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

browse

無料

Fast headless browser for QA testing and site dogfooding. Navigate any URL, interact with elements, verify page state, diff before/after actions, take annotated screenshots, check responsive layouts, test forms and uploads, handle dialogs, and assert element states. ~100ms per command. Use when you need to test a feature, verify a deployment, dogfood a user flow, or file a bug with evidence. Use when asked to "open in browser", "test the site", "take a screenshot", or "dogfood this".

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

debug

無料

Systematic debugging with root cause investigation. Four phases: investigate, analyze, hypothesize, implement. Iron Law: no fixes without root cause.

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

Design consultation: understands your product, researches the landscape, proposes a complete design system (aesthetic, typography, color, layout, spacing, motion), and generates font+color preview pages. Creates DESIGN.md as your project's design source of truth. For existing sites, use /plan-design-review to infer the system instead. Use when asked to "design system", "brand guidelines", or "create DESIGN.md".

日本語の概要は準備中です。原文の説明を表示しています。

WillNigri/Agentic-Tool-Optimization342026年9月8日 更新

WillNigri のスキルをすべて見る

このスキルの問題を報告する