本文へ移動
cccskills
無料GitHub で公開

bmad-run-inspector

Inspect and explain bmad-loop run artifacts under `.bmad-loop/runs/`. Use for live health checks ("is it stuck?", "is the loop alive?", "what is the agent doing?") and post-run forensics ("what happened?", "why did story X fail?", "where did the tokens go?"). Also triggers on "monitor/check/summarize the loop", "why did the run pause?", Vietnamese "theo dõi bmad loop", "phân tích run", "đọc log của run", and equivalent requests in other languages. Use only for bmad-loop's own run artifacts, not application runtime logs or CI-provider build logs.

インストール方法を見る

含まれるファイル(15)

  • SKILL.md14.9 KB
  • assets/templates/environment.md1.7 KB
  • assets/templates/environment.toml2.1 KB
  • assets/templates/project-readme.md1.8 KB
  • README.md6.4 KB
  • README.vi.md8.7 KB
  • README.zh.md7.0 KB
  • references/adapter-bootstrap.md10.7 KB
  • references/anomaly-triage.md18.2 KB
  • references/log-forensics.md11.3 KB
  • references/run-anatomy.md20.6 KB
  • scripts/bootstrap_adapter.py4.8 KB
  • scripts/extract_transcript.py10.8 KB
  • scripts/run_probe.py23.7 KB
  • tests/test_scripts.py7.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

BMAD Run Inspector

Verified against bmad-loop 0.11.1. The bundled scripts run on Python 3.9+. bmad-loop itself needs 3.11+, but it often lives in an isolated environment while the host's python3 is still stock macOS 3.9.

A run leaves a directory of evidence. This skill is about reading it honestly: saying what it proves, and refusing to say what it doesn't.

The one thing to get right

The .log file is a terminal redraw capture, not a transcript. The coding CLI paints a live TUI into a tmux or psmux pane. The capture records every repaint.

One adapter differs. opencode-http writes a real transcript instead. Read references/log-forensics.md before applying any of this to it.

Three consequences:

  • Each logical line appears hundreds of times, growing character by character. tail -50 returns fragments of one frame, not the last 50 things that happened.
  • Whitespace is dropped unpredictably. Done(10tooluses·99.2ktokens·1m0s) is one real line.
  • Long tool output is collapsed into … +N lines (ctrl+o to expand) and never painted. Those lines do not exist in the file.

The third one decides what you may claim. Test summaries sit inside collapsed blocks. Your test runner's pass/fail line is long output, so the TUI hides it. One measured 4.35 MB log had 58 collapse markers hiding roughly 2,463 lines. Its largest single block was 484 lines.

So never report that tests passed or failed based on the log. Report which test commands ran. The verdict comes from "Deciding whether verify passed" below.

The heuristics are stack-specific. They fail quietly. scripts/extract_transcript.py matches one coding CLI's TUI vocabulary: its spinner frames, its Done(N tool uses · T tokens · Ns) footer, its … +N lines (ctrl+o to expand) marker. Its results and errors sections match one test runner's banner, one type checker's error codes, one runtime's errno names.

Change the CLI or change the test stack and the script does not error. It matches fewer patterns and still returns a result. An empty --section errors or --section results on an unfamiliar CLI or stack is a miss, not good news.

So establish which CLI and which stack the repo runs before you read anything. The next section covers that. Run scripts/extract_transcript.py --collapsed to print how much is hidden, then quote that number instead of hedging.

Know the project before you read it

This skill ships generic. Everything that varies between repositories lives in a project adapter at _project/bmad-loop/. Read it before the first inspection. Bootstrap it when it is not there.

FileHolds
_project/bmad-loop/environment.tomlthe values: which coding CLI, which test runner and type checker, which multiplexer, whether the dev skill writes result.json, which verify steps are non-fatal, what feeds the backlog
_project/bmad-loop/environment.mdthe knowledge: a dated current-state snapshot, which of the extractor's constants match here and which don't, and the judgment calls a bare value can't carry

When environment.toml is missing:

python3 <skill>/scripts/bootstrap_adapter.py --repo-root /path/to/repo

It writes skeletons, never overwrites, and prints every value still marked TODO(confirm: …). Those TODOs are the point. Research each one from a real file or a real bmad-loop command. Then show the user the drafted values and where each came from, before you inspect anything.

An adapter of plausible guesses is worse than no adapter. A wrong coding_cli makes the extractor match nothing and return a result that reads clean. references/adapter-bootstrap.md sources each field, including the ones no one can answer until a run exists.

When the adapter and a live run disagree, the run wins. The disagreement is itself a finding about the adapter. Report it. Do not quietly override either one.

Check the CLI before reading the disk

bmad-loop answers some questions faster and more reliably than parsing artifacts. list, status, diagnose, validate, adapters and bare mux are safe on a live run. Try them first.

CommandWhat it gives you
list, statusRun state, but only for .bmad-loop/runs/. An archived run is invisible to both and returns no such run. Extract the tarball and read the artifacts by hand, starting at ## Workflow
adaptersWhich coding-CLI adapter each profile selects. This is where the adapter's orchestrator.coding_cli comes from. Re-run it when an extraction comes back suspiciously empty
validateLive host facts no run directory holds: multiplexer availability and version, whether the coding CLI is on PATH, hook registration and staleness, worktree cleanliness. Run it when a story fails for reasons that look environmental

Never mutate what you are observing. These commands read as harmless and are not:

  • mux set writes policy.toml.
  • confirm and decisions act on their target unless called with --list.
  • attach joins a live session. Any keystroke sent to it acts.
  • probe-adapter --probe launches a real CLI turn.
  • run, sweep, clean and cleanup act unless given --dry-run.
  • tui is not confirmed inert in every view. Treat it as not scriptable.

references/anomaly-triage.md has the full table.

Workflow

Live watch and post-hoc forensics use the same three steps. Only the leading question differs.

1. Probe the state

python3 <skill>/scripts/run_probe.py --project /path/to/repo

Prints health flags, per-task phase/attempt/review_cycle, heartbeats, log sizes, journal tail, ATTENTION metadata, any pending hard or graceful stop request, and a findings list.

It also writes .probe-snapshot.json into the run directory. The next probe reads that file and reports what changed. The delta is what separates "working" from "hung". It also separates a new ATTENTION notice from an unchanged append-only file. One reading alone answers neither.

Read thresholds from state.json's policy_snapshot, never from memory. Every project tunes max_dev_attempts and session_timeout_min differently. A remembered 2 becomes a false alarm on the next repo.

For a live watch, run this on an interval and compare against the previous probe. For forensics on a finished run, one probe is enough. Go straight to the flags.

2. Reconstruct the narrative

python3 <skill>/scripts/extract_transcript.py --section tools      # what it did, in order
python3 <skill>/scripts/extract_transcript.py --section subagents  # cost per delegated task
python3 <skill>/scripts/extract_transcript.py --collapsed          # how much is unreadable

The script strips escapes and rebuilds each logical line by keeping the longest variant seen. Sections: tools, subagents, errors, results, prose, progress.

progress reports the orchestrating session's own elapsed/token counter. Claude Code prints it as 50m 20s · ↓151.7k tokens. Another coding CLI prints it differently, or not at all.

A log that grows while this counter stands still is worth investigating. It is not proof of a hang. A session that delegates hands its footer to the subagents and stops painting its own counter. So check the prose and results tails for named subagent lines before you call a stall. references/log-forensics.md has the decision table.

3. Cross-check against the working tree

The log says what the agent tried. Git says what actually landed. Read tasks.<story>.baseline_commit from state.json, then:

git status --short
git diff --stat <baseline_commit>

This is where you catch the difference between an agent that wrote code and an agent that narrated writing code. It also catches partial work: five locale files touched and the sixth missed, a service added with no test beside it.

When scm.isolation is none, that diff sits in the user's live checkout. Say so. If the run fails, those changes stay there, and rollback_on_failure = false means nothing cleans them up.

Deciding whether verify passed

The log cannot tell you. Read journal.jsonl instead. It is authoritative and structured.

Do not use session-end.status as the verdict. It looks like one and is not. session-end.status is one of completed | stalled | timeout | crashed | over_budget | aborted, and it describes only whether the CLI session ended normally. It never says whether the work was accepted. A completed session can still be rejected. A crashed or timeout session can still be salvaged. Grepping this field and stopping there is the mistake.

Read the actual verdict in this order:

  1. dev-decision.action is the authoritative outcome for that attempt. One of proceed, retry, defer, pause, salvage.
  2. The terminal journal kind records where the story landed: story-done, story-deferred, story-escalated, story-awaiting-operator.
  3. tasks.<story>.phase in state.json should agree with whichever of the above fired.

A finally block writes session-end for every session, crashed ones included. So a launched session with no session-end is itself a finding, not a gap to explain away. Silence anywhere else carries no such guarantee: a journal with no failure entries proves only that nothing reportable has happened yet.

A story can land at story-awaiting-operator and stay there indefinitely. That is terminal, not stuck. It clears only when a human runs bmad-loop confirm <story-key>. Read references/anomaly-triage.md for the full handling. Do not improvise it here.

A story landing at story-escalated pauses the run. Its reason needs one extra step, because bmad-loop cuts the escalation text at 2000 characters and appends no marker. Five places carry byte-identical copies of that same cut: dev-decision.reason, story-escalated.reason, run-paused.reason, state.json's paused_reason, and the ATTENTION notice. Corroborating them against each other is circular. It proves nothing.

The uncut text is in the story spec's ## Auto Run Result section. tasks.<story>.spec_file names the file. Reading only the truncated copies is how a real blocker gets reported as a misclassification. references/anomaly-triage.md has the reading order and the matching care about which remedy to offer.

Watch the field names. session-end carries status. dev-decision carries a differently-named session_status. rc belongs to plugin-hook alone. The wrong key on the wrong kind returns a plausible-looking wrong answer.

Re-run the command yourself if the user needs the actual failing assertions. The journal gives the verdict, not the test output. Run the verify command from policy_snapshot.verify.commands directly and report that.

Name these two traps when you report:

  • || true swallows failures. A verify command ending in || true always exits 0. These are operator-authored, not shipped by bmad-loop. The adapter's verify.non_fatal_steps names the ones a given project made non-fatal, derived from policy_snapshot.verify.commands. Re-check there when the two disagree. The failure itself is invisible. The only symptom is downstream: a sprint backlog count that never moves although a story reached done.
  • Verify runs twice per story and discards output on timeout. A verify step with no output did not necessarily skip. It may have timed out and thrown the evidence away.

Reporting

Lead with the state, then the analysis, then the recommendation. On a live watch where nothing is wrong, one or two lines is the whole report. The user asked to be told when something is wrong. A wall of green text trains them to stop reading.

When something is wrong, use this shape:

<current state: story, phase, attempt, elapsed>
<what changed since last check>
<the finding, and the evidence for it>
<what to do — the exact command>

Anomalies fall into three tiers. references/anomaly-triage.md has the full table, the policy key behind each threshold, and the remedy.

  • Tier 1, needs a human now. crashed or crash_error set. paused_reason or paused_stage set. Engine pid dead while the run is unfinished. A new or unresolved ATTENTION notice. The run concluding. An ATTENTION file's existence alone is not enough, because the file is append-only. An escalation needs the extra step above before you explain the pause.
  • Tier 2, about to fail. attempt at the policy max. review_cycle not converging. stall_armed set or nudges sent. Stale heartbeat. Session budget nearly gone while still in dev.
  • Tier 3, silent rot. The ones nothing else catches. Log growing while the progress counter is frozen. Identical tool calls repeating across checks. Deferred-work ledger swelling while sweep.auto = "never". Backlog stuck despite stories completing.

Tier 3 is the reason this skill exists. bmad-loop tui already shows tiers 1 and 2. Tier 3 is visible only to someone who reads the artifacts and compares them over time.

Reference material

Read these when the question goes past the workflow above:

FileRead it when
references/run-anatomy.mdYou need the exact key that answers a question: which file, which field, what its values mean
references/log-forensics.mdThe reconstruction is losing something, or you need data the default sections drop
references/anomaly-triage.mdYou have a finding and need the threshold's source and the right remedy
references/adapter-bootstrap.md_project/bmad-loop/ is missing or incomplete, and you need where each field's value legitimately comes from

Honesty rules

These exist because the failure mode of this task is a confident, wrong, reassuring report.

  • Distinguish "I read this" from "I inferred this". The user acts on the difference.
  • Absence of error lines is not evidence of success. That holds double here, because the error lines are structurally absent from the capture.
  • Absence of a stated blocker is not evidence that there was no blocker. When a notice is truncated, say the text is partial. Go to the uncut source before concluding anything.
  • When a reading is ambiguous, name the extra command that would settle it, and offer to run it.
  • Never claim a story is done because the agent said it was done. Check the phase and the diff.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Coach structured brainstorming sessions on any topic: diverge on ideas together, then converge on priorities and next steps in a session document. Use whenever the user wants ideas or wants to think out loud, in any language ("brainstorm with me", "give me ideas for…", "help me come up with…", "let's explore options for…", "I'm stuck on the name / the features"; Vietnamese "brainstorm cùng tôi", "cho tôi ý tưởng", "nghĩ cùng tôi xem"), even when they never say "brainstorm". Also use when the user wants multiple perspectives or a devil's-advocate pass on their ideas ("give me different angles", "phản biện mấy ý tưởng này", "party mode"). Not for diagnosing why something is failing (problem-solver), user research (design-thinking), or company-level strategic bets (strategy-board).

日本語の概要は準備中です。原文の説明を表示しています。

tronghieu/agent-skills742026年10月10日 更新

Audit the reasoning of a document or argument (memo, proposal, investment analysis, board paper, article, the user's own draft): claims, evidence, unstated assumptions, logical gaps, fallacies, and what would falsify it, each anchored to quoted text. Use whenever the user wants reasoning examined, in any language ("audit this argument", "is this analysis sound", "poke holes in this proposal", "review the logic of my draft"; Vietnamese "phản biện giúp tôi", "soi lập luận này", "tài liệu này có lỗ hổng gì", "góp ý bản nháp của tôi"), even when they never say "critical thinking". Also use when the user drops a document and asks whether to trust or act on it, or asks to see their reasoning profile. Not for company-level strategic bets (strategy-board) or learning a topic through dialogue (socratic-questor).

日本語の概要は準備中です。原文の説明を表示しています。

tronghieu/agent-skills742026年10月10日 更新

cv-scorer

無料

Score candidate CVs on a 100-point scale against a Job Description. Use this skill when the user wants to evaluate, score, rank, or screen candidate CVs/resumes against a JD. Also trigger when the user mentions 'review CV', 'screen resume', 'rate candidates', 'shortlist', or any context involving matching resumes to job requirements.

日本語の概要は準備中です。原文の説明を表示しています。

tronghieu/agent-skills742026年10月10日 更新

Act as an end-to-end data scientist: turn business questions into defensible analysis, validated models, and decision-ready reports. Use whenever the user asks to analyze, explore, or profile a dataset or CSV/Parquet/Excel file; asks what drives a metric or why a number changed ("why did churn go up?"); wants to test whether a difference is real (A/B tests, experiments, "is this significant?", "how many samples do I need?"); wants a predictive model (churn, forecast, scoring, segmentation, classification, regression); asks to review an existing analysis, notebook, or model for flaws; or needs results written up for decision-makers. Triggers in any language ("phân tích dữ liệu", "xây model dự đoán", "kiểm định A/B"), even when they never say "data science" or "statistics".

日本語の概要は準備中です。原文の説明を表示しています。

tronghieu/agent-skills742026年10月10日 更新

Read, study, and answer questions about long books and papers without loading the whole text into one context window, keeping page-anchored notes as a reusable workspace. Use whenever the user asks to read, study, summarize, analyze, review, or answer questions about a long book, textbook, PDF, EPUB, thesis, dissertation, survey paper, or any document of roughly 50+ pages, in any language (Vietnamese "đọc sách", "tóm tắt sách", "phân tích luận án"). Also use when a `{slug}-notes/` workspace from a previous session sits next to a source file and the user asks a follow-up question about that book. Not for short documents (under ~30 pages) that fit in context; read those directly.

日本語の概要は準備中です。原文の説明を表示しています。

tronghieu/agent-skills742026年10月10日 更新

design-thinking

無料日本語概要

Facilitate a full design-thinking engagement (Empathize, Define, Ideate, Prototype, Test) grounded in real user evidence. Use whenever the user wants to understand users deeply and design a solution for them: "run design thinking", "understand our users", "design user interviews", "synthesize these interview notes", "build personas", "write How-Might-We questions", "prototype this concept", "design a usability test", "our users are churning and we don't know why", in any language ("tư duy thiết kế", "nghiên cứu người dùng", "phỏng vấn khách hàng", "デザイン思考", "设计思维"), even when they never say "design thinking". Also use when the user drops raw interview notes, transcripts, or survey exports and wants insights, personas, or product decisions. Not for "is there a market for X" desk questions (market-researcher).

tronghieu/agent-skills742026年10月10日 更新

tronghieu のスキルをすべて見る

このスキルの問題を報告する