Dispatch independent tasks to parallel workers or subagents without write collisions. Use when: running or planning agents in parallel, even two; check scopes before any launch.
日本語の概要は準備中です。原文の説明を表示しています。
Write or assess tests that prove behavior and would fail without the fix. Use when: writing tests, TDD, or asked whether a green test is enough.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Write or strengthen tests for a named behavior. Use existing tests directly when the task is only to run a known suite; this skill is not a required wrapper. A test is useful when it distinguishes an accepted outcome from a plausible failure, not merely when it executes the implementation. A green test is evidence for the caller's decision, never the test author's merge approval.
| Mode | Use when | Result |
|---|---|---|
generate | Existing behavior needs tests | Useful tests and focused/suite results |
coverage | The caller asks to find or fill gaps | Before/after coverage, valuable tests and remaining risks |
tdd | New behavior is being developed test first | Real expected RED, implementation, green and refactor |
strategy | The caller wants test design only | Prioritized risks and proposed checks in the existing discussion |
Default to generate; mode and scope are skill prompt choices, not invented
CLI flags. Coverage thresholds come from the caller or repository.
Prefer exact observable values or errors when known. Use properties or invariants when they express the contract more faithfully than one example. Differential agreement needs an independently credible reference. A smoke check proves only what it observes; it cannot establish an exact behavior by itself. Explain a material oracle limit in the native handoff, without creating a worksheet or mandatory report.
Establish that an important new behavioral check can catch the defect it claims to guard. An authentic pre-fix RED or reproduction is usually sufficient. If a regression test was written after the fix, run it against the pre-fix version (revert the fix in an isolated copy) or use a safe, targeted negative control. Mutate only when that would resolve real doubt about the oracle, then restore and verify the candidate. Do not demand one mutation experiment per table row or new test.
Confirm the runner completed, the intended tests actually ran, and assertions observe the promised behavior. Report crashes, truncation, unexpected skips or exclusions as gaps. When runner discovery or failure reporting changed, use a negative control through that same path before trusting green. No need to re-prove an unchanged healthy runner on each edit.
.feature file is optional;
if the repository already uses scenario-to-test annotations, maintain them
and use its scenario coverage checker.coverage
mode or under an existing repository requirement.tdd mode run it before implementation and require the expected
missing-behavior failure, then implement and refactor under green.Load only the guidance needed by the subject:
Tests belong in the repository's language-native locations. Check facts and limits go in the existing handoff:
tests: <file::name> -> <behavior and exact values it asserts>
proof: <each important new test> -> how it was shown to fail on its defect:
pre-fix RED | reverted-fix run | negative control | not shown
commands: <exact command> -> <result>
harness: <runner completed; how many ran; skips or exclusions>
defects: <discovered defect and reproducer, or none>
unchecked: <material behavior no test covers>
Persist coverage or other reports only when requested or required by a declared
consumer, at its selected destination; no automatic .agents/ output. Factual
green is input to the caller's merge or review decision, not the test author's
binding PASS.
Example: for a duplicate Job delivery, assert that the completed result is returned and the external side effect is called only once. A coverage increase without those assertions would not prove the behavior.
This guidance uses original examples informed by Matt Pocock's engineering skills, with AgentOps' existing acceptance and evidence boundaries.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Dispatch independent tasks to parallel workers or subagents without write collisions. Use when: running or planning agents in parallel, even two; check scopes before any launch.
日本語の概要は準備中です。原文の説明を表示しています。
Run a supplied task in headless AGY (Antigravity, Gemini) and collect its result. Use when: AGY, Antigravity or Gemini is requested by name; never a fallback.
日本語の概要は準備中です。原文の説明を表示しています。
Run one prompt through headless Claude with scoped permissions and a time bound. Use when: scripting or automating a `claude -p` call, even a simple one.
日本語の概要は準備中です。原文の説明を表示しています。
Run one prompt through headless Codex and capture the result. Use when: wanting a one-shot `codex exec` run or CI step. Not for batches or retries.
日本語の概要は準備中です。原文の説明を表示しています。
Compare independent opinions from several models or contexts without inflating agreement. Use when: wanting a second opinion or debate, or summarizing several reviewers' results.
日本語の概要は準備中です。原文の説明を表示しています。
Draft or lint a bounded long-running goal prompt with a finish line and hard limits. Use when: selected by name; one change goes to Plan.
日本語の概要は準備中です。原文の説明を表示しています。