Analyze graded `trial_result` JSONL outputs from the eval pipeline, including reliability and custom slicing.
日本語の概要は準備中です。原文の説明を表示しています。
Use test-driven development for behavior-changing feature or fix work, and whenever the user mentions TDD, test-first, red-green-refactor, tracer bullets, integration tests, or public-interface behavior tests. Skip for docs-only, path-only rename, formatting-only, or purely mechanical chores unless explicitly requested.
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Work in vertical red-green-refactor slices:
Do not write a batch of imagined tests before implementation. Each new test should respond to what the last cycle taught you about the actual behavior and interface.
bun test for targeted tests and bun --bun tsc --noEmit for the type gate.test, not it, and organize with describe.AGENTS.md rules for the touched area.Explore the codebase first. If the public interface, expected behavior, or risk boundary is discoverable from code, tests, or docs, use that evidence instead of asking the user.
Ask one concise question only when the decision cannot be discovered and a reasonable assumption would be risky. Include your recommended answer.
Before writing production code, identify:
For each behavior:
If the test passes before production code changes, it is not a valid RED signal. Tighten the test, choose a different behavior, or explain why existing coverage already proves the behavior.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Analyze graded `trial_result` JSONL outputs from the eval pipeline, including reliability and custom slicing.
日本語の概要は準備中です。原文の説明を表示しています。
Interview the user relentlessly about a plan or design until reaching shared understanding, resolving each branch of the decision tree. Use when user wants to stress-test a plan, get grilled on their design, or mentions "grill me".
日本語の概要は準備中です。原文の説明を表示しています。
Build command-only adapters for `eval` run mode using strict stdin/stdout JSON contracts.
日本語の概要は準備中です。原文の説明を表示しています。
End-to-end eval-suite orchestration with the `eval` command: run -> grade -> compare -> calibrate, using strict JSON contracts and schema discovery.
日本語の概要は準備中です。原文の説明を表示しています。