Audit Claude Code agents, skills, and commands for quality and production readiness. Use when evaluating skill quality, checking production readiness scores, or comparing agents against best-practice templates.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluate staged changes using LLM-as-a-Judge before committing
インストールする前に、エージェントに与えられる指示の中身を確認できます。
Evaluate staged git changes using the output-evaluator agent to catch issues before committing.
Run git diff --cached --stat to see what's staged. If nothing is staged, inform the user and exit.
Run git diff --cached to get the complete diff of all staged changes.
Use the Task tool to launch the output-evaluator agent with the diff:
Evaluate these staged changes for correctness, completeness, and safety.
Return a JSON verdict with scores and issues.
Changes:
[paste the git diff here]
Based on the evaluation result:
If APPROVE:
If NEEDS_REVIEW:
If REJECT:
If user confirms, create the commit using the standard commit flow.
/validate-changes
Output:
Evaluating 3 staged files...
VERDICT: NEEDS_REVIEW
Scores:
Correctness: 8/10
Completeness: 6/10
Safety: 9/10
Issues Found:
[MEDIUM] src/api/handler.ts:45
Missing error handling for network failures
[LOW] src/utils/format.ts:12
Consider adding input validation
Suggestion: Add try-catch around the fetch call in handler.ts
How would you like to proceed?
1. Fix issues and re-evaluate
2. Commit anyway (1 medium issue)
3. Abort
This command invokes an LLM evaluation, which uses API tokens:
For automatic evaluation on every commit, see pre-commit-evaluator.sh hook.
This command is the manual alternative when you want control over when evaluation runs.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Audit Claude Code agents, skills, and commands for quality and production readiness. Use when evaluating skill quality, checking production readiness scores, or comparing agents against best-practice templates.
日本語の概要は準備中です。原文の説明を表示しています。
Codebase health audit scoring 7 categories with progression plan
日本語の概要は準備中です。原文の説明を表示しています。
Autonomous improvement loop: scan codebase metrics, scaffold experiment files, run agent-driven iterations until metric improves
日本語の概要は準備中です。原文の説明を表示しています。
Generate bounded independent candidates, score them against a frozen rubric, and verify the selected result with a proof log.
日本語の概要は準備中です。原文の説明を表示しています。
Post-deploy monitoring: watch production after a deploy and alert on regressions
日本語の概要は準備中です。原文の説明を表示しています。
Restore context after /clear by summarizing recent work and project state
日本語の概要は準備中です。原文の説明を表示しています。