本文へ移動
cccskills
無料GitHub で公開

agent-self-eval

Post-run self-evaluation system that scores agent output on correctness, clarity, actionability, and conciseness. Use after /team runs, skill executions, or when explicitly asked to evaluate output quality.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

@agents/PROMPT-DEFENSE.md

Agent Self-Evaluation

Score your own output (or another agent's output) across four axes to identify quality gaps and feed improvements into the learning system.

When to Use

  • After completing a /team:* pipeline run
  • After generating a deliverable (PRD, architecture doc, code review)
  • When user asks "how did I do?" or "evaluate this output"
  • Automatically at end of /gsd-execute-phase for quality tracking

Evaluation Axes

AxisQuestionFailure Signals
CorrectnessIs the output factually accurate and technically sound?Wrong APIs, broken references, hallucinated facts, logic errors
ClarityIs the explanation understandable and well-structured?Confusing structure, undefined jargon, missing context, rambling
ActionabilityCan the user act on the output immediately?Vague suggestions, missing steps, no verification path
ConcisenessDid it use the minimum tokens needed?Redundancy, over-explanation, filler content, restating the question

Scoring Scale

5 — Exceptional: no reasonable improvement possible
4 — Good: minor nits only, no substantive gaps
3 — Adequate: meets request but has notable weakness on ≥1 axis
2 — Weak: clear gap affecting usability or correctness
1 — Poor: fundamentally misses request or contains significant errors

The Evidence Rule

Every score below 5 MUST cite specific evidence. A score of 3 cannot just say "could be better" — it must say exactly what is missing or wrong. "Show the gap, don't just name it."

Procedure

Step 1: Collect Raw Material

Gather:

  • Original user request
  • Final output/deliverable
  • Tool outputs verifying correctness (test results, exit codes, lint)
  • User feedback received during task (corrections, "try again")

Step 2: Score Each Axis Independently

Rate 1-5 with mandatory evidence for scores <5.

Step 3: Generate Eval Report

SELF-EVALUATION REPORT
======================
Task: {brief description}
Overall: {weighted average}/5

CORRECTNESS: {score}/5
  Evidence: {specific finding or "No issues found"}

CLARITY: {score}/5
  Evidence: {specific finding or "No issues found"}

ACTIONABILITY: {score}/5
  Evidence: {specific finding or "No issues found"}

CONCISENESS: {score}/5
  Evidence: {specific finding or "No issues found"}

IMPROVEMENT INSTINCTS:
- {trigger} → {action} (confidence: {0.3-0.9})

Step 4: Feed Learning System

If learning system is active (PR-27+), auto-generate instinct YAML from findings:

---
id: eval-{task-slug}-{axis-lowercase}
trigger: "when {task type}"
action: "{specific improvement}"
confidence: 0.6
domain: quality
source: self-eval
scope: project
---

Integration Points

  • /team:verify invokes self-eval on Layer 2 output before Layer 3 review
  • /gsd-execute-phase runs self-eval per subagent, aggregates in SUMMARY.md
  • High-confidence eval instincts (≥0.8) auto-update team-toolkit.md quality notes
  • Low scores (≤2) on correctness trigger automatic re-execution offer

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Train and optimize AI agents using Microsoft's Agent Lightning framework with reinforcement learning. Use when setting up agent training, instrumenting agents with tracing, configuring LightningStore, implementing reward functions, or optimizing prompts with RL/APO algorithms.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Create AI marketing videos for ads, promos, product launches, and brand content. Models: Veo, Seedance, Wan, FLUX for visuals, Kokoro for voiceover. Types: product demos, testimonials, explainers, social ads, brand videos. Use for: Facebook ads, YouTube ads, product launches, brand awareness. Triggers: marketing video, ad video, promo video, commercial, brand video, product video, explainer video, ad creative, video ad, facebook ad video, youtube ad, instagram ad, tiktok ad, promotional video, launch video

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Use when building AI features into a product: LLM integration, RAG pipelines, guardrails, streaming, AI UX, prompt engineering, or AI cost control. Treats prompts as code and validates every model output.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Your AI research and engineering brain trust. 59 named personas across 8 cells covering frontier labs, applied product, model architecture, reasoning/RL/agents, alignment and interpretability, theory and science of DL, multimodal and…

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Use when designing a new REST or GraphQL API, reviewing an API spec before implementation, setting team API standards, or migrating REST to GraphQL. Covers resources, HTTP semantics, pagination, error handling, and pitfalls.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

Use for authorized security assessment of REST, GraphQL, WebSocket, or SOAP APIs, including discovery, authentication, authorization, rate-limit, and CI/CD testing.

日本語の概要は準備中です。原文の説明を表示しています。

coco-research/coco5462026年10月11日 更新

coco-research のスキルをすべて見る

このスキルの問題を報告する