本文へ移動
cccskills
無料GitHub で公開

eval-output

Orchestrator for the eval-output skill suite — evaluate LLM and agent outputs for quality, accuracy, helpfulness, and safety using structured rubrics and LLM-as-judge techniques. Load when the user says "evaluate this output", "score this response", "run an eval", "LLM as judge", "evaluate agent output", "how good is this response", "rate this answer", "eval this", or provides an LLM output that should be assessed for quality. Single entry point for all output evaluation workflows.

インストール方法を見る

含まれるファイル(3)

  • SKILL.md7.6 KB
  • README.md765 B
  • references/examples.md1.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Eval Output

You are the orchestrator for the eval-output skill suite. You accept any LLM or agent output, classify the evaluation need, route to the correct sub-skill, and present a unified evaluation report. You are opinionated — you recommend the right evaluation approach based on what the user is actually trying to learn.

Hard Rules

  • No scoring without criteria. Every evaluation must use an explicit rubric — never score on vibes. If no rubric exists, route to eval-rubric-design first.
  • No single overall score. Always score dimensions independently. A single number hides tradeoffs and blocks root-cause analysis.
  • Justification before score. All LLM-as-judge scoring must require chain-of-thought reasoning before the numeric score — this improves reliability 15-25% (GER-Eval, arXiv:2602.08672).
  • Hard gates are pass/fail. Safety, compliance, and format requirements are binary — never averaged into a quality score.
  • Validate the judge itself before gating on its scores. The judge is a measurement instrument: calibrate it against a small human-labeled golden set using chance-corrected agreement (Cohen's κ) + failure-class recall — raw agreement overstates judge ability by 33–41pp, and a judge can score 20/21 on good outputs while missing 9/9 failures. Protocol: eval-judge/references/judge-calibration.md.
  • Max 1 clarifying question. If evaluation type is ambiguous, ask one question. Never two.

Workflow

Step 1 — Accept Input

Accept: LLM/agent output to evaluate, optional rubric, optional reference/expected output, optional context (prompt, retrieval context, conversation history).

Step 2 — Classify Evaluation Need

SignalRoutes to
User asks to create rubric, define criteria, design eval dimensionseval-rubric-design
User provides output + wants it scored/judged/comparedeval-judge
User wants to set up automated evals, CI integration, eval pipelineeval-pipeline
User provides output but no rubric existseval-rubric-design first → then eval-judge

If ambiguous: ask one question — "Do you want to (a) design evaluation criteria, (b) score a specific output, or (c) set up an automated eval pipeline?"

Step 3 — Route to Sub-Skill

Invoke the matched sub-skill with all available context.

Step 4 — Unified Report

Present the unified report (see Output Format). If blocked at rubric design, report why and stop.


Call Graph

eval-output (orchestrator)
|- eval-rubric-design  → produces rubric docs in docs/evals/
|- eval-judge          → scores outputs using rubrics (direct or pairwise)
\- eval-pipeline       → designs automated eval systems

Output Format

=== Eval Output Report ===
Target: [what was evaluated — output type, task, model]
Eval type: [rubric-design / direct-scoring / pairwise / pipeline-design]

=== Evaluation ===
[Sub-skill specific output]

=== Summary ===
[Key findings, recommendations, next steps]

Gotchas

  • An output that "sounds good" can still fail on accuracy, safety, or completeness — never skip structured evaluation because the output reads well.
  • If the user provides two outputs to compare, route to eval-judge in pairwise mode — not two separate direct scoring runs.
  • Rubrics drift over time as tasks and models evolve. Recommend periodic rubric review when eval results change unexpectedly.
  • Self-evaluation (model judging its own output) has known self-enhancement bias. Recommend a different model for judging when possible.
  • High aggregate scores can mask low business value. Weight rubric dimensions by business impact — a model scoring 48/100 overall can deliver more value than one scoring 62/100 if it wins on the dimensions that matter (AlphaEval 2026).
  • Long-form agent outputs contradict themselves. For any output >1 page, eval-judge runs an internal consistency check (Step 4b) — numeric, factual, and logical consistency across sections.
  • Multi-step agent pipelines need per-step evaluation. Cascade dependency is the #1 pipeline failure mode. eval-pipeline enforces per-step checkpoints before end-to-end eval.
  • Production evals run on traces. Instrument the shipped product via agent-observability first; runtime-learning-loop consumes eval scores and requires a quarantined held-out split — never let optimization touch it.
  • Inter-judge agreement can be illusory. Multiple judges agreeing proves shared surface heuristics, not correctness — rubric structure alone restores 62% of agreement. Ground rubrics in domain knowledge (eval-rubric-design) and validate against human labels before treating consensus as truth (arXiv:2603.11027, 2026).

Example

<examples> <example> <input>Evaluate this response my agent gave about database indexing</input> <output> === Eval Output Report === Target: Agent response on database indexing Eval type: Needs rubric first

No rubric found for this task. Routing to eval-rubric-design to create one.

[Invokes eval-rubric-design → user approves rubric → invokes eval-judge with rubric]

=== Evaluation === [Structured scores per dimension with justifications]

=== Summary === Overall: 3 of 5 dimensions scored 4+/5. Accuracy strong, completeness weak. Recommendation: Add coverage of partial indexes and composite index ordering. </output> </example> </examples>


Common Rationalizations

ExcuseReality
Judge without rubricRubric or dimensions required before scoring.
Single score, no rationaleEvery score needs cited evidence.
Skip bias mitigationPairwise needs position-swap or length check.
Eval once, never againPipeline skills define regression reruns.

Verification

  • Rubric or dimensions referenced
  • Scores tied to observable criteria
  • Bias mitigations applied for pairwise
  • Outputs under docs/evals/ when files written

Red Flags

  • Scoring performed without an explicit rubric
  • Single overall score hides dimensional tradeoffs
  • Numeric score assigned before written justification
  • Two outputs compared via duplicate solo evaluations

Prune Log

Last pruned: 2026-07-08

  • Added judge-validation hard rule + evaluation-illusion gotcha (2026 research pass — arXiv:2606.19544, arXiv:2603.11027)

Impact Report

After completing, always report:

Evaluation complete: [target description]
Eval type: [rubric-design / direct-scoring / pairwise / pipeline-design]
Sub-skill invoked: [name]
Dimensions scored: [N]
Hard gates: [N] pass, [N] fail
Key finding: [one-line summary]
Next step: [recommendation]

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Design execution structure for decomposed processes: single agent or multi-agent topology. Load when user says "design an agent for this", "what agent structure do I need", "architect this", "should this be multi-agent", "what's the right execution structure", "agent topology", "how should agents be organized". Takes process-decomposer output as primary input. If triggered directly without a process entry, calls process-decomposer first.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Internal skill. Called by setup-evaluation after a PASS. Launches agents from a validated architecture spec using Claude Code / Ampcode native parallelism (Task tool). Does NOT generate scripts or SDK code — it outputs structured spawn instructions that the platform executes natively. Never invoked directly by the user. Never launches without a setup-evaluation PASS.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Sync library skills from an agent-loom upstream repo into this project's .agents/skills while preserving project-local and forked skills. Load when the user asks to sync agent-loom, update skills from upstream, rsync from ../agent-loom, pull new library skills, upgrade installed skills, or refresh the .agents folder without losing custom project skills. Also triggers on "sync skills from agent-loom", "update my agent skills", "pull skill library updates", or "merge agent-loom improvements into this repo".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Instrument a shipped product's AI agents with tracing and observability so you can see what they did, why outputs happened, and what each run cost. Plain-language primer plus free-tier-first backend selection (Langfuse, Phoenix, LangSmith, Braintrust) and OpenTelemetry/OpenInference instrumentation. Load when the user asks to add observability, add tracing, instrument my agents, see what my agent is doing in production, set up Langfuse or Phoenix or LangSmith, debug why my agent gave a bad answer, or track LLM cost per request. Also fires when agent-system-architecture or setup-evaluation requires an observability plan for an agent-chain product. NOT for tracing the coding agent itself — that is run-trace. Precondition for runtime-learning-loop.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Run a structured retrospective after development-phase runs of your product's agents — interview the owner in plain language about what went well and poorly, draft ranked improvement hypotheses, then design and run small n=1/n=2 experiments with pre-declared success criteria, guardrails, stop conditions, and a cost/ROI kill-switch. Load when the user says how did that run go, retro this run, the agent output was bad, what should we improve, draft hypotheses, run a small experiment, or after repeated dev runs of an agentic system produce uneven quality. Priority: output quality over performance over cost, each with diminishing-returns stops. NOT a product A/B test (experimentation), NOT coding-agent harness repair (harness-evolution), NOT production-scale learning (runtime-learning-loop).

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

dvy1987 のスキルをすべて見る

このスキルの問題を報告する