本文へ移動
cccskills
無料GitHub で公開

eval-frameworks

Build evaluation frameworks for LLM systems — harness design, graders, datasets, regression tracking, and human-in-the-loop eval. Use when you need systematic measurement of model or agent quality.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Eval Frameworks

If you can't measure it, you can't improve it — and you'll ship regressions you never notice. An eval framework is the harness, datasets, graders, and processes that turn "does it work?" from a feeling into a number tracked over time.

Overview

Components: a task dataset (representative, versioned), a runner (executes the system on tasks, reproducibly), graders (deterministic checks, rubric judges, humans), metrics (aggregated, sliced by category), and tracking (scores over time, per change). The framework's value compounds: every eval run makes the next decision easier and the next regression harder to ship.

When to use

  • Any LLM feature heading to production: build the evals before launch.
  • Comparing models, prompts, or architectures: decide by numbers.
  • Regression prevention: catching quality drops from changes.
  • Tracking quality over time as models and data evolve.

Core concepts

  • Task datasets: real tasks from production (logs, tickets, user queries), categorized by type and difficulty. Versioned; refreshed as the product evolves. 50–200 tasks is a practical range.
  • Runner: executes tasks deterministically — fixed seeds where possible, recorded configs, parallel execution. Reproducibility is the point.
  • Graders: exact match / structured checks for deterministic outputs; rubric-based LLM judges for open-ended (validated against humans); human grading for the highest-stakes slices.
  • Metrics: aggregate scores plus slices — by task type, difficulty, and failure mode. The slices tell you what to fix; the aggregate tells you if you're improving.
  • Regression tracking: scores stored per run with the code/prompt/model version. Diffs on every change; regressions block or flag.
  • Human-in-the-loop: sampled human review calibrating automated graders and catching what they miss. Automation scales; humans ground truth.

Practical workflow

  1. Mine real tasks from production data; categorize by type and difficulty.
  2. Write graders: deterministic where possible; rubric judges where needed — and validate judges against human grades first.
  3. Build the runner: reproducible execution, parallel, with full per-task logging.
  4. Establish baselines: current system scores, sliced by category. This is your "before" picture.
  5. Integrate into the change process: evals run on every meaningful change; regressions investigated before shipping.
  6. Maintain: refresh tasks as the product evolves, re-validate judges periodically, review slices for new failure modes.
Eval framework anatomy:
TASKS/    versioned task sets, categorized
RUNNER/   reproducible execution + logging
GRADERS/  deterministic checks + validated judges
METRICS/  aggregates + slices by type/difficulty
HISTORY/  scores per version — the regression record
HUMAN/    sampled review calibrating automation

Common pitfalls

  • Synthetic tasks only: evals on invented tasks while production differs. Mine from real usage.
  • Unvalidated judges: LLM graders never checked against humans. Measure agreement before trusting.
  • Aggregate-only reporting: "82% pass" hiding a collapsed category. Slice always.
  • No versioning: tasks or graders changing silently. Version everything; diffs must be meaningful.
  • Evals as ceremony: run once, filed away. Value comes from repetition — integrate into the workflow.
  • Grader gaming: optimizing for the grader instead of the task. Rotate tasks; keep humans in the loop.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Master 1Password: vaults, Watchtower, passkeys, SSH agent, CLI, and family/team administration. Use when getting full value from 1Password personally or administering it for others.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

Create 3D visuals with modeling, texturing, lighting, rendering, and optimization for web and product.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

Create 3D web experiences: scene setup, models, materials, lighting, animation, scroll-driven scenes, and performance budgets. Use when adding 3D to websites beyond basic demos.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

Writing abstracts that get papers read — structured content, the 5-sentence core, and journal-specific constraints.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

Learn effectively from courses and academies: choosing programs, studying actively, and converting courses into skills. Use when investing time/money in structured learning.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

Audit designs for accessibility with WCAG checklists covering color, type, focus, motion, and content.

日本語の概要は準備中です。原文の説明を表示しています。

aicodedecode/awesome-muse-skills122026年10月10日 更新

aicodedecode のスキルをすべて見る

このスキルの問題を報告する