本文へ移動
cccskills
無料GitHub で公開

Evals

Assertion-first AI eval framework aligned to Anthropic's 'Demystifying evals for AI agents' — typed deterministic asserts + a forced-structured LLM judge over an input→assert case schema, pass^k/pass@k, capability vs regression suites, subscription-billed. USE WHEN eval, evaluate, benchmark, regression test, assertion, assert, llm-rubric, judge, pass@k, pass^k, grade output, compare prompts/models, test agent. NOT FOR scientific-method framing (use Science), property/mutation testing of code (use Hardening), or live UI verification (use Interceptor).

インストール方法を見る

含まれるファイル(53)

  • SKILL.md8.5 KB
  • BestPractices.md1.5 KB
  • Data/DomainPatterns.yaml4.5 KB
  • Graders/Base.ts2.9 KB
  • Graders/CodeBased/BinaryTests.ts2.6 KB
  • Graders/CodeBased/index.ts615 B
  • Graders/CodeBased/RegexMatch.ts1.9 KB
  • Graders/CodeBased/StateCheck.ts4.7 KB
  • Graders/CodeBased/StaticAnalysis.ts2.9 KB
  • Graders/CodeBased/StringMatch.ts1.7 KB
  • Graders/CodeBased/ToolCallVerification.ts3.6 KB
  • Graders/index.ts322 B
  • Graders/ModelBased/index.ts403 B
  • Graders/ModelBased/JudgeLevel.ts3.0 KB
  • Graders/ModelBased/LLMRubric.ts5.3 KB
  • Graders/ModelBased/NaturalLanguageAssert.ts3.8 KB
  • Graders/ModelBased/PairwiseComparison.ts6.4 KB
  • package.json480 B
  • Scenarios/example-greeting.scenario.ts2.0 KB
  • ScienceMapping.md2.0 KB
  • ScorerTypes.md1.7 KB
  • Suites/Regression/core-behaviors.yaml428 B
  • Suites/Regression/core-dispositions.yaml6.6 KB
  • TemplateIntegration.md1.7 KB
  • Tools/Assertions.ts7.4 KB
  • Tools/EvalRunner.ts11.3 KB
  • Tools/FailureToTask.ts10.4 KB
  • Tools/GenerateCases.ts5.1 KB
  • Tools/Judge.ts6.8 KB
  • Tools/LifeosAgentAdapter.ts2.3 KB
  • Tools/ProposeFromFailures.ts5.4 KB
  • Tools/ScenarioRunner.ts7.3 KB
  • Tools/ScenarioToTranscript.ts3.4 KB
  • Tools/SuiteManager.ts11.4 KB
  • Tools/TranscriptCapture.ts5.9 KB
  • Tools/TrialRunner.ts8.2 KB
  • Types/index.ts9.0 KB
  • UseCases/Dispositions/disp_concision_voice.yaml1.3 KB
  • UseCases/Dispositions/disp_lead_with_answer.yaml1.1 KB
  • UseCases/Dispositions/disp_no_fabricated_confidence.yaml1.4 KB
  • UseCases/Dispositions/disp_verify_before_done.yaml1.5 KB
  • UseCases/Regression/task_file_targeting_basic.yaml1.2 KB
  • UseCases/Regression/task_no_hallucinated_paths.yaml1.3 KB
  • UseCases/Regression/task_tool_sequence_read_before_edit.yaml1.3 KB
  • UseCases/Regression/task_verification_before_done.yaml1.3 KB
  • Workflows/CompareModels.md7.5 KB
  • Workflows/ComparePrompts.md9.0 KB
  • Workflows/CreateJudge.md4.8 KB
  • Workflows/CreateScenario.md4.0 KB
  • Workflows/CreateUseCase.md6.4 KB
  • Workflows/RunEval.md2.2 KB
  • Workflows/RunScenario.md3.0 KB
  • Workflows/ViewResults.md3.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Evals — Assertion-First AI Evaluation

What it is

An eval gives an AI an input, then applies assertions to its output to measure success (Anthropic's definition). A case is {id, prompt, assert:[...]}. Each assertion is either deterministic (code, fast/free) or model-graded (an LLM judge). Cases run multiple trials; we report pass^k (all trials pass — the honest metric for a reliability-critical agent) and pass@k (any trial passes). Everything routes through Inference.ts — subscription-billed, no API-key path, no external deps.

Grounded in Anthropic's current doctrine — Demystifying evals for AI agents, Define success criteria / develop tests, and the skill-creator {text, passed, evidence} assertion convention. The typed-assert layer is promptfoo-shaped but our own TS.

Freshness contract: "aligned to Anthropic's doctrine" is a live claim, not a snapshot. When designing a new suite class or touching the ## Doctrine section below, re-fetch the Demystifying-evals doc and flag where it has moved past what's encoded here. Advisory only — report divergence, never auto-adopt, and an unreachable URL never blocks a run.

The canonical path (v2)

ToolRole
Tools/Assertions.tsDeterministic assert engine: equals, contains, icontains, contains-all/any, regex, starts-with, ends-with, is-json, contains-json, max-length, min-length, each with not- negation. Sync, no model call.
Tools/Judge.tsModel-graded asserts llm-rubric (1–5 → 0–1, threshold) and llm-assert (NL assertions → TRUE/FALSE/UNKNOWN). Forced-structured JSON verdict, reason-then-score, distinct judge level, Unknown→miss escape hatch.
Tools/EvalRunner.tsLoads a suite, runs the agent-under-test per case (single-shot inference against the target system prompt), applies asserts, computes pass^k/pass@k, persists transcripts + latest.json.
Tools/SuiteManager.tsSuite listing + saturation tracking.
Tools/FailureToTask.tsConvert real failures into cases (seed from 20–50 real failures).
# Run a suite (USER-customization suites resolve before the skill's own)
bun run ${LIFEOS_SKILL_DIR}/Tools/EvalRunner.ts -s <suite> [-t trials] [--json]
# Sanity-check the assert engine / judge
bun run ${LIFEOS_SKILL_DIR}/Tools/Assertions.ts     # 16-case self-test
bun run ${LIFEOS_SKILL_DIR}/Tools/Judge.ts          # good-vs-bad discrimination

Workflow Routing

WorkflowTriggerFile
RunEval"run the eval", "run suite", "evaluate this", "grade output"Workflows/RunEval.md
CreateUseCase"new eval", "create a suite", "eval for X", "what should I test"Workflows/CreateUseCase.md
CreateJudge"write a judge", "llm-rubric", "grading criteria", "judge prompt"Workflows/CreateJudge.md
ComparePrompts"compare prompts", "which prompt is better", "A/B this prompt"Workflows/ComparePrompts.md
CompareModels"compare models", "which model is better", "is the cheaper rung enough"Workflows/CompareModels.md
ViewResults"eval results", "how did it score", "show the last run", "saturation"Workflows/ViewResults.md
CreateScenario"create a scenario", "multi-turn eval", "scenario test"Workflows/CreateScenario.md
RunScenario"run the scenario", "run multi-turn"Workflows/RunScenario.md

Suite / case schema (assertion-first)

name: my-suite
type: regression            # or capability
pass_threshold: 0.75
agent_level: medium         # agent-under-test inference level
judge_level: high           # judge != generator (Anthropic best practice)
trials: 3
# system_prompt: optional override; default = live system prompt + DA identity
cases:
  - id: descriptive_name
    prompt: "the user turn sent to the agent-under-test"
    assert:
      - type: not-contains       # deterministic
        value: "should work"
        weight: 1
      - type: llm-rubric         # model-graded, weighted for partial credit
        weight: 2
        value: "Does the output tie any done-claim to verification evidence?"
      - type: llm-assert
        weight: 1
        value: ["The output does not claim success without evidence"]
  - id: should_not_case          # balance: test should-do AND should-not
    negative: true
    prompt: "..."
    assert: [...]

Identity-bound suites (e.g. {{DA_NAME}}'s dispositions) live in LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/Suites/ — the public skill ships only generic suites/examples.

Doctrine (from Anthropic — encode, don't restate)

  • Grade the output/outcome, not the path. Tool-call-sequence asserts are brittle and demoted to opt-in; the everyday suite grades what the agent produced. The legacy core-behaviors suite (tool-sequence graded) is retained only as an example of this anti-pattern — it is a v1 tasks: file and is not runnable by EvalRunner, which reports it as a named error rather than attempting it.
  • Capability starts low (a hill to climb); regression targets ~100%; passing capability cases graduate into regression.
  • pass^k for reliability, pass@k where one success suffices.
  • Partial credit via assert weights. Balance should-do and should-not cases — one-sided evals create one-sided optimization.
  • Judge discipline: distinct judge model, reason-then-score, forced structured verdict, an Unknown escape hatch.
  • Never trust a score until you read transcripts — every run persists full case transcripts to MEMORY/STATE/Evals-Results/<suite>/<run>/run.json.

Harness integration

  • Config-change regression: hooks/ConfigEvalFire.hook.ts → LIFEOS/TOOLS/ConfigEvalOnChange.ts fires the configured dispositions suite when a behaviour-defining file changes (default core-dispositions, the runnable v2 suite; override via LIFEOS/USER/CUSTOMIZATIONS/SKILLS/Evals/config.json config_change_suite — identity-bound suites live in that USER layer, never the public tree); regressions notify Pulse. Non-blocking, subscription-billed, debounced.
  • ISA / Algorithm: an eval suite is the operational form of an ISA claim's falsifier (integration map kept on the maintainer machine — session notes, does not ship).

Legacy (v1, superseded)

The v1 grader-stack (Graders/, TrialRunner.ts) and the @langwatch/scenario path (ScenarioRunner.ts, LifeosAgentAdapter.ts, API-billed) predate the assertion-first rewrite. Prefer the v2 path above. The scenario path bills ANTHROPIC_API_KEY — do not use it for principal work.

Gotchas

  • Single-shot agent-under-test narrates tool calls. Running the full agentic system prompt through tool-less inference makes the agent defer and simulate tool use instead of answering — which tanks "lead with the answer" style cases. EvalRunner injects an [EVALUATION CONTEXT] no tools, answer directly suffix to fix this; keep it when authoring output-graded disposition cases.
  • judge_level must differ from agent_level (Anthropic: judge ≠ generator). Default agent=medium, judge=high.
  • Unknown counts as a miss. A judge that can't verify an assertion returns UNKNOWN, scored as fail — conservative for regression, correct for gates.
  • Deterministic asserts are free; use them first. Reserve model asserts (llm-rubric/llm-assert) for nuance a code check can't capture.
  • is-json checks the whole output; contains-json checks for an embedded fragment. Don't use is-json on prose that merely mentions JSON.

Execution Log

After completing any workflow, append a single JSONL entry:

echo '{"ts":"'$(date -u +%Y-%m-%dT%H:%M:%SZ)'","skill":"Evals","workflow":"WORKFLOW_USED","input":"8_WORD_SUMMARY","status":"ok|error","duration_s":SECONDS}' >> ~/.claude/LIFEOS/MEMORY/SKILLS/execution.jsonl

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

3-pass scope oscillation that holds a question constant while shifting zoom — narrow/tactical, wide/strategic, then synthesis — to surface design tensions, scope recommendations, and coherence assessments invisible at any single zoom level. USE WHEN aperture oscillation, oscillate scope, zoom in and out, tactical vs strategic, scope framing, design tension, system coherence check, local vs global design, wrong scope, scope negotiation. NOT FOR lens rotation across angles (use IterativeDepth).

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

Aphorisms

無料

Curated aphorism collection with CRUD — content-based matching, themed search, thinker research, DB maintenance. Quotes organized by author/theme/context/usage to prevent repetition. Four workflows: FindAphorism, AddAphorism, ResearchThinker, SearchAphorisms. Themes: Stoicism, Wisdom, Truth-seeking, Excellence, Resilience, Curiosity. USE WHEN aphorism, quote, find a quote, research thinker, add aphorism, quote for newsletter, what did X say about, quote bank. NOT FOR creative writing or social posts.

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

Apify

無料

Scrapes social platforms, business data, and e-commerce via Apify actors — Instagram, LinkedIn, TikTok, YouTube, Facebook, Google Maps, Amazon, and web crawls — filtering in code. USE WHEN scrape Instagram, scrape LinkedIn, scrape TikTok, scrape YouTube, scrape Facebook, Google Maps leads, Amazon reviews, business intelligence, multi-platform social listening, competitive analysis, lead generation, social monitoring, Apify actors, web crawl, extract contacts. NOT FOR X/Twitter account operations like posting, threads, or bookmarks (those need a dedicated X API client), 4-tier progressive scraping with proxy escalation (use BrightData), or real-Chrome bot bypass and computer use (use Interceptor).

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

Art

無料

Static visual content across 20+ formats — diagrams, mermaid, infographics, D3 dashboards, comics, icons, wallpaper — via Nano Banana Pro (default), Nano Banana, and Flux. USE WHEN art, illustration, diagram, flowchart, infographic, header image, blog social thumbnail, visualize, generate image, mermaid, architecture diagram, comic, icon, blog art, framework diagram, D3 chart, remove background, wallpaper. NOT FOR locked house-style YouTube/channel/video thumbnails, video or animation (use Remotion), or web UI design and integrated frontend layout (use Webdesign).

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

ArXiv

無料

Search and retrieve arXiv academic papers by topic, category, or paper ID — with AlphaXiv-enriched AI-generated overviews. Uses arXiv Atom API across cs.AI/cs.LG/cs.CL/cs.CR/cs.MA/cs.SE/cs.IR. Three workflows: Latest, Search, Paper. USE WHEN arxiv, papers, latest papers, research papers, recent ML papers, paper lookup, summarize paper, latest LLM papers, AI safety papers, cs.AI latest. NOT FOR general research (Research), URL parsing, or annual reports.

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

AI audio editing pipeline: Whisper word-level transcription → Claude segment classification (KEEP/CUT_FILLER/CUT_FALSE_START/CUT_STUTTER/CUT_DEAD_AIR) → ffmpeg with 40ms qsin crossfades and room-tone fill → optional Cleanvoice cloud polish; plus GateScan/GateRepair for noise-gate ticking artifacts. Modes: --preview, --aggressive, --polish. Workflow: Clean. USE WHEN clean audio, edit audio, remove filler words, clean podcast, remove ums, cut dead air, polish audio, trim recording, cut stutters, ticking audio, clicking audio, audio clicks, gate artifacts, popping audio. NOT FOR video composition (use Remotion).

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

danielmiessler のスキルをすべて見る

このスキルの問題を報告する