This skill should be used when the user asks to "check accessibility", "audit WCAG compliance", "scan HTML for a11y issues", "check color contrast", or "find accessibility violations in web pages".
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used when the user asks to "evaluate LLM output quality", "set up LLM-as-judge", "build an eval rubric", "compare model outputs pairwise", or "measure agent quality".
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Category: Engineering Domain: AI Engineering
Design and run trustworthy evaluations for LLM and agent outputs: pick the right grading method (programmatic check, LLM-as-judge, or human review), write a scoring rubric that judges can apply consistently, rank competing variants by pairwise comparison, and watch for the biases that quietly corrupt judge scores — position bias, verbosity bias, and self-preference. The goal is an eval that you can trust enough to ship on: calibrated against human labels, cheap enough to run on every change, and tracked alongside cost and latency so you never trade quality away by accident. This skill is model- and vendor-agnostic: it reasons about the evaluation method, not any one provider's API, and its scripts aggregate scores you have already collected — they never call a model.
Before designing or running an evaluation, confirm these inputs. If any is unknown or vague, ASK — do not assume:
criteria and weights)rubric_scorer.py absolute scoring vs pairwise_ranking.py comparison vs human-in-the-loop)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions.
cd engineering/agentic-evaluation-framework
# 1. Score outputs against a weighted rubric + check inter-rater agreement
python scripts/rubric_scorer.py --data rubric_scores.json
# 2. Rank competing variants from pairwise (A-vs-B) judgements
python scripts/pairwise_ranking.py --data pairwise_matches.json
# JSON output for piping into a dashboard or CI gate
python scripts/rubric_scorer.py --data rubric_scores.json --json
| Tool | Purpose | Key Flags |
|---|---|---|
scripts/rubric_scorer.py | Aggregate per-criterion scores into weighted totals, per-criterion means, pass/fail vs thresholds, and an inter-rater agreement metric | --data, --json |
scripts/pairwise_ranking.py | Turn head-to-head win/loss records into a ranking via Elo + Bradley-Terry, plus a win-rate matrix | --data, --k, --base, --json |
Both scripts: Python 3 standard library only, argparse CLI, --json and human-readable output. They compute over scores you provide and never call a model. Run --help for full usage.
references/llm-judge-methodology.md).rubric_scorer.py input JSON.rubric_scorer.py and read inter_rater_agreement: low agreement means the rubric is ambiguous, not that a grader is wrong — tighten the anchors and re-score before trusting any number.--json → pass/fail), and re-run agreement periodically to catch judge drift.pairwise_ranking.py matches, using "winner": "tie" for disagreements.pairwise_ranking.py to get Elo and Bradley-Terry rankings plus the win-rate matrix; Bradley-Terry is order-independent and preferred for a fixed batch, Elo for a streaming sequence of matches.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
This skill should be used when the user asks to "check accessibility", "audit WCAG compliance", "scan HTML for a11y issues", "check color contrast", or "find accessibility violations in web pages".
日本語の概要は準備中です。原文の説明を表示しています。
Design and analyze A/B tests: sample size, test duration, and statistical significance for conversion experiments. Use when setting up an A/B test, calculating sample size, designing an experiment, or analyzing results.
日本語の概要は準備中です。原文の説明を表示しています。
Design and run statistically rigorous A/B tests and experiments. Use when planning experiments, calculating sample sizes, designing test variants, selecting metrics, analyzing results, or when someone says "let's test that."
日本語の概要は準備中です。原文の説明を表示しています。
Sales execution across pipeline, discovery, demos, negotiation, and closing. Use when qualifying opportunities, running MEDDIC discovery, building account plans, handling objections, structuring proposals, or forecasting pipeline.
日本語の概要は準備中です。原文の説明を表示しています。
Design ad creative across Google, Meta, LinkedIn, Twitter/X, and TikTok with platform format specs, headline formulas, and A/B testing. Use when writing ad copy, generating headline variations, creating ad sets, or validating creative.
日本語の概要は準備中です。原文の説明を表示しています。
Answer Engine Optimization (AEO): optimize content to be cited by LLMs (ChatGPT, Claude, Perplexity, Gemini) in their answers. Use when designing content for LLM citation, auditing citability, or structuring Q&A schema.
日本語の概要は準備中です。原文の説明を表示しています。