本文へ移動
cccskills
無料GitHub で公開

write-code-eval

Write code evaluators for known failure modes with objective rules. Use when code can check the rule from a trace, with or without a reference answer. Use `write-judge-prompt` when the rule requires interpretation.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md1.5 KB
  • agents/openai.yaml264 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Write a code evaluator

Start with a failure mode found through error analysis. Write one check for that failure mode, much like a unit test that asserts what should hold for each trace.

  1. State the rule and identify the trace fields or reference data the check needs. If the rule requires interpretation, use write-judge-prompt.
  2. Implement the check in the project's language and eval framework. Return a result and a reason in the format that framework expects.
  3. Test known passes and failures, including borderline cases. Run the check on available traces and inspect mistakes. If the rule uses a proxy for human judgment, compare its results with human labels.

Examples

Failure modePossible check
Invalid output structureParse the output and check required fields
Missing or forbidden textMatch a string or pattern
Citation not in retrieved documentsCompare cited IDs with retrieved IDs
Bad tool callCheck arguments against the tool schema or run the call in a safe test environment
Wrong valueCompare the output with a reference value

Choose the check from the failure rule. For a failure with both objective and interpretive parts, check the objective part with code and use a judge for the rest.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Run error analysis on a dataset. Build a review UI, select diverse samples, monitor annotations, and organize failure modes.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-discovery, write-judge-prompt, or validate-evaluator instead).

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for something no other skill in this plugin matches. Do NOT use when a more specific skill in this plugin already matches; load that skill directly.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generation system, measuring retrieval quality, assessing generation faithfulness or relevance, generating synthetic QA pairs for retrieval testing, or optimizing chunking strategies.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Create diverse synthetic test inputs for LLM pipeline evaluation using dimension-based tuple generation. Use when bootstrapping an eval dataset, when real user data is sparse, or when stress-testing specific failure hypotheses. Do NOT use when you already have 100+ representative real traces (use stratified sampling instead), or when the task is collecting production logs.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

ai-evals-course のスキルをすべて見る

このスキルの問題を報告する