本文へ移動
cccskills
無料GitHub で公開

agent-evaluation

Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, objective assertions, rubric grading, trajectory analysis, cost/latency tracking, and regression comparison. Use when measuring agent quality, optimizing skill triggering, comparing prompts or models, or gating an AI feature release. Not for ordinary deterministic unit tests.

インストール方法を見る

含まれるファイル(3)

  • SKILL.md2.4 KB
  • agents/openai.yaml221 B
  • references/eval-schema.md1.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Agent Evaluation

Build a quality flywheel that can distinguish a real improvement from a lucky run.

Define the evaluation contract

Name the target behavior, users, risks, baseline, candidate, environment, stochastic settings, and decision threshold. Start with a few realistic cases, including a boundary or failure case. Split trigger-query optimization into a fixed training set and held-out validation set.

Use eval-schema.md for cases, assertions, timing, and result records.

Run isolated comparisons

  1. Snapshot the baseline before changing the candidate.
  2. Run baseline and candidate on identical inputs in fresh contexts with no leaked expected answer or previous trace.
  3. Capture final artifacts, public transcript/tool summaries, duration, token or request cost, and failures.
  4. Grade deterministic assertions first; use a blinded rubric or human review for qualities that cannot be measured mechanically.
  5. Repeat stochastic cases enough to expose variance. Do not hide flakiness by dropping inconvenient runs.

Measure task success, instruction adherence, tool selection and arguments, trajectory efficiency, grounding, safety, output quality, latency, and cost only when relevant. A single aggregate score must not hide a release-blocking metric.

Analyze and iterate

Cluster repeated failures by cause, change one owning layer, rerun the affected cases, then run the regression set. Compare candidate against baseline and reject improvements that regress a protected metric beyond its tolerance.

Use writing-skills for skill-specific authoring and release-engineering for production promotion. Never claim a score that was not read from an actual result artifact.

Completion condition

Cases, environment, baseline, candidate, artifacts, graders, costs, and limits are reproducible; the decision follows predefined thresholds rather than a post-hoc interpretation of the preferred result.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Designs or reviews HTTP, REST, GraphQL, RPC, CLI, webhook, event, and service interfaces with explicit inputs, outputs, errors, compatibility, idempotency, pagination, authentication, versioning, and observability. Use when introducing or changing an API or cross-component contract. Not for internal implementation details with no boundary or for debugging one API failure; use root-cause-debugging there.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

Reviews an existing codebase for structural friction, unclear ownership, leaky or shallow interfaces, excessive coupling, misplaced state, poor testability, and risky dependency direction, then prioritizes evidence-backed improvement candidates. Use for architecture audits, modularization, modernization, or recurring cross-cutting change pain. Not for designing one new interface, simplifying a local function, or fixing a reproduced bug.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

Validates a running application, CLI, API, service, or generated artifact as a user or operator against a prewritten observable behavior contract while remaining source-blind. Use for acceptance checks, runtime proof, anti-fake probes, release smoke tests, or an independent companion to code review. Not for source-quality findings, root-cause diagnosis, or visual design judgment outside the contract.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

Closes a completed development branch by checking the final diff and proof, presenting merge, pull-request, keep, or discard options, and cleaning up only after the user or repository workflow chooses a path. Use when feature work is complete and the branch must be integrated or retired. Not for claiming a feature is complete before verification.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

Tests a real web user flow with a browser by asserting semantic behavior, network and loading states, keyboard access, responsive layouts, and stable visual evidence. Use for browser bugs, end-to-end UI behavior, responsive or accessibility checks, and screenshot baselines. Not for static source review without a browser or for backend-only tests.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

Simplifies recently changed or explicitly scoped code while preserving observable behavior, error semantics, side effects, ordering, and public contracts. Use for readability cleanup, dead-code removal, reducing nesting, eliminating redundant wrappers, or making an implementation easier to maintain. Not for architecture redesign, new behavior, speculative cleanup, or a refactor without an adequate safety net.

日本語の概要は準備中です。原文の説明を表示しています。

thiientv/godmode962026年8月26日 更新

thiientv のスキルをすべて見る

このスキルの問題を報告する