本文へ移動
cccskills
無料GitHub で公開

phoenix-evals

Build and run evaluators for AI/LLM applications using Phoenix.

インストール方法を見る

含まれるファイル(35)

  • SKILL.md4.5 KB
  • references/axial-coding.md2.3 KB
  • references/common-mistakes-python.md7.0 KB
  • references/error-analysis-multi-turn.md1.3 KB
  • references/error-analysis.md4.2 KB
  • references/evaluate-dataframe-python.md4.4 KB
  • references/evaluators-code-python.md3.0 KB
  • references/evaluators-code-typescript.md1.2 KB
  • references/evaluators-custom-templates.md1.2 KB
  • references/evaluators-llm-python.md2.6 KB
  • references/evaluators-llm-typescript.md1.4 KB
  • references/evaluators-overview.md1.1 KB
  • references/evaluators-pre-built.md2.4 KB
  • references/evaluators-rag.md3.1 KB
  • references/experiments-datasets-python.md4.8 KB
  • references/experiments-datasets-typescript.md2.7 KB
  • references/experiments-overview.md1.7 KB
  • references/experiments-running-python.md2.9 KB
  • references/experiments-running-typescript.md3.1 KB
  • references/experiments-synthetic-python.md1.8 KB
  • references/experiments-synthetic-typescript.md2.2 KB
  • references/fundamentals-anti-patterns.md1.6 KB
  • references/fundamentals-model-selection.md1.4 KB
  • references/fundamentals.md1.8 KB
  • references/observe-sampling-python.md2.4 KB
  • references/observe-sampling-typescript.md3.9 KB
  • references/observe-tracing-setup.md3.3 KB
  • references/production-continuous.md3.6 KB
  • references/production-guardrails.md1.3 KB
  • references/production-overview.md2.3 KB
  • references/setup-python.md1.9 KB
  • references/setup-typescript.md981 B
  • references/validation-evaluators-python.md1.0 KB
  • references/validation-evaluators-typescript.md4.4 KB
  • references/validation.md1.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

TaskFiles
Setupsetup-python, setup-typescript
Decide what to evaluateevaluators-overview
Choose a judge modelfundamentals-model-selection
Use pre-built evaluatorsevaluators-pre-built
Build code evaluatorevaluators-code-python, evaluators-code-typescript
Build LLM evaluatorevaluators-llm-python, evaluators-llm-typescript, evaluators-custom-templates
Batch evaluate DataFrameevaluate-dataframe-python
Understand experimentsexperiments-overview
Run experimentexperiments-running-python, experiments-running-typescript
Create datasetexperiments-datasets-python, experiments-datasets-typescript
Generate synthetic dataexperiments-synthetic-python, experiments-synthetic-typescript
Validate evaluator accuracyvalidation, validation-evaluators-python, validation-evaluators-typescript
Sample traces for reviewobserve-sampling-python, observe-sampling-typescript
Analyze errorserror-analysis, error-analysis-multi-turn, axial-coding
RAG evalsevaluators-rag
Avoid common mistakescommon-mistakes-python, fundamentals-anti-patterns
Productionproduction-overview, production-guardrails, production-continuous

Workflows

Starting Fresh: observe-tracing-setup → error-analysis → axial-coding → evaluators-overview

Building Evaluator: fundamentals → common-mistakes-python → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Production: production-overview → production-guardrails → production-continuous

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill when the user explicitly asks to map, document, or onboard into an existing codebase. Trigger for prompts like "map this codebase", "document this architecture", "onboard me to this repo", or "create codebase docs". Do not trigger for routine feature implementation, bug fixes, or narrow code edits unless the user asks for repository-level discovery.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Run the AgentRC readiness assessment on the current repository and produce a static HTML dashboard at reports/index.html. Wraps `npx github:microsoft/agentrc readiness` and hands off rendering to the @ai-readiness-reporter custom agent. Supports policies (--policy) for org-specific scoring. Use when asked to assess, audit, or score the AI readiness of a repo.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Generate tailored AI agent instruction files via AgentRC instructions command. Produces .github/copilot-instructions.md (default, recommended for Copilot in VS Code) plus optional per-area .instructions.md files with applyTo globs for monorepos. Use after running /acreadiness-assess to close gaps in the AI Tooling pillar.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Help the user pick, write, or apply an AgentRC policy. Policies customise readiness scoring by disabling irrelevant checks, overriding impact/level, setting pass-rate thresholds, or chaining org baselines with team overrides. Use when the user asks about strict mode, AI-only scoring, custom weights, CI gating, or wants org-wide standardisation.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Use this skill when the user shares ad campaign performance data and asks what to cut, scale, or test. Trigger for prompts like "analyze my ad campaigns", "where am I wasting ad spend", "reallocate my ad budget", "which ads are actually working", or "ROAS analysis". Do not trigger for campaign planning or creative generation without performance data.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Add educational comments to the file specified, or prompt asking for file to comment if one is not provided.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

github のスキルをすべて見る

このスキルの問題を報告する