本文へ移動
cccskills

「eval」の検索結果

4,083 件 ・ 関連度順

概要と使いどころ

deepeval

無料

DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).

日本語の概要は準備中です。原文の説明を表示しています。

confident-ai/deepeval1.9万2026年10月10日 更新

Instrument an AI application with DeepEval's native tracing so its behavior is visible in Confident AI. TRIGGER when the user wants to add DeepEval tracing or @observe to an LLM app, agent, RAG pipeline, or chatbot; wire a framework, model-provider, or vector-database integration (LangGraph, LangChain, OpenAI Agents, LlamaIndex, Pydantic AI, CrewAI, and others); choose between a native integration and manual instrumentation; set span types, tags, or metadata; or send DeepEval-SDK traces to Confident AI's Observatory. DO NOT TRIGGER for building DeepEval pytest eval suites, datasets, goldens, metrics, or deepeval test run (use the `deepeval` skill), or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill). This skill is purely DeepEval-SDK instrumentation — producing well-formed traces, not running evals.

日本語の概要は準備中です。原文の説明を表示しています。

confident-ai/deepeval1.9万2026年10月10日 更新

Design automated evaluation pipelines for LLM and agent systems — combining deterministic checks, statistical metrics, and LLM-as-judge scoring into repeatable, CI-integrated eval suites. Load when the user asks to set up automated evals, design an eval pipeline, integrate evals into CI/CD, create an eval suite, do eval-driven development, or says "automate my evals", "CI eval integration", "evaluation pipeline", "continuous evaluation", "monitoring eval quality", "set up regression testing for my agent". Sub-skill of eval-output orchestrator.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Provides context about the Roo Code evals system structure in this monorepo. Use when tasks mention "evals", "evaluation", "eval runs", "eval exercises", or working with the evals infrastructure. Helps distinguish between the evals execution system (packages/evals, apps/web-evals) and the public website evals display page (apps/web-roo-code/src/app/evals).

日本語の概要は準備中です。原文の説明を表示しています。

foryourhealth111-pixel/Vibe-Skills3,6512026年8月31日 更新

evaluating-llms-harness

無料日本語概要

大規模言語モデルの知識・推論・コード生成を共通の課題で評価するスキル。複数モデルの比較表や学習途中の成績推移を作り、評価結果と個別の回答を保存します。

  • MMLU・GSM8Kでのモデル評価
  • 複数モデルの比較表作成
  • 学習途中のモデルの成績追跡
NousResearch/hermes-agent25.3万2026年10月11日 更新

INVOKE THIS SKILL for LLM-as-judge evaluation workflows on Arize: creating/updating evaluators, running evaluations on spans or experiments, tasks, trigger-run, column mapping, and continuous monitoring. Use when the user says: create an evaluator, LLM judge, hallucination/faithfulness/correctness/relevance, run eval, score my spans or experiment, ax tasks, trigger-run, trigger eval, column mapping, continuous monitoring, query filter for evals, evaluator version, or improve an evaluator prompt.

日本語の概要は準備中です。原文の説明を表示しています。

jcasnellie69/homelab-config22026年10月9日 更新

Export raw OpenTelemetry traces from an AI application to Confident AI's Observatory. TRIGGER when the user wants to send OpenTelemetry or OTLP traces/spans from an LLM app, agent, RAG pipeline, or chatbot to Confident AI; configure the Confident AI OTLP endpoint; set confident.span.* or confident.trace.* attributes; export AI-app traces without the deepeval Python package; wire an OTLPSpanExporter, OpenTelemetry Collector, or vendor-neutral OTel SDK to Confident AI; or pick the US vs EU OTLP endpoint. Language-agnostic: the mechanism is OTLP attribute keys plus an exporter endpoint. DO NOT TRIGGER for building DeepEval pytest eval suites, datasets, goldens, metrics, or deepeval test run (use the `deepeval` skill); for instrumenting with the DeepEval SDK's @observe decorator or framework integrations (use the `deepeval-tracing` skill); or for non-AI software such as web servers, CRUD backends, or infrastructure: the confident.* attributes describe AI components only.

日本語の概要は準備中です。原文の説明を表示しています。

confident-ai/deepeval1.9万2026年10月10日 更新

This skill should be used when the user wants to "run an evaluation", "evaluate my agent", "evaluate my ADK agent", "write an eval dataset", "analyze eval failures", "compare eval results", "optimize agent", or needs guidance on the Agent Platform eval methodology and the Quality Flywheel. Covers eval metrics, dataset schema, LLM-as-judge scoring, and common failure causes. Applies to any agents-cli project, whatever framework the agent is written in. Do NOT use for agent API code patterns (ADK: use google-agents-cli-adk-code), deployment (use google-agents-cli-deploy), or project scaffolding (use google-agents-cli-scaffold).

日本語の概要は準備中です。原文の説明を表示しています。

google/agents-cli6,0732026年10月6日 更新

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-discovery, write-judge-prompt, or validate-evaluator instead).

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generation system, measuring retrieval quality, assessing generation faithfulness or relevance, generating synthetic QA pairs for retrieval testing, or optimizing chunking strategies.

日本語の概要は準備中です。原文の説明を表示しています。

ai-evals-course/evals-skills1,4822026年9月25日 更新

Guide for diagnosing and improving MSBuild project evaluation performance. Only activate in MSBuild/.NET build context. USE FOR: builds slow before any compilation starts, high evaluation time in binlog analysis, expensive glob patterns walking large directories (node_modules, .git, bin/obj), deep import chains (>20 levels), preprocessed output >10K lines indicating heavy evaluation, property functions with file I/O ($([System.IO.File]::ReadAllText(...))), multiple evaluations per project. Covers the 5 MSBuild evaluation phases, glob optimization via DefaultItemExcludes, import chain analysis with /pp preprocessing. DO NOT USE FOR: compilation-time slowness (use build-perf-diagnostics), incremental build issues (use incremental-build), non-MSBuild build systems. INVOKES: binlog MCP server tools (evaluations, evaluation_global_properties, evaluation_properties, imports, properties); falls back to dotnet msbuild -pp:full.xml for preprocessing, /clp:PerformanceSummary.

日本語の概要は準備中です。原文の説明を表示しています。

bouclem/skills62026年5月31日 更新

Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "evaluation rubric for", "how do I measure quality of". Sub-skill of eval-output orchestrator.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1012026年10月10日 更新

Suede AI eval design and coverage audit: AI-SPEC, failure-mode rubric with severity scoring, concrete pass/fail eval cases, coverage and infrastructure scores, and mechanical acceptance gates. Use when a change ships LLM, RAG, agent, classifier, prompt, or generated-media behavior, or when asked to write evals for an AI feature, design test cases for a model surface, audit existing eval coverage, or judge whether AI behavior is safe to ship. No AI-SPEC means no eval plan, and no eval plan holds the recommended ship verdict. NOT FOR: reviewing or grading the implementation code behind the AI surface (use suede-code); wiring a passing suite into CI as a required check (use suede-ci-gate); UAT of the built feature beyond the eval suite (a private Suede Labs companion, not in this pack).

日本語の概要は準備中です。原文の説明を表示しています。

JasonColapietro/suede-creator-skills1272026年10月10日 更新

agentic-bench

無料日本語概要

Autonomous model validation and benchmarking. Investigates any ML model (LLM, image gen, TTS, time series, etc.), runs it on GPU cloud, evaluates quality and performance, and generates HTML reports. Use when user asks to verify, benchmark, evaluate, or test a model. Triggers on "verify model", "benchmark", "evaluate model", "test model", "run benchmark", "model evaluation", "モデルを検証", "ベンチマーク", "モデルを試して".

nyosegawa/agentic-bench52026年3月8日 更新

Orchestrator for the eval-output skill suite — evaluate LLM and agent outputs for quality, accuracy, helpfulness, and safety using structured rubrics and LLM-as-judge techniques. Load when the user says "evaluate this output", "score this response", "run an eval", "LLM as judge", "evaluate agent output", "how good is this response", "rate this answer", "eval this", or provides an LLM output that should be assessed for quality. Single entry point for all output evaluation workflows.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.

日本語の概要は準備中です。原文の説明を表示しています。

github/awesome-copilot4万2026年10月9日 更新

Capture and review Agent Skill observations, and create, run, diagnose, or iteratively improve Skill evaluations (evals) with the skill-up CLI. Use when the user asks to record explicitly attributed Skill usage or feedback; review observations or user-supplied traces/transcripts; turn an approved observation into a regression case; evaluate, test, regress, verify, fix, improve, iterate, or evolve a Skill; add or strengthen eval cases; write eval.yaml/case.yaml; run skill-up run/validate/list-cases/report/import/init; or migrate from Anthropic evals.json. Plugin observation capture/review require either the Codex skill-up plugin or an observer-enabled DSH skill-up plugin. User-supplied external evidence can be reviewed without a plugin; evaluation remains multi-engine.

日本語の概要は準備中です。原文の説明を表示しています。

alibaba/skill-up1,1742026年10月10日 更新

rag-eval

無料

Use when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues. Also use when the user mentions RAG evaluation, faithfulness checking, hallucination detection in RAG, retrieval quality, chunking optimization, or "is my RAG pipeline working." Outputs a diagnostic matrix that pinpoints whether problems are in retrieval or generation.

日本語の概要は準備中です。原文の説明を表示しています。

agentscope-ai/OpenJudge8712026年9月11日 更新

Use when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a labeled evaluation set. Also use when the user mentions test data design, eval coverage, difficulty stratification, synthetic data generation for eval, or "how to create good evaluation data." Outputs datasets in OpenJudge-compatible format.

日本語の概要は準備中です。原文の説明を表示しています。

agentscope-ai/OpenJudge8712026年9月11日 更新

Advanced Evaluation workflow skill. Use this skill when the user needs This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.

日本語の概要は準備中です。原文の説明を表示しています。

diegosouzapw/awesome-omni-skills1592026年7月8日 更新

Design eval harnesses — task schemas, metrics, dataset versioning, eval-as-code patterns. Use when asked to "build an eval harness", "set up eval-as-code", or "version our eval datasets".

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

Author and run offline evals for Agentforce agents: fixture format, scoring rubrics, hallucination checks, regression baselines, CI integration, prompt-change safety. Trigger keywords: agent eval, hallucination, doesn't hallucinate, agentforce offline evals, rubric for agent response. NOT for deciding which layers of testing the agent needs — use agentforce/agentforce-testing-strategy. NOT for Testing Center, AiEvaluationDefinition or sf agent test runs — use agentforce/agent-testing-and-evaluation.

日本語の概要は準備中です。原文の説明を表示しています。

PranavNagrecha/AwesomeSalesforceSkills192026年10月4日 更新

Run evaluations for Hugging Face Hub models using inspect-ai and lighteval on local hardware. Use for backend selection, local GPU evals, and choosing between vLLM / Transformers / accelerate. Not for HF Jobs orchestration, model-card PRs, .eval_results publication, or community-evals automation.

日本語の概要は準備中です。原文の説明を表示しています。

bg-szy/TOP-SKILLS62026年9月8日 更新