Guide for diagnosing and improving MSBuild project evaluation performance. Only activate in MSBuild/.NET build context. USE FOR: builds slow before any compilation starts, high evaluation time in binlog analysis, expensive glob patterns walking large directories (node_modules, .git, bin/obj), deep import chains (>20 levels), preprocessed output >10K lines indicating heavy evaluation, property functions with file I/O ($([System.IO.File]::ReadAllText(...))), multiple evaluations per project. Covers the 5 MSBuild evaluation phases, glob optimization via DefaultItemExcludes, import chain analysis with /pp preprocessing. DO NOT USE FOR: compilation-time slowness (use build-perf-diagnostics), incremental build issues (use incremental-build), non-MSBuild build systems. INVOKES: binlog MCP server tools (evaluations, evaluation_global_properties, evaluation_properties, imports, properties); falls back to dotnet msbuild -pp:full.xml for preprocessing, /clp:PerformanceSummary.
日本語の概要は準備中です。原文の説明を表示しています。
bouclem/skills☆ 62026年5月31日 更新
Advanced Evaluation workflow skill. Use this skill when the user needs This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
日本語の概要は準備中です。原文の説明を表示しています。
diegosouzapw/awesome-omni-skills☆ 1592026年7月8日 更新
Advanced Evaluation workflow skill. Use this skill when the user needs This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
日本語の概要は準備中です。原文の説明を表示しています。
diegosouzapw/awesome-omni-skills☆ 1592026年7月8日 更新
Guide for diagnosing and improving MSBuild project evaluation performance. USE FOR: builds slow before any compilation starts, high evaluation time in binlog analysis, expensive glob patterns walking large directories (node_modules, .git, bin/obj), deep import chains (>20 levels), preprocessed output >10K lines indicating heavy evaluation, property functions with file I/O ($([System.IO.File]::ReadAllText(...))), multiple evaluations per project. Covers the 5 MSBuild evaluation phases, glob optimization via DefaultItemExcludes, import chain analysis with /pp preprocessing. DO NOT USE FOR: compilation-time slowness (use build-perf-diagnostics), incremental build issues (use incremental-build), non-MSBuild build systems.
日本語の概要は準備中です。原文の説明を表示しています。
dotnet/skills☆ 5,5982026年10月11日 更新
Guide for diagnosing and improving MSBuild project evaluation performance. USE FOR: builds slow before any compilation starts, high evaluation time in binlog analysis, expensive glob patterns walking large directories (node_modules, .git, bin/obj), deep import chains (>20 levels), preprocessed output >10K lines indicating heavy evaluation, property functions with file I/O ($([System.IO.File]::ReadAllText(...))), multiple evaluations per project. Covers the 5 MSBuild evaluation phases, glob optimization via DefaultItemExcludes, import chain analysis with /pp preprocessing. DO NOT USE FOR: compilation-time slowness (use build-perf-diagnostics), incremental build issues (use incremental-build), non-MSBuild build systems.
日本語の概要は準備中です。原文の説明を表示しています。
managedcode/dotnet-skills☆ 4852026年10月10日 更新
Investigate AI observability evaluations of both types — `hog` (deterministic code-based) and `llm_judge` (LLM-prompt-based). Find existing evaluations, inspect their configuration, run them against specific generations, query individual pass/fail results, and generate AI-powered summaries of patterns across many runs. Use when the user asks to debug why an evaluation is failing, surface common failure modes, compare results across filters, dry-run a Hog evaluator, prototype a new LLM-judge prompt, or manage the evaluation lifecycle (create, update, enable/disable, delete).
日本語の概要は準備中です。原文の説明を表示しています。
0xAidan/polymarket-bot-test☆ 42026年9月4日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
sickn33/agentic-awesome-skills☆ 4.7万2026年10月10日 更新
[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude.
日本語の概要は準備中です。原文の説明を表示しています。
rlaope/oh-my-hermes☆ 3,2542026年10月10日 更新
Configures and runs LLM evaluation using Promptfoo framework. Use when setting up prompt testing, creating evaluation configs (promptfooconfig.yaml), writing Python custom assertions, implementing llm-rubric for LLM-as-judge, or managing few-shot examples in prompts. Triggers on keywords like "promptfoo", "eval", "LLM evaluation", "prompt testing", or "model comparison".
日本語の概要は準備中です。原文の説明を表示しています。
daymade/claude-code-skills☆ 1,4512026年10月10日 更新
Use when the user has evaluation principles or a dataset but needs help choosing the right graders, designing evaluation metrics, creating LLM-as-judge prompts, combining multiple metrics into a composite score, or building an automated evaluation pipeline. Also use when the user mentions grader selection, metric design, judge prompt engineering, rubric design, evaluation pipeline code, or "how to evaluate [X] automatically." Outputs executable OpenJudge pipeline code.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Use when the user has run multiple evaluation skills and wants a comprehensive analysis — maturity assessment, cross-skill signals, trends, prioritized actions, and an executive summary. Also use when the user mentions eval health check, evaluation audit, ship readiness, evaluation maturity, or "how good is my evaluation system itself." This is a read-only analysis skill.
日本語の概要は準備中です。原文の説明を表示しています。
agentscope-ai/OpenJudge☆ 8712026年9月11日 更新
Create and check long-running video material evaluation tasks. Use this skill when the user wants to submit videos for evaluation, check an existing video evaluation task list, or fetch the result of a previously created video evaluation task. This skill is not a general-purpose video upload skill because upload is allowed only as an internal step of task creation. Authentication uses an API key passed as an Authorization bearer token.
日本語の概要は準備中です。原文の説明を表示しています。
bytedance/agentkit-samples☆ 4702026年10月9日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
lingxling/awesome-skills-cn☆ 3032026年10月7日 更新
Agent Evaluation workflow skill. Use this skill when the user needs Testing and benchmarking LLM agents including behavioral testing, and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
日本語の概要は準備中です。原文の説明を表示しています。
diegosouzapw/awesome-omni-skills☆ 1592026年7月8日 更新
Agent Evaluation workflow skill. Use this skill when the user needs Testing and benchmarking LLM agents including behavioral testing, and the operator should preserve the upstream workflow, copied support files, and provenance before merging or handing off.
日本語の概要は準備中です。原文の説明を表示しています。
diegosouzapw/awesome-omni-skills☆ 1592026年7月8日 更新
Foundry hosted agent を Agent 365 Autopilot として発行し、自分のメールアドレス・予定表・権限で働く『デジタルな同僚』を Teams / Microsoft 365 Copilot に公開する。初回 scaffold で AI チームメイト評価Hub(Code Apps)、通常会話を採点する Python EvaluationWorker、回帰テストを同時生成し、publish 時に Foundry Monitor/Evaluations の日次評価も自動構成する。複数 Autopilot は共通 Dataverse Hub を agentkey で安全に共有する。メール応対 / Dataverse 検索 / Web 検索 / 定期実行 / コード実行 / 成果物共有 / Teams プレゼンスの機能ブロックに対応し、自己ホスト経路も互換用途として残す。CI/CD・レビューゲートは alm スキルに委譲する。
geekfujiwara/CodeAppsDevelopmentStandard☆ 722026年10月9日 更新
Use when testing, evaluating, or building regression suites for Agentforce agents: conversation testing in Agent Builder, topic (now subagent) coverage and utterance testing, Testing API and AiEvaluationDefinition metadata, Agentforce DX CLI test runs (sf agent generate test-spec, sf agent test create/run/resume/results/list), evaluation metrics (containment rate, escalation rate, CSAT, topic activation accuracy), and post-deploy analytics via Enhanced Event Logs. Triggers: 'how do I test my Agentforce agent', 'agent routes to wrong subagent', 'write utterance tests', 'regression test after topic change', 'measure agent quality', 'agent containment rate', 'run agent tests from the CLI'. NOT for agent creation, topic design, or action contract design — use agentforce/agentforce-agent-creation, agentforce/agent-topic-design, or agentforce/agent-actions respectively.
日本語の概要は準備中です。原文の説明を表示しています。
PranavNagrecha/AwesomeSalesforceSkills☆ 192026年10月4日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
sinhoneyy/master-skills☆ 142026年9月5日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
JantonioFC/skillsbank☆ 92026年8月4日 更新
採用・人材戦略の専門スキル。職務記述書(JD)作成、採用計画策定、コンピテンシーモデル設計、 面接評価基準設計、オンボーディング計画をサポート。採用要件の明確化から内定後のフォローまで 一貫した人材獲得プロセスを支援。日英両言語のテンプレートを提供し、グローバル採用にも対応。 Use when: creating job descriptions, designing competency models, planning recruitment strategies, developing interview evaluation criteria, or creating onboarding plans. Triggers: "JD作成", "採用計画", "面接評価", "オンボーディング", "コンピテンシー", "人材獲得", "job description", "recruitment plan", "interview evaluation", "hiring", "talent acquisition"
takusaotome/claude-skills-library☆ 92026年10月5日 更新
Evaluate recommender system quality for the CSX4207 Vinyl Record Store. Use whenever you must compute or report recommender metrics (Precision@k, Recall@k, HitRate@k, MRR, MAP@k, NDCG@k, coverage, diversity, novelty, serendipity, personalization), design an evaluation protocol/split, run a baseline comparison, or write the evaluation section of a course deliverable. Covers formulas, JavaScript reference implementations, and the reporting checklist.
日本語の概要は準備中です。原文の説明を表示しています。
bg-szy/TOP-SKILLS☆ 62026年9月8日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
bouclem/skills☆ 62026年5月31日 更新
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
lucaspmarie-a11y/claude-skills-vault☆ 52026年6月7日 更新