Evaluates factor crowding with positioning analysis, valuation spread monitoring, and unwind risk assessment. Use when analyzing factor crowding, assessing unwind risk, or monitoring factor valuation extremes.
日本語の概要は準備中です。原文の説明を表示しています。
3,141 件 ・ 関連度順
概要と使いどころ
Evaluates factor crowding with positioning analysis, valuation spread monitoring, and unwind risk assessment. Use when analyzing factor crowding, assessing unwind risk, or monitoring factor valuation extremes.
日本語の概要は準備中です。原文の説明を表示しています。
"M&A integration playbook covering eight modules: strategic rationale, target screening, due diligence (financial/legal/commercial), valuation with valuation bridge, synergy analysis, deal structuring (stock vs. asset, cash vs. equity, earn-out), SPA key clauses, and post-merger integration (PMI). Use for deal evaluation, valuation disputes, structure design, earn-out design, synergy breakdown, integration risk, or hostile takeover defense. Triggers: 『併購』『收購』『M&A』『盡職調查』『DD』『估值橋』『綜效』『earn-out』『換股比例』『PMI』『買殼』『借殼上市』『敵意併購』『交易結構』. For Taiwan EMBA 財管組 case studies and term reports (台大/政大/陽交). Complements Asgard `biz-dcf` and `fin-modeling` (valuation tools) plus `biz-corporate-governance` (governance layer) by providing the transaction-layer framework.".
日本語の概要は準備中です。原文の説明を表示しています。
Use when testing, evaluating, or building regression suites for Agentforce agents: conversation testing in Agent Builder, topic (now subagent) coverage and utterance testing, Testing API and AiEvaluationDefinition metadata, Agentforce DX CLI test runs (sf agent generate test-spec, sf agent test create/run/resume/results/list), evaluation metrics (containment rate, escalation rate, CSAT, topic activation accuracy), and post-deploy analytics via Enhanced Event Logs. Triggers: 'how do I test my Agentforce agent', 'agent routes to wrong subagent', 'write utterance tests', 'regression test after topic change', 'measure agent quality', 'agent containment rate', 'run agent tests from the CLI'. NOT for agent creation, topic design, or action contract design — use agentforce/agentforce-agent-creation, agentforce/agent-topic-design, or agentforce/agent-actions respectively.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluate recommender system quality for the CSX4207 Vinyl Record Store. Use whenever you must compute or report recommender metrics (Precision@k, Recall@k, HitRate@k, MRR, MAP@k, NDCG@k, coverage, diversity, novelty, serendipity, personalization), design an evaluation protocol/split, run a baseline comparison, or write the evaluation section of a course deliverable. Covers formulas, JavaScript reference implementations, and the reporting checklist.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluate LLM outputs using frontier models as judges. Use for pairwise model comparison, quality scoring with custom rubrics, and automated evaluation pipelines. Covers position bias mitigation, statistical significance, and generating preference data for DPO/RLHF.
日本語の概要は準備中です。原文の説明を表示しています。
論文や研究計画を共通の評価基準で読み、研究方法、根拠、引用、文章構成を点検するスキル。評価の理由と修正の優先順位、次に直す内容を整理します。
アプリ開発を計画・実装・評価のエージェントに分け、動作中の画面を検証しながら、設定した品質基準に向けて修正を繰り返す開発支援スキル。
工場やオフィスの電力・天然ガス調達を、料金体系、使用量、契約条件から検討するスキル。供給会社の比較、ピーク料金対策、再生可能エネルギー契約や予算作成を支援します。
将来のキャッシュフローを現在価値に換算し、企業価値と1株当たりの評価額を算出するExcelを作成します。前提を変えた再計算やシナリオ比較、感度分析も扱います。
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates an ML GPU cluster for a cloud trial or acceptance test: environment dump, isolated newest PyTorch, matmul FLOPS (MAMF/MSMF) on every GPU while the others compute, intra-node all-reduce bandwidth and per-call latency of every collective, the same inter-node on every node you were given (omit those sections if there is only one node), fio on local disk and shared FS, dated markdown report. Use when the user asks to evaluate a cluster, kick the tires on trial nodes, run cluster acceptance, or measure GPU/network/storage. Canonical copy: https://github.com/stas00/ml-engineering/blob/master/skills/evaluate-cluster/SKILL.md
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used when writing, enhancing, or evaluating the launch prompt for a long-running autonomous agent or a parallel multi-agent orchestration attacking a hard problem: pseudo-formal task briefs that define terms and an exact success predicate linguistically, enumerate non-counting outcomes, set persistence rules with explicit stop and return conditions and effort floors, manage a diverse portfolio of parallel approaches with an approach registry and blocked-route bookkeeping, and gate the return on adversarial audit. Route agent topology and coordination protocols to multi-agent-patterns, runtime control surfaces and loop governance to harness-engineering, evaluator and quality-gate construction to evaluation, judge design to advanced-evaluation, and compaction or memory mechanics to context-compression and memory-systems.
日本語の概要は準備中です。原文の説明を表示しています。
Runs the CARLA Leaderboard evaluator over a set of routes with your agent — routes/routes-subset/repetitions/track selection, checkpoint and live-results endpoints, resume after a crash, recording, and the debug levels — for leaderboard 1.0, 2.0 or 2.1. Preflights the whole stack (version pairing, maps, ports, agent class, sensor budget) before spending hours, and hands the world back in async mode. Use when the user asks to "run the leaderboard", "evaluate my agent", "run routes_training", "resume an evaluation", or "reproduce my submission locally".
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
日本語の概要は準備中です。原文の説明を表示しています。
Make a custom AI agent — Python or TypeScript/JavaScript, on a framework or hand-built — report what it did to Failproof AI, and run your own evaluator worker (the "eval pod") that scores those runs. Reach for it on vague phrasing too: "add observability to my agent", "why isn't my agent showing up?", "run an LLM judge on our own infra". Trigger when the user wants to: • plan an integration — which points in the agent loop to record; • instrument — add `failproofai-sdk` (Python) or `@failproofai/sdk` (Node, Bun, Deno, Next.js): turn on an adapter (LangChain/LangGraph, CrewAI, LlamaIndex, Pydantic AI, Vercel AI SDK, Mastra) or wire a hand-built loop; • verify — confirm events are written, or debug an integration that produces nothing; • evaluate — write, deploy or debug an Evaluator worker in Python or TypeScript. NOT for reading telemetry or scores that already landed (that's `fp-cloud-cli`), or deciding what is worth evaluating (that's `failproofai-eval-brainstorm`).
日本語の概要は準備中です。原文の説明を表示しています。
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
日本語の概要は準備中です。原文の説明を表示しています。
Analyze how a private SaaS company's ARR valuation multiple changed across funding rounds, and attribute the compression or expansion to rate cycles and macro selloffs, growth deceleration, narrative shifts (including an AI premium), competition, and investor demand, benchmarked against private-market medians and peers. Use this skill whenever the user asks about valuation compression, ARR multiples, round-to-round valuation or multiple changes, down rounds, or wants to compare a VC-backed software company's funding rounds. Research the rounds rather than answering from memory.
日本語の概要は準備中です。原文の説明を表示しています。
[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude.
日本語の概要は準備中です。原文の説明を表示しています。
Use this skill when a user wants to create, run, or analyze evaluation suites for Microsoft 365 Copilot declarative agents with the public @microsoft/m365-copilot-eval CLI. Trigger on intents such as "evaluate my agent", "test my agent", "run my evals", "create eval prompts", "add multi-turn tests", "tune evaluator thresholds", "why is my agent failing", or "set up eval environment variables".
日本語の概要は準備中です。原文の説明を表示しています。
A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill before publishing, updating, or deploying. Covers two hard veto gates (structural redlines + research integrity redlines), static quality scoring across 25 criteria (ISO 25010 + OpenSSF + Agent), dynamic test input generation, multi-mode execution testing, multi-layer output evaluation with five specialized category rubrics (Evidence Insight / Protocol Design / Data Analysis / Academic Writing / Other), a Research Veto that applies to all four research categories, human eval viewer generation, actionable P0/P1/P2 optimization recommendations, and automatic skill improvement that outputs a polished, production-ready SKILL.md. Also use whenever a user says "audit my skill", "evaluate my skill", "improve my skill", or wants a corrected version after evaluation.
日本語の概要は準備中です。原文の説明を表示しています。
Configures and runs LLM evaluation using Promptfoo framework. Use when setting up prompt testing, creating evaluation configs (promptfooconfig.yaml), writing Python custom assertions, implementing llm-rubric for LLM-as-judge, or managing few-shot examples in prompts. Triggers on keywords like "promptfoo", "eval", "LLM evaluation", "prompt testing", or "model comparison".
日本語の概要は準備中です。原文の説明を表示しています。