本文へ移動
cccskills

「model quality」の検索結果

316 件 ・ 関連度順

概要と使いどころ

agentic-bench

無料日本語概要

Autonomous model validation and benchmarking. Investigates any ML model (LLM, image gen, TTS, time series, etc.), runs it on GPU cloud, evaluates quality and performance, and generates HTML reports. Use when user asks to verify, benchmark, evaluate, or test a model. Triggers on "verify model", "benchmark", "evaluate model", "test model", "run benchmark", "model evaluation", "モデルを検証", "ベンチマーク", "モデルを試して".

nyosegawa/agentic-bench52026年3月8日 更新

Cross-model benchmark for gstack skills. Runs the same prompt through Claude, GPT (via Codex CLI), and Gemini side-by-side — compares latency, tokens, cost, and optionally quality via LLM judge. Answers "which model is actually best for this skill?" with data instead of vibes. Separate from /benchmark, which measures web page performance. Use when: "benchmark models", "compare models", "which model is best for X", "cross-model comparison", "model shootout". (gstack) Voice triggers (speech-to-text aliases): "compare models", "model shootout", "which model is best".

日本語の概要は準備中です。原文の説明を表示しています。

mostafasudo/appilot32026年9月14日 更新

This skill should be used whenever the user mentions a "semantic model", "data model", or "dataset", or asks to "build", "model", "design", "optimize", "review", or "audit" one, or to "add a measure", "add a relationship", "create a role" / "set up RLS", "add a calculation group", "set up incremental refresh", "fix a star schema", "reduce model size", "prepare a model for Copilot / AI", or "check model quality". Covers the full lifecycle (design, build, refresh, review) and drives every operation through the `te` CLI first, then TOM (connect-pbid) or a model MCP, then TMDL authoring (the tmdl skill). Not for report visuals (use pbir-cli) or isolated DAX query tuning (use the dax skill).

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

aws-ai-ml

無料

Selects, deploys, and customizes AI models on Amazon SageMaker. Training or Processing jobs, fine-tuning (SFT/DPO/RLVR/RLAIF), model selection, dataset preparation, evaluation, SageMaker or Bedrock deployment, inference optimization and endpoint diagnostics. Covers the full lifecycle from planning through production. Use when fine-tuning models on SageMaker, choosing which base model to customize, fine-tune, or deploy from SageMaker JumpStart or Hub, SageMakerPublicHub, or the SageMaker public model catalog, transforming or validating training data, evaluating model quality, deploying or optimizing endpoints, configuring IAM/S3 for training, or managing SageMaker Managed MLflow. Use for endpoint health, failures, latency, logs, metrics, errors. Covers Serverless Model Customization, Nova and OSS deployment paths, and PySDK v3. NOT for Ground Truth labeling, Feature Store, or general-purpose AWS infrastructure.

日本語の概要は準備中です。原文の説明を表示しています。

aws/agent-toolkit-for-aws2,8432026年10月10日 更新

Expert-level air quality engineering covering pollutant sources, atmospheric dispersion, air quality standards, emission controls, and air quality modeling.

日本語の概要は準備中です。原文の説明を表示しています。

luokai0/ai-agent-skills-by-luo-kai122026年5月6日 更新

Role-based multi-model strategy. The model catalog lives in a local configuration file instead of the skill text -- roles (thinker/decision-maker, long-running worker, draft generator, reviewer/acceptance, escalation), reasoning effort as an independent axis, mandatory review for draft models, provider switching for independent reviews, budgets, delegation flow, and cost/quality of small workers.

日本語の概要は準備中です。原文の説明を表示しています。

ellmos-ai/skills72026年10月12日 更新

Use when the user wants to send the same prompt to several LLMs and compare the answers - sequentially (stopwatch, clean per-lane timing) or in parallel (true race), with optional repetitions per model, judged by the starting model across quality, correctness, completeness, instruction fidelity and latency (time is only one dimension). Triggers on /compare-race, 'model race', 'vergleiche modelle', 'gleicher prompt an mehrere modelle'.

日本語の概要は準備中です。原文の説明を表示しています。

ellmos-ai/skills72026年10月12日 更新

mle-workflow

無料日本語概要

機械学習モデルを本番で使うために、データの仕様、再現可能な学習、品質評価、配備、監視、切り戻しを整理し、実装計画やレビュー項目にまとめるスキル。

  • 実験コードを学習パイプラインにしたいとき
  • モデル公開前の品質基準づくり
  • データ漏洩や前処理の不一致の調査
affaan-m/ECC27.7万2026年10月12日 更新

Identify poisoned training data and backdoored ML models across the pipeline using IBM's Adversarial Robustness Toolbox (activation clustering, spectral signatures, trigger reconstruction), Cleanlab for label-quality issues, and supply-chain checks like weight-hash verification and safetensors enforcement. Use before training or deploying on third-party/user-contributed data or downloaded checkpoints, during ML supply-chain reviews, or when investigating model misbehavior tied to specific inputs (suspected backdoor trigger).

日本語の概要は準備中です。原文の説明を表示しています。

mukul975/Anthropic-Cybersecurity-Skills3.4万2026年8月31日 更新

Use when the user wants to migrate code that calls OpenAI, Gemini/Google AI, or the Anthropic API to Amazon Bedrock — a pure model/SDK rewrite. End-to-end: assesses the codebase, rewrites SDK calls, evaluates output quality against Bedrock, and delivers a ready-to-merge git branch. Also has an access-only mode for checking or enabling Bedrock model access for specific models (e.g. 'request access to Claude on Bedrock', 'which models do I turn on') with no code migration — it checks each model and walks the right access step, then stops. Not for agent runtime selection or agent migration planning (use agent-advisor), nor standalone cost estimates or infra-only migration. The full migration REQUIRES gcp-to-aws OR azure-to-aws installed alongside this skill (gcp-to-aws preferred when both are present); access-only mode needs neither.

日本語の概要は準備中です。原文の説明を表示しています。

aws/agent-toolkit-for-aws2,8432026年10月10日 更新

Resolve model profile (quality/balanced/budget) at orchestration start and map agents to specific models. Enables cost/quality tradeoffs by selecting appropriate AI models for each agent role.

日本語の概要は準備中です。原文の説明を表示しています。

a5c-ai/babysitter1,8412026年9月17日 更新

Collect and assess quality metrics for Simulink models using the metric.Engine API. Use when evaluating model design quality (maintainability, complexity, architecture) or testing completeness (unit testing, SIL, PIL, coverage, pass rates) after edits or test runs. Covers metric discovery, collection, report generation, and diagnostics. Triggers on: "collect metrics", "quality metrics", "metric.Engine", "testing metrics", "maintainability metrics", "SIL metrics", "PIL metrics", "what metrics are available", "assess quality", "coverage metrics", "generate metric report", "diagnose metric errors".

日本語の概要は準備中です。原文の説明を表示しています。

matlab/simulink-agentic-toolkit1,2142026年10月8日 更新

Use when comparing Scenario models on the same brief through MCP: a bake-off with shared prompts and inputs, cost, generation time and quality measured side by side, choosing a default model, testing a new release against the incumbent, or finding which model handles a style, edit, or reference image best. Keywords: model comparison, benchmark, bake-off, A/B test, contact sheet, cost per asset, latency.

日本語の概要は準備中です。原文の説明を表示しています。

scenario-labs/skills9642026年10月10日 更新

Run AI agent and LLM evaluations in CI/CD pipelines — automated quality gates that fail the build when AI output quality drops. Use when someone asks to "test my AI agent", "add evals to CI", "catch prompt regressions", "compare models", "evaluate LLM output quality", "set up AI quality gates", or "benchmark my agent before deploying". Covers eval frameworks (Cobalt, Promptfoo, Braintrust), LLM-as-judge scoring, threshold-based assertions, and GitHub Actions integration.

日本語の概要は準備中です。原文の説明を表示しています。

TerminalSkills/skills1632026年10月4日 更新

obliteratus

無料日本語概要

重みが公開されたLLMの回答拒否に関わる仕組みを分析し、再学習なしで重みを変更する実験を支援。計算環境に合うモデルや手法の選択、変更前後の品質評価まで扱います。

  • LLMの回答拒否の仕組みを調べたいとき
  • 再学習せず重み変更を実験したいとき
  • 計算環境に合うモデルと手法を選びたいとき
NousResearch/hermes-agent25.3万2026年10月11日 更新

saelens

無料日本語概要

言語モデル内部の反応を少数の特徴に分解し、学習した概念を調べます。学習済みSAEの分析、自作SAEの訓練、特徴を使った出力調整と品質評価を案内します。

  • 学習済みSAEで特徴を調べたいとき
  • 特定のモデル層にSAEを訓練する
  • 入力文ごとの特徴の反応を比較する
NousResearch/hermes-agent25.3万2026年10月11日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.

日本語の概要は準備中です。原文の説明を表示しています。

foryourhealth111-pixel/Vibe-Skills3,6522026年8月31日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

foryourhealth111-pixel/Vibe-Skills3,6522026年8月31日 更新

Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation. Use when analyzing ChangeNet model failures, investigating poor recall / FAR / PASS-NO_PASS metrics, auditing visual inspection pipeline quality, or running an RCA report for an AOI defect-detection model. Trigger phrases include "RCA on my ChangeNet model", "why is my AOI model failing", "audit ChangeNet predictions", "investigate FAR regressions", "root cause analysis on visual-changenet".

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5602026年10月9日 更新

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

日本語の概要は準備中です。原文の説明を表示しています。

OpenRaiser/NanoResearch1,3402026年10月9日 更新