本文へ移動
cccskills

「inference」の検索結果

654 件 ・ 関連度順

概要と使いどころ

Lists all inference providers offered during NemoClaw onboarding. Use when explaining which providers are available, what the onboard wizard presents, or how inference routing works. Trigger keywords - nemoclaw inference options, nemoclaw onboarding providers, nemoclaw inference routing, switch nemoclaw inference model, change inference runtime, nemoclaw local inference, ollama nemoclaw, vllm nemoclaw, local model server, openai compatible endpoint.

日本語の概要は準備中です。原文の説明を表示しています。

composio-community/nemoclaw-composio22026年4月30日 更新

Debug inference clients that use an attached provider and its native endpoint, including hosted APIs and host-local Ollama, vLLM, SGLang, TRT-LLM, LM Studio, or NIM. Use for provider attachment, endpoint policy, credential substitution, topology, and migration from the removed managed inference endpoint. Trigger keywords - debug inference, managed inference endpoint, local inference, ollama, lm studio, vllm, sglang, trtllm, NIM, inference failing, model server unreachable, credential_endpoint_mismatch, host.openshell.internal.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/OpenShell1.6万2026年10月10日 更新

Detect MITRE ATLAS AML.T0024 attacks (model stealing, inversion, membership inference) performed via inference-API abuse, by monitoring per-principal query volume/distribution, rate-limiting and perturbing outputs, and red-teaming your model's extractability. Use for a public or partner inference API needing cloning/inversion/membership-inference detection, or a pre-deployment red-team exercise to measure extraction risk.

日本語の概要は準備中です。原文の説明を表示しています。

mukul975/Anthropic-Cybersecurity-Skills3.4万2026年8月31日 更新

Guides the migration of existing AI workloads (Cloud Run, Gemini API, Gemini Enterprise Agent Platform) to self-hosted GKE inference using gcloud and kubectl. Use when the user has an existing AI inference workload (on Cloud Run, the Gemini API, Gemini Enterprise Agent Platform, or a custom VM) and wants to move it to self-hosted inference on GKE, or asks follow-up questions during such a migration (hardware sizing, model staging, manifest generation, validation, traffic cutover). DO NOT use for brand new GKE inference deployments with no existing workload to migrate (use gke-inference instead). DO NOT use if the user intends to automate the migration via the Gemini Cloud Assist MCP server.

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Conducts privacy auditing of AI models including training data extraction testing, membership inference attacks, model inversion testing, and attribute inference assessment. Uses ML Privacy Meter and related tools to quantify privacy leakage. Keywords: model audit, membership inference, privacy meter, model inversion, training data extraction.

日本語の概要は準備中です。原文の説明を表示しています。

mukul975/Privacy-Data-Protection-Skills3022026年3月17日 更新

Managing privacy risks from AI-driven inferences about individuals including derived data classification, profiling under GDPR Art. 22, inference accuracy obligations, and controlling automated personality/behaviour predictions. Keywords: AI inference, derived data, profiling, automated predictions, GDPR.

日本語の概要は準備中です。原文の説明を表示しています。

mukul975/Privacy-Data-Protection-Skills3022026年3月17日 更新

Guides the migration of existing AI workloads (Cloud Run, Gemini API, Gemini Enterprise Agent Platform) to self-hosted GKE inference using gcloud and kubectl. Use when the user has an existing AI inference workload (on Cloud Run, the Gemini API, Gemini Enterprise Agent Platform, or a custom VM) and wants to move it to self-hosted inference on GKE, or asks follow-up questions during such a migration (hardware sizing, model staging, manifest generation, validation, traffic cutover). DO NOT use for brand new GKE inference deployments with no existing workload to migrate (use gke-inference instead). DO NOT use if the user intends to automate the migration via the Gemini Cloud Assist MCP server.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

This skill should be used when the user asks to "implement a DiD regression", "run a staggered difference-in-differences", "set up an event study", "implement IV / 2SLS", "run a regression discontinuity design", "build synthetic control", "do propensity score matching", "test parallel trends", "run Honest DiD", "wild cluster bootstrap", "Callaway-Sant'Anna", "Sun-Abraham", "Bacon decomposition", "double machine learning", "debiased machine learning", "DML / DoubleML", "cross-fitting", "causal forest", "heterogeneous treatment effects", "ATE with ML controls", "which causal design fits my data", or needs causal-inference code, diagnostics, or identification writing in Python (StatsPAI), R, or Stata. Based on Scott Cunningham's Causal Inference: The Mixtape, with runnable validated templates, ML-based causal inference patterns, and Journal of Finance applications.

日本語の概要は準備中です。原文の説明を表示しています。

Jill0099/causal-inference-mixtape172026年8月12日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for migrating existing AI workloads to GKE (use google-cloud-solution-guided-gke-ai-migration), GKE RAG with Cloud SQL/AlloyDB (use google-cloud-solution-rag-enterprise-search-gke-sqldb), or batch/HPC (use gke-batch-hpc).

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

This skill should be used when the user asks to "implement a DiD regression", "write a causal inference pipeline", "set up an event study", "implement instrumental variables", "run a regression discontinuity design", "build a synthetic control model", "implement propensity score matching", "write parallel trends test", "implement Bacon decomposition", or needs code templates for causal inference methods in Python, R, or Stata. Based on Scott Cunningham's Causal Inference: The Mixtape.

日本語の概要は準備中です。原文の説明を表示しています。

brycewang-stanford/Auto-Empirical-Research-Skills4,5732026年10月5日 更新

predictive-modeling-diagnostics

無料日本語概要

予測モデリング(教師あり機械学習 / 予測モデル / 精度評価 / 交差検証 / リーク、決定木 / ランダムフォレスト / GBDT / LightGBM / XGBoost / CatBoost、モデル解釈 / 特徴量重要度 / SHAP / PDP / ICE、時系列 / 季節性 / ARIMA / SARIMA / Prophet / VAR / 変化点 / 時系列予測)を実行したら必ずセットで出す図と値のルーター。train_test_split, cross_val_score, GridSearchCV, Pipeline, TimeSeriesSplit, DecisionTree, RandomForest, lgb., xgb., CatBoost, dtreeviz, shap., permutation_importance, PartialDependenceDisplay, statsmodels.tsa, SARIMAX, STL, adfuller, plot_acf, Prophet, ruptures がコードに現れたとき、または ユーザーが「予測モデルを作って」「精度を出して」「どの変数が効いているか」「来月を予測して」「季節性を見て」と 言ったときに使う。診断・リーク・ベースライン・残差に言及がなくても適用する。SKILL.md のルーティング表で手法を特定し、 対応する references/<手法>.md を読んでから実行する。回帰係数の推論・検定は statistical-inference-diagnostics、 効果推定は causal-inference-diagnostics を使う。

atsushi-green/ds-ai-coding-skills892026年10月4日 更新

Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation, audio transcription, translation, summarization, and Inference Endpoints

日本語の概要は準備中です。原文の説明を表示しています。

agents-inc/skills242026年9月8日 更新

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for generic batch jobs or HPC task queues (use gke-batch-hpc instead).

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Detect model stealing, model inversion, and membership inference performed through inference-API abuse by monitoring query patterns, applying output perturbation, and red-teaming your own model's extractability.

日本語の概要は準備中です。原文の説明を表示しています。

andycungkrinx91/konoha92026年10月9日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Run 150+ AI apps via inference.sh CLI (infsh) — image generation, video creation, LLMs, search, 3D, social automation. Uses the terminal tool. Triggers: inference.sh, infsh, ai apps, flux, veo, image generation, video generation, seedream, seedance, tavily

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

ito-inference

無料日本語概要

Itôで予約済みのGPUを使ったモデル配信の相談に対し、現状の未対応範囲と将来必要な確認事項を整理するスキル。現在は配信を開始せず、制約を案内します。

  • Itô GPU予約後の配信対応を確認したいとき
  • OpenAI互換の接続先の制約を知りたいとき
  • 将来のモデル配信設定を確認したいとき
affaan-m/ECC27.7万2026年10月10日 更新

flash-attention

無料日本語概要

長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。

  • 長い入力での学習を高速化したいとき
  • 長文推論のメモリ使用量を抑えたいとき
  • 処理速度とモデル出力の比較
NousResearch/hermes-agent25.3万2026年10月11日 更新

inference-sh-cli

無料日本語概要

画像・動画の生成やAI検索など、150以上のクラウドAIアプリを共通のコマンドで探して実行します。ローカルの画像や音声を使う加工・生成にも対応します。

  • 文章から画像を生成したいとき
  • 手元の画像を動画にしたいとき
  • 写真と音声でアバター動画を作りたいとき
NousResearch/hermes-agent25.3万2026年10月11日 更新