Install Triton + SageAttention to accelerate ComfyUI (the sageattn attention_mode and inductor torch.compile used by WanVideoWrapper / many video graphs). Windows-first (triton-windows + woct0rdho prebuilt SageAttention wheels matched to torch/CUDA/python into the RIGHT python), plus Linux (official triton + build) and Mac (N/A → sdpa/MPS). Also covers the SAFE sdpa / no-compile fallback so an example that assumes sageattn + torch.compile still runs when these aren't installed (video-extend TRAP 5). Use when a loader crashes with "No module named 'sageattention'" or reports triton unavailable, when asked to speed up Wan/video workflows, or when deciding whether to install acceleration vs. fall back.
日本語の概要は準備中です。原文の説明を表示しています。
artokun/comfyui-mcp☆ 8072026年10月5日 更新
Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the user does not name it: verify operator step sequence, detect missing or out-of-order SOP steps, score factory/work-cell video for procedure compliance, run VLM-based SOP checking on industrial cameras, or call /v1/chat/completions with a file, RTSP, or Basler camera. Also trigger for its internals: SOPVideoProcessor, DeepStream GEBD model (e.g. DDM) via Triton CAPI, nvds_custom_postprocess, Cosmos Reason 1/2 vLLM, SSE streaming, Kafka NvProto/JSON output, Basler/Pylon camera + emulation, Docker compose, chunk-level latency. Do NOT trigger for generic DeepStream pipelines, object detection/tracking, NIM imports, or video summarization.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
算子级开发与性能优化:Triton / Ascend C / Catlass / PyPTO / TileLang 算子 编写、精度对齐与性能调优。优先路由到外部 cannbot-skills 技能库(不重复其内容), 本 skill 只保留场景 → skill 映射与 MindIE-SD 特有补充。 当用户需要新增/优化算子、定位算子性能或精度问题时使用此 skill。 即使用户只提到"写个 triton kernel"或"这个算子怎么加速"而未说算子,也应触发; pattern 融合/编译后端见 pattern-dev,算子基准选型/接入测试见 benchmark-dev。 由 dev-workflow 或 pattern-dev 的 replacement kernel 场景触发。
日本語の概要は準備中です。原文の説明を表示しています。
Ascend/MindIE-SD☆ 152026年10月10日 更新
Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. Includes canary deployments, autoscaling, model versioning, A/B testing, and GPU resource management for production model serving.
日本語の概要は準備中です。原文の説明を表示しています。
BagelHole/DevOps-Security-Agent-Skills☆ 1,1542026年5月22日 更新
Provides guidance for writing, optimizing, and benchmarking Triton kernels for Intel XPU GPUs (Battlemage/Arc Pro B50) using the Xe-Forge optimization framework. Includes an LLM-driven trial-loop workflow (analyze, validate, benchmark, profile, finalize), XPU-specific patterns (tensor descriptors, GRF mode, tile swizzling), KernelBench fused kernels, and Flash Attention.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
Provides guidance for writing and benchmarking optimized Triton kernels for AMD GPUs (MI355X, R9700) on ROCm, targeting HuggingFace diffusers (LTX-Video, SD3, FLUX) and transformers. Core kernels: RMSNorm, RoPE 3D, GEGLU, AdaLN. Includes XCD swizzle, autotune, diffusers integration patterns, and LTX-Video pipeline injection.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
LLM and ML model deployment for inference. Use when serving models in production, building AI APIs, or optimizing inference. Covers vLLM (LLM serving), TensorRT-LLM (GPU optimization), Ollama (local), BentoML (ML deployment), Triton (multi-model), LangChain (orchestration), LlamaIndex (RAG), and streaming patterns.
日本語の概要は準備中です。原文の説明を表示しています。
ancoleman/ai-design-components☆ 5252025年12月11日 更新
Pre-execution Solana transaction streaming via Jito ShredStream, Shyft RabbitStream, and Triton Deshred
日本語の概要は準備中です。原文の説明を表示しています。
agiprolabs/claude-trading-skills☆ 4102026年9月3日 更新
Think and work like an expert MLOps Engineer. Use when a task calls for MLOps Engineer judgment. Reasons from data contracts, feature parity, evaluation gates, and rollback-readiness through MLflow/W&B registries, Feast feature stores, KServe/Triton serving, Great Expectations/TFDV validation, and Evidently PSI/KS drift monitors while treating train-serve skew, data leakage, silent degradation, and schema/concept drift as first-class failure modes.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agents☆ 1992026年10月3日 更新
Compile, profile, diagnose, optimize, and compare Ascend NPU operators across CANN versions and A2/A3/A5 for Ascend C, CATLASS, Triton-Ascend, TileLang-Ascend, PyPTO, and SHMEM/MC2. Use when an operator already has a runnable implementation and the user asks for msOpProf collection, bottleneck analysis, source-level tuning, or a reproducible before/after performance report. Do not use as the primary workflow for operator creation, migration, or unresolved correctness failures.
日本語の概要は準備中です。原文の説明を表示しています。
kali20gakki/msAgent☆ 322026年10月10日 更新