Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. Use for edge deployment, M1/M2/M3 Macs, AMD/Intel GPUs, or when CUDA is unavailable. Supports GGUF quantization (1.5-8 bit) for reduced memory and 4-10× speedup vs PyTorch on CPU.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
huang-sh/DeepScience☆ 42026年7月15日 更新
Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs. Use when deploying or serving AI/ML models, running GPU-accelerated workloads (training, fine-tuning, inference), serving web endpoints, scheduling batch jobs, or scaling Python code to cloud containers with the Modal SDK.
日本語の概要は準備中です。原文の説明を表示しています。
K-Dense-AI/scientific-agent-skills☆ 4.8万2026年10月5日 更新
Pick the right serving container for a SageMaker model deployment and find its current image URI. Use this skill whenever about to deploy a model to a SageMaker endpoint and an image URI needs to be chosen — including when the user says "deploy this LLM", "host this HuggingFace model", "serve this fine-tuned model", "deploy this embedding model", "host a reranker", "serve a sentence-transformers model", or when about to hardcode any container URI in deployment code. HuggingFace-curated Deep Learning Containers are ALWAYS preferred: HuggingFace vLLM (LLMs and generative rerankers), HuggingFace vLLM-Omni (multimodal), TEI (embeddings/cross-encoder rerankers), HF Inference Toolkit (other transformers). Generic images (AWS vLLM, DJL-LMI, SGLang) are used only when no HuggingFace image is compatible — never merely because they carry a newer version. Never hardcode a container URI from memory and never default to TGI. Prevents stale-image failures and wrong-region URIs.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/skills☆ 1.1万2026年10月9日 更新
0G Compute Network guide for decentralized AI inference, fine-tuning, and GPU services. Covers chatbots, image generation, speech-to-text, SDK integration (0g-serving-broker), processResponse API, broker.inference methods, CLI commands (0g-compute-cli), and account management. Use this skill for any 0G compute, 0G AI, or decentralized GPU question.
日本語の概要は準備中です。原文の説明を表示しています。
internet-court/internet-court-skill☆ 6,7182026年8月20日 更新
Implements the NOWAIT technique for efficient reasoning in R1-style LLMs. Use when optimizing inference of reasoning models (QwQ, DeepSeek-R1, Phi4-Reasoning, Qwen3, Kimi-VL, QvQ), reducing chain-of-thought token usage by 27-51% while preserving accuracy. Triggers on "optimize reasoning", "reduce thinking tokens", "efficient inference", "suppress reflection tokens", or when working with verbose CoT outputs.
日本語の概要は準備中です。原文の説明を表示しています。
foryourhealth111-pixel/Vibe-Skills☆ 3,6532026年8月31日 更新
Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
LLM serving for latency, batching, caching, quantization, and routing. Use when tuning vLLM or SGLang, sizing KV cache, adding speculative decoding, or cutting serving cost.
日本語の概要は準備中です。原文の説明を表示しています。
vasilyu1983/AI-Agents-public☆ 912026年10月5日 更新
MLOps and the production ML lifecycle -- model packaging and serving, CI/CD for ML, experiment tracking, model registries, reproducibility, production monitoring for data and concept drift, retraining pipelines, A/B and shadow deployment, and rollback. Covers batch vs online/real-time inference, REST endpoints, feature stores, data and version pinning, deterministic pipelines, performance-decay detection, and retraining triggers. Use when deploying models to production, serving predictions, monitoring for data or concept drift, setting up ML CI/CD, tracking experiments, managing a model registry, or planning retraining, shadow rollout, and rollback.
日本語の概要は準備中です。原文の説明を表示しています。
Tibsfox/gsd-skill-creator☆ 712026年7月20日 更新
Executa inferência de LLM em CPU, Apple Silicon e GPUs consumer sem hardware NVIDIA. Use para edge deployment, Macs M1/M2/M3, GPUs AMD/Intel ou quando CUDA não está disponível. Suporta quantização GGUF (1,5-8 bits) para redução de memória e aceleração de 4-10× vs PyTorch em CPU.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Otimiza inferência de LLM com NVIDIA TensorRT para máxima vazão e latência mínima. Use para implantação em produção em GPUs NVIDIA (A100/H100), quando você precisa de inferência 10-100x mais rápida que PyTorch, ou para servir modelos com quantização (FP8/INT4), batching em voo e escalabilidade multi-GPU.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Serve LLMs com alta throughput usando PagedAttention do vLLM e continuous batching. Use ao fazer deploy de APIs LLM em produção, otimizar latência/throughput de inferência, ou servir modelos com memória GPU limitada. Suporta endpoints compatíveis com OpenAI, quantização (GPTQ/AWQ/FP8) e tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Use when working with Bentoml — bentoML model serving and packaging management. Covers service management, model packaging, deployment status, API testing, Bento building, and runner configuration. Use when packaging ML models for serving, deploying BentoML services, debugging inference issues, or managing model artifacts.
日本語の概要は準備中です。原文の説明を表示しています。
cloudthinker-ai/CloudSkills☆ 62026年4月5日 更新
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新
为单实例推理服务比较批处理、请求调度、KV 缓存、卸载和推测解码。用于容量、吞吐、延迟和每合格请求成本的取舍。
日本語の概要は準備中です。原文の説明を表示しています。
bojieli/ai-infra-book☆ 6,2262026年10月1日 更新
Convert existing PyTorch Lightning training code into an NVFLARE federated job using the Lightning Client API patch, local validation, and job export; use only when the request names federated/NVFLARE conversion or asks multiple sites to train collaboratively while keeping each site's data local, and either names PyTorch Lightning or preliminary source inspection identifies one Lightning owner; do not use for non-federated Lightning work such as DDP, profiling, inference serving, or training-loop changes, nor for plain PyTorch, TensorFlow/Keras, other frameworks, deployment, POC/production lifecycle, or experiment workflows.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. WHEN: pod crashes or Pending, CrashLoopBackOff, OOMKilled, ImagePullBackOff, node NotReady, DNS or ingress failure, connectivity timeout, network policy, SNAT exhaustion, node-pool scaling blocked by QuotaExceeded or InsufficientVCPUQuota, upgrade stuck, spot or zone disruption, a bare VMExtensionProvisioningError or AllocationFailed wrapper, an uncataloged capacity symptom, or 'investigate my AKS cluster'. DO NOT USE FOR: packet capture (use aks-network-capture); GPU or model-serving issues (use aks-gpu-inference); cluster creation or provisioning (use azure-kubernetes); cost (use cost-analysis from the optional azure-cost plugin); pod rightsizing (use azure-kubernetes); a fully qualified documented AKS signature with every required nested qualifier (use aks-known-issues); standalone failures on non-AKS Azure resources (use azure-diagnostics). Unqualified errors and open-ended incidents stay here only for AKS.
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/skills☆ 3,1012026年10月10日 更新
Set up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN: "setup AI Runway", "onboard AKS cluster", "install AI Runway", "airunway setup", "deploy model to AKS", "GPU inference on AKS", "KAITO setup on AKS", "run LLM on AKS", "vLLM on AKS", "set up model serving on AKS", "AI Runway controller".
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/skills☆ 3,1012026年10月10日 更新
Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. Use together with the dstack skill, and only when the user explicitly asks to create a preset or manage existing presets, not for deploying or serving a model.
日本語の概要は準備中です。原文の説明を表示しています。
dstackai/dstack☆ 2,2772026年10月11日 更新
Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. Implements model serving, feature engineering, A/B testing, and monitoring. Use PROACTIVELY for ML model deployment, inference optimization, or production ML infrastructure.
日本語の概要は準備中です。原文の説明を表示しています。
rmyndharis/antigravity-skills☆ 1,7372026年10月1日 更新
Debug live Azure Kubernetes Service (AKS) incidents with a read-only, evidence-first investigation. WHEN: pod crashes or Pending, CrashLoopBackOff, OOMKilled, ImagePullBackOff, node NotReady, DNS or ingress failure, connectivity timeout, network policy, SNAT exhaustion, node-pool scaling blocked by QuotaExceeded or InsufficientVCPUQuota, upgrade stuck, spot or zone disruption, a bare VMExtensionProvisioningError or AllocationFailed wrapper, an uncataloged capacity symptom, or 'investigate my AKS cluster'. DO NOT USE FOR: packet capture (use aks-network-capture); GPU or model-serving issues (use aks-gpu-inference); cluster creation or provisioning (use azure-kubernetes); cost (use cost-analysis from the optional azure-cost plugin); pod rightsizing (use azure-kubernetes); a fully qualified documented AKS signature with every required nested qualifier (use aks-known-issues); standalone failures on non-AKS Azure resources (use azure-diagnostics). Unqualified errors and open-ended incidents stay here only for AKS.
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/azure-skills☆ 1,5582026年10月10日 更新
Set up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN: "setup AI Runway", "onboard AKS cluster", "install AI Runway", "airunway setup", "deploy model to AKS", "GPU inference on AKS", "KAITO setup on AKS", "run LLM on AKS", "vLLM on AKS", "set up model serving on AKS", "AI Runway controller".
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/azure-skills☆ 1,5582026年10月10日 更新
Deploy and manage vLLM for high-throughput LLM inference. Configure continuous batching, tensor parallelism, quantization, and OpenAI-compatible API endpoints for production LLM serving.
日本語の概要は準備中です。原文の説明を表示しています。
BagelHole/DevOps-Security-Agent-Skills☆ 1,1542026年5月22日 更新