本文へ移動
cccskills

「tensorrt-llm」の検索結果

12 件 ・ 関連度順

概要と使いどころ

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Otimiza inferência de LLM com NVIDIA TensorRT para máxima vazão e latência mínima. Use para implantação em produção em GPUs NVIDIA (A100/H100), quando você precisa de inferência 10-100x mais rápida que PyTorch, ou para servir modelos com quantização (FP8/INT4), batching em voo e escalabilidade multi-GPU.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

dgx-spark

無料日本語概要

NVIDIA DGX Spark User Guide リファレンス。GB10 Grace Blackwell Superchip ハードウェア仕様、first-boot / UEFI 設定、Spark Stacking クラスタ、 DGX OS / NGC / container runtime ソフトウェアスタック、 PXE / cloud-init プロビジョニング、fleet lifecycle 運用、 apt / fwupdmgr 更新、system recovery、PSIRT サポート、 vLLM / TensorRT-LLM / Ollama / ComfyUI / NIM / NeMo fine-tune / NCCL 等の build.nvidia.com/spark playbook、複数 Spark 接続・Tailscale・VS Code、 NVIDIA Developer Forums 由来の known issues / troubleshooting。

Fandhe-AI/agent-reference-skills42026年10月9日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

High-throughput LLM inference on NVIDIA GPUs.

日本語の概要は準備中です。原文の説明を表示しています。

NousResearch/hermes-agent25.3万2026年10月10日 更新

Unified LLM torch-profiler triage skill for `sglang`, `vllm`, `TensorRT-LLM`, and `TokenSpeed`. Use it to inspect an existing `trace.json(.gz)` or profile directory, or to drive live profiling against a running server when supported and return one three-table report with kernel, overlap-opportunity, and fuse-pattern tables.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月11日 更新

LLM and ML model deployment for inference. Use when serving models in production, building AI APIs, or optimizing inference. Covers vLLM (LLM serving), TensorRT-LLM (GPU optimization), Ollama (local), BentoML (ML deployment), Triton (multi-model), LangChain (orchestration), LlamaIndex (RAG), and streaming patterns.

日本語の概要は準備中です。原文の説明を表示しています。

ancoleman/ai-design-components5252025年12月11日 更新

Cross-engine decision rubric for self-hosting or recommending an LLM serving stack. Picks among vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX as a function of (hardware × workload × constraint), not "which is fastest". Activates whenever a coder-agent must choose, defend, or migrate a serving runtime.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新

Decision SOP for serving LLMs with vLLM. Covers PagedAttention mental model, quantization/parallelism/batching tradeoffs, OOM triage, and when NOT to use vLLM. Activates when a coder-agent is choosing or tuning an inference engine, debugging vLLM throughput/latency/OOM, or comparing vLLM against TGI/SGLang/TensorRT-LLM/llama.cpp.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4372026年10月9日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新