本文へ移動
cccskills
無料GitHub で公開日本語紹介

tensorrt-llm

NVIDIA GPU上で大規模言語モデルの応答生成を高速化し、APIとして提供する設定を支援します。量子化や一括処理、複数GPUへの分散も扱います。

原文High-throughput LLM inference on NVIDIA GPUs.

インストール方法を見る

こんなときに便利

  • 言語モデルをチャットAPIで提供したいとき
  • 多数のプロンプトを一括処理したいとき
  • 量子化でメモリ使用量を抑えたいとき
  • 大きなモデルを複数GPUで動かしたいとき

日本語での紹介

できること

NVIDIA GPU上で大規模言語モデルの推論、つまり入力に対する応答生成を高速化し、サービスとして提供する設定を支援します。TensorRT-LLMを使ったモデルの読み込み、生成条件の設定、trtllm-serveによるAPI起動を扱います。数値の精度を落としてメモリを節約する量子化や、複数の入力をまとめるバッチ処理、複数GPUへの分散も対象です。

こんなときに便利

応答の待ち時間を抑えたいアプリや、多数の入力を処理するモデル運用に向いています。大きなモデルを複数GPUで動かしたい場合や、量子化モデルの設定を検討する際にも使えます。生成中の処理をまとめる仕組みや、生成に使うメモリの管理も扱います。

使い方の例

  • 「trtllm-serveでモデルを起動し、チャットAPIから呼び出せるようにして」
  • 「多数のプロンプトをまとめて処理するコードを作って」
  • 「モデルを複数GPUに分散し、FP8で動かす設定を作って」

注意点

対応環境はLinuxとNVIDIA GPUです。依存ライブラリはtensorrt-llmとtorchで、資料ではCUDA 13.2.1、TensorRT 10.x、Python 3.10〜3.12が必要とされています。DockerイメージはNGCで配布されています。CPUやApple Silicon向けの運用は対象外です。

この紹介文は、公開されている SKILL.md をもとに AI(Claude Haiku)が作成しました。正確な仕様は下の原文を確認してください。

含まれるファイル(4)

  • SKILL.md4.9 KB
  • references/multi-gpu.md6.5 KB
  • references/optimization.md5.5 KB
  • references/serving.md9.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

TensorRT-LLM

NVIDIA's open-source library for optimizing LLM inference with high performance on NVIDIA GPUs.

When to use TensorRT-LLM

Use TensorRT-LLM when:

  • Deploying on NVIDIA GPUs (A100, H100, GB200)
  • Need maximum throughput (24,000+ tokens/sec on Llama 3)
  • Require low latency for real-time applications
  • Working with quantized models (FP8, INT4, FP4)
  • Scaling across multiple GPUs or nodes

Use vLLM instead when:

  • Need simpler setup and Python-first API
  • Want PagedAttention without TensorRT compilation
  • Working with AMD GPUs or non-NVIDIA hardware

Use llama.cpp instead when:

  • Deploying on CPU or Apple Silicon
  • Need edge deployment without NVIDIA GPUs
  • Want simpler GGUF quantization format

Quick start

Installation

# Docker (recommended) — images are on NGC (nvcr.io), not Docker Hub.
# Replace x.y.z with the desired version (e.g. 1.2.1). Browse tags on NGC:
# https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags
docker pull nvcr.io/nvidia/tensorrt-llm/release:x.y.z

# pip install (current stable GA)
pip install tensorrt_llm

# Requires CUDA 13.2.1, TensorRT 10.x, Python 3.10-3.12

Basic inference

from tensorrt_llm import LLM, SamplingParams

# Initialize model
llm = LLM(model="meta-llama/Meta-Llama-3-8B")

# Configure sampling
sampling_params = SamplingParams(
    max_tokens=100,
    temperature=0.7,
    top_p=0.9
)

# Generate
prompts = ["Explain quantum computing"]
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.text)

Serving with trtllm-serve

# Start server (automatic model download and compilation)
# Tensor parallelism across 4 GPUs
trtllm-serve meta-llama/Meta-Llama-3-8B \
    --tp_size 4 \
    --max_batch_size 256 \
    --max_num_tokens 4096

# Client request
curl -X POST http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Meta-Llama-3-8B",
    "messages": [{"role": "user", "content": "Hello!"}],
    "temperature": 0.7,
    "max_tokens": 100
  }'

Key features

Performance optimizations

  • In-flight batching: Dynamic batching during generation
  • Paged KV cache: Efficient memory management
  • Flash Attention: Optimized attention kernels
  • Quantization: FP8, INT4, FP4 for 2-4× faster inference
  • CUDA graphs: Reduced kernel launch overhead

Parallelism

  • Tensor parallelism (TP): Split model across GPUs
  • Pipeline parallelism (PP): Layer-wise distribution
  • Expert parallelism: For Mixture-of-Experts models
  • Multi-node: Scale beyond single machine

Advanced features

  • Speculative decoding: Faster generation with draft models
  • LoRA serving: Efficient multi-adapter deployment
  • Disaggregated serving: Separate prefill and generation

Common patterns

Quantized model (FP8)

from tensorrt_llm import LLM

# Load FP8 quantized model (2× faster, 50% memory)
llm = LLM(
    model="meta-llama/Meta-Llama-3-70B",
    dtype="fp8",
    max_num_tokens=8192
)

# Inference same as before
outputs = llm.generate(["Summarize this article..."])

Multi-GPU deployment

# Tensor parallelism across 8 GPUs
llm = LLM(
    model="meta-llama/Meta-Llama-3-405B",
    tensor_parallel_size=8,
    dtype="fp8"
)

Batch inference

# Process 100 prompts efficiently
prompts = [f"Question {i}: ..." for i in range(100)]

outputs = llm.generate(
    prompts,
    sampling_params=SamplingParams(max_tokens=200)
)

# Automatic in-flight batching for maximum throughput

Performance benchmarks

Meta Llama 3-8B (H100 GPU):

  • Throughput: 24,000 tokens/sec
  • Latency: ~10ms per token
  • vs PyTorch: 100× faster

Llama 3-70B (8× A100 80GB):

  • FP8 quantization: 2× faster than FP16
  • Memory: 50% reduction with FP8

Supported models

  • LLaMA family: Llama 2, Llama 3, CodeLlama
  • GPT family: GPT-2, GPT-J, GPT-NeoX
  • Qwen: Qwen, Qwen2, QwQ
  • DeepSeek: DeepSeek-V2, DeepSeek-V3
  • Mixtral: Mixtral-8x7B, Mixtral-8x22B
  • Vision: LLaVA, Phi-3-vision
  • 100+ models on HuggingFace

References

Resources

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

1password

無料日本語概要

1Passwordに保存したパスワードやAPIキーを、コマンド実行や設定テンプレートで使うためのスキル。CLIの導入、認証、秘密情報の取得・受け渡しを案内します。

  • 1Password CLIの導入と設定
  • 保存済みAPIキーでコマンド実行
  • 設定テンプレートへの秘密情報の埋め込み
NousResearch/hermes-agent25.3万2026年10月11日 更新

3-statement-model

無料日本語概要

損益計算書・貸借対照表・キャッシュフロー計算書を数式で連動させるExcelモデルを作ります。実績と予測の入力、前提を変えた比較、財務三表の整合性確認を段階的に進めます。

  • 財務テンプレートへの実績入力
  • 財務三表が連動する予測の作成
  • 基本・上振れ・下振れシナリオの比較
NousResearch/hermes-agent25.3万2026年10月11日 更新

Run PyTorch training across GPUs with minimal changes.

日本語の概要は準備中です。原文の説明を表示しています。

NousResearch/hermes-agent25.3万2026年10月11日 更新

Set up Actual Computer (actual.inc) inference in Hermes.

日本語の概要は準備中です。原文の説明を表示しています。

NousResearch/hermes-agent25.3万2026年10月11日 更新

adversarial-ux-test

無料日本語概要

技術が苦手で不満を抱きやすい利用者になりきってアプリを試し、操作の分かりにくさや離脱の原因を発見。実際の改善課題と個人的な不満を分け、修正案をまとめます。

  • 公開やデモ前に使いやすさを確認したいとき
  • 新規ユーザーの体験を点検したいとき
  • 登録・課金でのつまずきを探したいとき
NousResearch/hermes-agent25.3万2026年10月11日 更新

agent-merge-conflict-arbiter

無料日本語概要

2つのエージェントが別々に加えた変更のGit競合を、双方の差分と意図から中立的に解決。競合箇所ごとの判断理由を記録し、ビルドやテストで統合結果を確認します。

  • 2つのエージェントのブランチ競合を解消
  • 並行したworktreeの変更を統合
  • 競合する設計判断の理由を確認したいとき
NousResearch/hermes-agent25.3万2026年10月11日 更新

NousResearch のスキルをすべて見る

このスキルの問題を報告する