長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。
- 長い入力での学習を高速化したいとき
- 長文推論のメモリ使用量を抑えたいとき
- 処理速度とモデル出力の比較
41 件 ・ 関連度順
概要と使いどころ
長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。
Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention
日本語の概要は準備中です。原文の説明を表示しています。
Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.
日本語の概要は準備中です。原文の説明を表示しています。
量化格式契约与位级对齐:从设备字节反推量化器的**编码公式、舍入模式、scale 粒度、退化块规则**, 并据此重实现/融合量化器(MXFP8 e8m0、int8 perblock、fp8 perchannel 等), 用「同进程同数据 + 独立参考实现 + 中点与退化输入全覆盖」做到**逐字节精确**。 当自研量化 kernel 与框架算子**对不上**、需要逆向某个量化器的数值契约、需要判断 「量化残差算不算可接受」、或要在融合内核里复现 aclnn 量化语义时使用此 skill。 即使用户只说「量化结果和参考不一致」「这个 scale 怎么算出来的」「融合量化器精度对不上」 而未提"契约",也应触发。**注意**:单纯选量化档位/配置(该不该开量化、开哪一档)走 dit-perf-opt;算子级性能调优与 DSL 选型走 operator-dev; 并行作用域导致的静默失效(DiT 侧「改了并行但不报错也没生效」的判定)走 dit-parallel-opt。
日本語の概要は準備中です。原文の説明を表示しています。
Fornece orientação para treinamento RL de nível empresarial usando miles, um fork pronto para produção do slime. Use ao treinar grandes modelos MoE com FP8/INT4, necessitando alinhamento treino-inferência ou exigindo RL especulativo para máxima taxa de transferência.
日本語の概要は準備中です。原文の説明を表示しています。
Otimiza inferência de LLM com NVIDIA TensorRT para máxima vazão e latência mínima. Use para implantação em produção em GPUs NVIDIA (A100/H100), quando você precisa de inferência 10-100x mais rápida que PyTorch, ou para servir modelos com quantização (FP8/INT4), batching em voo e escalabilidade multi-GPU.
日本語の概要は準備中です。原文の説明を表示しています。
Otimiza atenção em transformers com Flash Attention para ganho de 2-4x em velocidade e redução de 10-20x em memória. Use ao treinar/executar transformers com sequências longas (>512 tokens), ao encontrar problemas de memória GPU com atenção, ou quando precisa de inferência mais rápida. Suporta SDPA nativo do PyTorch, biblioteca flash-attn, H100 FP8 e sliding window attention.
日本語の概要は準備中です。原文の説明を表示しています。
Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
日本語の概要は準備中です。原文の説明を表示しています。
Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA GPU上で大規模言語モデルの応答生成を高速化し、APIとして提供する設定を支援します。量子化や一括処理、複数GPUへの分散も扱います。
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.
日本語の概要は準備中です。原文の説明を表示しています。
Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.
日本語の概要は準備中です。原文の説明を表示しています。
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.
日本語の概要は準備中です。原文の説明を表示しています。
Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.
日本語の概要は準備中です。原文の説明を表示しています。
Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.
日本語の概要は準備中です。原文の説明を表示しています。