本文へ移動
cccskills

「fp8」の検索結果

41 件 ・ 関連度順

概要と使いどころ

flash-attention

無料日本語概要

長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。

  • 長い入力での学習を高速化したいとき
  • 長文推論のメモリ使用量を抑えたいとき
  • 処理速度とモデル出力の比較
NousResearch/hermes-agent25.3万2026年10月11日 更新

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Provides guidance for enterprise-grade RL training using miles, a production-ready fork of slime. Use when training large MoE models with FP8/INT4, needing train-inference alignment, or requiring speculative RL for maximum throughput.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

量化格式契约与位级对齐:从设备字节反推量化器的**编码公式、舍入模式、scale 粒度、退化块规则**, 并据此重实现/融合量化器(MXFP8 e8m0、int8 perblock、fp8 perchannel 等), 用「同进程同数据 + 独立参考实现 + 中点与退化输入全覆盖」做到**逐字节精确**。 当自研量化 kernel 与框架算子**对不上**、需要逆向某个量化器的数值契约、需要判断 「量化残差算不算可接受」、或要在融合内核里复现 aclnn 量化语义时使用此 skill。 即使用户只说「量化结果和参考不一致」「这个 scale 怎么算出来的」「融合量化器精度对不上」 而未提"契约",也应触发。**注意**:单纯选量化档位/配置(该不该开量化、开哪一档)走 dit-perf-opt;算子级性能调优与 DSL 选型走 operator-dev; 并行作用域导致的静默失效(DiT 侧「改了并行但不报错也没生效」的判定)走 dit-parallel-opt。

日本語の概要は準備中です。原文の説明を表示しています。

Ascend/MindIE-SD152026年10月10日 更新

Fornece orientação para treinamento RL de nível empresarial usando miles, um fork pronto para produção do slime. Use ao treinar grandes modelos MoE com FP8/INT4, necessitando alinhamento treino-inferência ou exigindo RL especulativo para máxima taxa de transferência.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Otimiza inferência de LLM com NVIDIA TensorRT para máxima vazão e latência mínima. Use para implantação em produção em GPUs NVIDIA (A100/H100), quando você precisa de inferência 10-100x mais rápida que PyTorch, ou para servir modelos com quantização (FP8/INT4), batching em voo e escalabilidade multi-GPU.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Otimiza atenção em transformers com Flash Attention para ganho de 2-4x em velocidade e redução de 10-20x em memória. Use ao treinar/executar transformers com sequências longas (>512 tokens), ao encontrar problemas de memória GPU com atenção, ou quando precisa de inferência mais rápida. Suporta SDPA nativo do PyTorch, biblioteca flash-attn, H100 FP8 e sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

tensorrt-llm

無料日本語概要

NVIDIA GPU上で大規模言語モデルの応答生成を高速化し、APIとして提供する設定を支援します。量子化や一括処理、複数GPUへの分散も扱います。

  • 言語モデルをチャットAPIで提供したいとき
  • 多数のプロンプトを一括処理したいとき
  • 量子化でメモリ使用量を抑えたいとき
NousResearch/hermes-agent25.3万2026年10月11日 更新

Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8. Use after a checkpoint passes promotion, when choosing a quantization format for a target device, or when an exported model fails its smoke test.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

Use when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月11日 更新

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GPU memory. Supports OpenAI-compatible endpoints, quantization (GPTQ/AWQ/FP8), and tensor parallelism.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Simplest distributed training API. 4 lines to add distributed support to any PyTorch script. Unified API for DeepSpeed/FSDP/Megatron/DDP. Automatic device placement, mixed precision (FP16/BF16/FP8). Interactive config, single launch command. HuggingFace ecosystem standard.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新