本文へ移動
cccskills

「flash attention」の検索結果

18 件 ・ 関連度順

概要と使いどころ

flash-attention

無料日本語概要

長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。

  • 長い入力での学習を高速化したいとき
  • 長文推論のメモリ使用量を抑えたいとき
  • 処理速度とモデル出力の比較
NousResearch/hermes-agent25.3万2026年10月11日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Otimiza atenção em transformers com Flash Attention para ganho de 2-4x em velocidade e redução de 10-20x em memória. Use ao treinar/executar transformers com sequências longas (>512 tokens), ao encontrar problemas de memória GPU com atenção, ou quando precisa de inferência mais rápida. Suporta SDPA nativo do PyTorch, biblioteca flash-attn, H100 FP8 e sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Operate Flash Linear Attention package workflows: setup, kernels, layers/models, KDA/context parallel, and benchmarking.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/RuView9.7万2026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/ruflo7.4万2026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/RuVector4,5562026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/agentic-flow8172026年10月10日 更新

Provides guidance for writing, optimizing, and benchmarking Triton kernels for Intel XPU GPUs (Battlemage/Arc Pro B50) using the Xe-Forge optimization framework. Includes an LLM-driven trial-loop workflow (analyze, validate, benchmark, profile, finalize), XPU-specific patterns (tensor descriptors, GRF mode, tile swizzling), KernelBench fused kernels, and Flash Attention.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/kernels7652026年10月10日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/ruv-FANN3852026年8月9日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

spencermarx/open-code-review3712026年7月28日 更新

Think and work like an expert Deep Learning Scientist. Use when a task calls for Deep Learning Scientist judgment. Reasons from CNN/Transformer inductive bias, Li et al. loss landscapes, grokking/mode connectivity, and Kaplan/Chinchilla scaling (~20 tokens/param); designs ResNet/ViT/DiT/MoE/FlashAttention stacks with FLOPs-matched ablations; trains AdamW+cosine/WSD via Megatron-FSDP/DeepSpeed; evaluates FID/MMLU-Pro/MMLU-CF with lm-eval decontamination and Pineau/NeurIPS reproducibility checklists.

日本語の概要は準備中です。原文の説明を表示しています。

K-Dense-AI/scientific-agents1992026年10月3日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/marketing1292026年5月23日 更新

Speed up long-sequence transformer training and inference.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新