本文へ移動
cccskills

「attention optimization」の検索結果

24 件 ・ 関連度順

概要と使いどころ

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

flash-attention

無料日本語概要

長い入力を扱うTransformerの学習・推論で、注意機構の計算を高速化し、GPUメモリ使用量を抑えるスキル。導入から性能測定、出力の比較まで案内します。

  • 長い入力での学習を高速化したいとき
  • 長文推論のメモリ使用量を抑えたいとき
  • 処理速度とモデル出力の比較
NousResearch/hermes-agent25.3万2026年10月11日 更新

Otimiza atenção em transformers com Flash Attention para ganho de 2-4x em velocidade e redução de 10-20x em memória. Use ao treinar/executar transformers com sequências longas (>512 tokens), ao encontrar problemas de memória GPU com atenção, ou quando precisa de inferência mais rápida. Suporta SDPA nativo do PyTorch, biblioteca flash-attn, H100 FP8 e sliding window attention.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

This skill should be used to explain or reason about the foundational concepts of context engineering: what context is, the anatomy of a context window, how attention mechanics work, the U-shaped attention curve, why context quality matters more than quantity, and the mental models needed to interpret every other context-engineering decision. Use this for conceptual explanation, onboarding, and background reading. Route operational work to the specialized skills: debugging attention failures goes to context-degradation, token-efficiency work goes to context-optimization, conversation summarization goes to context-compression, and project-shape decisions go to project-development.

日本語の概要は準備中です。原文の説明を表示しています。

muratcankoylan/Agent-Skills-for-Context-Engineering1.8万2026年10月1日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/RuView9.7万2026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/ruflo7.4万2026年10月11日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques. Use when optimizing inference speed (1.5-3.6× speedup), reducing latency for real-time applications, or deploying models with limited compute. Covers draft models, tree-based attention, Jacobi iteration, parallel token generation, and production deployment strategies.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/RuVector4,5552026年10月11日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/agentic-flow8172026年10月10日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/ruv-FANN3852026年8月9日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

spencermarx/open-code-review3712026年7月28日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ruvnet/marketing1292026年5月23日 更新

DiT 计算模块(L3):把**已定位的 DiT 计算瓶颈**落成特性级选档与实施—— 量化档(W8A16 / W4A16 / W8A8 系列 / W4A4 / MXFP8 / FA 量化)、稀疏(rf_v2 / ada_bsa)、 缓存(DiTCache / AttentionCache / 时间步优化)、编译启用(MindieSDBackend / Pattern 融合 / ACLGraph) 的**开不开、开哪一档、怎么开、怎么复验**;依据是 `docs/zh/features/*`(特性真源)+ framework-integration/references/framework-support-matrix.md(支持状态)。 即使用户只说"这个模型怎么加速""量化/稀疏/Cache 怎么选怎么开""要不要开量化、开哪一档""这个档位开了有没有效果" 而未提 profiling,也应触发。 **入口条件**:瓶颈点已明确(用户带一句实测锚点,或编排层交付标签)时由域入口 `performance-optimization` 按标签分发到本技能;**瓶颈未明("怎么加速 / 跑通 / 采 profile")先走 `model-auto-optimization` 定位**,不在本技能内做占比分析。 near-miss:多卡并行形态 / 通信掩盖 / TP·offload 选型 → `dit-parallel-opt`;VAE 解码段与 host 固定开销 → 各自模块(VAE / host);单算子实现级实测选型(mindie_bench)→ `benchmark-dev`; 需要新增 pattern / 算子才能落地本档 → `pattern-dev` / `operator-dev`;框架侧开关与使能验证 → `framework-integration`;量化器位级契约与精度对齐(编码公式 / 舍入 / scale 粒度)→ `quantization-dev`;精度验收判据 → `accuracy-gate`;数字入库口径 → `perf-gate`。

日本語の概要は準備中です。原文の説明を表示しています。

Ascend/MindIE-SD152026年10月10日 更新

deepspeed

無料

Expert guidance for distributed training with DeepSpeed - ZeRO optimization stages, pipeline parallelism, FP16/BF16/FP8, 1-bit Adam, sparse attention

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Achieve aggressive v3 performance targets: 2.49x-7.47x Flash Attention speedup, 150x-12,500x search improvements, 50-75% memory reduction. Comprehensive benchmarking and optimization suite.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Write landing page copy with attention to the hero, value proposition, social proof, objection handling, and conversion-focused CTAs. Use this skill whenever the user wants to write a landing page, sales page, hero section, or any conversion-focused web copy. Triggers on landing page, sales page, hero copy, value proposition, headline, subheadline, hero section, CTA copy, conversion copy, opt-in page, squeeze page. Also triggers when the user has a marketing campaign or product launch needing dedicated conversion copy. Use `cro-optimization` instead when the page already exists and the ask is to lift its conversion rate through testing rather than to write the copy.

日本語の概要は準備中です。原文の説明を表示しています。

rampstackco/claude-skills9462026年10月7日 更新

Provides guidance for writing, optimizing, and benchmarking Triton kernels for Intel XPU GPUs (Battlemage/Arc Pro B50) using the Xe-Forge optimization framework. Includes an LLM-driven trial-loop workflow (analyze, validate, benchmark, profile, finalize), XPU-specific patterns (tensor descriptors, GRF mode, tile swizzling), KernelBench fused kernels, and Flash Attention.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/kernels7652026年10月10日 更新

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供 model inference performance optimization 的脱敏案例。纯 API 调用、提示词或 token 用量优化,以及电脑卡顿,不进入本技能的模型推理实验流程。

日本語の概要は準備中です。原文の説明を表示しています。

majiayu000/spellbook2872026年10月8日 更新

Acelere a inferência de LLMs usando especulative decoding, múltiplas cabeças Medusa e técnicas de lookahead decoding. Use ao otimizar velocidade de inferência (aceleração de 1,5-3,6×), reduzir latência em aplicações em tempo real ou fazer deploy de modelos com recursos computacionais limitados. Cobre modelos draft, atenção em árvore, iteração de Jacobi, geração paralela de tokens e estratégias de deploy em produção.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新