本文へ移動
cccskills

「cuda」の検索結果

95 件 ・ 関連度順

概要と使いどころ

nvidia-cuda

無料日本語概要

NVIDIA CUDA GPU 並列コンピューティングリファレンス。 CUDA C++ / CUDA Python, kernel, nvcc, Unified Memory, CUDA Graphs, Cooperative Groups, Driver API, マルチ GPU。 PTX ISA 命令セット, state space, MMA 命令。 Blackwell チューニング, Streaming Multiprocessor, NVLink。 CUTLASS / CuTe DSL / CuTe C++ GEMM, Tensor Core, tcgen05, Operator API, nsight-compute, compute-sanitizer。

Fandhe-AI/agent-reference-skills42026年10月11日 更新

Use this skill when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via doca_gpu_eth_rxq / doca_gpu_eth_txq, standing up the per-CUDA-device doca_gpu context, designing the persistent CUDA kernel that drains the GPU-visible queue, running the dual capability check (DOCA cap-query plus cudaGetDeviceProperties), registering cudaMalloc pools via doca_buf_arr_create_*, or debugging DOCA_ERROR_* returns from the GPUNetIO API. Trigger even when the user does not explicitly mention "DOCA GPUNetIO" or "persistent kernel" — typical implicit phrasings include "CUDA kernel reading packets directly from the NIC", "GPU-initiated networking on BlueField", "DOCA_ERROR_DRIVER on doca_gpu_create", "nvidia_peermem not loaded", "kernel-per-packet is too slow", or "which GPU supports GPU-side packet I/O". Refuse and route elsewhere for general CUDA programming, DOCA Ethernet queue bring-up, DOCA DPA, or DOCA install — those belong to other skills.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

doca-gpi

無料

Use this skill for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. Covers picking GPI vs doca-gpunetio, the doca_gpi / domain / channel object model, the GPU-side handle handoff (doca_gpu_gpi_channel*), attaching GPU memory to a GPI domain, the domain and channel attribute objects, and debugging DOCA_ERROR_* from doca_gpi_* calls. Trigger even when the user does not explicitly mention "DOCA GPI" — implicit phrasings include "my CUDA kernel needs to post RDMA directly from GPU memory", "DOCA_ERROR_* from doca_gpi_gpu_channel_get", "how do I hand a GPU handle to my CUDA kernel", "how many channels can a GPI domain hold", or "GPU kernel driving RDMA without the host CPU on the path". Refuse and route elsewhere for the doca-gpunetio Send/Receive surface, the doca-rdma queue lifecycle, DPA-side initiation (doca-rdmi), or the CUDA programming model — those belong to other skills.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use this skill when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCA_EXPERIMENTAL-marked C entry points that doca-bench loads to measure a workload class its built-in modes do not cover, with doca_bench_cuda as the shipped reference exemplar. Trigger even when the user does not say "doca-bench-extension" or "doca_bench_cuda" — typical implicit phrasings include "no built-in doca-bench mode fits my workload", "how do I benchmark a CUDA GPUNetIO RX/TX kernel", "doca-bench cannot find or load my custom .so", "extension exported symbols do not match what the parent expects", "soversion mismatch after a DOCA upgrade", or "my GPU kernel hangs because stop_flag was never set". Refuse and route elsewhere for questions about which built-in doca-bench mode to pick, DOCA GPUNetIO programming semantics, CUDA toolkit installation, or contributor work on in-tree extensions — those belong to other skills.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use when benchmarking, profiling, or A/B-testing CUDA kernels or end-to-end perf in the imp inference engine on RTX 5090 (sm_120), including refreshing tests/perf_baseline.json or publishing numbers to docs/BENCHMARKS.md and the README. Triggers on "benchmark kernel", "profile cuda", "ncu", "nsys", "kernel timing", "kernel sum", "occupancy", "bandwidth bound", "compute bound", "roofline", "perf baseline", "is this regression real", "decode dropped", "aggregate throughput", "two-image A/B", "prefill kernel A/B". Do NOT use for writing/optimizing kernel code (sm120-cuda-expert) or output-quality checks (check-degeneration).

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when writing, reviewing, or optimizing CUDA kernels targeting sm_120a (RTX 5090 / Consumer Blackwell, GB202) in the imp inference engine. Triggers on CUDA/PTX kernel code, shared-memory layout, bank conflicts, tensor-core MMA (mxf4nvf4, tf32, 3xTF32/3xFP16), GEMV/GEMM/attention/quantization/GDN-scan kernels, decode tok/s under expected, kernel emitting HMMA instead of mxf4nvf4, occupancy, register pressure, __launch_bounds__, spills, PDL. Pair with `benchmark-cuda` for measurement and `check-degeneration` after hot-path changes.

日本語の概要は準備中です。原文の説明を表示しています。

kekzl/imp442026年10月11日 更新

Use when debugging CUDA with cuda-gdb or Compute Sanitizer, reading GPU core dumps, using device printf, or triaging error codes 700, 701, 702, and 719. Not for performance: use cuda-profiling.

日本語の概要は準備中です。原文の説明を表示しています。

OutlineDriven/odin-claude-plugin372026年9月29日 更新

colab-video

無料日本語概要

窓際族物語の動画生成工程(/local-videoのステップ7)だけをGoogle Colab(CUDA GPU)で実行する派生ワークフロー。ローカルMacにCUDA GPUが無くMiniMax H3を実行できないとき、ユーザーから「Colabで動画を作って」「H3をColabで回して」「LTXで動画を作って」と指示されたときに必ず使用する。台本・音声・キーフレーム生成・画像検証(ステップ1〜6)と最終結合(ステップ8〜9)はローカルで行い、動画生成だけを同梱ノートブック(H3=h3_colab.ipynb / LTX-2.5=ltx25_colab.ipynb)でColabに切り出す。セリフ(wav駆動リップシンク)はH3、セリフなしI2VチャプターはLTX-2.5も選べる。無料T4は配管検証用、本番生成はL4/A100(Pay As You Go / Colab Pro)。

sobaya-0141/Seedance_Madogiwa42026年10月9日 更新

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. Use for CUDA/GPU optimization; CPU-bound NumPy, SciPy, pandas, scikit-learn, NetworkX, scikit-image, vector-search, image-processing, graph, simulation, or file-I/O workloads; CuPy, cuDF, cuML, cuGraph, cuVS, cuCIM, KvikIO, Warp, Newton, Numba-CUDA, or RAFT questions; and profiling, memory-transfer, kernel, or multi-GPU bottlenecks. Also use when large data-parallel Python code is slow and GPU acceleration is a plausible option, even if the user does not name CUDA.

日本語の概要は準備中です。原文の説明を表示しています。

K-Dense-AI/scientific-agent-skills4.8万2026年10月5日 更新

Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment. Use for Conda installs; redirect CUDA 12 hosts to container/wheel.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use when porting circuits from another framework (e.g. Qiskit) into CUDA-Q kernels while preserving the source algorithm and validation fidelity.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Optimize MATLAB design files for GPU Coder to generate faster CUDA code. Iteratively profiles, rewrites, and benchmarks until performance targets are met or diagnostics are resolved. Use when asked to: optimize for GPU Coder, improve GPU codegen performance, profile generated GPU/CUDA code, profile GPU MEX, fix gpuPerformanceAnalyzer diagnostics, speed up GPU MEX, reduce GPU memory transfers, improve kernel parallelism, rewrite MATLAB for CUDA, or run gpuPerformanceAnalyzer.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1502026年10月9日 更新

Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. Use when the user wants to integrate an AI model into an application with code generation as the end goal — generating MEX, CUDA MEX, static library, dynamic library, or executable — or using the model in Simulink for simulation and code generation. Covers PyTorch ExportedProgram (.pt2) via loadPyTorchExportedProgram and LiteRT (.tflite) via loadLiteRTModel (R2026a+). Keywords: PyTorch, torch, .pt2, ExportedProgram, loadPyTorchExportedProgram, invoke, codegen, MEX, CUDA, GPU, C, C++, deploy, AI model, deep learning model, LiteRT, TFLite, TensorFlow Lite, Simulink, slbuild, PyTorch ExportedProgram block, MATLAB Function block, dlosslib, loadLiteRTModel.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1502026年10月9日 更新

cuda

無料

Use when writing CUDA kernels, managing the thread, block, and grid hierarchy, tiling shared memory, using streams, setting nvcc flags, or using Thrust. Not for kernel debugging: use cuda-debugging.

日本語の概要は準備中です。原文の説明を表示しています。

OutlineDriven/odin-claude-plugin372026年9月29日 更新

Use when profiling CUDA with Nsight Systems or Nsight Compute, reading roofline and occupancy metrics, or annotating phases with NVTX. Not for correctness: use cuda-debugging.

日本語の概要は準備中です。原文の説明を表示しています。

OutlineDriven/odin-claude-plugin372026年9月29日 更新

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). Use when installing PyTorch/Unsloth/TRL/vLLM on DGX Spark, hitting libcudart or wheel-ABI errors on aarch64, or choosing between NGC containers and bare pip installs.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

日本語の概要は準備中です。原文の説明を表示しています。

sgl-project/sglang3.7万2026年10月12日 更新

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use this skill when the user is building, running, or interpreting the doca/tools/gpunetio_ib_write_bw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests through the doca-gpunetio device-side surface to measure sustained GPU-driven WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user does not explicitly mention "doca-gpunetio-ib-write-bw" or "GPUNetIO" — typical implicit phrasings include "measure WRITE BW when the GPU posts the WRs", "BW swings between runs on the same flags", "is the NIC saturated or am I CPU-bound on the CUDA kernel", "meson compile fails for the GPUNetIO bw tool", "nvidia_peermem isn't picking up my GPU buffer", or "GPU-initiated WRITE throughput vs CPU-initiated perftest". Refuse and route elsewhere for general doca-gpunetio library work, DOCA install, the GPU-initiated WRITE latency analog, the CPU-initiated upstream perftest, or application-level end-to-end throughput — those belong to other skills.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Use this skill when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the `gpunetio_ib_write_lat` client + server pair under `doca/tools/gpunetio_ib_write_lat/`, checking GPU-NIC pairing, reading the half-iter / full-iter / CUDA-side usec columns, characterizing median / p99 / jitter for a real-time control loop, picking GPUNetIO vs GPI vs CPU-initiated `perftest`, or weighing the latency-vs-batching trade-off. Trigger even without 'GPUNetIO' or 'ib_write_lat': 'GPU kernel RDMA latency benchmark', 'how fast can a CUDA kernel post a WRITE', 'p99 RDMA latency on H100 + ConnectX', 'kernel-launched WR tail latency', or 'compare GPU-init vs CPU-init perftest'. Route elsewhere for bandwidth runs (doca-gpunetio-ib-write-bw), the GPI surface (doca-gpi), library debugging (doca-gpunetio), or DOCA install.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI). Use for solver internals, PRs, DCO, and code conventions.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5612026年10月10日 更新

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia), Python, and PyTorch. Use when checking component versions, verifying CUDA/driver compatibility, detecting version mismatches across nodes, planning upgrades, documenting cluster configuration, or troubleshooting version-related issues on HyperPod. Triggers on requests about versions, compatibility, component checks, or upgrade planning for HyperPod clusters.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Provides guidance for writing and benchmarking optimized CUDA kernels for NVIDIA GPUs (H100, A100, T4) targeting HuggingFace diffusers and transformers libraries. Kernels must be kernel-builder/ABI3-compliant: no pybind11, no setup.py, TORCH_LIBRARY_EXPAND bindings only. Supports models like LTX-Video, Stable Diffusion, LLaMA, Mistral, and Qwen. Includes integration with HuggingFace Kernels Hub (get_kernel) for loading pre-compiled kernels. Includes benchmarking scripts to compare kernel performance against baseline implementations.

日本語の概要は準備中です。原文の説明を表示しています。

huggingface/kernels7652026年10月10日 更新