Provides guidance for writing and benchmarking optimized CUDA kernels for NVIDIA GPUs (H100, A100, T4) targeting HuggingFace diffusers and transformers libraries. Kernels must be kernel-builder/ABI3-compliant: no pybind11, no setup.py, TORCH_LIBRARY_EXPAND bindings only. Supports models like LTX-Video, Stable Diffusion, LLaMA, Mistral, and Qwen. Includes integration with HuggingFace Kernels Hub (get_kernel) for loading pre-compiled kernels. Includes benchmarking scripts to compare kernel performance against baseline implementations.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
Apply the SGLang kernels RFC when adding, moving, splitting, or reviewing kernel APIs, registry metadata, kernel tests, benchmarks, and model-specific implementations. Use with add-jit-kernel, add-sgl-kernel, and write-sglang-test for placement and migration checks.
日本語の概要は準備中です。原文の説明を表示しています。
sgl-project/sglang☆ 3.7万2026年10月11日 更新
MNN CPU 后端 kernel 开发分支(`skills/cpu/` 下,另一分支是 `cpu/optimize` 性能归因)。覆盖标量 oracle → C++ SIMD → intrinsic → 汇编的四级实现阶梯、pack/kernel ABI 契约(tile、cell stride、后处理参数)、CoreFunctions 派发表注册与二级表安全构造、跨 ISA × 精度的正确性门禁,以及 AArch64 / x86_64 / RISC-V 三份实现参考。为 CPU 后端新写或移植 kernel(NEON / SSE / AVX / RVV intrinsic 或 .S 汇编)、新增一层 ISA、改 pack mode 或 tile 参数、把 kernel 挂进函数表时使用。已定位到 kernel 才进本分支。
日本語の概要は準備中です。原文の説明を表示しています。
alibaba/MNN☆ 1.6万2026年10月10日 更新
Use this skill when the user is doing hands-on DOCA DPA host-side work on a BlueField — creating the `doca_dpa` Core context, loading a DPACC-compiled DPA app image (`doca_dpa_app`), creating DPA threads, launching kernels via `doca_dpa_kernel_launch_update_*`, draining `doca_dpa_completion`, running `doca_dpa_cap_*` discovery, choosing between the DPA comm component (inter-DPA messaging) and the DPA verbs component (in-kernel RDMA), or debugging `DOCA_ERROR_*` from `doca_dpa_*`. Trigger even without "DOCA DPA" or "Data-Path Accelerator": "run compute on the DPA from my host", "DPA kernel hangs, no completion", "DOCA_ERROR_DRIVER on launch", "DOCA/DPACC version skew", or "does this BlueField expose a DPA". Route elsewhere for DPA-side kernel programming itself, DPACC compiler internals, host↔DPU messaging (doca-comch), host-side RDMA (doca-rdma), and GPU-initiated networking (doca-gpunetio).
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Use this skill when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via doca_gpu_eth_rxq / doca_gpu_eth_txq, standing up the per-CUDA-device doca_gpu context, designing the persistent CUDA kernel that drains the GPU-visible queue, running the dual capability check (DOCA cap-query plus cudaGetDeviceProperties), registering cudaMalloc pools via doca_buf_arr_create_*, or debugging DOCA_ERROR_* returns from the GPUNetIO API. Trigger even when the user does not explicitly mention "DOCA GPUNetIO" or "persistent kernel" — typical implicit phrasings include "CUDA kernel reading packets directly from the NIC", "GPU-initiated networking on BlueField", "DOCA_ERROR_DRIVER on doca_gpu_create", "nvidia_peermem not loaded", "kernel-per-packet is too slow", or "which GPU supports GPU-side packet I/O". Refuse and route elsewhere for general CUDA programming, DOCA Ethernet queue bring-up, DOCA DPA, or DOCA install — those belong to other skills.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Provides guidance for writing, optimizing, and benchmarking C++ CPU kernels with SIMD intrinsics (AVX2/AVX512) for the Hugging Face kernels ecosystem. Includes a two-phase workflow: Phase 1 correctness (generic → AVX2) and Phase 2 performance exploration (AVX512 with branching trial loop), runtime CPU dispatch, OpenMP threading, and brgemm integration for GEMM-heavy kernels.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
Provides guidance for writing, optimizing, and benchmarking Triton kernels for Intel XPU GPUs (Battlemage/Arc Pro B50) using the Xe-Forge optimization framework. Includes an LLM-driven trial-loop workflow (analyze, validate, benchmark, profile, finalize), XPU-specific patterns (tensor descriptors, GRF mode, tile swizzling), KernelBench fused kernels, and Flash Attention.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU kernel wiki. Use when asked to distill an optimization run, a kernel_opt trace directory or a version ladder into wiki records; to report what such a run actually achieved; to build, extend, re-run or validate the staging store behind those records; or to explain how a trace-derived record's number, snippet or provenance was established.
日本語の概要は準備中です。原文の説明を表示しています。
alibaba/atrex-kernel-agent☆ 1672026年10月10日 更新
Detect kernel-level rootkits in Linux memory dumps using Volatility3 linux plugins (check_syscall, lsmod, hidden_modules), rkhunter system scanning, and /proc vs /sys discrepancy analysis to identify hooked syscalls, hidden kernel modules, and tampered system structures.
日本語の概要は準備中です。原文の説明を表示しています。
mukul975/Anthropic-Cybersecurity-Skills☆ 3.4万2026年8月31日 更新
Use this skill for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. Covers picking GPI vs doca-gpunetio, the doca_gpi / domain / channel object model, the GPU-side handle handoff (doca_gpu_gpi_channel*), attaching GPU memory to a GPI domain, the domain and channel attribute objects, and debugging DOCA_ERROR_* from doca_gpi_* calls. Trigger even when the user does not explicitly mention "DOCA GPI" — implicit phrasings include "my CUDA kernel needs to post RDMA directly from GPU memory", "DOCA_ERROR_* from doca_gpi_gpu_channel_get", "how do I hand a GPU handle to my CUDA kernel", "how many channels can a GPI domain hold", or "GPU kernel driving RDMA without the host CPU on the path". Refuse and route elsewhere for the doca-gpunetio Send/Receive surface, the doca-rdma queue lifecycle, DPA-side initiation (doca-rdmi), or the CUDA programming model — those belong to other skills.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Use this skill when the user is doing hands-on DOCA Flow DPA Provider work — exporting a `doca-flow` pipe or external resource (index-selector/memory) into BlueField DPA address space so a DPACC-built kernel can read counters, mutate hash-pipe entries, and update/read memory or index-selector resources inline with Flow. Covers per-port `doca_flow_dpa_ctx`, three queue types (general/resources-write/resources-read), the order-sensitive export handshake (`_export_prepare` → add entries → `_export` → `_get_device_addr`), and DPA-side device API. Trigger even when the user does not say "DOCA Flow DPA Provider" — implicit phrasings include "DPA kernel never sees entries in the exported pipe", "BAD_STATE from `_pipe_export`", "how do I disable a hash entry from a DPA kernel", "DPA memory read returns no value", or "DPA-side post keeps returning AGAIN". Refuse and route elsewhere for `doca-flow` pipe construction, generic host-side DPA (`doca-dpa`), or DPA-side kernel-writing — those belong to other skills.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Use this skill when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the `gpunetio_ib_write_lat` client + server pair under `doca/tools/gpunetio_ib_write_lat/`, checking GPU-NIC pairing, reading the half-iter / full-iter / CUDA-side usec columns, characterizing median / p99 / jitter for a real-time control loop, picking GPUNetIO vs GPI vs CPU-initiated `perftest`, or weighing the latency-vs-batching trade-off. Trigger even without 'GPUNetIO' or 'ib_write_lat': 'GPU kernel RDMA latency benchmark', 'how fast can a CUDA kernel post a WRITE', 'p99 RDMA latency on H100 + ConnectX', 'kernel-launched WR tail latency', or 'compare GPU-init vs CPU-init perftest'. Route elsewhere for bandwidth runs (doca-gpunetio-ib-write-bw), the GPI surface (doca-gpi), library debugging (doca-gpunetio), or DOCA install.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Use this skill when the user runs doca_dpa_hl_tracer to capture/decode DPA-side traces at the programming-events layer (kernel entry/exit, sync points, comm primitive calls, RDMA WR submission, completion drain) — picking TRACE vs CRIT, tuning the JSON config (file-size limits + file_size_limit_policy, thread priorities/cores), decoding against the matching DPA-side ELF, or diagnosing empty/noisy captures. Trigger even when the user does not explicitly mention "DOCA DPA tracer" or "high-level tracer" — typical implicit phrasings include "DPA kernel returns wrong result but host completions look clean", "kernel-entry to first-comm latency is huge", "RDMA WR to drain gap on the DPA", "trace file truncated mid-run", "TRACE doubled my DPA latency", or "tracer wrote a file but parser shows zero events". Refuse and route elsewhere for writing DPA kernels, DPA-Comms/DPA-Verbs programming, raw per-cycle DPA profiling, host-side doca-dpa debugging, or production DPA telemetry — those belong to other skills.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Linux kernel exploitation playbook. Use when exploiting kernel vulnerabilities (UAF, OOB, race condition, type confusion) for privilege escalation via commit_creds, modprobe_path overwrite, or kernel ROP chains in CTF and real-world scenarios.
日本語の概要は準備中です。原文の説明を表示しています。
yaklang/hack-skills☆ 2,4262026年9月13日 更新
Provides guidance for writing and benchmarking optimized Triton kernels for AMD GPUs (MI355X, R9700) on ROCm, targeting HuggingFace diffusers (LTX-Video, SD3, FLUX) and transformers. Core kernels: RMSNorm, RoPE 3D, GEGLU, AdaLN. Includes XCD swizzle, autotune, diffusers integration patterns, and LTX-Video pipeline injection.
日本語の概要は準備中です。原文の説明を表示しています。
huggingface/kernels☆ 7652026年10月10日 更新
Kernel-version skew check (ADR-027). Reports manifest surface + manifest kernel + installed kernel + verdict (match/patch-diff/minor-diff/major-diff). Exits 1 on minor/major skew with a copy-pasteable `npm install @metaharness/kernel@X.Y.Z` next step. Exits 2 if no .harness/manifest.json at path.
日本語の概要は準備中です。原文の説明を表示しています。
ruvnet/metaharness☆ 6962026年10月10日 更新
Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel. Use for device-level bottleneck diagnosis, kernel-internal critical-path questions, or evidence that genuinely needs both; do not require all three modes.
日本語の概要は準備中です。原文の説明を表示しています。
alibaba/atrex-kernel-agent☆ 1672026年10月10日 更新
Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific kernel-internal timing question.
日本語の概要は準備中です。原文の説明を表示しています。
alibaba/atrex-kernel-agent☆ 1672026年10月10日 更新
Use when writing, optimizing, or benchmarking a C++ CPU kernel with AVX2 or AVX512 intrinsics for the Hugging Face kernels ecosystem. Not for CUDA kernels: use cuda.
日本語の概要は準備中です。原文の説明を表示しています。
OutlineDriven/odin-claude-plugin☆ 372026年9月29日 更新
Linux kernel exploitation playbook. Use when exploiting kernel vulnerabilities (UAF, OOB, race condition, type confusion) for privilege escalation via commit_creds, modprobe_path overwrite, or kernel ROP chains in CTF and real-world scenarios.
日本語の概要は準備中です。原文の説明を表示しています。
ShulkwiSEC/bb-huge☆ 242026年7月11日 更新
Analyzes malicious and vulnerable Windows kernel drivers (.sys) by parsing the PE for the native subsystem, identifying DriverEntry/IRP dispatch and IOCTL handlers, and flagging BYOVD and kernel-callback abuse. Activates for requests to analyze a Windows driver, examine a .sys sample, or assess a BYOVD/kernel driver threat.
日本語の概要は準備中です。原文の説明を表示しています。
meltedinhex/analyst-ai-pack☆ 212026年7月7日 更新
Analyzes bootkit and rootkit samples by identifying boot-process tampering (MBR/VBR/ UEFI), kernel-mode components, and stealth hooking techniques from static indicators. Activates for requests to analyze a bootkit or rootkit, examine MBR/UEFI tampering, or identify kernel-mode stealth components.
日本語の概要は準備中です。原文の説明を表示しています。
meltedinhex/analyst-ai-pack☆ 212026年7月7日 更新
Detect kernel-level rootkits in Linux memory dumps using Volatility3 linux plugins (check_syscall, lsmod, hidden_modules), rkhunter system scanning, and /proc vs /sys discrepancy analysis to identify hooked syscalls, hidden kernel modules, and tampered system structures.
日本語の概要は準備中です。原文の説明を表示しています。
andycungkrinx91/konoha☆ 92026年10月9日 更新
Detect kernel-level rootkits in Linux memory dumps using Volatility3 linux plugins (check_syscall, lsmod, hidden_modules), rkhunter system scanning, and /proc vs /sys discrepancy analysis to identify hooked syscalls, hidden kernel modules, and tampered system structures.
日本語の概要は準備中です。原文の説明を表示しています。
micsapp/micstec-skills☆ 42026年3月20日 更新