Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
Field-tested methodology and concrete recipes for training and operating large-scale LLM/VLM/multi-modal models end to end - choosing and benchmarking accelerators, storage and network; SLURM/Kubernetes orchestration; maximizing training throughput and fitting models in memory; diagnosing and surviving training instabilities, NaN/Inf, and hardware/job failures; checkpointing and fault tolerance; inference performance and memory; debugging multi-node/ multi-GPU hangs; and writing/running tests. Use when the user is training or fine-tuning large models, hits low TFLOPS/MFU, OOM, slow dataloading, a loss spike/divergence, a NCCL/InfiniBand or multi-node hang, node/GPU failures, checkpoint or preemption problems, storage/network bottlenecks, or needs to pick GPUs/cloud/file-systems or size inference latency/throughput. Distilled from "Machine Learning Engineering", the latest version of which can be found at https://github.com/stas00/ml-engineering The latest SKILL.md version can be found at https://github.com/stas00/ml-engineering/blob/master/skills/ml-engineering/SKILL.md
日本語の概要は準備中です。原文の説明を表示しています。
stas00/ml-engineering☆ 1.9万2026年10月8日 更新
Comprehensive guide for writing Apache Cassandra in-JVM distributed tests (dtests). Use when creating tests that simulate multi-node Cassandra clusters within a single JVM for faster integration testing. Covers cluster creation (single-node, multi-node, multi-datacenter), configuration (all cassandra.yaml parameters, features, network topology), instance lifecycle (startup/shutdown/restart), query execution, message filtering for failure scenarios, running code on instances, ClusterUtils utilities, and debugging classloader-related issues (serialization failures, same-class-different-classloader problems).
日本語の概要は準備中です。原文の説明を表示しています。
apache/cassandra☆ 1万2026年10月11日 更新
Start (or restart) a local multi-node Temps cluster using Docker-in-Docker — one control plane + 3 worker nodes, each a privileged DinD container running its own dockerd + `temps agent`, wired with the real multi-host overlay (VXLAN, compute_cidr allocation) via `tools/dev-cluster/` in whichever checkout/worktree you run it from. Invoke when the user says "start the temps cluster", "spin up a multi-node dev cluster", "test this with multiple workers", "bring up worker nodes locally", "docker in dind cluster", or wants to verify cross-node behavior (node targeting, cluster DNS, overlay networking, node join/mTLS) that a single-node `start-temps` server cannot exercise. Distinct from `start-temps` (single-node, native binary, port-slot based) and `start-temps-ee` (single-node EE binary) — this is the only path that actually has more than one node.
日本語の概要は準備中です。原文の説明を表示しています。
gotempsh/temps☆ 8332026年10月11日 更新
Advanced RuView capabilities — RuvSense multistatic sensing (attention-weighted fusion, geometric diversity, persistent field model), cross-viewpoint fusion across multiple nodes, RF tomography (ISTA L1 solver, voxel grids), longitudinal biomechanics drift, pre-movement intention signals, adversarial signal detection, and multistatic mesh security hardening. Use for research-grade or multi-node deployments.
日本語の概要は準備中です。原文の説明を表示しています。
ruvnet/RuView☆ 9.7万2026年10月11日 更新
Configure RuView — ESP32 sdkconfig variants, NVS provisioning, WiFi channel / MAC filter overrides (ADR-060), edge intelligence modules (ADR-041), sensing-server flags, multi-node mesh, and Cognitum Seed integration. Use when adjusting how a deployed RuView system behaves without changing code.
日本語の概要は準備中です。原文の説明を表示しています。
ruvnet/RuView☆ 9.7万2026年10月11日 更新
Internal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for reverse proxies, Host headers, forwarded headers, vhost routing, websocket upgrades, path-prefix rewriting, base-URL derivation, and multi-node route resolution. Use when the user asks which host or container serves a route, why a public-looking domain still belongs to the sandbox, how headers or proxies change behavior, or how a route resolves across proxy, container, and worker boundaries. Use only after `$ctf-sandbox-orchestrator` has already established sandbox assumptions and routed here.
日本語の概要は準備中です。原文の説明を表示しています。
zhaoxuya520/reverse-skill☆ 4.1万2026年9月22日 更新
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after recipe-runner brings a deployment up (especially disagg/multi-node) to confirm the KV transport is correct; use troubleshoot for diagnosing already-failed pods.
日本語の概要は準備中です。原文の説明を表示しています。
openai/plugins☆ 7,3922026年10月8日 更新
Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. Use after recipe-runner brings a deployment up (especially disagg/multi-node) to confirm the KV transport is correct; use troubleshoot for diagnosing already-failed pods.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5602026年10月10日 更新
Use when erlang distributed systems including node connectivity, distributed processes, global name registration, distributed supervision, network partitions, and building fault-tolerant multi-node applications on the BEAM VM.
日本語の概要は準備中です。原文の説明を表示しています。
benchflow-ai/skillsbench☆ 1,8372026年7月24日 更新
Distributed consensus algorithms and logical time for cloud and multi-node systems. Covers Lamport clocks, vector clocks, FLP impossibility, Paxos (basic, multi, fast), Raft, Viewstamped Replication, Byzantine fault tolerance basics, quorum reads/writes (N/R/W), leader election, and TLA+ specification style. Use when designing replicated state machines, picking a consensus protocol, reasoning about split-brain and quorum loss, or writing formal specs for distributed coordination.
日本語の概要は準備中です。原文の説明を表示しています。
Tibsfox/gsd-skill-creator☆ 712026年7月20日 更新
Internal downstream skill for ctf-sandbox-orchestrator. CTF-sandbox workflow for reverse proxies, Host headers, forwarded headers, vhost routing, websocket upgrades, path-prefix rewriting, base-URL derivation, and multi-node route resolution. Use when the user asks which host or container serves a route, why a public-looking domain still belongs to the sandbox, how headers or proxies change behavior, or how a route resolves across proxy, container, and worker boundaries. Use only after `$ctf-sandbox-orchestrator` has already established sandbox assumptions and routed here.
日本語の概要は準備中です。原文の説明を表示しています。
zhaoxuya520/AI-Fullstack-Delivery-Workflow☆ 532026年5月19日 更新
Instâncias GPU em nuvem reservadas e sob demanda para treinamento e inferência de ML. Use quando você precisar de instâncias GPU dedicadas com acesso SSH simples, sistemas de arquivos persistentes ou clusters multi-node de alto desempenho para treinamento em larga escala.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Orquestração de treinamento distribuído em clusters. Escala PyTorch/TensorFlow/HuggingFace do laptop para milhares de nós. Ajuste de hiperparâmetros integrado com Ray Tune, tolerância a falhas, escalabilidade elástica. Use ao treinar modelos massivos em múltiplas máquinas ou executar varreduras distribuídas de hiperparâmetros.
日本語の概要は準備中です。原文の説明を表示しています。
artubss/SKILLS-CLAUDE-CODE☆ 112026年5月17日 更新
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
日本語の概要は準備中です。原文の説明を表示しています。
ibragimov-oasis/vibe-coder☆ 22026年6月24日 更新
Reserved and on-demand GPU cloud instances for ML training and inference. Use when you need dedicated GPU instances with simple SSH access, persistent filesystems, or high-performance multi-node clusters for large-scale training.
日本語の概要は準備中です。原文の説明を表示しています。
Lord1Egypt/awesome-skill-forge☆ 22026年6月10日 更新