本文へ移動
cccskills

スキルを探す

2 件(stas00 のリポジトリ) ・ 人気順

概要と使いどころ

Field-tested methodology and concrete recipes for training and operating large-scale LLM/VLM/multi-modal models end to end - choosing and benchmarking accelerators, storage and network; SLURM/Kubernetes orchestration; maximizing training throughput and fitting models in memory; diagnosing and surviving training instabilities, NaN/Inf, and hardware/job failures; checkpointing and fault tolerance; inference performance and memory; debugging multi-node/ multi-GPU hangs; and writing/running tests. Use when the user is training or fine-tuning large models, hits low TFLOPS/MFU, OOM, slow dataloading, a loss spike/divergence, a NCCL/InfiniBand or multi-node hang, node/GPU failures, checkpoint or preemption problems, storage/network bottlenecks, or needs to pick GPUs/cloud/file-systems or size inference latency/throughput. Distilled from "Machine Learning Engineering", the latest version of which can be found at https://github.com/stas00/ml-engineering The latest SKILL.md version can be found at https://github.com/stas00/ml-engineering/blob/master/skills/ml-engineering/SKILL.md

日本語の概要は準備中です。原文の説明を表示しています。

stas00/ml-engineering1.9万2026年10月8日 更新

Evaluates an ML GPU cluster for a cloud trial or acceptance test: environment dump, isolated newest PyTorch, matmul FLOPS (MAMF/MSMF) on every GPU while the others compute, intra-node all-reduce bandwidth and per-call latency of every collective, the same inter-node on every node you were given (omit those sections if there is only one node), fio on local disk and shared FS, dated markdown report. Use when the user asks to evaluate a cluster, kick the tires on trial nodes, run cluster acceptance, or measure GPU/network/storage. Canonical copy: https://github.com/stas00/ml-engineering/blob/master/skills/evaluate-cluster/SKILL.md

日本語の概要は準備中です。原文の説明を表示しています。

stas00/ml-engineering1.9万2026年10月8日 更新