本文へ移動
cccskills
無料GitHub で公開

tinygrad

Deep learning framework development with tinygrad - a minimal tensor library with autograd, JIT compilation, and multi-device support. Use when writing neural networks, training models, implementing tensor operations, working with UOps/PatternMatcher for graph transformations, or contributing to tinygrad internals. Triggers on tinygrad imports, Tensor operations, nn modules, optimizer usage, schedule/codegen work, or device backends.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md4.6 KB
  • README.md729 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

tinygrad

A minimal deep learning framework focused on beauty and minimalism. Every line must earn its keep.

Quick Reference

from tinygrad import Tensor, TinyJit, nn, dtypes, Device, GlobalCounters

# Tensor creation
x = Tensor([1, 2, 3])
x = Tensor.rand(2, 3)
x = Tensor.kaiming_uniform(128, 784)

# Operations are lazy until realized
y = (x + 1).relu().sum()
y.realize()  # or y.numpy()

# Training context
with Tensor.train():
  loss = model(x).sparse_categorical_crossentropy(labels).backward()
  optim.step()

Architecture Pipeline

  1. Tensor (tinygrad/tensor.py) - User API, creates UOp graph
  2. UOp (tinygrad/uop/ops.py) - Unified IR for all operations
  3. Schedule (tinygrad/engine/schedule.py) - Converts tensor UOps to kernel UOps
  4. Codegen (tinygrad/codegen/) - Converts kernel UOps to device code
  5. Runtime (tinygrad/runtime/) - Device-specific execution

Training Loop Pattern

from tinygrad import Tensor, TinyJit, nn
from tinygrad.nn.datasets import mnist

X_train, Y_train, X_test, Y_test = mnist()
model = Model()
optim = nn.optim.Adam(nn.state.get_parameters(model))

@TinyJit
@Tensor.train()
def train_step():
  optim.zero_grad()
  samples = Tensor.randint(512, high=X_train.shape[0])
  loss = model(X_train[samples]).sparse_categorical_crossentropy(Y_train[samples]).backward()
  return loss.realize(*optim.schedule_step())

for i in range(100):
  loss = train_step()

Model Definition

Models are plain Python classes with __call__. No base class required.

class Model:
  def __init__(self):
    self.l1 = nn.Linear(784, 128)
    self.l2 = nn.Linear(128, 10)
  def __call__(self, x):
    return self.l1(x).relu().sequential([self.l2])

Available nn modules: Linear, Conv2d, BatchNorm, LayerNorm, RMSNorm, Embedding, GroupNorm, LSTMCell

Optimizers: SGD, Adam, AdamW, LARS, LAMB, Muon

State Dict / Weights

from tinygrad.nn.state import safe_save, safe_load, get_state_dict, load_state_dict, get_parameters

# Save/load safetensors
safe_save(get_state_dict(model), "model.safetensors")
load_state_dict(model, safe_load("model.safetensors"))

# Get all trainable params
params = get_parameters(model)

JIT Compilation

TinyJit captures and replays kernel graphs. Input shapes must be fixed.

@TinyJit
def forward(x):
  return model(x).realize()

# First call captures, subsequent calls replay
out = forward(batch)

Device Management

from tinygrad import Device
print(Device.DEFAULT)  # Auto-detected: METAL, CUDA, AMD, CPU, etc.

# Force device
x = Tensor.rand(10, device="CPU")
x = x.to("CUDA")

Environment Variables

VariableValuesDescription
DEBUG1-7Increasing verbosity (4=code, 7=asm)
VIZ1Graph visualization
BEAM#Kernel beam search width
NOOPT1Disable optimizations
SPEC1-2UOp spec verification

Debugging

# Visualize computation graph
VIZ=1 python -c "from tinygrad import Tensor; Tensor.ones(10).sum().realize()"

# Show generated code
DEBUG=4 python script.py

# Run tests
python -m pytest test/test_tensor.py -xvs

UOp and PatternMatcher (Internals)

UOps are immutable, cached graph nodes. Use PatternMatcher for transformations:

from tinygrad.uop.ops import UOp, Ops
from tinygrad.uop.upat import UPat, PatternMatcher, graph_rewrite

pm = PatternMatcher([
  (UPat(Ops.ADD, src=(UPat.cvar("x"), UPat.cvar("x"))), lambda x: x * 2),
])
result = graph_rewrite(uop, pm)

Key UOp properties: op, dtype, src, arg, tag

Define PatternMatchers at module level - they're slow to construct.

Style Guide

  • 2-space indentation, 150 char line limit
  • Prefer readability over cleverness
  • Never mix functionality changes with whitespace changes
  • All functionality changes must be tested
  • Run pre-commit run --all-files before commits

Testing

python -m pytest test/test_tensor.py -xvs
python -m pytest test/unit/test_schedule_cache.py -x --timeout=60
SPEC=2 python -m pytest test/test_something.py  # With spec verification

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use when the user requests integration testing, feature validation, or test plan execution

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

anneal

無料

Use when the user wants to systematically fix AI code slop — duplicated logic, over-engineering, silent error swallowing, convention drift, cargo-cult patterns, and other LLM-introduced architectural decay — over a specified duration

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

Produce a researched long-form article from a topic prompt via an orchestrated pipeline - research agent (first-person sources, working-definition gate), narrative-architecture outline, writer/cold-reviewer loop with an explicit ACCEPT/REVISE verdict contract, then a catalog-deslop pass with a regression gate. The orchestrator dispatches subagents only; the writer never judges its own draft. Use when the user says "article factory", "write an article about X", "run the article pipeline", or asks for a researched long-form piece produced end-to-end. For essays and micro posts in the user's own voice without a research stage, use the prose skill instead.

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

Runs autonomous keep/discard experiments on a codebase to optimize a single metric for a fixed duration, in the style of karpathy/autoresearch. Use when the user says "autoresearch" (optionally with a focus, e.g. "autoresearch the optimizer"), asks to run experiments on a repo overnight, to hill-climb or optimize a metric autonomously, or points at a repo with a karpathy-style program.md.

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

Create custom modules for [Harbor Boost](https://github.com/av/harbor/tree/main/boost), an optimizing LLM proxy. Use when building Python modules that intercept/transform LLM chat completions—reasoning chains, prompt injection, structured outputs, artifacts, or custom workflows. Triggers on requests to create Boost modules, extend LLM behavior via proxy, or implement chat completion middleware.

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

bugbash

無料

Systematically explore and test any software project (CLI, API, Backend, Library, etc.) to find bugs, usability issues, and edge cases. Produces a structured report with full reproduction evidence (exact commands, inputs, logs, and tracebacks) for every issue.

日本語の概要は準備中です。原文の説明を表示しています。

av/skills192026年10月9日 更新

av のスキルをすべて見る

このスキルの問題を報告する