Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when a task touches a Jupyter or marimo notebook, asks to productionise or schedule notebook logic, or when you are tempted to do the work in a notebook — extract logic into tested modules, keep the notebook thin and reproducible, and commit it the way the repo does
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Notebooks are good for looking at data and bad at being software: hidden state from out-of-order cells, no tests, diffs full of JSON and base64 images. In this role a notebook is never the deliverable. Logic moves into an importable module with tests; the notebook imports it and shows results (rule no-notebook-only-logic).
Core principle: if a test, a pipeline or a service needs it, it lives in src/; the notebook only calls it.
Find out whether the notebook works today, on a fresh kernel, top to bottom:
mkdir -p /tmp/tt-<task key>
uv run jupyter nbconvert --to notebook --execute notebooks/churn.ipynb --output-dir /tmp/tt-<task key>/
# or, when the notebook has a parameters cell
uv run papermill notebooks/churn.ipynb /tmp/tt-<task key>/churn.out.ipynb -p sample_frac 0.01
Record the result: runs, fails at cell N, or cannot run here (needs warehouse credentials, a GPU, a file not in the repo). If it could not be executed, say so explicitly in the closing message — never describe a notebook as working because it looks right.
For each cell that computes something (load, clean, feature, train, evaluate):
# ❌ notebook cell — nothing can test or reuse this
df = pd.read_csv("/Users/ana/data/orders.csv")
df = df[df.status != "test"]
df["value"] = df.qty * df.price
monthly = df.groupby(df.ordered_at.str[:7]).value.sum()
# ✅ src/shop/metrics.py
def monthly_revenue(orders: pd.DataFrame) -> pd.Series:
real = orders[orders["status"] != "test"]
value = real["qty"] * real["price"]
month = real["ordered_at"].dt.to_period("M")
return value.groupby(month).sum().rename("revenue")
# ✅ tests/test_metrics.py
def test_monthly_revenue_excludes_test_orders_and_groups_by_month() -> None:
orders = pd.DataFrame({
"status": ["ok", "test", "ok"],
"qty": [2, 1, 1],
"price": [5.0, 100.0, 3.0],
"ordered_at": pd.to_datetime(["2026-01-31", "2026-01-15", "2026-02-01"]),
})
assert monthly_revenue(orders).to_dict() == {
pd.Period("2026-01", "M"): 10.0,
pd.Period("2026-02", "M"): 3.0,
}
# ✅ notebook cell afterwards
orders = load_orders(Settings())
monthly_revenue(orders).plot.bar()
Bugs you find while extracting (a string-sliced month, a filter that drops nulls) get their own failing test first, and a line in the closing message — the numbers the notebook used to show change.
!pip install, no %cd, no absolute paths, no credentials — configuration comes from settings or environment variables.| Repo uses | Do |
|---|---|
| nbstripout (git filter or pre-commit hook) | commit with outputs stripped; never bypass the hook |
| jupytext pairing | edit either side, then jupytext --sync notebooks/churn.ipynb; review the .py diff |
| marimo | notebooks are plain .py files: marimo edit, run as a script with python notebooks/churn.py, convert with marimo convert churn.ipynb > churn.py only if the task asks |
| nothing yet | strip outputs before committing and say so; adding a notebook tool is a separate decision (rule repo-conventions-win) |
If a notebook must run on a schedule, prefer turning it into a pipeline entrypoint that calls the module (data-pipeline-orchestration). Where the repository already runs notebooks with papermill, keep a single cell tagged parameters, keep the logic in the package, and write outputs to a path derived from the parameters.
When the repository executes notebooks in CI (pytest --nbmake notebooks/ or an nbconvert step), keep that green: the notebook must run on the fixture or a tiny sample, fast, without credentials. Otherwise the notebook's logic is covered by the module tests, and the notebook itself is checked by running it once as in section 1.
%run.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when the diff adds or changes an endpoint, resolver, RPC, job or query that takes an object id, a role check, a request binding or a tenant filter - BOLA/IDOR, function-level authorization, mass assignment and tenant scoping
日本語の概要は準備中です。原文の説明を表示しています。
Use on every UI change - semantic HTML, labels for controls, keyboard-navigable dialogs/menus, visible focus, and never color as the only signal
日本語の概要は準備中です。原文の説明を表示しています。
Use when a task changes any screen, form, dialog, menu or control - Lighthouse/axe scan of the changed screens, a keyboard walk, and the thresholds that fail a task
日本語の概要は準備中です。原文の説明を表示しています。
How to work a task returned with review, QA or UAT findings. Use when a task is in need_revision or PR review comments are in your context.
日本語の概要は準備中です。原文の説明を表示しています。
Use when deciding whether a request needs an analiz task before implementation - the conditions that require the architect's analysis versus going straight to implementation
日本語の概要は準備中です。原文の説明を表示しています。