Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when writing tests for data or ML code — what to test at each layer (transforms, schemas, properties, pipeline smoke, model-quality thresholds, determinism and parity), how to build small synthetic fixtures, and what not to test
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Data code is easy to "test" badly: a test that re-runs the pipeline and asserts the metric equals last week's 0.8312 breaks on every library upgrade and catches nothing. Good data tests pin behaviour on tiny, explicit fixtures where you know the right answer by hand, and assert model quality as a threshold on planted signal.
Core principle: test your logic, not the library; derive expected values by hand on a fixture small enough to read; assert ranges for anything learned.
| Layer | What it proves | Share of tests |
|---|---|---|
| Pure transforms | A cleaning rule, a feature, a metric gives the hand-derived answer | most |
| Schemas / contracts | Bad input is rejected with the right column named | one per rule |
| Properties (Hypothesis) | Invariants hold for many generated inputs | a few, on core transforms |
| Pipeline smoke | Entry point runs end-to-end on a fixture, writes a loadable artifact | one per entrypoint |
| Model quality | Beats the baseline by a margin on planted signal, fixed seed | one or two per model |
| Determinism and parity | Same seed → same output; saved/loaded and served predictions equal training-time ones | one each |
| Warehouse | dbt data and unit tests (sql-analytics-and-dbt) | per model |
Build them in the test or conftest.py. Never a production extract (PII, size, drift) and never the network.
def test_order_value_sums_lines_and_ignores_cancelled() -> None:
lines = pd.DataFrame({
"order_id": ["o1", "o1", "o2", "o3"],
"amount": [10.0, 5.0, 7.0, 3.0],
"status": ["ok", "ok", "cancelled", "ok"],
})
expected = pd.Series({"o1": 15.0, "o3": 3.0}, name="order_value")
expected.index.name = "order_id"
pd.testing.assert_series_equal(order_value(lines), expected)
Every fixture for a transform carries the cases that break data code: a null, a duplicate key, an unseen category, an empty frame, a single row, a boundary timestamp, a timezone-aware value. Each earns its own named test.
For model tests, plant a known signal with a seeded generator:
@pytest.fixture
def churn_fixture() -> tuple[pd.DataFrame, np.ndarray]:
rng = np.random.default_rng(7)
n = 400
X = pd.DataFrame({
"days_since_order": rng.integers(0, 120, n),
"orders_90d": rng.poisson(3, n),
"plan": rng.choice(["free", "pro"], n),
})
logit = 0.04 * X["days_since_order"] - 0.5 * X["orders_90d"] - 1.0
y = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int).to_numpy()
return X, y
Use them for invariants that hold for every input: totals are preserved, no NaN is introduced, outputs stay in range, a cleaning step is idempotent, a row-wise transform keeps the row count.
from hypothesis import given, strategies as st
from hypothesis.extra.pandas import column, data_frames, range_indexes
order_lines = data_frames(
columns=[
column("order_id", elements=st.sampled_from(["o1", "o2", "o3"])),
column("amount", elements=st.floats(0, 1e6)),
column("status", elements=st.sampled_from(["ok", "cancelled"])),
],
index=range_indexes(max_size=50),
)
@given(order_lines)
def test_order_value_total_equals_sum_of_ok_lines(lines: pd.DataFrame) -> None:
result = order_value(lines)
assert result.sum() == pytest.approx(lines.loc[lines["status"] == "ok", "amount"].sum())
assert (result >= 0).all()
Keep CI deterministic with a profile (settings.register_profile("ci", derandomize=True, deadline=None)) if the repository has none; follow its profile if it does.
def test_model_beats_prevalence_baseline_on_planted_signal(churn_fixture) -> None:
X, y = churn_fixture
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
score = cross_val_score(build_model(seed=0), X, y, cv=cv, scoring="average_precision").mean()
assert score > y.mean() + 0.15
def test_train_writes_a_loadable_model(tmp_path, churn_fixture) -> None:
X, y = churn_fixture
path = train(X, y, out_dir=tmp_path, seed=0)
model = load_model(path)
assert model.predict_proba(X.head(5)).shape == (5, 2)
def test_training_is_deterministic_for_a_seed(churn_fixture) -> None:
X, y = churn_fixture
a = build_model(seed=0).fit(X, y).predict_proba(X)
b = build_model(seed=0).fit(X, y).predict_proba(X)
np.testing.assert_array_equal(a, b)
Also: saved-then-loaded model predicts exactly what the in-memory one did; the serving endpoint returns what pipeline.predict_proba returns for the same fixture row (model-serving-and-monitoring); a PyTorch model overfits one tiny batch (deep-learning-pytorch).
pytest.approx, np.testing.assert_allclose(actual, expected, rtol=1e-6), pd.testing.assert_frame_equal(..., check_exact=False, rtol=1e-6).assert_frame_equal, not a handful of cells; decide check_dtype and check_like (column order) deliberately.> baseline + margin, 0 <= p <= 1), never exact values.StandardScaler scales or pandas' groupby groups — test your use of them.The default suite runs in seconds: fixtures of tens to a few hundred rows, n_estimators and epochs turned down through the same config the code reads, n_jobs=1 where parallel start-up dominates. A slow suite gets skipped, and a skipped suite catches nothing.
score == 0.8312.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when the diff adds or changes an endpoint, resolver, RPC, job or query that takes an object id, a role check, a request binding or a tenant filter - BOLA/IDOR, function-level authorization, mass assignment and tenant scoping
日本語の概要は準備中です。原文の説明を表示しています。
Use on every UI change - semantic HTML, labels for controls, keyboard-navigable dialogs/menus, visible focus, and never color as the only signal
日本語の概要は準備中です。原文の説明を表示しています。
Use when a task changes any screen, form, dialog, menu or control - Lighthouse/axe scan of the changed screens, a keyboard walk, and the thresholds that fail a task
日本語の概要は準備中です。原文の説明を表示しています。
How to work a task returned with review, QA or UAT findings. Use when a task is in need_revision or PR review comments are in your context.
日本語の概要は準備中です。原文の説明を表示しています。
Use when deciding whether a request needs an analiz task before implementation - the conditions that require the architect's analysis versus going straight to implementation
日本語の概要は準備中です。原文の説明を表示しています。