Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when dataframe code is slow or runs out of memory, when data outgrows pandas, or when writing pandas 3, Polars or DuckDB code — vectorisation, dtypes, Copy-on-Write, lazy and out-of-core execution, and measuring before and after
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Most slow data code is Python looping over rows, reading columns it never uses, or holding strings as generic objects. Fix those in the library the repository already uses before reaching for a new one; when the data truly outgrows memory, DuckDB and Polars process it lazily and out of core.
Core principle: measure first, pin behaviour with a test, optimise, measure again — and report both numbers.
start = time.perf_counter()
result = build_features(sample)
print(f"{time.perf_counter() - start:.2f}s, {result.memory_usage(deep=True).sum() / 1e6:.0f} MB")
Measure on a realistic sample (not 20 rows, not the full warehouse). Before rewriting a slow function, make sure a fixture test pins its output; the faster version must pass the same test. Put "before → after" timings and memory in your closing message.
pandas 3 makes Copy-on-Write the only mode and a dedicated string dtype the default. Code written for pandas 1.x/2.x can silently stop working:
# ❌ chained assignment — never updates df under Copy-on-Write
df["discount"][df["plan"] == "pro"] = 0.1
# ❌ inplace on a selected column — never updates df
df["age"].fillna(0, inplace=True)
# ✅ one indexing operation on the object you mean to change
df.loc[df["plan"] == "pro", "discount"] = 0.1
df["age"] = df["age"].fillna(0)
sub = df[mask]; sub["x"] = 1 changes sub only, with no SettingWithCopyWarning and no .copy() needed.str dtype (Arrow-backed when PyArrow is installed; missing values are NaN). df["s"].dtype == object is now False — use pd.api.types.is_string_dtype.# ❌ Python function per row
df["tier"] = df.apply(lambda r: "high" if r.spend > 1000 else ("mid" if r.spend > 100 else "low"), axis=1)
# ✅ vectorised
df["tier"] = np.select([df["spend"] > 1000, df["spend"] > 100], ["high", "mid"], default="low")
# ✅ or as an ordered categorical
df["tier"] = pd.cut(df["spend"], bins=[-np.inf, 100, 1000, np.inf], labels=["low", "mid", "high"])
| Instead of | Use |
|---|---|
iterrows, itertuples, apply(axis=1) | column arithmetic, np.where, np.select, pd.cut |
| a loop doing dict lookups | series.map(mapping) or a merge |
| groupby, then merge the result back | groupby(...)[col].transform("sum") |
| Python string functions per value | the .str accessor (fast on the Arrow-backed dtype) |
pd.concat inside a loop | collect frames in a list, concat once |
pd.read_parquet("events.parquet", columns=["user_id", "ts", "amount"], filters=[("ts", ">=", start)])
pd.read_csv(path, usecols=["user_id", "plan"], dtype={"plan": "category"})
category for low-cardinality strings; float32 for model features where the precision is enough.df.memory_usage(deep=True) before and after.con = duckdb.connect()
con.execute("SET memory_limit = '4GB'")
daily = con.execute(
"""
select user_id, date_trunc('day', ts) as day, sum(amount) as spend
from read_parquet('data/events/*/*.parquet', hive_partitioning = true)
where ts >= ?
group by all
""",
[start],
).df()
.df() returns pandas, .pl() Polars; COPY (select ...) TO 'out.parquet' (FORMAT parquet) writes without a round-trip through Python.? parameters; EXPLAIN ANALYZE shows where time goes.spend = (
pl.scan_parquet("data/events/*.parquet")
.filter(pl.col("ts") >= start)
.group_by("user_id")
.agg(pl.col("amount").sum().alias("spend"), pl.len().alias("n_events"))
.collect(engine="streaming")
)
scan_* + collect() lets the optimiser push filters and column selection into the read; .explain() prints the plan.pl.when(...).then(...).otherwise(...) and .over("user_id") replace row loops and groupby-merge-back.map_elements runs Python per value — the same trap as apply(axis=1).polars<2 stays there; the code above is written against the 1.x API — check the changelog before using it on 2.x.| Situation | Do |
|---|---|
| Repo uses pandas, data fits in memory | optimise the pandas code (sections 2–4) |
| Repo already uses DuckDB or Polars | use it for the heavy step |
| Data outgrows memory and the repo has neither | say so in a comment and propose DuckDB or Polars; do not add a second dataframe library in a feature task (rule stack-choice) |
| Heavy aggregation already in the warehouse | push it into SQL or dbt instead of pulling raw rows |
Convert between libraries once, at a boundary — not back and forth inside a loop.
apply(axis=1) "because it is readable".object dtype columns full of numbers.ChainedAssignmentError warning in the test output.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases
日本語の概要は準備中です。原文の説明を表示しています。
Use when the diff adds or changes an endpoint, resolver, RPC, job or query that takes an object id, a role check, a request binding or a tenant filter - BOLA/IDOR, function-level authorization, mass assignment and tenant scoping
日本語の概要は準備中です。原文の説明を表示しています。
Use on every UI change - semantic HTML, labels for controls, keyboard-navigable dialogs/menus, visible focus, and never color as the only signal
日本語の概要は準備中です。原文の説明を表示しています。
Use when a task changes any screen, form, dialog, menu or control - Lighthouse/axe scan of the changed screens, a keyboard walk, and the thresholds that fail a task
日本語の概要は準備中です。原文の説明を表示しています。
How to work a task returned with review, QA or UAT findings. Use when a task is in need_revision or PR review comments are in your context.
日本語の概要は準備中です。原文の説明を表示しています。
Use when deciding whether a request needs an analiz task before implementation - the conditions that require the architect's analysis versus going straight to implementation
日本語の概要は準備中です。原文の説明を表示しています。