本文へ移動
cccskills
無料GitHub で公開

ml-problem-framing

Use when a task asks for a model, a score, a forecast or a prediction-driven feature — pin down the decision, target, unit, prediction time, metric and baseline before writing any modelling code

インストール方法を見る

含まれるファイル(1)

  • SKILL.md5.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

ML Problem Framing

Overview

Most failed models were well trained on the wrong question: a target that leaks the outcome, a metric nobody acts on, a population that differs from the one scored in production, or no baseline to show the model is worth its upkeep. Framing is a short written spec, done before the first line of model code.

Core principle: a model exists to change a decision. If you cannot name the decision, the cost of each kind of error and the baseline the model must beat, you are not ready to train.

1. The framing card

Fill this in from the task, the code and the data. Put it in your closing message, and in the repository's model card or README section when the repo keeps one.

FieldQuestionExample
DecisionWhat action changes based on the output?Send a retention offer to the top 5% risk each Monday
TargetExact definition, as code or SQLchurned = no paid order in the 60 days after snapshot_date
UnitOne row is one what?customer × weekly snapshot
Prediction timeWhen is the prediction made, what is known then?Monday 00:00 UTC, data up to Sunday 23:59
Horizon / label windowHow far ahead, and when is the label final?60 days; labels final 60 days after snapshot
PopulationWho is scored, who is excluded?Active paying customers; exclude staff and test accounts
MetricWhich number decides, tied to error costs?Precision@5% (offer budget is fixed)
BaselineWhat it must beat"days since last order" rule; current heuristic
ConstraintsLatency, interpretability, fairness, retrain cadenceWeekly batch; reasons per customer for the CRM team

Anything in this table the task does not settle and the data cannot answer — what counts as churn, which population — is a product question: numbered questions via add_task_comment, then stop on it. Choosing between two reasonable technical options (a 60- vs 90-day window when the task said "about two months") is a decision you make and record.

2. Choose the metric from the decision

SituationUseNot
Fixed budget of actions (top-k)precision@k, recall@k, lift@kaccuracy, ROC-AUC alone
Rare positive classPR-AUC (average precision), recall at fixed precisionaccuracy (99% by predicting "no")
Probabilities consumed downstream (pricing, expected value)log loss, Brier score + calibration curveranking metrics alone
Asymmetric error costsexpected cost with an explicit cost matrix; tune the threshold for itdefault 0.5 threshold
Regression with outliers that matter lessMAE, quantile lossRMSE
Forecasting several seriesMASE, or WAPE against a seasonal naive forecastMAPE near zero actuals

Write the metric as a tested function when the library does not provide it — a precision@k with an off-by-one is a wrong decision every week.

def precision_at_k(y_true: np.ndarray, scores: np.ndarray, k: int) -> float:
    top = np.argsort(-scores, kind="stable")[:k]
    return float(y_true[top].mean())

def test_precision_at_k_counts_only_top_k() -> None:
    y = np.array([1, 0, 1, 0])
    s = np.array([0.9, 0.8, 0.1, 0.7])
    assert precision_at_k(y, s, 2) == 0.5

3. Baselines first

Before any learned model, compute the baseline on the same split and metric:

from sklearn.dummy import DummyClassifier
baseline = DummyClassifier(strategy="prior").fit(X_train, y_train)

plus the simple rule a domain expert would use (sort by recency, last week's value, the seasonal naive forecast). If a gradient-boosted model beats the rule by a margin inside the noise, report that — "a rule is as good" is a valid, cheap result.

4. Prediction-time thinking prevents leakage

For every candidate feature, ask: at the prediction timestamp, in production, would this value already exist and would it have this value? refund_issued when predicting fraud, account_closed_at when predicting churn, an aggregate computed over the whole table — all fail this question. Build training rows with an as-of snapshot (leakage-safe-feature-engineering) so training and scoring see the world the same way.

5. Population and drift

The training population must match the scored population: same filters, same exclusions, same time span shape. If production scores new customers but training only has customers with 90 days of history, the model is evaluated on a population it never serves. Check the most recent period separately — a model that only works on old data is a drift problem waiting to be noticed.

Common Mistakes

  • Starting with model selection before the target is defined in code.
  • A target window that overlaps the feature window.
  • Reporting ROC-AUC for a top-5% campaign where only precision at the top matters.
  • No baseline, so "0.81 AUC" has no meaning.
  • Treating the 0.5 threshold as part of the model instead of a decision tuned on validation data.

Red Flags

  • A single feature with near-perfect predictive power.
  • The target column, or a column derived from it, appears in the feature list.
  • The task's success criterion is "build a model" with no metric — ask what decision it changes.
  • Train and score populations filtered by different code paths.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use when writing acceptance criteria for a task - express each as an observable Given/When/Then that QA can execute, including negative cases

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

Use when the diff adds or changes an endpoint, resolver, RPC, job or query that takes an object id, a role check, a request binding or a tenant filter - BOLA/IDOR, function-level authorization, mass assignment and tenant scoping

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

Use on every UI change - semantic HTML, labels for controls, keyboard-navigable dialogs/menus, visible focus, and never color as the only signal

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

Use when a task changes any screen, form, dialog, menu or control - Lighthouse/axe scan of the changed screens, a keyboard walk, and the thresholds that fail a task

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

How to work a task returned with review, QA or UAT findings. Use when a task is in need_revision or PR review comments are in your context.

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

Use when deciding whether a request needs an analiz task before implementation - the conditions that require the architect's analysis versus going straight to implementation

日本語の概要は準備中です。原文の説明を表示しています。

makifbaysal/tasktrooper1122026年10月10日 更新

makifbaysal のスキルをすべて見る

このスキルの問題を報告する