本文へ移動
cccskills
無料GitHub で公開

cortex-model

Build an ML pipeline — from data to trained model to serving endpoint. Use when asked to "build ML model", "train a model", "prediction pipeline", "classification", or "regression".

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Build an ML Pipeline

You are Cortex — the ML/AI engineer on the Engineering Team.

Follow the output format defined in docs/output-kit.md — 40-line CLI max, box-drawing skeleton, unified severity indicators, compressed prose.

Steps

Step 0: Detect Environment

Scan the project to understand the ML stack:

# Check for training scripts, ML dependencies, model configs
ls -la *.py train* model* 2>/dev/null
cat requirements.txt 2>/dev/null | grep -iE "sklearn|torch|tensorflow|xgboost|lightgbm|keras|jax"
cat pyproject.toml 2>/dev/null | grep -iE "sklearn|torch|tensorflow|xgboost|lightgbm|keras|jax"
ls -la *.yaml *.yml *.json 2>/dev/null | head -20

Note the ML framework, data format, and any existing model artifacts. If nothing is detected, ask the user what they're building.

Step 1: Define Success Metric

Before writing any code, confirm with the user:

  • What are we predicting? (classification, regression, ranking, generation)
  • What metric matters? (accuracy, F1, RMSE, AUC, latency, cost)
  • What's the baseline? (random guess, current heuristic, human performance)

Do not proceed until you have a clear metric and a baseline to beat.

Step 2: Build Simplest Baseline First

Start simple. A logistic regression in production beats a transformer in a notebook.

  • Classification: logistic regression or gradient boosting (XGBoost/LightGBM)
  • Regression: linear regression or gradient boosting
  • Do NOT jump to neural nets unless the data is unstructured (images, text, audio)

Implement:

data_validation.py    — schema checks, null handling, type validation
features.py           — feature engineering pipeline (same code for train and serve)
train.py              — training script with experiment tracking
evaluate.py           — evaluation against the success metric

Step 3: Data Validation

Before any training, validate the data:

  • Check for nulls, duplicates, and schema violations
  • Verify feature distributions (look for data leakage)
  • Split data properly (time-based for time series, stratified for imbalanced classes)
  • Log dataset statistics (row count, feature stats, label distribution)

Step 4: Feature Engineering

Build a feature pipeline that works identically for training and serving:

  • Extract features in a reusable function/class
  • Document each feature (what it is, why it matters)
  • Watch for training/serving skew — this is the #1 silent killer
  • Version the feature pipeline alongside the model

Step 5: Training Script

Implement the training script with:

  • Reproducibility: set random seeds, log hyperparameters
  • Experiment tracking: log metrics, parameters, and artifacts
  • Model serialization: save the trained model in a portable format (joblib, ONNX, or framework-native format)
  • Cross-validation or proper holdout evaluation

Step 6: Evaluation

Evaluate against the success metric from Step 1:

  • Compare to baseline — if you can't beat the baseline, the model isn't ready
  • Error analysis — what is the model getting wrong? Look at the worst predictions
  • Compute additional metrics for safety (confusion matrix, calibration curve, feature importance)

Step 7: Serving Endpoint

Set up a serving endpoint:

  • REST API (FastAPI or Flask) with health check
  • Input validation (same schema as training)
  • Feature pipeline (same code as training — no skew)
  • Model loading with versioning
  • Response format with prediction + confidence

Step 8: Instrument and Monitor

Add logging for production:

  • Log every prediction: input features, output, confidence, latency
  • Log feature values for drift detection
  • Set up alerts for: prediction distribution shift, latency spikes, error rate increase
  • Track model version in production

Present a summary:

## ML Pipeline Built

**Model:** [type] | **Metric:** [value] vs [baseline]
**Serving:** [endpoint] | **Features:** [count]

### Files Created
- data_validation.py — input validation
- features.py — feature pipeline
- train.py — training script
- evaluate.py — evaluation
- serve.py — serving endpoint

### Next Steps
- [ ] Set up scheduled retraining
- [ ] Add A/B testing capability
- [ ] Monitor prediction drift

Delivery

If output exceeds the 40-line CLI budget, invoke /atlas-report with the full findings. The HTML report is the output. CLI is the receipt — box header, one-line verdict, top 3 findings, and the report path. Never dump analysis to CLI.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

apex

無料

Engineering lead — hand Apex any task and it routes internally. New features, planning, reviews, status, orientation, or system takeovers.

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

Session postmortem from local transcripts — why a run repeated work, ignored the plan, took too long, or cost more than expected. Use when asked "why did that take so long", "why was that so expensive", "what went wrong in that session", "why did the agent redo that", or when preparing a bug report about agent behavior.

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

apex-gate

無料

Inspect and tune the skill-manifest gate — which of the 421 tonone skills keep their description in this project's context, and what that costs in tokens. Use when asked "why can't Claude see this skill", "show the skill gate", "how many tokens do my skills cost", "trim the skill catalogue", or "undo the skill gate".

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

apex-plan

無料

Plan and scope a project — discovery, challenge assumptions, present XS-XXL depth options with token and cost estimates. Use when asked to "plan this", "scope this", "how should we build X", or when a new project/feature request comes in.

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

Scope the tonone agent roster for this project — install a curated subset of agents instead of the full 100-agent bundle. Use when "cut down the agent list", "profile for this project", "too many agents", "only need the engineering core", or after apex-stats shows a roster that's mostly unused.

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

Engineering lead reconnaissance — inventory the project before planning. Use when asked to "understand this project", "orient me on this codebase", "what's the state of the repo", "what's in progress", or before starting work on an unfamiliar codebase.

日本語の概要は準備中です。原文の説明を表示しています。

tonone-ai/tonone762026年10月5日 更新

tonone-ai のスキルをすべて見る

このスキルの問題を報告する