本文へ移動
cccskills
無料GitHub で公開

harness-stripping

Systematically remove one harness component at a time and measure impact, killing scaffolding that no longer earns its complexity.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md4.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Harness Stripping

Every harness component was added to compensate for a specific model failure. Models improve. Components don't retire themselves. The scaffolding that saved you on Sonnet 4.5 may be dead weight — or actively harmful — on Opus 4.6. Strip it deliberately, one piece at a time, and let evals tell you what still earns its keep.

Inspired by Prithvi's March 2026 harness post on evaluator-generator separation and the general "re-test your assumptions each model bump" discipline.

When to apply

  • A model upgrade just landed and your harness was tuned for the previous generation.
  • A component's justification is "we added this because the model used to do X" — and you haven't checked whether it still does X.
  • The harness has accreted over months and nobody remembers what half the machinery is for.
  • Cost or latency is climbing and you suspect redundant belt-and-suspenders layers.

Procedure

  1. Inventory the components. List every distinct piece of scaffolding: prompt sections, tool wrappers, post-hoc validators, retry loops, evaluator personas, structured-output enforcers, sandbox rules. One row per component. Note the failure mode each was added to prevent.

  2. Rank by suspicion. Put the components most likely to be obsolete at the top: anything added before the last two model bumps, anything targeting a failure mode you haven't seen recently, anything whose original justification is now folklore.

  3. Pick a baseline eval. You need a repeatable metric before you touch anything. Reuse an existing eval set if you have one; otherwise pick 20–60 tasks representative of production work. Record baseline score, cost, and wall-clock.

  4. Strip one component. Only one. Comment it out or gate it behind a flag — don't delete yet. Re-run the eval.

  5. Compare against baseline.

    • Score within noise, cost/latency down → the component is dead weight. Delete.
    • Score drops measurably → the component still earns its complexity. Restore and note what failure mode returned.
    • Score improves → the component was actively harmful. Delete and investigate why (often: over-constraining a now-capable model).
  6. Commit the delta. Land the strip (or the restore-with-notes) as its own commit. Do not batch multiple strips into one change — you lose the ability to attribute the score movement.

  7. Repeat for the next component. Re-establish baseline from the new state each round, not the original. Compounding strips have compounding effects.

Anti-patterns

  • Stripping two components at once — you can't tell which one mattered. Halve the signal, double the confusion.
  • Skipping the eval "because it's obviously safe to remove" — the harness accreted for reasons. Some are still real. Measure.
  • Deleting instead of gating on the first pass — you will want to A/B mid-review. Flag first, delete after the eval confirms.
  • Trusting anecdotes over the eval — "it feels better without it" is how load-bearing components get removed. If the eval doesn't show it, it isn't there.
  • Stripping components that guard safety, sandboxing, or cost caps — those aren't compensating for model weakness. Leave them.
  • Doing this on prod traffic — run against an eval set, not real users. The failure modes you're re-probing are exactly the ones that hurt users.

What to strip first

Highest yield in practice:

  • Structured-output enforcers layered on top of models that now emit valid JSON natively.
  • Multi-step "plan then execute" wrappers on tasks the model now one-shots.
  • Retry loops around tool calls that no longer flake.
  • Evaluator personas whose critiques the generator now anticipates on its own.
  • Verbose "remember to do X" prompt sections where X is now default behavior.

When NOT to apply

Don't strip mid-project on a live long-running run — you'll perturb sessions in flight. Do it between projects, or on a forked branch. Also skip if you don't have an eval you trust; stripping without measurement is guessing.

Related

  • [[shift-notes]] — record which components were stripped and when, so the next audit doesn't re-strip and re-restore the same piece.
  • [[adversarial-verify]] — the evaluator-generator pattern that may itself be a strip candidate on newer models.
  • [[broken-window-check]] — if you strip a component and the eval regresses in a specific way, that's your new broken window to hunt.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

a11y-pass

無料

Catch the accessibility failures that ship in almost every AI-built UI. Use after building any interactive component.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable). Read it alongside claude-progress.txt at session start — prose is for humans, JSON is for the loop.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Review a diff against the goal spec assuming the code is BROKEN. The reviewer that lives in the maker's head always agrees with itself — this pulls review into a hostile, separate pass. Invoke after every code change before marking work done.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Verify that an endpoint checks ownership, not just authentication. Use on any handler that reads or mutates user data.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Find the exact commit that introduced a bug. Use when something worked before and broke, and you don't know which change did it.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Before picking new work, smoke-test the last "completed" feature. If it's broken, revert and re-open it before touching anything else. Kills the "looks shipped, isn't shipped" bug across sessions.

日本語の概要は準備中です。原文の説明を表示しています。

Archive228/loopkit7552026年7月15日 更新

Archive228 のスキルをすべて見る

このスキルの問題を報告する