本文へ移動
cccskills
無料GitHub で公開

harness-evolution

Improve agent reliability over time — diagnose why agents fail and fix the setup. Triggers on: agent keeps failing, same mistake again, agent not improving, make agent smarter, agent quality plateau, agents ignore skills, agent skips tests, fix agent behavior, agent unreliable, improve agent setup, self-improving harness, agents worse over time, tune agent instructions, agent going in circles, agent ignores AGENTS.md, repeated agent errors. Requires harness v0 and eval harness. AUTO-ROUTED from harness-engineering on symptoms. Not first setup — harness-generation first.

インストール方法を見る

含まれるファイル(4)

  • SKILL.md7.0 KB
  • references/diagnosis-etclovg.md3.5 KB
  • references/evolution-loop.md4.2 KB
  • references/examples.md1.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Harness Evolution

You close the harness improvement loop: execute → trace → diagnose layer → propose minimal edit → regression validate → promote or reject. Model weights are out of scope.

Hard Rules

Never run an evolution round without harness vN manifest and operational eval harness. Never propose an edit without trace evidence tied to a failure cluster. Never accept an edit without regression gate: held-in Δ ≥ 0 AND held-out Δ ≥ 0 AND max(Δ) > 0 (Self-Harness). Never make outcome-only edits — attribute failure to an ETCLOVG layer first (HarnessFix). Never allow evolve agent to modify verifier config, eval held-out tasks, or LLM API keys (AHE sandbox). Never promote prompt-only changes when tools/middleware/skills are the diagnosed layer (AHE ablation). Never bypass file-scope guard — edits only to paths declared in manifest allowed_write.


Workflow

Step 0 — Preconditions (mandatory)

Verify:

  1. docs/harness/manifest.json exists (else → harness-generation).
  2. docs/harness/eval-interface.md + regression task set defined (else → eval-rubric-design → eval-pipeline).
  3. Held-out split documented — never fed to proposer (Self-Harness, Meta-Harness).

FAIL fast with specific route if any missing.

Step 1 — Capture traces

Collect from: benchmark runs, docs/memory/agent-handoffs.md, session logs, or docs/harness/runs/iteration_NNN/. Distill to layered digest per AHE experience observability — raw millions of tokens are not fed to the proposer.

Step 2 — Diagnose (ETCLOVG + HTIR)

Per references/diagnosis-etclovg.md:

  • Normalize traces to step-level nodes (HarnessFix HTIR pattern).
  • Attribute each failure cluster to one primary layer: Execution, Tooling, Context, Lifecycle, Observability, Verification, Governance.
  • Consolidate recurring flaws into actionable records — one mechanism per record.

Step 3 — Propose diverse-minimal candidates

Generate K candidate edits (default K=3), each:

  • Tied to one failure mechanism (Self-Harness).
  • Scoped to manifest allowed_write paths (metaharness scope guard).
  • Documented with evidence quad (AHE): failure evidence, root cause, targeted fix, predicted impact.

Write docs/harness/evolve/change_manifest.json before evaluation.

Step 4 — Regression validate

Invoke eval-pipeline (harness regression mode) on held-in + held-out splits.

GateRule
Self-Harness acceptanceheld-in Δ ≥ 0 AND held-out Δ ≥ 0 AND improvement > 0
ScopeNo files outside allowed_write
No-changeZero file changes → inherit parent scores, do not promote (metaharness)
pass@1Optimize pass@1, not pass@k flaky strategies (AHE)

Step 5 — Promote or reject

Accept: bump manifest version to vN+1, update hashes, archive run under docs/harness/runs/. Reject: log predicted-vs-actual in manifest; if same flaw persists 2+ rounds at same layer → rollback component and pivot layer (AHE).

Optional label-free path (RHO): when no labeled eval exists, use self-consistency + pairwise self-preference among candidates — still require positive mean score before promote.

Step 6 — Memory + handoff

On promote: memory-capture with harness version, delta metrics, and changed components. Append docs/skill-outputs/SKILL-OUTPUTS.md.


Gotchas

  • Compressed feedback loses credit assignment — never reduce traces to scalar score only (Meta-Harness).
  • Runtime supervision patches suppress errors without fixing harness flaws — reject as edits (HarnessFix).
  • Self-attribution misses regressions — manifest must list risk_tasks predicted to break (AHE).
  • Generic prompt bloat — every instruction must map to a diagnosed failure cluster.
  • Label-free RHO is fallback — prefer verifier-backed regression when labels exist.

Output Format

Harness evolution — round [N]
Diagnosed layer: [ETCLOVG]
Candidates: [K] | Accepted: [id or none]
Held-in Δ: [x] | Held-out Δ: [y]
Promoted: v[N] → v[N+1] | [rejected — reason]
Changed components: [list]
Next: [another round | reality-check claim audit]

Example

<examples> <example> <input>Agent keeps retrying the same failing tool call — improve the harness.</input> <output> Harness evolution — round 1 Diagnosed layer: Tooling (F6 tool-use loop) Candidates: 3 | Accepted: candidate-2 (middleware retry cap + alternate tool path) Held-in Δ: +2 | Held-out Δ: +1 Promoted: v0 → v1 | Changed: docs/harness/middleware.md, tool descriptions </output> </example> </examples>

Common Rationalizations

ExcuseReality
"Skip eval — vibes say it's better"No regression gate = unfalsifiable (reality-check)
"Fix in the prompt only"AHE: prompt-only regresses; fix diagnosed layer
"Use test failures as held-out"Contaminates proposer — splits are sacred
"One big harness rewrite"Diverse-minimal proposals beat monolithic edits
"Evolve without manifest"No attribution, no rollback

Verification

  • Preconditions verified (manifest + eval harness + held-out split)
  • Failure attributed to ETCLOVG layer with trace refs
  • change_manifest.json with evidence quad
  • Regression run via eval-pipeline
  • Promotion only if dual-split rule passes

Red Flags

  • Evolution round without eval harness
  • Held-out tasks leaked to proposer
  • Scope violations in edited files
  • Prompt bloat without failure mapping

Prune Log

Last pruned: 2026-07-05

  • Deep learn-from: evolution-loop, diagnosis-etclovg, examples L3 (5 papers + 5 repos)

Impact Report

Harness evolution round [N]: [accepted|rejected]
Layer: [ETCLOVG] | v[N]→v[N+1]
Held-out Δ: [x] | Components changed: [list]
eval-pipeline: [run id]

Reference Files

  • references/evolution-loop.md — full loop, auto-harness 3-step gate, filesystem artifact store
  • references/diagnosis-etclovg.md — HTIR nodes, layer attribution, flaw records
  • references/examples.md — accept, reject, and RHO fallback examples

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Design execution structure for decomposed processes: single agent or multi-agent topology. Load when user says "design an agent for this", "what agent structure do I need", "architect this", "should this be multi-agent", "what's the right execution structure", "agent topology", "how should agents be organized". Takes process-decomposer output as primary input. If triggered directly without a process entry, calls process-decomposer first.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Internal skill. Called by setup-evaluation after a PASS. Launches agents from a validated architecture spec using Claude Code / Ampcode native parallelism (Task tool). Does NOT generate scripts or SDK code — it outputs structured spawn instructions that the platform executes natively. Never invoked directly by the user. Never launches without a setup-evaluation PASS.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Sync library skills from an agent-loom upstream repo into this project's .agents/skills while preserving project-local and forked skills. Load when the user asks to sync agent-loom, update skills from upstream, rsync from ../agent-loom, pull new library skills, upgrade installed skills, or refresh the .agents folder without losing custom project skills. Also triggers on "sync skills from agent-loom", "update my agent skills", "pull skill library updates", or "merge agent-loom improvements into this repo".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Instrument a shipped product's AI agents with tracing and observability so you can see what they did, why outputs happened, and what each run cost. Plain-language primer plus free-tier-first backend selection (Langfuse, Phoenix, LangSmith, Braintrust) and OpenTelemetry/OpenInference instrumentation. Load when the user asks to add observability, add tracing, instrument my agents, see what my agent is doing in production, set up Langfuse or Phoenix or LangSmith, debug why my agent gave a bad answer, or track LLM cost per request. Also fires when agent-system-architecture or setup-evaluation requires an observability plan for an agent-chain product. NOT for tracing the coding agent itself — that is run-trace. Precondition for runtime-learning-loop.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Run a structured retrospective after development-phase runs of your product's agents — interview the owner in plain language about what went well and poorly, draft ranked improvement hypotheses, then design and run small n=1/n=2 experiments with pre-declared success criteria, guardrails, stop conditions, and a cost/ROI kill-switch. Load when the user says how did that run go, retro this run, the agent output was bad, what should we improve, draft hypotheses, run a small experiment, or after repeated dev runs of an agentic system produce uneven quality. Priority: output quality over performance over cost, each with diminishing-returns stops. NOT a product A/B test (experimentation), NOT coding-agent harness repair (harness-evolution), NOT production-scale learning (runtime-learning-loop).

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

dvy1987 のスキルをすべて見る

このスキルの問題を報告する