本文へ移動
cccskills
無料GitHub で公開

harness-engineering

Orchestrator for agent harness work — the setup that makes AI agents follow project rules and improve when they fail. FIRES PROACTIVELY when agents misbehave, repeat mistakes, ignore instructions, skip skills, or when AGENTS.md exists but docs/harness/manifest.json is missing. Also triggers on: harness engineering, agent scaffold, agent keeps failing, agent not following instructions, make agents reliable, agents going off rails, agent forgot context, improve agent setup, self-improving agents, agents keep making mistakes, why is my agent bad, agent quality, agent setup broken, agents ignore skills, same mistake again, fix agent behavior, tune agent instructions, set up agent infrastructure, after project setup agents still bad. Routes bootstrap vs evolution. Not multi-agent topology — agent-builder.

インストール方法を見る

含まれるファイル(4)

  • SKILL.md6.5 KB
  • references/examples.md857 B
  • references/harness-readiness-gate.md2.9 KB
  • references/routing.md2.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Harness Engineering

You are the harness orchestrator. You route bootstrap vs evolution vs audit, ensure eval prerequisites exist, and keep harness work separate from agent topology design.

Trigger Discipline

Proactive — user need not say "harness." Fire when: agent misbehavior symptoms; AGENTS.md without docs/harness/manifest.json; post-project-setup backfill; self-improvement claims. Full symptom table: references/harness-readiness-gate.md.

Hard Rules

Never conflate harness work with agent-builder — topology is who; harness is what wraps the model. Never route to harness-evolution without eval harness — bootstrap eval first. Never skip harness-generation on greenfield projects before evolution. Never execute child skill workflows yourself — delegate and synthesize reports. Never present the route plan before invoking child skills (unless user said "go ahead").


Workflow

Step 0 — Harness readiness scan (mandatory)

Read references/harness-readiness-gate.md. Silent: manifest exists? eval interface? symptoms? No manifest → bootstrap. Manifest + symptoms → evolution. Self-improvement claim → reality-check.

Step 1 — Classify intent

SignalRoute
New project / no docs/harness/manifest.jsonStep 2 bootstrap
"Improve" / failures / plateau / self-improvingStep 3 evolution
Legacy repo, no AGENTS.mdretroactive-project-setup → bootstrap
"Is it self-improving?" / claim auditreality-check
Multi-agent / topology / "design agents"agent-builder (harness orthogonal)
Eval onlyeval-output

Read references/routing.md for disambiguation vs project-setup and project-orchestrator.

Step 2 — Bootstrap path

project-setup (if no AGENTS.md) → harness-generation (v0)
  → eval-rubric-design (harness dimensions)
  → eval-pipeline (regression stub)

Skip project-setup if populated AGENTS.md exists — merge via harness-generation only.

Step 3 — Evolution path

Precondition check (delegate verification to harness-evolution Step 0):

  • manifest exists
  • eval harness operational
  • held-out split defined

If missing: run bootstrap substeps, then harness-evolution.

Step 4 — Agent-chain coordination

When user is building multi-agent systems:

  1. harness-generation if no v0 (parallel-safe with process-decomposer).
  2. agent-builder for topology.
  3. setup-evaluation must PASS harness + eval checks before agent-launcher.

Step 5 — Unified report

Present child skill outputs in one summary (see Output Format).


Gotchas

  • Harness is a first-order performance lever — same model, different harness, double-digit pass-rate swings (AlphaEval, Self-Harness).
  • Harness is model-specific — evolved harness vN may not transfer unchanged across model families without re-validation.
  • PROGRAM.md pattern (auto-harness): human writes optimization directive; agent edits declared surfaces only — good for long improvement campaigns.
  • Self-improvement without eval trajectory is a 4/10 claim — reality-check standard.

Output Format

=== Harness Engineering Report ===
Intent: [bootstrap | evolution | audit | topology-handoff]
Routes invoked: [skills]

=== Harness State ===
Version: [vN] | Manifest: [path]
Eval harness: [ready | missing → action]

=== Child Results ===
[per-skill summaries]

=== Next ===
[recommended follow-up]

Example

<examples> <example> <input>Improve my agent harness — it keeps failing on lint steps.</input> <output> === Harness Engineering Report === Intent: evolution Routes: harness-evolution (round 1), eval-pipeline (regression)

=== Harness State === Version: v0 → v1 | Eval harness: ready

=== Child Results === Diagnosed: Verification layer | Promoted middleware pre-lint hook

=== Next === Run round 2 only if held-out plateaus </output> </example> </examples>

Common Rationalizations

ExcuseReality
"agent-builder handles harness"Topology only — run harness-generation
"Skip eval for quick fix"harness-evolution hard-fails — route eval first
"project-setup is enough"One-shot AGENTS.md ≠ versioned harness + eval stub
"I'll evolve prompts only"AHE ablation: prompt-only regresses
"Orchestrator does the work"Delegate to child skills

Verification

  • Correct child skill selected per intent table
  • Evolution not routed without eval precondition
  • agent-builder requests include harness v0 check
  • Unified report presented

Red Flags

  • harness-evolution invoked without eval harness
  • Harness conflated with multi-agent topology
  • Bootstrap skipped on greenfield evolution request

Prune Log

Last pruned: 2026-07-05

  • Deep learn-from: routing.md Pareto frontier; INGEST-QUEUE pairwise compares cleared

Impact Report

Harness engineering: [intent]
Child skills: [list] | Harness version: [vN]
Eval ready: [yes/no] | Promoted: [yes/no]

Reference Files

  • references/routing.md — disambiguation matrix, lifecycle map, PROGRAM.md pattern
  • references/harness-readiness-gate.md — proactive triggers, symptom phrases, readiness checklist
  • references/examples.md — bootstrap, evolution, and agent-chain coordination examples

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Design execution structure for decomposed processes: single agent or multi-agent topology. Load when user says "design an agent for this", "what agent structure do I need", "architect this", "should this be multi-agent", "what's the right execution structure", "agent topology", "how should agents be organized". Takes process-decomposer output as primary input. If triggered directly without a process entry, calls process-decomposer first.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Internal skill. Called by setup-evaluation after a PASS. Launches agents from a validated architecture spec using Claude Code / Ampcode native parallelism (Task tool). Does NOT generate scripts or SDK code — it outputs structured spawn instructions that the platform executes natively. Never invoked directly by the user. Never launches without a setup-evaluation PASS.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Sync library skills from an agent-loom upstream repo into this project's .agents/skills while preserving project-local and forked skills. Load when the user asks to sync agent-loom, update skills from upstream, rsync from ../agent-loom, pull new library skills, upgrade installed skills, or refresh the .agents folder without losing custom project skills. Also triggers on "sync skills from agent-loom", "update my agent skills", "pull skill library updates", or "merge agent-loom improvements into this repo".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Instrument a shipped product's AI agents with tracing and observability so you can see what they did, why outputs happened, and what each run cost. Plain-language primer plus free-tier-first backend selection (Langfuse, Phoenix, LangSmith, Braintrust) and OpenTelemetry/OpenInference instrumentation. Load when the user asks to add observability, add tracing, instrument my agents, see what my agent is doing in production, set up Langfuse or Phoenix or LangSmith, debug why my agent gave a bad answer, or track LLM cost per request. Also fires when agent-system-architecture or setup-evaluation requires an observability plan for an agent-chain product. NOT for tracing the coding agent itself — that is run-trace. Precondition for runtime-learning-loop.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Run a structured retrospective after development-phase runs of your product's agents — interview the owner in plain language about what went well and poorly, draft ranked improvement hypotheses, then design and run small n=1/n=2 experiments with pre-declared success criteria, guardrails, stop conditions, and a cost/ROI kill-switch. Load when the user says how did that run go, retro this run, the agent output was bad, what should we improve, draft hypotheses, run a small experiment, or after repeated dev runs of an agentic system produce uneven quality. Priority: output quality over performance over cost, each with diminishing-returns stops. NOT a product A/B test (experimentation), NOT coding-agent harness repair (harness-evolution), NOT production-scale learning (runtime-learning-loop).

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

dvy1987 のスキルをすべて見る

このスキルの問題を報告する