本文へ移動
cccskills
無料GitHub で公開

run-trace

Append structured execution traces across operational, cognitive, and contextual surfaces with minimal overhead. Load when inspecting agent runs, logging tool calls and observations, enabling post-run debugging, or pairing with structured-planning step IDs. Also triggers on "trace this run", "log execution", "agent observability", "run log", or when fault-localize needs evidence. Default-on during multi-step plans. Traces live at .agent-loom/traces/ — git-ignored by default.

インストール方法を見る

含まれるファイル(4)

  • SKILL.md3.6 KB
  • references/examples.md614 B
  • references/TRACE-SCHEMA.md1.4 KB
  • scripts/trace_query.py2.5 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Run Trace

You append low-overhead, structured records per meaningful step — tool call, observation, reflection, error. Three surfaces: operational, cognitive, contextual. Never block execution to log.

Hard Rules

Append to .agent-loom/traces/<run-id>.jsonl — one JSON object per line per references/TRACE-SCHEMA.md. Reuse step_id from active structured-planning plan when present. Never persist secrets — refs only in input_ref / output_ref. Logging is append-only — never rewrite prior lines. Ensure .agent-loom/traces/ is gitignored in consumer projects.


Workflow

Step 1 — Allocate run_id

Format: YYYY-MM-DDTHH:MM:SSZ-<slug> at run start. Create empty jsonl file.

Step 2 — Append per step

After each meaningful event, append one record with correct surface:

EventSurface
Tool/command executedoperational
Plan/route/reflectioncognitive
cwd, branch, versionscontextual

Step 3 — On error

Set error field on the operational record. Continue tracing if run continues.

Step 4 — Query when needed

python3 .agents/skills/run-trace/scripts/trace_query.py <path> timeline
python3 .agents/skills/run-trace/scripts/trace_query.py <path> errors

Step 5 — Hand off

On failed run, pass trace path to fault-localize.


Gotchas

  • Pretty-printed JSON breaks JSONL — single-line objects only.
  • Logging raw env vars may leak secrets — whitelist keys only.
  • High-frequency polling loops — batch or sample; don't log every poll.

Output Format

## Run trace — [run_id]

Trace file: `.agent-loom/traces/[run_id].jsonl`
Records: N | Errors: N

Latest:
- [ts] [surface] [step_id] [action]

Query: trace_query.py [path] timeline

Examples

Teaser: 5-step plan run → 12 operational + 3 cognitive records → S3 error captured with exit code.

Full pairs: references/examples.md


Common Rationalizations

ExcuseReality
"Tracing is too heavy"One JSON line per step is cheap.
"I'll remember what happened"You won't — especially after revert.
"Chat history is enough"Not structured; can't query errors programmatically.
"Secrets in output are fine locally"Traces get committed, shared, uploaded — refs only.
"Skip cognitive surface"Route decisions are the hardest to debug.

Verification

  • Trace file exists and is valid JSONL
  • step_ids align with plan when planning active
  • No secrets in trace payloads
  • traces/ gitignored

Red Flags

  • Trace file with invalid JSON lines
  • Missing error field on failed tool steps
  • Secrets written to jsonl

Prune Log

Last pruned: 2026-07-05

  • Initial release from high-leverage skill spec (Skill 4 family)

Impact Report

Trace: [run_id] | Records: N | Errors: N | Path: .agent-loom/traces/...

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Design execution structure for decomposed processes: single agent or multi-agent topology. Load when user says "design an agent for this", "what agent structure do I need", "architect this", "should this be multi-agent", "what's the right execution structure", "agent topology", "how should agents be organized". Takes process-decomposer output as primary input. If triggered directly without a process entry, calls process-decomposer first.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Internal skill. Called by setup-evaluation after a PASS. Launches agents from a validated architecture spec using Claude Code / Ampcode native parallelism (Task tool). Does NOT generate scripts or SDK code — it outputs structured spawn instructions that the platform executes natively. Never invoked directly by the user. Never launches without a setup-evaluation PASS.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Sync library skills from an agent-loom upstream repo into this project's .agents/skills while preserving project-local and forked skills. Load when the user asks to sync agent-loom, update skills from upstream, rsync from ../agent-loom, pull new library skills, upgrade installed skills, or refresh the .agents folder without losing custom project skills. Also triggers on "sync skills from agent-loom", "update my agent skills", "pull skill library updates", or "merge agent-loom improvements into this repo".

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Instrument a shipped product's AI agents with tracing and observability so you can see what they did, why outputs happened, and what each run cost. Plain-language primer plus free-tier-first backend selection (Langfuse, Phoenix, LangSmith, Braintrust) and OpenTelemetry/OpenInference instrumentation. Load when the user asks to add observability, add tracing, instrument my agents, see what my agent is doing in production, set up Langfuse or Phoenix or LangSmith, debug why my agent gave a bad answer, or track LLM cost per request. Also fires when agent-system-architecture or setup-evaluation requires an observability plan for an agent-chain product. NOT for tracing the coding agent itself — that is run-trace. Precondition for runtime-learning-loop.

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

Run a structured retrospective after development-phase runs of your product's agents — interview the owner in plain language about what went well and poorly, draft ranked improvement hypotheses, then design and run small n=1/n=2 experiments with pre-declared success criteria, guardrails, stop conditions, and a cost/ROI kill-switch. Load when the user says how did that run go, retro this run, the agent output was bad, what should we improve, draft hypotheses, run a small experiment, or after repeated dev runs of an agentic system produce uneven quality. Priority: output quality over performance over cost, each with diminishing-returns stops. NOT a product A/B test (experimentation), NOT coding-agent harness repair (harness-evolution), NOT production-scale learning (runtime-learning-loop).

日本語の概要は準備中です。原文の説明を表示しています。

dvy1987/agent-loom32026年8月8日 更新

dvy1987 のスキルをすべて見る

このスキルの問題を報告する