本文へ移動
cccskills
無料GitHub で公開

graph-evolution

Compares Trailmark code graphs at two source code snapshots (git commits, tags, or directories) to surface security-relevant structural changes. Detects new attack paths, complexity shifts, blast radius growth, taint propagation changes, and privilege boundary modifications that text diffs miss. Use when comparing code between commits or tags, analyzing structural evolution, detecting attack surface growth, reviewing what changed between audit snapshots, or finding security-relevant changes that text diffs miss.

インストール方法を見る

含まれるファイル(6)

  • SKILL.md13.2 KB
  • agents/openai.yaml246 B
  • assets/trail-of-bits-mark.svg3.0 KB
  • references/evolution-metrics.md5.1 KB
  • references/report-format.md4.2 KB
  • scripts/graph_diff.py7.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Graph Evolution

Builds Trailmark code graphs at two source snapshots and computes a structural diff. Surfaces security-relevant changes that text-level diffs miss: new attack paths, complexity shifts, blast radius growth, taint propagation changes, and privilege boundary modifications.

When to Use

  • Comparing two git refs to understand what structurally changed
  • Auditing a range of commits for security-relevant evolution
  • Detecting new attack paths created by code changes
  • Finding functions whose blast radius or complexity grew silently
  • Identifying taint propagation changes across refactors
  • Pre-release structural comparison (tag-to-tag or branch-to-branch)

When NOT to Use

  • Line-level code review (use differential-review for text-diff analysis)
  • Single-snapshot analysis (use the trailmark skill directly)
  • Diagram generation from a single snapshot (use the diagramming-code skill)
  • Mutation testing triage (use the genotoxic skill)

Rationalizations to Reject

RationalizationWhy It's WrongRequired Action
"We just need the structural diff, skip pre-analysis"Without pre-analysis, you miss taint changes, blast radius growth, and privilege boundary shiftsRun engine.preanalysis() on both snapshots
"Text diff covers what changed"Text diffs miss new attack paths, transitive complexity shifts, and subgraph membership changesUse structural diff to complement text diff
"Only added nodes matter"Removed security functions and shifted privilege boundaries are equally dangerousReview removals and modifications, not just additions
"Low-severity structural changes can be ignored"INFO-level changes (dead code removal) can mask removed security checksClassify every change, review removals for replaced functionality
"One snapshot's graph is enough for comparison"Single-snapshot analysis can't detect evolution — you need both before and afterAlways build and export both graphs
"Tool isn't installed, I'll compare manually"Manual comparison misses what graph analysis catchesInstall trailmark first
"The diff came back empty, so nothing changed structurally"trailmark diff defaults --language to python and exits 0 with empty arrays on any other target, so an empty diff reads identically whether the code is unchanged or the language was wrongPass --language explicitly and re-run before concluding no change

Prerequisites

trailmark must be installed. If uv run trailmark fails, run:

uv tool install trailmark
# Python snippets: uv run --with trailmark python -   (a tool env is not importable)

DO NOT fall back to "manual comparison" or reading source files as a substitute for running trailmark. The tool must be installed and used programmatically. If installation fails, report the error.


Quick Start

# Compare two git refs (e.g., tags, branches, commits)
# 1. Build graphs at each snapshot
# 2. Run pre-analysis on both
# 3. Compute structural diff
# 4. Generate report

# Step-by-step: see Workflow below

Decision Tree

├─ Need to understand what each metric means?
│  └─ Read: references/evolution-metrics.md
│
├─ Need the report output format?
│  └─ Read: references/report-format.md
│
├─ Already have two graph JSON exports?
│  └─ Jump to Phase 3 (run native diff + graph_diff.py)
│
└─ Starting from two git refs?
   └─ Start at Phase 1

Workflow

Graph Evolution Progress:
- [ ] Phase 1: Create snapshots (git worktrees)
- [ ] Phase 2: Build graphs + pre-analysis on both snapshots
- [ ] Phase 3: Compute structural diff
- [ ] Phase 4: Interpret diff and generate report
- [ ] Phase 5: Clean up worktrees

Phase 1: Create Snapshots

Use git worktrees to get clean copies of each ref without disturbing the working tree.

# Create temp directories for worktrees
BEFORE_DIR=$(mktemp -d)
AFTER_DIR=$(mktemp -d)

# Create worktrees (run from repo root)
git worktree add "$BEFORE_DIR" {before_ref}
git worktree add "$AFTER_DIR" {after_ref}

If comparing two directories instead of git refs, skip this phase and use the directory paths directly in Phase 2.

Phase 2: Build Graphs and Run Pre-Analysis

Build Trailmark graphs for both snapshots and run pre-analysis on each. Pre-analysis computes blast radius, taint propagation, privilege boundaries, and entrypoint enumeration.

from trailmark.query.api import QueryEngine

def build_and_export(target_dir, output_path, language="auto"):
    """Build graph, run pre-analysis, export JSON."""
    engine = QueryEngine.from_directory(target_dir, language=language)
    engine.preanalysis()
    json_str = engine.to_json()
    with open(output_path, "w") as f:
        f.write(json_str)
    return engine.summary()

import tempfile, os
work_dir = tempfile.mkdtemp(prefix="trailmark_evolution_")
before_json = os.path.join(work_dir, "before_graph.json")
after_json = os.path.join(work_dir, "after_graph.json")

before_summary = build_and_export(
    "{before_dir}", before_json
)
after_summary = build_and_export(
    "{after_dir}", after_json
)

Verify both graphs built successfully by checking the summary output. If either fails, rerun with an explicit language or comma-separated list instead of auto.

Phase 3: Compute Structural Diff

Run both:

  1. Trailmark's native structural diff for nodes, edges, and entrypoints
  2. The plugin's graph_diff.py helper for subgraph membership changes

Use the same work_dir from Phase 2, and pass the same --language value Phase 2 built with. trailmark diff defaults that flag to python, so on any other target the default exits 0 and writes empty arrays rather than reporting a mismatch.

trailmark diff --json --language auto "{before_dir}" "{after_dir}" > "{work_dir}/trailmark_diff.json" || \
  uv run trailmark diff --json --language auto "{before_dir}" "{after_dir}" > "{work_dir}/trailmark_diff.json"

uv run {baseDir}/scripts/graph_diff.py \
    --before "{before_json}" \
    --after "{after_json}" > "{work_dir}/subgraph_diff.json"

If Phase 2 needed an explicit language or a comma-separated list instead of auto, use that same value here.

If either diff command fails or writes an empty JSON file, stop and report the error instead of continuing to Phase 4.

A trailmark_diff.json whose nodes, edges, and entrypoints arrays are all empty means either nothing changed structurally or both snapshots parsed to (near-)empty graphs. Decide which using Phase 2's graph summaries: if either snapshot's node count is zero or implausibly small for the target, the parse missed the code — name the language set explicitly (rust, solidity, python,rust) and re-run. Healthy node counts on both snapshots plus an empty diff is genuine structural stability.

The native Trailmark diff contains:

KeyContents
summary_deltaChanges in node/edge/entrypoint counts
nodes.addedNew functions, classes, methods
nodes.removedDeleted functions, classes, methods
nodes.modifiedFunctions with changed CC, params, line span
edges.addedNew call/inheritance/import relationships
edges.removedDeleted relationships
entrypointsAdded, removed, and modified entrypoints

The subgraph diff contains:

KeyContents
subgraphsPer-subgraph membership changes (tainted, high_blast_radius, etc.)

Phase 4: Interpret Diff and Generate Report

Read both diff JSON files and generate a security-focused markdown report. See references/report-format.md for the full template.

Interpretation priorities (highest to lowest):

  1. New tainted paths — nodes entering the tainted subgraph, especially if they also appear in added edges targeting sensitive functions
  2. Privilege boundary changes — new or removed trust transitions from the native entrypoint/edge diff plus the subgraph diff
  3. Attack surface growth — new entrypoints, especially untrusted_external, from trailmark_diff.json
  4. Blast radius increases — nodes entering high_blast_radius
  5. Complexity spikes — CC increases > 3 on tainted or entrypoint-reachable nodes
  6. Structural additions — new nodes and edges (review needed)
  7. Structural removals — verify removed security functions were replaced

Cross-reference structural changes with git diff {before_ref}..{after_ref} to add source-level context to findings.

Severity classification:

SeverityStructural Signal
CRITICALNew tainted path to sensitive function, removed auth boundary
HIGHNew entrypoint + high blast radius, large CC increase on tainted node
MEDIUMNew trust-boundary-crossing edges, moderate CC increase
LOWAdded nodes without entrypoint reachability
INFODead code removal, complexity reductions

For detailed metric definitions, see references/evolution-metrics.md.

Phase 5: Clean Up

Remove git worktrees after the report is written:

git worktree remove "{before_dir}"
git worktree remove "{after_dir}"

Diff Reference

trailmark diff --json --language auto BEFORE AFTER
uv run {baseDir}/scripts/graph_diff.py [OPTIONS]

trailmark diff --language defaults to python. On a target in any other language that default still exits 0, emitting well-formed JSON with empty nodes, edges, and entrypoints arrays, so always pass the flag: auto detects and merges every supported language found under the target, and a single name (rust, solidity) or comma-separated list (python,rust) pins an explicit set. auto fails loudly with No supported languages detected under <path> when a snapshot holds nothing it can parse, which is the outcome you want. Confirm the language first; only then can an empty diff count as evidence that nothing changed.

Use trailmark diff for:

  • Node/edge changes
  • Added/removed/modified entrypoints
  • Human-readable structural diff reports

Use graph_diff.py for:

  • Subgraph membership changes derived from engine.preanalysis()
  • tainted, high_blast_radius, privilege_boundary, and related sets
ArgumentDefaultDescription
--beforerequiredPath to the "before" graph JSON
--afterrequiredPath to the "after" graph JSON
--indent2JSON output indentation

graph_diff.py input format: Trailmark JSON exports from engine.to_json(). graph_diff.py output: JSON structural diff for nodes, edges, and subgraphs.


Quality Checklist

Before delivering the report:

  • Both graphs built successfully (check summaries)
  • Pre-analysis ran on both snapshots
  • Native Trailmark diff computed (trailmark_diff.json); if it is empty, both snapshots' Phase 2 node counts were non-zero, so empty means stable
  • Subgraph diff computed and non-empty (subgraph_diff.json)
  • All subgraph changes interpreted (tainted, blast radius, etc.)
  • Critical findings include evidence (node IDs, edge diffs)
  • Severity levels assigned to all findings
  • Source-level context added via git diff cross-reference
  • Worktrees cleaned up (or temp dirs removed)
  • Report written to GRAPH_EVOLUTION_*.md

Integration

trailmark skill: Phase 2 uses the trailmark API for graph building and pre-analysis. All trailmark query patterns work on either snapshot's engine.

differential-review skill: Use graph-evolution for structural analysis, differential-review for line-level code review. The two are complementary — graph-evolution finds attack paths that text diffs miss, while differential-review provides git blame context and micro-adversarial analysis.

trailmark-review-gate skill: Use trailmark-review-gate after graph-evolution when a branch, pull request, fix commit, or release diff needs a PASS/WARN/FAIL/UNKNOWN structural review packet. The gate applies deterministic review rules to graph-evolution output; it does not replace human review.

genotoxic skill: If graph-evolution reveals new high-CC tainted nodes, feed them to genotoxic for mutation testing triage.

diagramming-code skill: Generate before/after diagrams to visualize structural changes. Use call-graph or data-flow diagrams focused on changed nodes.


Supporting Documentation

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Builds and runs code under AddressSanitizer to catch buffer overflows, use-after-free, and other memory errors during fuzzing or tests. Covers -fsanitize=address builds, ASAN_OPTIONS, reading the crash report, LeakSanitizer, and the overhead and platform trade-offs. Use when fuzzing C/C++ or Rust that has unsafe blocks or FFI, when debugging a memory corruption crash, or when reading an ASan stack trace.

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

aflpp

無料

Sets up and runs AFL++ for multi-core fuzzing of C/C++ projects built with afl-clang-fast or afl-gcc-fast. Covers instrumentation modes, parallel main and secondary campaigns, persistent mode, corpus minimization, and crash triage. Use when scaling fuzzing across cores, fuzzing a mature C/C++ codebase, reading the afl-fuzz status screen, or moving on after libFuzzer has plateaued.

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

Audits GitHub Actions workflows for security vulnerabilities in AI agent integrations including Claude Code Action, Gemini CLI, OpenAI Codex, and GitHub AI Inference. Detects attack vectors where attacker-controlled input reaches AI agents running in CI/CD pipelines, including env var intermediary patterns, direct expression injection, dangerous sandbox configurations, and wildcard user allowlists. Use when reviewing workflow files that invoke AI coding agents, auditing CI/CD pipeline security for prompt injection risks, or evaluating agentic action configurations.

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

Scans Algorand smart contracts for 11 common vulnerabilities including rekeying attacks, unchecked transaction fees, missing field validations, and access control issues. Use when auditing Algorand projects (TEAL/PyTeal).

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

atheris

無料

Sets up and runs Atheris, the coverage-guided Python fuzzer built on libFuzzer. Covers TestOneInput harnesses, FuzzedDataProvider, instrumenting both pure Python and native C extensions, and running under AddressSanitizer. Use when fuzzing a Python package, hunting memory corruption in a Python C extension, or choosing between Atheris and Hypothesis for a Python target.

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

Augments Trailmark code graphs with external audit findings from SARIF static analysis results, weAudit annotation files, and version-gated Trailmark 0.4.x binary-analysis graph exports. Maps findings to graph nodes by file and line overlap, creates severity-based subgraphs, and enables cross-referencing findings with pre-analysis data (blast radius, taint, etc.). Use when projecting SARIF results onto a code graph, overlaying weAudit annotations, importing binary graph findings, cross-referencing Semgrep, CodeQL, or binary-analysis findings with call graph data, or visualizing audit findings in the context of code structure.

日本語の概要は準備中です。原文の説明を表示しています。

trailofbits/skills7,4762026年10月10日 更新

trailofbits のスキルをすべて見る

このスキルの問題を報告する