本文へ移動
cccskills
無料GitHub で公開

run-summary

Summarize the architecture of code generated by a single retort run. Produces module-level structure, interfaces, and control flow in a form suitable for cross-run comparison — not a full codebase-summary.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md5.1 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Run Summary

Overview

Retort runs produce small, single-task codebases (one REST API, one CLI, one service). Full codebase-summary treatment is overkill — but a consistent lightweight summary per run is exactly what makes cross-run comparison useful.

This skill is an intentionally scoped-down adaptation of pourpoise's codebase-summary. It runs in seconds, not minutes, and emits only what's useful when comparing 48+ small generated projects.

Parameters

  • codebase_path (required): The run directory (same as evaluate-run's run_dir)
  • output_dir (optional, default: {codebase_path}/summary): Where to write the summary files

Steps

1. Infer the surface

Read TASK.md to understand what the code is supposed to do — the "surface" of the project. One paragraph, no judgments.

2. Map the modules

Generate {output_dir}/modules.md:

# Modules

| Path | Purpose | Entry points |
|------|---------|--------------|
| src/app.py | HTTP server, route handlers | `app`, `create_app()` |
| src/models.py | SQLAlchemy models | `Book`, `Base` |
| src/db.py | Connection + migrations | `get_engine()` |
| tests/test_app.py | API integration tests | 8 test functions |

Constraints:

  • You MUST list every non-generated source file (skip node_modules, target, __pycache__, .git, lock files, build artifacts).
  • The "Purpose" column is one line extracted from the code, not invented.
  • The "Entry points" column is the publicly-named functions, classes, or exported symbols — not every local helper.

3. Describe the interfaces

Generate {output_dir}/interfaces.md covering what the code exposes:

  • HTTP routes (method, path, short description)
  • CLI commands (subcommand + flags)
  • Library API (exported classes/functions)
  • Data schemas (tables, message formats)

Example:

# Interfaces

## HTTP routes

| Method | Path | Returns | Handler |
|--------|------|---------|---------|
| GET | /books | `[Book]` | `app.py:list_books` |
| POST | /books | `Book` | `app.py:create_book` |
| GET | /books/{id} | `Book \| 404` | `app.py:get_book` |

## Data schema

`books` table: id (int, pk), title (str), author (str), year (int).

Constraints:

  • You MUST grep / static-analyze the code rather than execute it to discover interfaces.
  • You MUST NOT invent endpoints the code doesn't actually declare.
  • If the code has none of the above categories, write (none) under that heading.

4. Trace the dominant control flow

Generate {output_dir}/flow.md with one Mermaid diagram showing the happy-path request/response for the main feature, plus a one-paragraph narration.

# Flow

```mermaid
sequenceDiagram
    Client->>app.py: GET /books
    app.py->>db.py: get_session()
    db.py-->>app.py: Session
    app.py->>Book: query.all()
    Book-->>app.py: [Book]
    app.py-->>Client: 200 {json}

A request to GET /books opens a DB session via db.py:get_session(), queries all Book rows, and returns them as JSON. No pagination, no filtering.


Constraints:
- You MUST pick the single most representative flow — the one a user of the generated code would hit first.
- You MUST note deviations from common patterns ("no input validation", "no error handling", "synchronous DB access in async handler").
- The narration MUST be factual, not prescriptive.

### 5. Write the index

Generate `{output_dir}/index.md` linking to the other three files and giving a 3–4 bullet summary:

```markdown
# Summary: {cell_name} · rep {replicate}

- **Shape:** {one-line description — "Flask REST API with SQLAlchemy", "Go net/http CRUD with in-memory store", etc.}
- **Structure:** {n} modules, {n} test files
- **Interfaces:** {n} HTTP routes / {n} CLI commands / {n} exported functions
- **Notable:** {what stands out — simplest/most-complex approach seen, unusual library choice, etc.}

See [modules.md](modules.md), [interfaces.md](interfaces.md), [flow.md](flow.md).

Constraints Summary

  • You MUST finish in under 90 seconds wall-clock. This is the fast-path summary.
  • You MUST NOT read anything under node_modules/, target/, __pycache__/, .git/, dist/, build/.
  • You MUST write exactly four files: index.md, modules.md, interfaces.md, flow.md. No more, no less.
  • You MUST keep descriptions factual — no quality judgments (that's evaluate-run's job).
  • Output files MUST be valid markdown that renders correctly in GitHub's viewer (Mermaid in a ```mermaid fenced block).

Troubleshooting

Code is unparseable / generated output is garbage

  • Write the index.md with **Shape:** unparseable — agent output did not produce a valid project.
  • Leave the other files near-empty with a one-line explanation.
  • Exit 0 so evaluate-run can still complete.

Too many files to summarize

  • This shouldn't happen in retort's small-task workspaces. If it does, cap the modules table at 50 rows and note the truncation in index.md.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

beads

無料

Use when working in a repository that uses bd or Beads for durable project task tracking, issue dependencies, blocker management, multi-session handoff, or shared work memory. Trigger when the user asks to find ready work, claim or close tasks, create follow-up work, inspect blockers, recover project context, or choose between local planning and persistent project tracking.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

Compare evaluated runs in a retort experiment along factor dimensions. Surfaces effects of each factor, aggregates across replicates, and highlights cells that diverge qualitatively — complementing (not replacing) retort's ANOVA analysis.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

Determine the TRUE cause of a failed retort run before attributing it. Ground-truth every failure (run its tests, read its agent logs, inspect its workspace) and classify it as an infrastructure false-fail, a genuine model miss, or an environment issue — never trust the gate verdict or a log signature alone. Use whenever a run is recorded failed (test_coverage=0 / gate fail), a pass rate looks low, or you're deciding whether a "failure" is real before reporting or proceeding.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

Evaluate a single retort experiment run. Score the generated code against the task's TASK.md requirements, run its build and tests, compute metrics, and emit a structured evaluation report plus a machine-readable findings file.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

Refresh the data tables in optimal-blog.md from master.db. Checks the data for integrity problems FIRST, then runs the generator that picks per-language winners and splices every GEN-marked table, then reconciles the surrounding prose. Use after new experiment results land, or when the optimal-blog numbers are stale.

日本語の概要は準備中です。原文の説明を表示しています。

adrianco/retort2082026年10月10日 更新

adrianco のスキルをすべて見る

このスキルの問題を報告する