本文へ移動
cccskills
無料GitHub で公開

agent-eval

Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md3.3 KB
  • corpus.json21.7 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

CodeGraph Quality Audit

Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/.

Prerequisites

  • tmux 3+, a logged-in claude CLI, node, git (macOS/Linux).
  • Run from the codegraph repo root.

Workflow

Copy this checklist:

- [ ] 1. Pick version (local or npm)
- [ ] 2. Pick language
- [ ] 3. Pick repo by size
- [ ] 4. Pick harness (headless / tmux / both)
- [ ] 5. Run audit.sh in the background
- [ ] 6. Report results

Step 1 — version. Ask with AskUserQuestion: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10). Map the answer to a VERSION token:

  • "Local dev build" → local
  • "Latest published" → latest
  • a typed version → that string (e.g. 0.7.10)

Step 2 — language. Read .claude/skills/agent-eval/corpus.json. Ask with AskUserQuestion which language to test, listing the languages that have entries.

Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files). Each entry carries the repo URL and a representative question.

Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token:

  • "Headless" → headless — claude -p with stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY).
  • "Interactive (tmux)" → tmux — drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower).
  • "Both" → all — headless + interactive (4 runs).

Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes):

scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>

Step 6 — report. When the job finishes, read the log and report per arm:

  • Headless (parse-run.mjs): total tool calls, file Reads, Grep/Bash, codegraph-tool calls, duration, total cost.
  • Interactive (parse-session.mjs): the VERDICT: codegraph_explore used Nx | Read N | Grep/Bash N and TOKENS: lines.

Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer.

Notes

  • The index is rebuilt every run (audit.sh wipes .codegraph) — different versions extract differently, so an index must be served by the same binary that built it.
  • audit.sh temporarily mutates the global codegraph install for the test, then restores your dev link via local-install.sh.
  • Corpus repos are cloned to /tmp/codegraph-corpus (reused if already present).
  • Add or edit repos in corpus.json (fields: name, repo, size, files, question).

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

使用 AACT (Aggregate Analysis of ClinicalTrials.gov) PostgreSQL 数据仓库进行批量、历史、聚合性临床试验数据挖掘。Use this skill when the user requests bulk SQL analysis over the full clinical trials data warehouse — historical trial trends, disease landscapes, similar-design matching, or multi-year aggregations across hundreds of thousands of NCT records. 触发场景包括:AACT 查询、临床试验批量分析、PostgreSQL 试验数据、全量 NCT 检索、试验数据挖掘、历史试验分析、clinical trials data warehouse、SQL trials、bulk trial analysis、disease landscape、试验设计相似性匹配、跨年度聚合、sponsor/phase/country 多维统计。**与 clinical-trials-v2 差异**:本 skill 走批量 SQL · 离线大数据(PostgreSQL);v2 走实时 API · 单查询。两者互补:单条 NCT 实时状态用 v2,百万级历史挖掘用本 skill。支持云端公共 PostgreSQL(aact-db.ctti-clinicaltrials.org · 零部署)和每日 dump 本地还原(高性能 · 离线)两种连接方式,自动检测优先用本地。跨平台(macOS/Linux/Windows)参数化 SQL 防注入,read-only 强制保护。

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

使用 AACT (Aggregate Analysis of ClinicalTrials.gov) PostgreSQL 数据库进行大批量临床试验历史分析与数据挖掘。触发场景包括:AACT 查询、临床试验批量分析、PostgreSQL 试验数据、全量 NCT 检索、试验数据挖掘、clinical trials data warehouse、疾病领域全景分析、设计相似试验匹配、跨年度试验趋势聚合。本 skill 通过 SQL 接口处理百万级试验记录,支持云端公共 PostgreSQL 服务(aact-db.ctti-clinicaltrials.org)和每日 dump 本地还原两种连接方式,自动检测并优先使用本地高性能模式。

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

Use when work produces a deliverable, factual claim, dataset, code change, decision, publication, deployment, automation, or irreversible action whose failure would matter; when the user asks for a quality gate, acceptance criteria, QA, completeness, validation, audit, evidence, preflight, release readiness, or a definition of done; or before claiming completion on medium- or high-risk work. Automatically decide whether a formal gate is warranted, derive task-specific pass/fail criteria, gather evidence, and block unsupported completion. Make sure to use this skill even when the user does not say "quality gate" if consequential work needs acceptance criteria or completion evidence. Skip formal gating for trivial, reversible, low-impact requests.

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

Create a changeset file in .changeset/ that describes a user-facing change for the release notes

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

Add a new competitor to the AI Visibility Tool Directory — researches the tool, generates data, takes a screenshot, and inserts into the codebase

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

add-lang

無料

Add tree-sitter language support to codegraph end-to-end — wire the grammar + extractor, write tests, then benchmark extraction quality and retrieval value on 3 popular real-world repos. Use when the user runs /add-lang <language> or asks to add/support a new language (e.g. Lua, Elixir, Zig, OCaml) in codegraph.

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

EthanYoQ のスキルをすべて見る

このスキルの問題を報告する