本文へ移動
cccskills
無料GitHub で公開

execution-grounded-selection

Pick among candidate outputs (code, configs, plans) by running them on diverse inputs and clustering by behavioural fingerprint, rather than by textual aggregation or log-probability. Activates when an executor returns multiple plausible candidates that need disambiguation, when output-majority voting would be the default choice, or when reviewing generated code that has not yet been validated. The 2026 evidence (Semantic Voting, arxiv 2605.08680v1) is that any execution-based selector dominates output-majority voting by 19-52pp; sketch-generated inputs beat random fuzz by 11.3pp. Triggers: "pick the best candidate", "majority vote on code", "select from N samples", "validate the generated output", "behavioural verification".

インストール方法を見る

含まれるファイル(1)

  • SKILL.md3.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Execution-Grounded Selection

Why

Output-majority voting (pick the most common string output) is dominated by any selector that actually runs the candidates. Semantic Voting (arxiv 2605.08680v1) shows 19-52pp improvement over output voting across multiple code benchmarks. The specific aggregation rule (majority, weighted, MBR-Exec) is statistically indistinguishable once execution evidence is present — execution is the dominant signal, aggregation is the residual.

This is the code-domain analogue of the noise-as-exploration Rosetta concept: execution diversity is the exploration mechanism; behavioural fingerprint is the equilibrium signal.

How

The Semantic Voting pipeline:

  1. Sample N candidates — generate at temperature > 0 (typical N = 5-10).
  2. Generate diverse inputs — sketch-generated inputs (derived from candidate population structure) beat random fuzz by ~11pp. If sketch generation is infeasible, fall back to LLM-generated test inputs > random fuzz.
  3. Execute each candidate on each input — collect (candidate_i, input_j, output_ij) tuples. Crash counts as a distinct fingerprint, not a discard.
  4. Cluster by fingerprint — equivalence on [output_ij for j in inputs] defines the cluster.
  5. Pick the largest cluster — break ties by candidate self-confidence or by Pareto on execution cost.

When to skip

  • Generation is deterministic (temperature = 0) — there's nothing to select among.
  • Execution is expensive or has side effects (touches the network, modifies state) — fall back to dry-run / static analysis.
  • The output isn't executable (e.g., natural-language summary) — use paired-trace audit instead.

Integration

  • wrap:verify and gsd-verify-work — pre-merge gate when verification produces multiple candidate fixes.
  • gsd-code-fixer — when the fixer proposes more than one fix per finding.
  • code-review — adds a "did you actually run it?" subsection to the review rubric.

Cross-references

  • Rosetta concept #9 (Execution-Grounded Selection) — canonical definition
  • College: agent-systems / agentic-code-generation / agent-execution-grounded-selection
  • Related skills: test-generator (generates the inputs; pair with this skill for the full loop)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Provides web accessibility best practices for semantic HTML, ARIA, keyboard navigation, color contrast, and screen reader patterns. Use when building UI components, reviewing accessibility, or when user mentions 'a11y', 'accessibility', 'ARIA', 'screen reader', 'keyboard navigation', 'WCAG'.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Active listening techniques for effective communication. Covers attending behaviors, paraphrasing, reflective listening, clarifying questions, empathic response, barriers to listening, listening in conflict, and cross-cultural listening. Use when building listening skills, improving understanding in conversation, mediating disputes, or analyzing communication breakdowns.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Adversarial spec-compliance PR review — cross-references diffs against approved specs, verifies runtime claims against source, detects competing PRs, audits scope/convention compliance. Use before merging.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Provides best practices for AI agent orchestration including MCP servers, A2A protocol, multi-agent coordination, and swarm architectures. Use when designing agent systems, configuring MCP servers, setting up agent teams, or when user mentions 'MCP', 'A2A', 'agent orchestration', 'multi-agent', 'swarm', 'agent team', 'LangGraph', 'CrewAI', 'AutoGen'.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Symbolic manipulation, equation solving, and algebraic structures for mathematical reasoning. Covers distributive law, factoring, completing the square, linear through polynomial equation solving, systems of equations (substitution, elimination, Gaussian elimination, matrix methods), algebraic structures (groups, rings, fields), modular arithmetic, polynomial theory, and inequalities. Use when solving equations, simplifying expressions, working with algebraic structures, or performing symbolic manipulation.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Understanding how algorithmic systems shape what users see, know, and do -- from recommendation feeds to search ranking to credit scoring to hiring software. Covers the mechanics of recommendation systems, algorithmic bias and its sources, personalization's effects on information diets, opacity and accountability, AI limitations (hallucination, confident wrongness), and the human-in-the-loop question. Use when a learner needs to think critically about why particular content reached them.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator712026年7月20日 更新

Tibsfox のスキルをすべて見る

このスキルの問題を報告する