Skill Creator
Create or improve a reusable capability, verify the user's actual outcome, and
deliver the declared source, package or installed result. Start at the unfinished
stage. Preserve an existing skill's name, supported work and user-owned data.
Use the current user's authorization and actual host capabilities before these
defaults. A tool's availability, a long conversation or the word “optimize” does
not authorize history access, new capabilities, agent fan-out or publication.
Resolve the task before acting
- Identify the result, supported inputs, triggers, authorized deliverables and
stopping condition from the conversation. Ask only for information or a real
unresolved choice that changes the work; reuse answers and existing authorization.
Distinguish creating a skill, editing one, optimizing this creator and doing a
one-off task. Do not turn the last into reusable tooling without authorization.
- Find the existing owner before creating or reallocating a capability. For that
decision, search current project and installed/source skills by capability
vocabulary, including project skills and profiles; verify coverage before
claiming none exist. Read the ownership decision.
For a narrow edit with unchanged ownership, retain the verified mapping and
proceed. Reuse the maintained execution owner and keep
our verified increment with the workflow that needs it; preserve the dependency
update path. Keep distinct permission or lifecycle boundaries distinct.
- Establish canonical source ownership before writing; see
Edit Skills at Source Location.
For an existing skill, freeze its complete bundle and inventory its jobs,
interfaces, failure/recovery cases and variants before the first edit; see
Step 4.
- List the proposed changes, then select the lowest verification tier that can
falsify them. Keep evaluation authorization separate from risk classification.
Record the tier, reason and default evidence in the existing plan/TodoList when
available. Add generic eval JSON/viewer tasks only after that pipeline is authorized.
Continue with the applicable source workflow below, retaining the common gates.
Verification depth router (run before choosing any workflow)
Read change verification before classifying a
change or launching evidence work. Apply its complete boundaries:
| Change | Minimum evidence path |
|---|
| Existing skill: spelling/format only, authoritative factual correction, or deterministic repair of an explicit existing contract; no changed capability/branch/interface/permission | Tier 1: validation, diff, migration gate and the matching direct check |
| Existing skill: bounded agent behavior change, no new capability/interface/script/dependency/permission; 1–2 named examples exercise the entire change | Tier 2: those output-level replays, narrow deterministic checks and required independent review |
| New skill, broad rewrite/methodology, new or materially changed capability/trigger family/branch/output/script/dependency/permission, high-risk automation or 3+ changed prompt classes | Tier 3: deterministic gates first, then evidence for the named unresolved failure axes |
Materially adding or rewriting a reference changes the runtime loading surface:
start at Tier 2, escalating for a Tier-3 trigger. A long file, subjective output,
uncertain scope or an additional benchmark request does not determine the tier.
Paired evaluation is separately authorized. Run it when the user explicitly
requests A/B, baselines, benchmarking, repeated trials, a viewer or multi-agent
evaluation; otherwise explain the decision it would resolve and obtain opt-in.
Declare distinct roles, necessary isolated arms/shards, total units and capped
concurrency before launching. Do not pre-launch downstream grading/viewer work.
Cancellation stops the heavy pipeline and remains in force until reauthorized;
retain necessary safety, preservation and the independently required review.
Select source workflows without losing the lifecycle
These are source and execution adaptations, not mutually exclusive products.
Classify verification and establish ownership before any specialized exit. Return
from the selected workflow to compatible editing, validation and delivery gates;
record which generic mechanics its verification protocol replaces.
| Source or task | Read and follow before acting |
|---|
| Current conversation or ordinary skill creation/editing | Read Capture Intent when inputs or intended behavior remain unresolved, then only the authoring sections needed by the actual draft; reuse already resolved intent/evidence for bounded corrections. Extract verified knowledge into references and needed parameterized helpers into scripts, after checking existing owners |
| Actual third-party installation/debugging work being distilled | Wrapper workflow; verify installation state rather than force incompatible file assertions |
| Explicitly approved earlier local sessions, with known session/time/scope boundaries | Conversation mining; retrieve through the owning index and exact reader, then use manifest → discover → redact → chunk for selected corpus distillation. For cross-task improvement, load its matched-case route before selecting changes; agents receive only prepared, redacted evidence |
| User-endorsed finished artifacts used to extract preferences | Artifact corpus distillation; extract evidence-backed decision rules, not merely a sample catalog |
| Authorized old/new or no-skill comparison | Paired evaluation and schemas; use the original immutable old skill for existing-skill comparisons |
| Real missed/misfired trigger requiring description tuning | Description triggering; first verify a working probe, inspect every skill call, and distinguish invocation from output correctness |
| Claude.ai, Cowork or a host missing a pipeline capability | Host adaptations; use actual tools and disclose omitted comparisons rather than simulate independent evidence |
An ordinary optimization does not authorize earlier-history mining. Do not open
raw transcripts in this context or send raw paths/content to agents. No history
access is a valid route when the live conversation and current bundle suffice.
For approved artifacts, admission requires the user's endorsement; “registered”
is not “distilled,” and approval of facts is not approval of a reusable template.
Ground the draft and its execution
Before adding technical or methodology claims, read
grounding these claims.
Verify landed claims against the actual target ref, not its commit message.
Retrieve the current official authoring practices
and read the applicable phases of development methodology
before authoring. Retrieve established domain methods as well as infrastructure
prior art; preserve the user's private methodology. Verify technical assertions
against observed behavior/current authority, or mark their evidence limits.
Treat “unsupported” as a hypothesis and investigate alternative documented paths.
- For external facts, code/config/SOP changes or validation warnings, follow
knowledge grounding and document alignment.
Update directly affected examples/help/docstrings before the first freeze;
live, replay and synthetic evidence support different claims.
- For operational first-use/setup/recovery, follow
first use and recovery before drafting.
Probe readiness, execute authorized preparation, resume from the first unmet
condition and verify the original usable result. Skip this route for
self-sufficient transforms and reference-only skills.
- For persisted formats/partial updates, permission-sensitive checks or large
command output, follow the affected stateful verification recipe.
A truncated response is not a complete read; use its complete-read recipe.
For provider/default-route/enable-switch changes, use its Operational route
changes recipe: preserve the original user result and exercise the actual
downstream consumer, including asynchronous/device removal and human-review
exits. A reachable pointer alone does not prove the handoff runs.
- For any necessary review/eval input copies, declare exact immutable inputs,
cumulative bytes, reserve and session owner, then use
materialization prepare → run → finish.
Finish after success, failure or interruption; preserve modified inputs/evidence.
- For overlapping installed skills, follow
precedence and coexistence.
Run the official-creator presence check there; absent/already configured means
continue silently. Installing a global routing hook requires authorization.
Use imperative instructions, clear failure/recovery exits and observable command
outcomes. Keep one canonical definition per contract; place conditional detail
under its owner with an action-time loading pointer. Keep user-mutable data outside
update-owned bundles. Match instructions and verification to the actual host and
production reader, rather than assume one host's tools or renderer apply everywhere.
CRITICAL: Edit Skills at Source Location
Read source location and activation
before source checks or availability claims. Treat user skill directories, symlinks
and plugin caches as installed entries until ownership proves otherwise.
Use the fixed command entry
for bundled tools from another directory, with absolute artifact paths:
python3 <skill-creator-dir>/scripts/creator.py source_contract check-path <skill-dir> \
--phase create --repo <source-worktree> --scope marketplace
Stop before writes on invalid or unknown; do not guess an owner. Linked
worktrees must name their own repository root while retaining the registered Git
identity. Use --scope project for project-local skills. Source, installed identity
and a fresh host's actual consumption remain separate observations.
Run bundled Python tools in the creator's locked uv project with uv run --frozen;
check prerequisites for the phases actually selected.
Reuse the shared cache and project-local environment; do not add per-call overlays
for locked dependencies or clean caches as part of ordinary skill work.
Step 4: Edit the Skill
For an existing skill, read existing-skill migration
before the first edit. Freeze the entire auditable old bundle from a full immutable
ref or provenance-bearing snapshot; preserve its inclusion policy using
snapshot archives.
Classify representation changes separately from retirement, narrowing, redesign,
factual repair and lossy summarization. Compression permits relocation/deduplication,
not silent behavior removal. Preserve each old scenario's trigger, decision inputs,
supported action, stop/confirmation, recovery boundary and verification path.
Retirement or changed boundaries require traceable authorization.
Run compare → classify → verify. Review every unmatched old unit with the owning
current file, locatable evidence and semantic reason; surviving only in tests,
evals or an unreachable reference is a gap. Regenerate stale reviews after edits.
Inspect changed prose pointers and adjacent rule clauses with the migration
guide's reference/self-application procedure; text survival alone does not prove
the section's behavior survived.
Finish this edit only after the
static preservation and required task-result gates
both clear.
Validate each entry edit immediately:
uv run --frozen python -m scripts.quick_validate <skill-dir>
Write descriptions as YAML block scalars. Keep distinct trigger jobs in the
description; body-only words cannot restore discovery. Bundle helpers/components
only within authorized capability work; user-approved fragments need a frozen
behavior/provenance contract. Do not turn a narrow edit into a new pipeline.
Review the changed behavior
Use assertions against the user's result, not a literal tool path. Match the actual
reader/production engine. Before writing or running a check, read the check-design section of
check design. Calibrate on known healthy and failing inputs;
cover every clause and examined-item count. False green, false alarms and vacuous
checks require different probes. Deterministic criteria use code; subjective
outputs use appropriately calibrated independent/human judgment. Read actual
outputs and traces before trusting a benchmark or reviewer finding.
Validate aggregate and standalone viewer grades through the shared
grading validator. Bind populated canonical
eval-directory assertions by original text and multiplicity; keep missing or
empty targets explicitly unbound for legacy/preparation observations.
Before shipping a new skill or a changed rule/contract/number, follow
independent review. Freeze the exact
artifact, reader spec, blast radius, failure axes and terminal condition. Use
fresh context and evidence outside the author's changes. Reproduce hypotheses;
after substantive fixes recheck the failed axes plus preservation with fresh
context. Persist and commit the current review in the private review archive,
not public source or disposable eval scratch. Stop at the declared boundary.
Step 5: Sanitization Review
Before public delivery, read publishing and packaging
and sanitization. Check destination
visibility. Semantically read shipped examples/identifiers as well as running
scanners; green scans do not establish privacy. Private findings are information
for the owner, not permission to replace working configuration.
Use the shared packaging policy for the shipping
set. Packaging an existing skill re-verifies its completed regression review;
a marker or commit alone is not authority. Clear scan/runtime errors, keep scan
and review evidence bound to current content, and bump the registered plugin's
release identity for shipped changes. Preserve unrelated registry entries.
For repository publication, follow release readiness
with committed exact-candidate independent evidence. A package-only/source-only
request does not acquire installation or publication scope. When local availability
is requested, invoke skill-governance and the source-sync owner, preserve host
disables, read the actual consumed file and verify the fresh host. A registered or
merged source is not proof of installation or business success.
Show the result, not just the work
Follow delivery identity at delivery. Show the
verified outcome, relevant evidence and limits using the qualified source identity
and plugin version; separate source, publication, installation and task-result claims.
Use the existing eval viewer for authorized paired outputs. For another result
whose evidence/options need visual inspection, hand off conditionally to
report-with-html; do not invent its unavailable template or generate HTML for
every small edit. For recurring customer reports, follow the approved-template
contract in report-template approval: show a real
verified report first, obtain separate approval of its future form, and store the
data-free template in stable user-local storage outside the skill package.
Iterate from actual evidence and user corrections. Fix the artifact and reusable
rule together within authorization, retaining exact private calibration words
where appropriate to their destination. Re-run only affected checks; do not restart
cancelled evaluation or optimize a description without real trigger evidence.
Finish when the declared outcome and evidence gates are satisfied.