Audit GitHub Actions that run AI agents for prompt injection, unsafe interpolation, sandbox gaps, and permissive actor rules. Use for agentic CI workflows, not general application code review.
日本語の概要は準備中です。原文の説明を表示しています。
Create new skills (SKILL.md files), modify and improve existing skills, and design skill descriptions for accurate triggering. Use when the user wants to create a new skill from scratch, edit an existing skill, optimize a skill's description, or convert a workflow they just demonstrated into a reusable skill.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
A skill for creating new skills and iteratively improving them.
At a high level, the process of creating a skill goes like this:
SKILL.md)Your job when using this skill is to figure out where the user is in this process and help them progress through the stages. If they say "I want to make a skill for X", help narrow down what they mean, write a draft, try a few realistic prompts, and iterate. If they already have a draft, jump straight to testing and iterating.
If the user just says "vibe with me, no formal evals", do that.
Skill-creator is liable to be used by people across a wide range of familiarity with coding jargon. Pay attention to context cues:
It is fine to briefly clarify a term when in doubt.
Start by understanding what the user wants. The current conversation may already contain the workflow to capture (e.g., they say "turn this into a skill"). If so, extract answers from the conversation history first — the tools used, the sequence of steps, corrections made, the input/output formats observed. The user can fill the gaps and confirm.
Ask:
Proactively ask about edge cases, input/output formats, example files, success criteria, and dependencies. Wait to write test prompts until this is ironed out.
If the platform supports parallel sub-tasks, research in parallel (search docs, find similar skills, check best practices).
Based on the interview, fill in:
name — the skill identifier (must match the directory name; lowercase alphanumeric with single hyphens, regex ^[a-z0-9]+(-[a-z0-9]+)*$).description — when to trigger and what the skill does. This is the primary triggering mechanism. Include both what the skill does AND specific contexts for when to use it. All "when to use" info goes here, not in the body. Skills tend to under-trigger, so make descriptions slightly "pushy" — e.g. instead of "Build a fast dashboard", write "Build a fast dashboard. Make sure to use this skill whenever the user mentions dashboards, data visualization, internal metrics, or wants to display any kind of company data, even if they do not explicitly ask for a 'dashboard.'"license (optional) — the license under which the skill is distributed.Then write the body.
skill-name/
├── SKILL.md (required)
│ ├── YAML frontmatter (name, description required)
│ └── Markdown instructions
└── Bundled resources (optional)
├── scripts/ — Executable code for deterministic / repetitive tasks
├── references/ — Docs loaded into context as needed
└── assets/ — Files used in output (templates, icons, fonts)
Skills use a three-level loading system:
SKILL.md body — in context whenever the skill triggers; aim for under 500 lines.Key patterns:
SKILL.md under 500 lines. If approaching the limit, add a layer of hierarchy with clear pointers to follow-up files.SKILL.md with guidance on when to read them.Domain organization: when a skill supports multiple domains/frameworks, organize by variant:
cloud-deploy/
├── SKILL.md (workflow + selection)
└── references/
├── aws.md
├── gcp.md
└── azure.md
The model reads only the relevant reference file.
Skills must not contain malware, exploit code, or any content that could compromise system security. A skill's contents should not surprise the user given its description. Do not create misleading skills or skills designed to facilitate unauthorized access, data exfiltration, or other malicious activities. Roleplay-style skills are fine.
Defining output formats — use a clear template:
## Report structure
ALWAYS use this exact template:
# [Title]
## Executive summary
## Key findings
## Recommendations
Examples pattern — small, concrete examples help:
## Commit message format
**Example 1:**
Input: Added user authentication with JWT tokens
Output: feat(auth): implement JWT-based authentication
After drafting, come up with 2–3 realistic test prompts — the kind of thing a real user would actually say. Share them with the user: "Here are a few test cases I'd like to try. Do these look right, or do you want to add more?"
Try the skill on each prompt. Read the transcripts (not just the final outputs) to see whether the skill is causing the model to waste effort on unhelpful steps.
scripts/ and tell the skill to use it.After improving the skill, re-run the test prompts and check that the issues are resolved without breaking earlier behavior.
The description field is the primary mechanism that determines whether the model invokes a skill. To optimize it:
This is a modified port of anthropics/skills/skills/skill-creator. The upstream version ships bundled scripts and an HTML eval viewer (scripts/aggregate_benchmark.py, eval-viewer/generate_review.py, etc.) that are not included here. The reviewed upstream snapshot is pinned in UPSTREAMS.json.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Audit GitHub Actions that run AI agents for prompt injection, unsafe interpolation, sandbox gaps, and permissive actor rules. Use for agentic CI workflows, not general application code review.
日本語の概要は準備中です。原文の説明を表示しています。
Audit and improve project-rules files (AGENTS.md, CLAUDE.md, .agents/instructions, local overrides) so the agent keeps accurate project context. Use when the user asks to check, audit, review, update, improve, or fix their AGENTS.md or CLAUDE.md, mentions "project rules maintenance" or "agent context optimization", or when the codebase has changed enough that the rules file may be stale. Scans the repository for every rules file, grades each against a quality rubric, outputs a quality report, and applies targeted edits only after user approval.
日本語の概要は準備中です。原文の説明を表示しています。
Capture learnings from the current session into the project-rules file (AGENTS.md, CLAUDE.md, or local override) so future sessions benefit. Use when the user says "revise the rules", "update AGENTS.md / CLAUDE.md with what we just learned", "save this to project memory", "remember this for next time", or at the end of a productive session when valuable context has emerged that is not yet documented. This complements agents-md-improver — improver audits, while this one captures.
日本語の概要は準備中です。原文の説明を表示しています。
Operational rubric that turns "don't make AI slop" into observable properties, severity levels, evidence requirements, and repair actions for interface design. Use as the reference rubric when building or reviewing marketing sites, product interfaces, dashboards, portfolios, or e-commerce pages, especially alongside frontend-design.
日本語の概要は準備中です。原文の説明を表示しています。
Design a feature architecture by analyzing existing codebase patterns and conventions, then provide a comprehensive implementation blueprint with specific files to create or modify, component designs, data flows, and a build sequence. Use this skill when the user asks for an architecture design, an implementation plan for a non-trivial feature, or when dispatched as a sub-task during feature-dev architecture phase.
日本語の概要は準備中です。原文の説明を表示しています。
Deeply analyze an existing codebase feature by tracing execution paths, mapping architecture layers, understanding patterns and abstractions, and documenting dependencies. Use this skill when you need to understand how a feature works before modifying or extending it, when dispatched as a sub-task during feature-dev exploration, or when the user asks "how does X work in this codebase".
日本語の概要は準備中です。原文の説明を表示しています。