Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when designing production-LLM prompts — few-shot, chain-of-thought, system prompts, templates, self-verification — distinct from prompt-optimizer and refine-prompt.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Production patterns for LLM prompts: few-shot, chain-of-thought, system-prompt design, templating, self-verification. Distinct surface from sibling skills:
prompt-optimizer — polishes a single end-user prompt for ChatGPT / Claude / Gemini.refine-prompt — refines a free-form work prompt into engine-ready acceptance criteria.Do NOT use when:
prompt-optimizer.refine-prompt.Start at Level 1; only escalate when measurement says you must.
Level 1 Direct instruction "Summarize this article."
Level 2 + constraints (length, format, focus) "...in 3 bullets, key findings only."
Level 3 + reasoning scaffold "Read first, identify findings, then summarize."
Level 4 + few-shot examples "Like these examples: ..."
Level 5 + self-verification step "...then check answer against criteria; revise if fails."
Escalating without evidence is over-engineering. Each level adds tokens, latency, and a maintenance surface.
Fixed instruction hierarchy — every production prompt fills these slots in order:
[System context] role, expertise, constraints, safety
[Task instruction] what to do, in one sentence
[Examples] few-shot demonstrations (optional)
[Input data] the user-supplied content
[Output format] schema, length, citation rules
Stable slots (system, task, format) belong in cached prompt prefixes; volatile slots (examples, input) belong in the per-call portion.
Examples are uniform and small (< 20) → embed all of them; deterministic.
Examples are large or diverse → semantic-similarity retrieval per call.
Edge cases dominate → diversity-sampled examples (cluster + pick one per cluster).
Token budget tight → fewer, higher-quality examples beats many mediocre.
Examples drift with the data → regenerate from a labeled corpus on a schedule, not hand-edited.
Bad examples are worse than no examples — the model imitates structure.
CoT improves accuracy on multi-step reasoning, hurts on classification and lookup. Decision rule:
Task is multi-step / arithmetic / multi-hop → add CoT (zero-shot "let's think step by step", or few-shot CoT).
Task is single-step extraction / classify → CoT adds tokens without lift; skip.
You haven't measured → measure first, decide second.
Self-consistency needed (high-stakes answers) → sample N reasoning paths, majority vote.
Production prompts handle their own failure cases:
prompt-optimizer, refine-prompt, mcp-builder, async-python-patterns.agents/settings/contexts/skills-provenance.yml (entry: prompt-engineering-patterns).verify-before-complete, skill-quality, non-destructive-by-default.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Use when reviewing UI for accessibility — WCAG 2.2 AA, keyboard nav, focus, ARIA, contrast, screen-reader semantics — even on 'is this a11y-OK?' or 'mach das barrierefrei'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when defining or auditing the activation event — aha-moment selection, retention correlation, falsifiable definition. Triggers on 'what is our aha moment', 'redefine activation'.
日本語の概要は準備中です。原文の説明を表示しています。
Use when capturing an architectural decision — file naming, next ADR number, Status / Context / Decision / Consequences, index regen; fires even without saying 'ADR'.
日本語の概要は準備中です。原文の説明を表示しています。
Adversarial critique — devil's advocate, stress-test, honest teardown ('poke holes', 'be brutal', 'was hältst du davon'); explicit request only. Routine code or design review → code-review.
日本語の概要は準備中です。原文の説明を表示しています。
Use when reading, creating, or updating agent documentation, module docs, roadmaps, or AGENTS.md. Understands the full .augment/, agents/, and copilot-instructions structure.
日本語の概要は準備中です。原文の説明を表示しています。
Use for an adversarial red-team / blue-team / auditor review of an AI agent's CONFIG + behaviour (rules, skills, MCP, hooks, permissions) — attack-chain → defensive-gap list, not a code audit.
日本語の概要は準備中です。原文の説明を表示しています。