本文へ移動
cccskills
無料GitHub で公開

postmortem

Write a blameless postmortem after an incident. Use when the user says "postmortem", "post-incident review", "PIR", "what happened in that incident", "incident review", "blameless postmortem", "5 whys", "how do we prevent this again", or needs to document learnings from a production incident - even if they don't explicitly say "postmortem".

インストール方法を見る

含まれるファイル(1)

  • SKILL.md5.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Overview

Based on "The Field Guide to Understanding Human Error" by Sidney Dekker and the Google SRE Book. Dekker's fundamental insight: human error is never the cause of an incident - it is a symptom of a system that put people in a position where mistakes were likely. A blameless postmortem asks not "who made the mistake?" but "what conditions made this mistake likely, and how do we change those conditions?"

Google SRE's postmortem culture: incidents are opportunities to improve systems, not to assign blame. Engineers who feel safe reporting mistakes surface problems before they become incidents.

Workflow

Step 1: Schedule within 48 hours (SEV1) or 1 week (SEV2)

Postmortems decay in value quickly. Memories fade, context is lost. Set the meeting within 48 hours of a SEV1 resolution.

Attendees: IC, Operations Lead, on-call engineer(s), and anyone who can contribute to understanding what happened. No executives in the room - their presence changes behavior.

Step 2: Build the timeline

Reconstruct what happened in chronological order. Be precise about times.

Timeline:
[HH:MM] - [what happened - system event, human action, or observation]
[HH:MM] - [alert fired: name]
[HH:MM] - [on-call engineer paged]
[HH:MM] - [first action taken]
[HH:MM] - [mitigation applied]
[HH:MM] - [service restored]

Include: system events, alerts, human decisions, and the reasoning people had at the time.

Step 3: Identify contributing factors (not root cause)

Dekker's principle: complex systems rarely have a single root cause. They have contributing factors that aligned to create conditions for failure.

For each contributing factor, ask: "If this had been different, would the incident have been less likely or less severe?"

Common contributing factor categories:

  • Technical: insufficient alerting, missing circuit breakers, inadequate testing
  • Process: no deployment checklist, unclear escalation path, missing runbook
  • Knowledge: alert with no runbook, undocumented dependency, new team member
  • Capacity: system at limits, team understaffed, alert fatigue

Step 4: Apply the 5 Whys

For the primary contributing factor, drill down:

  • Why did X happen? - Because Y
  • Why did Y happen? - Because Z
  • Why did Z happen? - Because...

Stop when you reach a level where an action item can change the system. "Because the engineer made a mistake" is never the stopping point - that's where the analysis begins.

Step 5: Write action items

Every postmortem must produce action items. Without them, it's just storytelling.

Each action item:

Action: [specific thing to do]
Owner: [named person, not "the team"]
Due: [specific date]
Priority: [P1 - fix before next deploy / P2 - this sprint / P3 - this quarter]
Type: [Prevention / Detection / Mitigation / Process]

Action types:

  • Prevention - stops this failure mode from occurring
  • Detection - makes the problem visible sooner
  • Mitigation - reduces impact when it does occur
  • Process - changes how the team responds

Step 6: Write the postmortem document

Postmortem: [Incident title]
Date: [incident date]
Authors: [IC + contributors]
Severity: [SEV1/SEV2]
Duration: [start to resolution]
Impact: [users affected, features degraded]

Summary
[2-3 sentences: what happened, why it mattered, how it was resolved]

Timeline
[full timeline from Step 2]

Contributing Factors
[bullet list from Step 3 - no blame language]

5 Whys Analysis
[drill-down from Step 4]

What Went Well
[what helped contain or resolve the incident faster]

Action Items
[table from Step 5]

Lessons Learned
[2-3 insights that are generalizably useful beyond this specific incident]

Step 7: Share broadly

Postmortems are most valuable when shared. Publish to an internal wiki, engineering all-hands, or team newsletter. Other teams often have the same latent failure modes.

Anti-Patterns

1. Blame language Bad: "The engineer deployed without testing." Good: "The deployment process did not have a required staging environment check, making it possible to deploy untested code." Dekker: blame stops investigation. "Human error" as a cause explains nothing and changes nothing.

2. No action items Bad: Postmortem documents what happened but ends without concrete next steps. Good: Every postmortem produces at least 2 action items with owners and dates. Track them to completion.

3. Postmortem after the deadline Bad: "We'll get to it when things calm down." (more than 1 week for SEV1) Good: Timebox it. A 60-minute postmortem within 48 hours is worth more than a 2-hour postmortem 3 weeks later.

4. Only focusing on what went wrong Bad: Pure failure analysis, nothing about what was done right. Good: Include "What went well." The practices that helped contain the incident should be reinforced, not just the failures fixed.

Quality Checklist

  • Scheduled within 48h (SEV1) or 1 week (SEV2)
  • Timeline is precise (specific times, not "around noon")
  • Contributing factors identified with no blame language
  • 5 Whys drilled to a system-level cause, not a person
  • Every action item has: owner (named), due date, priority, type
  • "What went well" section included
  • Document shared with broader engineering team
  • Action items tracked to completion in next retrospective

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Rate any product, feature, or experience on the 11-star scale (Brian Chesky's Airbnb thought experiment). Use when user says "rate this experience", "11-star", "star rating", "experience audit", "how good is this", "experience rating", "product audit", "quality assessment", or wants to evaluate product quality and identify improvement paths. Also trigger when user wants to benchmark a feature, assess where a product stands, or map out what "great" looks like - even if they don't explicitly say "11-star".

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

Generate A/B copy variants grounded in Eugene Schwartz's awareness and market-sophistication stages and John Caples' tested-headline discipline. Every variant declares its angle, the awareness stage it targets, and a falsifiable hypothesis, and the skill recommends which pair to test first and what delta counts as signal. Use when the user says "A/B test this", "write variants", "give me options", "test different angles", "copy variants", "ab copy", "test this headline", "multiple versions", "which version is better", or wants to test copy before publishing. For reviewing copy that already exists, see copy-review. For sample size and stopping rules, see growth-experiment. For diagnosing a page rather than its words, see page-cro. For full paid ad sets by platform, see ad-campaign-writer.

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

Write measurable acceptance criteria for requirements or user stories. Use when the user says "write acceptance criteria", "definition of done", "how do we test this", "fit criteria", "given when then", "Gherkin scenarios", "what does done look like", "testable conditions", "how will we know this works" - even if they don't explicitly say "acceptance criteria".

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

Build a strategic account plan for winning, retaining, or expanding a named account. Use when the user says "build an account plan", "strategic account planning", "how do I grow this account", "account expansion strategy", "whitespace analysis", "QBR prep", "key account review", "map the stakeholders in this account", "land and expand plan", "how do I break into [company]", or wants a structured strategy for a specific account.

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

Write paid ad copy for LinkedIn, Google, Meta, and YouTube. Produces multiple variants per platform, each targeting a different Eugene Schwartz awareness level and Cialdini persuasion principle. Use when the user asks for ad copy, ad variants, "write LinkedIn ads", Google ads, Meta ads, paid social copy, ad headlines, or wants creative for a paid campaign. Reads brand voice, ICP, and positioning from knowledge/. For organic social posts, see linkedin-post or social-calendar. For the landing page the ad points at, see landing-page-writer. For testing variants, see ab-copy-writer.

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

Design clean and consistent APIs. Use when the user says "design an API", "API design review", "design these endpoints", "REST API for X", "GraphQL schema", "API contract", "how should this endpoint work", "review my API design", or wants to create or improve an API interface - even if they don't explicitly say "API design".

日本語の概要は準備中です。原文の説明を表示しています。

qa-aman/claude-skills202026年9月10日 更新

qa-aman のスキルをすべて見る

このスキルの問題を報告する