本文へ移動
cccskills
無料GitHub で公開

codexkit-incident-postmortem

Write blameless incident postmortems following Google SRE methodology. Covers timeline reconstruction, root cause analysis (5 Whys), impact assessment, and action items with owners. Use after production incidents, outages, or significant bugs to prevent recurrence.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md5.5 KB
  • agents/openai.yaml166 B
  • examples/common-mistakes.md794 B
  • examples/good-output.md958 B
  • verification/checklist.md1.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Incident Postmortem

When to Use

  • After any production incident that affected users
  • After near-misses that revealed systemic weaknesses
  • When establishing a blameless postmortem culture
  • When leadership requests incident RCA (Root Cause Analysis)

Procedure

Step 1 — Incident Summary

Write a 2–3 sentence summary:

  • What happened, when, and how long
  • Impact in business terms (users affected, revenue lost, SLA breach)
  • Current status (resolved, monitoring, partially mitigated)

Severity classification:

SeverityCriteria
SEV-1Complete outage or data loss affecting all users
SEV-2Major feature unavailable or significant degradation
SEV-3Minor feature impact, workaround available
SEV-4Cosmetic or non-user-facing issue

Step 2 — Timeline

Build a minute-by-minute timeline:

Time (UTC)EventSource
14:02Deploy v2.3.1 to productionCI/CD
14:05Error rate spikes to 15%Datadog
14:08PagerDuty alerts on-call engineerPagerDuty
14:12On-call begins investigationSlack thread
14:25Root cause identified: DB migration timeoutLogs
14:30Rollback initiatedCI/CD
14:35Service fully recoveredMonitoring

Time to Detect (TTD): 3 min | Time to Resolve (TTR): 33 min

Step 3 — Root Cause Analysis (5 Whys)

  1. Why did the service fail? → Database queries timed out
  2. Why did queries time out? → Migration locked a critical table for 8 minutes
  3. Why was the table locked? → ALTER TABLE ran without CONCURRENTLY flag
  4. Why wasn't it concurrent? → Migration script didn't follow the safe-migration checklist
  5. Why wasn't the checklist followed? → No automated check in CI pipeline

Root cause: Missing CI check for safe migration patterns.

Step 4 — Impact Assessment

MetricValue
Duration33 minutes
Users affected~12,000 (8% of DAU)
Requests failed~45,000 (500 errors)
Revenue impact~$2,400 estimated
SLA impact99.92% (target: 99.95%) — SLA breached

Step 5 — What Went Well / What Didn't

What Went WellWhat Didn't Go Well
Fast detection (3 min TTD)No pre-deploy migration testing
Clear rollback procedureMigration script wasn't reviewed
Team communicated via war roomPagerDuty escalation was slow

Step 6 — Action Items

#ActionTypeOwnerPriorityDeadline
1Add CI check for safe migration patternsPreventPlatformP1Sprint 5
2Add migration dry-run to staging deployDetectDevOpsP1Sprint 5
3Improve PagerDuty escalation timingProcessSREP2Sprint 6

Action types: Prevent (stop recurrence), Detect (catch faster), Mitigate (reduce impact)

Inputs

InputRequiredFormat
Incident descriptionYesWhat happened, when
Timeline / log dataYesTimestamps with events
Impact dataYesUsers affected, duration, revenue
Team notesRecommendedSlack threads, war room notes

Output

## Postmortem — [Incident Title]

**Date:** 2024-03-15 | **Severity:** SEV-2 | **Duration:** 33 min
**Author:** [On-call engineer] | **Reviewers:** [Team leads]

### Summary
Production database queries timed out for 33 minutes due to a table-locking
migration, affecting ~12,000 users and breaching our 99.95% SLA target.

### Timeline
[Minute-by-minute timeline]

### Root Cause
[5 Whys chain → root cause]

### Impact
[Quantified impact table]

### Lessons Learned
[What went well / what didn't]

### Action Items
[Prioritized actions with owners and deadlines]

Definition of Done

  • Blameless language throughout (no individual blame)
  • Timeline with timestamps and sources
  • 5 Whys reaching a systemic root cause
  • Impact quantified (users, duration, revenue, SLA)
  • Action items typed (prevent/detect/mitigate) with owners

Quality Criteria

  • Steps are executable in sequence without external context
  • Decision points have clear if/then branching
  • Rollback or abort procedures are documented for risky steps
  • Expected duration or time-per-step is estimated

Verification (4C)

CheckQuestion
CorrectnessDo the steps execute correctly in the order specified?
CompletenessAre decision points, error handling, and escalation paths all documented?
Context-fitCould someone with the right access but no prior context complete this runbook?
ConsequenceIf Step N fails and the operator skips to Step N+1, what breaks?

Edge Cases

  • Steps require access the operator doesn't have — Document exact access requirements upfront. Include escalation contact for emergency access.
  • Environment differs from documented state — Add a pre-flight check as Step 0 to verify prerequisites before starting.
  • Runbook is triggered during off-hours — Document who to contact and which steps can be safely deferred to business hours.

Changelog

  • v1.0.0 — Initial release

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Design rigorous A/B test plans with hypothesis, sample size calculation, Minimum Detectable Effect (MDE), randomization strategy, and decision rules. Includes guardrail metrics and rollout playbook. Use when planning product experiments, conversion optimization, or data-driven feature decisions.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

Review REST and GraphQL API designs for consistency, usability, and best practices. Covers naming conventions, versioning strategy, error format, pagination, authentication patterns, and breaking change detection. Use when reviewing API specs, designing new APIs, or auditing existing endpoints.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

Write Architecture Decision Records (ADRs) following the Michael Nygard format. Captures context, options considered, decision rationale, and consequences. Use when making technology choices, framework selections, or any architectural decision that future developers need to understand.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

Assess organizational readiness for financial audits (internal or external). Map assertions to account balances, check evidence completeness, score readiness using a Red/Amber/Green framework, and generate a remediation timeline. Aligned with SOX, IFRS, and GAAP audit standards. Use before scheduled audits or when preparing for first-time compliance.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

Design safe recurring Codex automations with clear prompts, outputs, schedules, and gating rules.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

Refine Product Backlog Items to meet INVEST criteria. Write User Stories with Acceptance Criteria in Given/When/Then format, estimate with Story Points, and flag dependencies. Use before sprint planning when backlog items need grooming. Do not use to prioritize the backlog — that is the Product Owner's decision.

日本語の概要は準備中です。原文の説明を表示しています。

hoavdc/CodexKit252026年10月11日 更新

hoavdc のスキルをすべて見る

このスキルの問題を報告する