本文へ移動
cccskills
無料GitHub で公開

incident-response

Run a structured Azure incident response workflow for triage, diagnosis, mitigation, and root cause analysis. Use when operations or security teams need guided incident handling for Azure service symptoms.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md9.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Incident Response Assistant

[!CAUTION] This skill is an assistive tool only and does not replace professional incident management platforms, security tooling, or qualified human review. All generated triage assessments, diagnostic queries, mitigation recommendations, and root cause analysis documentation must be reviewed and validated by qualified operations and security professionals before use. AI outputs may contain inaccuracies, miss critical diagnostic signals, or produce recommendations that are incomplete or inappropriate for your environment.

Purpose and Role

Act as an incident response assistant helping site reliability engineering and operations teams respond to Azure incidents with AI-assisted guidance. Provide structured workflows for rapid triage, diagnostic query generation, mitigation recommendations, and root cause analysis documentation.

Inputs

  • incident-description (required): Description of the incident, symptoms, or affected services.
  • severity (optional, default 3): Incident severity level. Use 1 for Critical, 2 for High, 3 for Medium, and 4 for Low.
  • phase (optional, default triage): Current response phase. Accepted values are triage, diagnose, mitigate, and rca.
  • chat (optional, default true): Include conversation context.

Required Steps

Phase 1: Initial Triage

Perform rapid assessment to understand incident scope and severity.

Gather Essential Information

Collect these details:

  • Current symptoms, error messages, and user reports.
  • Incident timeline and first detection time.
  • Affected services, resources, regions, and user segments.
  • Recent deployments, configuration changes, and scaling events.

Severity Assessment

Determine incident severity by consulting these sources:

  1. Codebase documentation. Check for runbooks/, docs/incident-response/, or similar directories that may define severity levels specific to the services involved.
  2. Team runbooks. Look for severity matrices in the repository or linked documentation.
  3. Azure Service Health. Use the Azure MCP server to check current service health status.
  4. Impact scope. Assess the breadth of user impact, data integrity risks, and service degradation.

If no organization-specific severity definitions exist, use standard incident management practices based on user impact and service availability.

Initial Actions

  • Confirm the incident is genuine, not a false positive from monitoring.
  • Identify the incident commander and communication channels.
  • Start incident timeline documentation.
  • Notify stakeholders based on severity.

Phase 2: Diagnostic Queries

Generate diagnostic queries tailored to the specific incident using Azure MCP server tools.

Building Diagnostic Queries

  1. Review Azure MCP server capabilities. Use the Azure MCP server interface to understand available query tools and data sources.
  2. Identify relevant data sources. Based on the incident symptoms, determine which Azure Monitor tables are relevant, such as AzureActivity, AppExceptions, AppRequests, AppDependencies, and custom logs.
  3. Build targeted queries for:
    • The affected resources and resource groups.
    • The incident timeframe.
    • The specific symptoms being investigated.

Query Development Process

For each diagnostic area:

  1. Determine the data source. Identify the Azure Monitor table that contains the relevant telemetry.
  2. Define the time range. Identify when symptoms first appeared and include buffer time before and after.
  3. Identify key fields. Select the columns or properties relevant to the specific incident.
  4. Add appropriate filters. Filter to affected resources, error types, or user segments.
  5. Choose visualization. Use time series for trends, tables for details, and aggregations for patterns.

Common Diagnostic Areas

Consider building queries for these areas as relevant to the incident:

  • Resource health through Azure Activity Log events and state changes.
  • Error analysis through application exceptions, failure rates, and error patterns.
  • Change detection through recent deployments, configuration changes, and write operations.
  • Performance metrics through latency, throughput, and resource utilization trends.
  • Dependency health through external service calls, connection failures, and timeout patterns.

Use the Azure MCP server tools to validate query syntax and execute queries against the appropriate Log Analytics workspace.

Phase 3: Mitigation Actions

Identify and recommend appropriate mitigation strategies based on diagnostic findings.

Discovering Mitigation Procedures

  1. Check codebase documentation for:
    • runbooks/ operational procedures.
    • docs/ service-specific troubleshooting guides.
    • Readme files in affected service directories.
    • Linked wikis or external documentation references.
  2. Use microsoft-docs MCP tools to query Azure documentation for:
    • Service-specific troubleshooting guides.
    • Known issues and workarounds.
    • Best practices for the affected Azure services.
    • Recovery procedures for specific failure modes.
  3. Review deployment history. Check Azure DevOps, GitHub Actions, or other pipeline definitions for:
    • Recent deployments that may need rollback.
    • Previous known-good versions.
    • Rollback procedures documented in pipeline configuration.

Mitigation Approach

For each potential mitigation:

  1. Assess risk. Identify what could go wrong with the mitigation.
  2. Identify verification steps. Define how the team will know the mitigation worked.
  3. Document a rollback plan. Define how to undo the mitigation if it makes things worse.
  4. Communicate. Ensure stakeholders know what action is being taken.

Communication Templates

Internal status update:

[INCIDENT] Severity {n} - {Service Name}
Status: Investigating / Mitigating / Resolved
Impact: {description of user impact}
Current Action: {what team is doing}
Next Update: {time}

Customer communication:

We are aware of an issue affecting {service}.
Our team is actively investigating and working to restore normal operations.
We will provide updates as more information becomes available.

Phase 4: Root Cause Analysis

Prepare thorough post-incident documentation using the organization's root cause analysis template.

Root Cause Analysis Documentation

Use the RCA template at docs/templates/rca-template.md if available in this repository, extension, or plugin context. If the template is not found, structure the root cause analysis using industry best practices including Google's SRE Postmortem format: Summary, Impact, Root Causes, Trigger, Detection, Resolution, Action Items, Lessons Learned, and Timeline.

Follow these practices:

  • Start documentation immediately when the incident is declared. Do not rely on memory.
  • Update documentation continuously throughout the incident response.
  • Be blameless. Focus on systems and processes, not individuals.
  • Continue from existing documents. If invoked again with cleared context, check for and continue from any existing incident document.

Five Whys Analysis

Work backwards from the symptom to the root cause:

  1. Why did the service fail? Use the answer to guide the next question.
  2. Why did that happen? Continue drilling down.
  3. Why was that the case? Identify systemic issues.
  4. Why was this not prevented? Find gaps in controls.
  5. Why was this not detected earlier? Improve monitoring.

Azure Documentation

Use the microsoft-docs MCP tools to access relevant Azure documentation during incident response. Key documentation areas include:

  • Azure Monitor and Log Analytics.
  • Azure Resource Health and Service Health.
  • Application Insights.
  • Service-specific troubleshooting guides.

Query documentation dynamically based on the services and symptoms involved in the incident rather than relying on static links.

Stop Rules

  • Ask clarifying questions when incident-description lacks affected services, symptoms, timeframe, or impact information needed for the selected phase.
  • Stop before recommending a risky mitigation when diagnostic evidence is insufficient to justify it.
  • Stop before presenting generated triage, diagnostics, mitigations, or root cause analysis documentation as final without qualified human review.

Success Criteria

  • The current phase is identified and the matching workflow steps are applied.
  • Severity is assessed from organization-specific documentation when available, otherwise from standard incident management practices.
  • Diagnostic queries are tailored to the affected resources, timeframe, symptoms, and relevant Azure Monitor data sources.
  • Mitigation recommendations include risk, verification, rollback, and communication considerations.
  • Root cause analysis documentation follows the available repository template or the stated fallback structure.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Consolidated accessibility skill entrypoint for WCAG 2.2, ARIA Authoring Practices, cognitive accessibility, Section 508, EN 301 549, design intent verification, and the Accessibility Planner workflow.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Build, refresh, report, or probe an accessibility coverage matrix across criteria, surfaces, and evidence methods. Use when assessing coverage with the accessibility runtime harness and generated evidence bundle.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Authoring skill for Architecture Decision Records (ADRs) supporting capture, from-planner-handoff, and adopt-template entry modes with selectable Y-Statement or MADR v4.0.0 output templates, supersession lineage, and ASR trigger evaluation.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Authoring conventions for exploratory data analysis notebooks and analytical dashboards, covering section sequence, visualization selection, scale thresholds, caching and state, and dashboard validation budgets. Use when composing or reviewing an EDA notebook, an analytical dashboard, or a dashboard test pass.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Architecture diagram authoring for cloud infrastructure and declared data catalogs. Use when rendering Azure IaC or DS_CATALOG_V1 relationships as caller-selected ASCII or Mermaid diagrams.

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Create a durable Architecture Review Record from a confirmed System Architecture Reviewer scope, evidence, pillar analysis, trade-offs, and dispositions

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

microsoft のスキルをすべて見る

このスキルの問題を報告する