本文へ移動
cccskills
無料GitHub で公開

infrastructure-drift-detection

Drift is the divergence between declared infrastructure and live infrastructure.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md10.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Infrastructure Drift Detection (Verbose)

Core Patterns

Where Drift Actually Comes From

Drift is the divergence between declared infrastructure and live infrastructure. Naming its sources matters because they call for different remedies.

SourceExampleRemedy
Emergency console changeSecurity group widened during an incidentRemove standing write access; codify after
Convenience console change"Just bumping the instance size"Same — plus a faster pipeline
Provider defaultsTags, ARNs, or KMS keys injected at createCodify the default, or ignore narrowly
Controllers and autoscalersdesired_count, node group sizeNarrow ignore_changes
Out-of-band automationA Lambda that rotates a ruleBring it into code or exclude explicitly
Manual state surgerystate rm after a botched applyReview process; import rather than remove

The dominant source, by a wide margin, is the first two. Someone with production console access fixes a real problem in ninety seconds and does not come back to write it down — because the incident is over and the fix is working.

What makes this expensive is the delay. The infrastructure is now correct and the code is now wrong, and nothing reports the discrepancy. Weeks later an unrelated pull request adds a tag; the plan includes a quiet ~ line narrowing that security group back to its declared value; the apply reproduces the original outage with no apparent cause. The eventual timeline reconstruction (see incident-timeline-creation) will show a deploy that "changed nothing relevant".

The Detection Mechanism

terraform plan refreshes state against the provider APIs and diffs. The flag that makes it automatable is -detailed-exitcode:

terraform init -input=false
terraform plan -detailed-exitcode -lock-timeout=5m -out=drift.tfplan
Exit codeMeaningJob behaviour
0No changesSuccess, no notification
1Error — credentials, lock contention, API failurePage the platform team
2Changes presentDrift finding: report and route

Collapsing 1 and 2 into "non-zero means drift" is the most common implementation bug: an expired credential then looks identical to a clean run of findings, and a persistently broken job reads as persistent drift until someone stops looking.

One important caveat: a scheduled plan against main reports both true drift and merged-but-unapplied code. Run the check against the ref that is actually deployed — tag it on apply and check out that tag — or you will spend your first month triaging your own backlog.

For a read-only check without producing a plan artifact, newer Terraform supports terraform plan -refresh-only -detailed-exitcode, which reports only changes the provider made outside Terraform and excludes changes originating in your code. That is usually the more precise drift signal.

Scheduling

Drift detection belongs on a timer, not in the deploy pipeline. Checking at deploy time is checking at the moment it is most expensive to act: the finding lands in front of an engineer who is trying to ship something unrelated, mixed into a plan they are about to approve.

name: drift-detection
on:
  schedule:
    - cron: "0 */6 * * *"
  workflow_dispatch: {}

permissions:
  id-token: write
  contents: read

jobs:
  detect:
    strategy:
      fail-fast: false
      matrix:
        stack: [network, data, platform, apps]
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ env.DEPLOYED_REF }}
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::111122223333:role/tf-drift-readonly
          aws-region: us-east-1
      - run: terraform -chdir=infra/${{ matrix.stack }}/prod init -input=false
      - id: plan
        continue-on-error: true
        run: |
          terraform -chdir=infra/${{ matrix.stack }}/prod \
            plan -refresh-only -detailed-exitcode -lock-timeout=5m -out=drift.tfplan
      - if: steps.plan.outcome == 'failure' && steps.plan.conclusion != 'success'
        run: ./scripts/report-drift.sh infra/${{ matrix.stack }}/prod/drift.tfplan

Cadence guidance:

EnvironmentIntervalRationale
Regulated / PCI / HIPAAHourlyEvidence requirement; short unauthorized-change window
Production4–6 hoursCatches same-day drift while the context is fresh
StagingDailyLower stakes, still worth trending
EphemeralNeverRebuilt from code anyway

Two hard rules for the scheduled job:

Read-only credentials. The role should hold Describe*/Get*/List* and nothing else. This is not merely defence in depth — it makes the next rule structurally impossible to violate.

Never auto-apply. An auto-remediating drift job reverts the mitigation an on-call engineer put in place, unattended, in the middle of the night, with no human in the loop. The correct output of a drift job is a notification.

Reporting

Raw plan output is unreadable at 200 lines. Parse the JSON:

terraform show -json drift.tfplan | jq -r '
  .resource_changes[]
  | select(.change.actions | inside(["no-op"]) | not)
  | "\(.change.actions | join(",")) \(.address)"
'

For attribute-level detail:

terraform show -json drift.tfplan | jq -r '
  .resource_changes[]
  | select(.change.actions[0] == "update")
  | .address as $a
  | .change.before as $b
  | .change.after  as $x
  | ($b | keys[]) as $k
  | select($b[$k] != $x[$k])
  | "\($a).\($k): \($b[$k]) -> \($x[$k])"
'

A good report answers three questions: which resource, which attribute, and who changed it. The third comes from the audit log, not from Terraform:

aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=ResourceName,AttributeValue=sg-0abc123 \
  --start-time "$(date -u -d '7 days ago' +%Y-%m-%dT%H:%M:%SZ)" \
  --query 'Events[].[EventTime,Username,EventName]' --output table

Route by owning team using resource tags. A shared #infra-alerts channel that receives every finding for every stack is a channel people mute within a fortnight.

Triage: Codify, Revert, or Ignore

Every finding resolves to exactly one of three outcomes, and leaving it unresolved is what makes drift chronic.

Codify when the manual change was correct. Write it into the module, open a PR, let the plan confirm it produces no diff, apply. This is the common case for incident mitigations, and it is the step teams skip.

Revert when the change was unauthorized or accidental. Apply the declared configuration — after confirming with the owning team that nothing depends on the live value. Then ask how the write happened; an unexplained production change is a security finding, not just a hygiene one.

Ignore when the attribute is legitimately owned by something else:

resource "aws_ecs_service" "api" {
  # ...
  lifecycle {
    ignore_changes = [
      desired_count,                 # managed by application autoscaling
      task_definition,               # managed by the deploy pipeline
    ]
  }
}

Name the attributes. ignore_changes = all converts the resource into a permanent blind spot — it will never report drift again, including a security group opened to the internet.

Complementary Controls

Terraform plan sees only what Terraform manages. Two gaps need separate coverage:

Unmanaged resources. Anything created outside code is invisible to plan entirely. Cloud-native config tooling — AWS Config rules, Azure Policy, GCP Organization Policy — evaluates every resource regardless of provenance, and is the right place for absolute rules like "no security group allows 0.0.0.0/0 on 22".

Preventive guardrails. Service control policies and IAM boundaries that deny the mutating action outright beat detecting it afterwards. Drift detection is a backstop for what the guardrails allow through.

Track findings over time. A resource that drifts on the same attribute every week is telling you the code is wrong, or the team's real workflow does not route through the pipeline. That is a process defect — feed it into the improvement loop (see continuous-improvement), because remediating it individually forever is not a solution.

Common Anti-Patterns

❌ Detecting drift only during a deploy — the finding arrives when it is most disruptive and least likely to be triaged properly. ✅ Scheduled detection every few hours, decoupled from releases.

❌ Treating exit codes 1 and 2 as the same — an expired credential is indistinguishable from clean findings, and a broken job looks like a busy one. ✅ Branch on the exit code: 1 pages, 2 reports.

❌ Auto-applying to "self-heal" — unattended reversion of an on-call engineer's mitigation, usually reproducing the original incident. ✅ Report only; humans decide codify, revert, or ignore. Read-only credentials make this enforceable.

❌ ignore_changes = all — silences the alarm rather than fixing the noise, including for changes that matter. ✅ Ignore specific attributes, with a comment naming what owns them.

❌ Planning against main and calling unapplied code "drift" — floods the report with your own pending changes. ✅ Check against the deployed ref, or use -refresh-only to isolate external changes.

❌ Standing production console write access for the whole team — this is the drift source; detection alone cannot outrun it. ✅ Read-only by default, time-boxed break-glass that alerts on assumption.

❌ Findings that are noted and forgotten — the same drift reappears every scheduled run and everyone learns to ignore the report. ✅ Every finding gets an owner and a resolution: codified, reverted, or ignored with justification.

❌ Assuming plan covers everything — resources created outside Terraform never appear in a plan at all. ✅ Pair with AWS Config / Azure Policy / GCP Org Policy for provenance-independent evaluation.

Drift Detection Checklist

  • Scheduled drift job on production, every 4–6 hours or better
  • Job runs against the deployed ref, not the tip of main
  • -detailed-exitcode used, with 1 and 2 handled differently
  • -refresh-only used where the goal is external-change detection
  • Detection role is strictly read-only
  • Job never applies
  • Findings parsed to resource and attribute level, not raw plan text
  • Reports routed to owning teams by tag
  • Audit log correlation identifies who made each change
  • Every finding closed as codify / revert / ignore-with-reason
  • ignore_changes scoped to named attributes; no all
  • Humans have no standing production write access
  • Cloud-native policy evaluation covers unmanaged resources
  • Recurring drift on the same resource escalated as a process defect

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill whenever the user wants to create or improve a presentation for an academic context — conference papers, seminar talks, thesis defenses, grant briefings, lab meetings, invited lectures, or any presentation where the audience will evaluate reasoning and evidence. Triggers include: 'conference talk', 'seminar slides', 'thesis defense', 'research presentation', 'academic deck', 'academic presentation'. Also triggers when the user asks to 'make slides' in combination with academic content (e.g., 'make slides for my paper on X', 'create a presentation for my dissertation defense', 'build a deck for my grant proposal'). This skill governs CONTENT and STRUCTURE decisions. For the technical work of creating or editing the .pptx file itself, also read the pptx SKILL.md.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

adp

無料

Redpanda's Agentic Data Plane: governance infrastructure for building, running, and governing AI agents and MCP servers, plus a proxying AI Gateway for LLM providers, operated via `rpk ai` and the ADP API. Use when creating or managing AI agents (managed or self-managed) via `rpk ai agent` or `AgentRegistryService`; configuring MCP servers (remote or managed catalog, code mode, auth); setting up LLM providers or querying models via `rpk ai llm`/`rpk ai model` or the AI Gateway proxy; or configuring budgets, guardrails, or Cedar access-control policies through the governance APIs. Also covers reading agent transcripts and spending insights, and wiring OAuth clients or providers to the aigw Authorization Server. For the separate rpk cloud mcp control-plane MCP server, see `/redpanda:rpk-cloud`.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads

無料

Operate professional paid advertising across Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, and X. Use for account intake, source-grounded audits, strategy, budget and measurement planning, creative production, experiments, reporting, monitoring, and explicitly approved campaign changes. Also trigger on PPC, paid social, retail media, attribution, tracking, landing pages, cross-platform conversion totals, negative keywords or search terms, beta-feature scoring, stale platform claims, API-token or credential setup, campaign deletion, and safe Claude Ads installation or uninstall.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads-apple

無料

Audit Apple Ads measurement, AdServices and AdAttributionKit, campaign and keyword structure, Search Match, App Store placements, custom product pages, bidding, budgets, MMP reconciliation, and policy. Use for Apple Ads, Apple Search Ads, App Store ads, Search Match, custom product pages, AdServices, or Apple app-install campaigns.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

Research competitor paid-ad presence, messaging, creative, formats, landing pages, keyword and auction signals, transparent ad libraries, and strategic gaps across supported platforms. Use for competitor ads, ad libraries, ad spy, competitive PPC analysis, competitor creative, Google Ads Transparency, Meta Ad Library, or paid-media competitor research.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads-dna

無料

Extract a public-safe brand and offer profile for paid advertising from an authorized website and operator input. Triggers on: brand DNA, brand profile, brand identity, brand style, brand colors, brand voice, visual identity, style guide, website brand analysis.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

skillmds のスキルをすべて見る

このスキルの問題を報告する