本文へ移動
cccskills
無料GitHub で公開

gcp-rightsizing

Analyze Compute Engine VM, Cloud SQL, Persistent Disk, and serverless (Cloud Functions/Cloud Run) utilization to identify right-sizing opportunities. Uses Cloud Monitoring metrics with anti-hallucination rules for E2 shared-core instances, sole-tenant nodes, preemptible/spot VMs, SUD eligibility, peak vs average analysis, CUD-aware savings, and estimated monthly savings calculations.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md14.6 KB
  • scripts/get_rightsizing_gcp.sh27.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

GCP Rightsizing Skill

Analyze resource utilization and identify right-sizing opportunities with anti-hallucination guardrails and reusable Cloud Monitoring functions.

Relationship to other GCP skills:

  • gcp-rightsizing/ -> "What is oversized" (utilization thresholds, downsize recommendations)
  • gcp-idle-resources/ -> "What is wasting money" (idle detection rules, cost estimation)
  • gcp/ -> "How to execute" (parallel patterns, monitoring aligners, billing/pricing scripts)

CRITICAL: Rightsizing Rules (Anti-Hallucination)

These rules are MANDATORY when analyzing resource utilization. Violating them produces incorrect recommendations that can cause outages.

Rule 1: Minimum 14-Day Observation Window

Short observation windows miss weekly patterns (batch jobs, weekend traffic, month-end spikes). Default --days 14, allow up to 90.

WRONG:  --days 3   -> Misses weekend batch jobs
WRONG:  --days 1   -> Captures only one day's pattern
CORRECT: --days 14 -> Captures at least 2 full weekly cycles

Rule 2: E2 Shared-Core Instances Have Hard CPU Limits

E2 shared-core VMs (e2-micro, e2-small, e2-medium) have CPU time limits, NOT burstable credits like AWS t-family. An e2-micro gets 2 vCPUs but is capped at 12.5% (0.25 vCPU equivalent) sustained. They cannot exceed their allocation even if idle beforehand.

MANDATORY: Do NOT treat E2 shared-core like AWS burstable. High CPU % on shared-core means the workload is hitting its hard cap.

WRONG:  e2-micro avg CPU 12% -> "underutilized, this is a 2-vCPU machine at only 12%"
CORRECT: e2-micro avg CPU 12% -> "near cap (12.5% limit). Consider upgrading to e2-small (25% cap)"
CORRECT: e2-small avg CPU 5% -> "using 5% of 25% cap, underutilized"

CPU cap reference:

Machine TypevCPUsCPU CapCap as %
e2-micro20.25 vCPU12.5%
e2-small20.50 vCPU25%
e2-medium21.00 vCPU50%

Rule 3: Sole-Tenant Nodes -- Report Node Utilization, Not VM

On sole-tenant nodes, VMs share a physical host. Rightsizing individual VMs without considering overall node fill rate is misleading.

MANDATORY: Flag sole-tenant VMs separately. The optimization is node fill rate, not individual VM utilization.

WRONG:  "vm-abc on sole-tenant node is at 5% CPU, downsize"
CORRECT: "vm-abc is on sole-tenant node node-group-xyz. Individual VM rightsizing deferred -- optimize node fill rate instead."

Rule 4: Peak vs Average -- Never Downsize on Average Alone

An instance with avg CPU 10% but max CPU 95% is a bursty workload. Downsizing would cause failures during peaks.

MANDATORY: Report BOTH Average AND Maximum statistics. Only flag for downsizing if max < threshold too.

WRONG:  avg CPU 10% -> "downsize"
CORRECT: avg CPU 10%, max CPU 22% -> "downsize candidate (both avg and max are low)"
CORRECT: avg CPU 10%, max CPU 95% -> "bursty workload, do NOT downsize"

Rule 5: Preemptible/Spot VMs -- Skip Rightsizing

Preemptible and Spot VMs already run at 60-91% discount. Rightsizing savings are marginal and the workload is already optimized for cost.

MANDATORY: Exclude preemptible/spot VMs from rightsizing analysis.

WRONG:  "preemptible vm-abc is underutilized, downsize"
CORRECT: "vm-abc is preemptible/spot -- skipped (already cost-optimized)"

Rule 6: Savings Estimates Must Caveat CUD/SUD Coverage

GCP applies Sustained Use Discounts (SUDs) automatically (up to 30% for N1/N2) and Committed Use Discounts (CUDs) contractually. Downsizing a CUD-covered VM does NOT immediately save money -- the commitment continues.

MANDATORY: Caveat all savings estimates. Query CUD coverage when possible.

WRONG:  "Downsize to save $70/mo"
CORRECT: "Estimated on-demand savings: $70/mo. Note: if CUD-covered, savings may not apply until commitment expires. SUD automatically adjusts."

Rule 7: Cloud SQL activationPolicy + HA Doubles Compute

Cloud SQL with availabilityType=REGIONAL (HA) doubles compute cost (standby replica). A db-custom-4-16384 HA instance costs 2x the single-zone price.

MANDATORY: Check availabilityType before estimating Cloud SQL costs.

WRONG:  db-custom-4-16384 -> $165/mo
CORRECT: db-custom-4-16384, availabilityType=REGIONAL -> $330/mo (HA doubles compute)
CORRECT: db-custom-4-16384, availabilityType=ZONAL -> $165/mo

Rule 8: Persistent Disk Type Migration (pd-standard -> pd-balanced)

pd-standard (HDD) is $0.04/GiB/mo but limited to 0.75 IOPS/GiB. pd-balanced (SSD) is $0.10/GiB/mo with 6 IOPS/GiB baseline. A 200GiB pd-standard has only 150 IOPS baseline.

MANDATORY: Compare actual IOPS usage against baseline before recommending type changes.

WRONG:  "pd-standard is slower, just switch to pd-balanced"
CORRECT: "200GiB pd-standard: 150 IOPS baseline, actual avg 40 IOPS (27%). Right-sized for IOPS."
CORRECT: "200GiB pd-standard: 150 IOPS baseline, actual avg 140 IOPS (93%). Upgrade to pd-balanced for 1200 IOPS baseline."

Rule 9: Cloud Functions/Cloud Run -- Flag Waste, Don't Prescribe Values

Serverless resource configuration is highly workload-specific. Flag obvious waste (allocated 4GiB, uses 200MiB) but do NOT recommend specific memory/CPU values. Let the user benchmark.

WRONG:  "Cloud Function allocates 2GiB, reduce to 256MiB"
CORRECT: "Cloud Function fn-abc allocates 2048MiB but peak memory usage is 180MiB. Candidate for memory reduction (benchmark required)."

Rule 10: SUD Eligibility Varies by Machine Family

N1 and N2 families get automatic SUDs (up to 30%). E2 and T2D do NOT get SUDs. C3 and M3 do NOT get SUDs (use CUDs instead).

MANDATORY: Include SUD eligibility when showing savings.

WRONG:  "All VMs get 30% sustained use discount"
CORRECT: "n2-standard-4: SUD-eligible (up to 30% automatic discount)"
CORRECT: "e2-standard-4: NOT SUD-eligible (use CUDs for discounts)"

Mandatory Pre-Analysis Checklist

Before writing ANY rightsizing analysis, verify ALL of the following:

  • Observation window >= 14 days (see Rule 1)
  • E2 shared-core instances analyzed against their CPU cap, not raw vCPU count (see Rule 2)
  • Sole-tenant VMs identified and handled separately (see Rule 3)
  • Both Average AND Maximum statistics reported for all metrics (see Rule 4)
  • Preemptible/Spot VMs excluded (see Rule 5)
  • Savings caveated as "based on on-demand rates; CUD/SUD may apply" (see Rule 6)
  • Cloud SQL HA flag checked (see Rule 7)
  • All cost figures include currency unit (USD)
  • Parallel execution used for all Cloud Monitoring queries (see gcp/SKILL.md)

Rightsizing Script (get_rightsizing_gcp.sh)

DO NOT read or modify the script file. Only source and call the functions.

SETUP (at the start of your script):

source ./_skills/connections/gcp/gcp-rightsizing/scripts/get_rightsizing_gcp.sh

All functions enforce anti-hallucination rules: 14-day minimum window, parallel Cloud Monitoring queries, TOON output format.

FUNCTION REFERENCE:

FunctionPurposeSignature
gcp_rightsizing_vmsCompute Engine VM CPU analysis with shared-core, sole-tenant, preemptible handling[--days N] [--project PROJECT]
gcp_rightsizing_cloudsqlCloud SQL CPU + connections with HA cost awareness[--days N] [--project PROJECT]
gcp_rightsizing_disksPersistent Disk IOPS/throughput utilization, type migration candidates[--days N] [--project PROJECT]
gcp_rightsizing_serverlessCloud Functions + Cloud Run memory/CPU waste detection[--days N] [--project PROJECT]
gcp_rightsizing_summaryRun all checks, output unified summary[--days N] [--project PROJECT]

RECOMMENDED WORKFLOW (every rightsizing analysis):

  1. Always run gcp_rightsizing_vms first -- Compute Engine is typically the largest compute cost
  2. Run gcp_rightsizing_cloudsql for database layer analysis
  3. Run gcp_rightsizing_disks for storage optimization (pd-standard -> pd-balanced candidates)
  4. Run gcp_rightsizing_serverless if Cloud Functions/Cloud Run is a significant cost driver
  5. Or run gcp_rightsizing_summary for a unified view across all resource types

Examples:

source ./_skills/connections/gcp/gcp-rightsizing/scripts/get_rightsizing_gcp.sh

# VM utilization analysis (14-day default)
gcp_rightsizing_vms

# VM analysis with 30-day window in specific project
gcp_rightsizing_vms --days 30 --project my-project-id

# Cloud SQL utilization
gcp_rightsizing_cloudsql --days 14

# PD rightsizing (type migration candidates)
gcp_rightsizing_disks --days 14

# Cloud Functions + Cloud Run analysis
gcp_rightsizing_serverless --days 14

# Full summary across all resource types
gcp_rightsizing_summary --days 30 --project my-project-id

GCP Machine Type Quick Reference

Approximate monthly on-demand costs in USD (us-central1, tax-exclusive). Use to validate savings estimates.

FamilyMachine TypevCPUMemory~Monthly USDSUD Eligible
E2e2-micro2 (shared 0.25)1 GB$6.11No
E2e2-small2 (shared 0.50)2 GB$12.23No
E2e2-medium2 (shared 1.00)4 GB$24.46No
E2e2-standard-228 GB$48.92No
E2e2-standard-4416 GB$97.83No
E2e2-standard-8832 GB$195.67No
N2n2-standard-228 GB$56.82Yes (up to 30%)
N2n2-standard-4416 GB$113.63Yes
N2n2-standard-8832 GB$227.26Yes
C3c3-standard-4416 GB$120.37No (CUD only)
N1n1-standard-113.75 GB$24.27Yes (up to 30%)
N1n1-standard-227.5 GB$48.55Yes

Downsizing savings: Moving one size down within a family typically saves ~50% (e.g., e2-standard-8 $196 -> e2-standard-4 $98 = $98/mo savings).

For machine types not listed above, use get_gcp_cost from the gcp/ skill.


Cloud SQL Tier Reference

TiervCPUMemory~Monthly USD (Zonal)~Monthly USD (HA)
db-f1-microshared0.6 GB$10.80N/A
db-g1-smallshared1.7 GB$36.00N/A
db-custom-1-384013.75 GB$49.64$99.29
db-custom-2-768027.5 GB$99.29$198.58
db-custom-4-15360415 GB$198.58$397.15
db-custom-8-30720830 GB$397.15$794.30

Persistent Disk Performance Reference

Disk Type$/GiB/moIOPS/GiBMax IOPSThroughput/GiB
pd-standard$0.040.75 R / 1.5 W7,5000.12 MB/s
pd-balanced$0.10680,0000.28 MB/s
pd-ssd$0.1730100,0000.48 MB/s
pd-extreme$0.125 + IOPSconfigurable120,0001.2 GB/s

Cloud Monitoring Metrics Used

ResourceMetricNamespaceAligner
VM CPUinstance/cpu/utilizationcompute.googleapis.comALIGN_MEAN, ALIGN_MAX
Cloud SQL CPUdatabase/cpu/utilizationcloudsql.googleapis.comALIGN_MEAN, ALIGN_MAX
Cloud SQL connectionsdatabase/network/connectionscloudsql.googleapis.comALIGN_MEAN
PD Read IOPSinstance/disk/read_ops_countcompute.googleapis.comALIGN_RATE
PD Write IOPSinstance/disk/write_ops_countcompute.googleapis.comALIGN_RATE
Cloud Function executionsfunction/execution_countcloudfunctions.googleapis.comALIGN_RATE
Cloud Function memoryfunction/user_memory_bytescloudfunctions.googleapis.comALIGN_MAX
Cloud Run request countrequest_countrun.googleapis.comALIGN_RATE
Cloud Run memory utilcontainer/memory/utilizationsrun.googleapis.comALIGN_MAX

CRITICAL: alignment-period MUST be >= 60 seconds when using aligners other than ALIGN_NONE.


Common Errors

ErrorCauseSolution
PERMISSION_DENIED on monitoringMissing monitoring.timeSeries.list permissionCheck service account roles
No data for VM CPUInstance was recently created or restartedExtend observation window
E2 shared-core appears underutilizedComparing against 2 vCPUs instead of CPU capUse the CPU cap table (Rule 2)
Savings estimate seems too highHA Cloud SQL not accounted forCheck availabilityType, double cost if REGIONAL (Rule 7)
Preemptible VM flagged for rightsizingScript didn't filter scheduling policyCheck scheduling.preemptible field (Rule 5)
Cloud SQL shows as idle but has replicasRead replicas are separate instancesCheck for read replica relationships
Cloud Function memory unclearMemory metric shows allocated, not peakUse function/user_memory_bytes with ALIGN_MAX

Output Format

Present results as a structured report:

Gcp Rightsizing Report
══════════════════════
Resources discovered: [count]

Resource       Status    Key Metric    Issues
──────────────────────────────────────────────
[name]         [ok/warn] [value]       [findings]

Summary: [total] resources | [ok] healthy | [warn] warnings | [crit] critical
Action Items: [list of prioritized findings]

Target ≤50 lines of output. Use tables for multi-resource comparisons.

Counter-Rationalizations

ShortcutCounterWhy
"I'll skip discovery and check known resources"Always run Phase 1 discovery firstResource names change, new resources appear — assumed names cause errors
"The user only asked for a quick check"Follow the full discovery → analysis flowQuick checks miss critical issues; structured analysis catches silent failures
"Default configuration is probably fine"Audit configuration explicitlyDefaults often leave logging, security, and optimization features disabled
"Metrics aren't needed for this"Always check relevant metrics when availableAPI/CLI responses show current state; metrics reveal trends and intermittent issues
"I don't have access to that"Try the command and report the actual errorAssumed permission failures prevent useful investigation; actual errors are informative

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Cloudflare GraphQL Analytics for zone traffic, firewall events, Workers metrics, and schema exploration. Use when querying Cloudflare analytics data or exploring the GraphQL API.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Use when working with Alloydb — google AlloyDB instance analysis, query insights, columnar engine optimization, maintenance windows, and cluster health.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Use when working with Aqua — aqua Security platform analysis. Covers container runtime protection, image assurance policies, compliance frameworks, vulnerability management, workload protection, and registry scanning. Use when analyzing container security posture, reviewing image compliance, investigating runtime alerts, or auditing security policies.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Use when working with Bigquery — google BigQuery job analysis, slot utilization, cost analysis, dataset management, and query optimization.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Use when working with Cassandra — apache Cassandra keyspace analysis, compaction strategies, repair status, nodetool operations, and cluster health monitoring.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

Use when working with Checkov — checkov infrastructure-as-code security scanning. Covers Terraform, CloudFormation, Kubernetes, and Dockerfile scanning, policy management, custom checks, compliance frameworks, and suppression management. Use when scanning IaC for security misconfigurations, evaluating compliance, managing custom policies, or reviewing scan results.

日本語の概要は準備中です。原文の説明を表示しています。

cloudthinker-ai/CloudSkills62026年4月5日 更新

cloudthinker-ai のスキルをすべて見る

このスキルの問題を報告する