本文へ移動
cccskills
無料GitHub で公開

google-cloud-waf-performance-optimization

Generates performance-focused guidance for Google Cloud workloads based on the design principles and recommendations in the Performance Optimization pillar of the Google Cloud Well-Architected Framework (WAF). Use this skill to evaluate a workload, identify performance requirements, and provide actionable recommendations for resource allocation, modular design, and elasticity.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md7.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Google Cloud Well-Architected Framework skill for the Performance Optimization pillar

Overview

The Performance Optimization pillar of the Google Cloud Well-Architected Framework provides principles and recommendations to help you design, build, and operate high-performing workloads. It focuses on efficiently allocating resources, leveraging modular architectures, and using data-driven insights to continuously monitor and improve performance as your business needs evolve.

Core principles

The recommendations in the performance optimization pillar of the Well-Architected Framework are aligned with the following core principles:

Relevant Google Cloud products

The following are examples of Google Cloud products and features that are relevant to performance optimization:

  • Compute and scaling

    • Compute Engine (MIGs): Managed instance groups that support autoscaling and load balancing for VM-based workloads.
    • Google Kubernetes Engine (GKE): Provides container orchestration with horizontal and vertical pod autoscaling.
    • Cloud Run: A fully managed serverless platform that automatically scales containers to zero or up based on traffic.
  • Data and caching

    • Cloud CDN: Low-latency content delivery network to cache static and dynamic content closer to end-users.
    • Memorystore: Managed in-memory data store for Valkey and Redis to provide sub-millisecond data access.
    • Bigtable: NoSQL database service for analytical and operational workloads requiring low latency and high throughput.
    • Spanner: RDBMS that provides global consistency, high availability, and horizontal scaling for mission-critical transactional applications.
  • Performance analysis and monitoring

    • Cloud Trace: Distributed tracing system that helps identify latency bottlenecks.
    • Cloud Profiler: Continuous CPU and memory profiling to identify resource-heavy application code.
    • Cloud Monitoring: Provides dashboards and alerts based on performance KPIs like latency and throughput.

Workload assessment questions

Ask appropriate questions to understand the performance-related requirements and constraints of the workload and the user's organization. Choose questions from the following list:

  • Plan resource allocation

    • When initially provisioning compute resources for a new application, which approach do you use to determine the required capacity for expected peak loads?
    • Which caching strategies (browser, in-memory, CDN, database) do you utilize to improve performance and responsiveness?
    • How do you optimize the performance of your data storage solutions (e.g., SSD vs HDD, storage classes) for your applications?
  • Promote modular design

    • Which architectural patterns (microservices, asynchronous messaging, stateless servers) do you employ to enhance performance and resilience?
    • How do you design your application to minimize the impact of failures in one part of the system on other parts?
  • Continuously monitor and improve performance

    • How frequently do you review and analyze the performance of your production applications and infrastructure?
    • Which tools or techniques (APM, distributed tracing, load testing) do you use to proactively identify and diagnose performance bottlenecks?
    • How do you incorporate performance considerations into your software development lifecycle (SDLC)?
  • Take advantage of elasticity

    • Which methods do you use to manage and optimize the cost of your cloud resources while maintaining performance?
    • How do you typically handle sudden spikes in traffic or workload on your applications?

Validation checklist

Use the following checklist to evaluate the architecture's alignment with performance optimization recommendations:

  • Resource allocation

    • Initial provisioning is based on load testing or historical data rather than general estimates.
    • Caching is implemented at multiple layers (CDN, in-memory, or browser) to offload backend systems.
    • Storage types (SSD/HDD) and classes are selected based on the specific I/O requirements of the workload.
  • Modular design

    • The architecture uses microservices or decoupled components to allow independent scaling.
    • Circuit breakers or bulkheads are implemented to isolate failures and prevent performance degradation across the system.
  • Monitoring and continuous improvement

    • Automated dashboards and alerts are configured for key performance indicators (KPIs).
    • Distributed tracing and profiling tools are used to identify code-level bottlenecks.
    • Performance testing (unit and integration) is integrated into the software development lifecycle.
  • Elasticity

    • Auto-scaling rules are configured and validated to handle variable demand.
    • The architecture leverages serverless or managed services to dynamically match capacity to load.
    • Resource utilization is reviewed regularly to eliminate idle overhead and balance cost with performance.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for standard infrastructure monitoring unrelated to AI agents, or when the agent is not instrumented with OpenTelemetry (for Reliability, Cost, Safety, Security alerts). NOTE: Reliability, Cost, Safety, and Security alerts use generic OTel metrics and work across runtimes (such as Cloud Run, Vertex AI). Quality alerts rely on Vertex AI Online Monitors and are strictly bound to Vertex AI deployments.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available models, check if a specific model is deployable (`gcloud ai model-garden models list-deployment-config`), query deployment cost, troubleshoot deployment errors (like quota limits), or undeploy/clean up endpoints. Also use when copying and deploying a 1P Tuned Model. Don't use for pure listing/discovery questions of the form "is X deployed?", "list my endpoints", or "which regions have models running?" — for those use `agent-platform-endpoint-management`. Don't use for public Vertex AI deployments (use the `vertex-deploy` skill) or for running model evaluations (use the `agent-platform-eval-flywheel` skill).

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for running model evaluations.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results before and after a fix, or when guidance is needed on Agent Platform eval methodology — including dataset schema, LLM-as-judge scoring, and common failure causes. For fine-tuning, use agent-platform-tuning. For general production deployment, use agent-platform-deploy.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate code for calling Gemini or OpenMaaS models, authenticate with GenAI SDK, OpenAI SDK, or legacy Agent Platform SDK, configure base URLs and global/regional endpoints, or troubleshoot 429 Resource Exhausted (DSQ), 400 User Validation, or 404 Not Found errors. Don't use for deploying models to endpoints or for running model evaluations.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry).

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

vaila-multimodaltoolbox のスキルをすべて見る

このスキルの問題を報告する