Generates reliability-focused guidance for Google Cloud workloads based on the design principles and recommendations in the Google Cloud Well-Architected Framework. Use this skill to evaluate a workload, identify reliability requirements, and provide actionable recommendations for build, deploy, and manage the workload reliably in Google Cloud.
日本語の概要は準備中です。原文の説明を表示しています。
tensology/decisionsai☆ 172026年10月10日 更新
Knowledge base from the NIST/SEMATECH e-Handbook of Statistical Methods (NIST HB 151) — the practical statistics reference for engineering, metrology, and quality. Use for: exploratory data analysis and the four univariate assumptions (4-plot); measurement process characterization including bias/precision, calibration designs, gauge R&R, and ISO/GUM uncertainty budgets; production process characterization (stability vs. capability); process modeling and regression (LS/WLS/NLS/LOESS); design of experiments (screening, fractional/factorial, response-surface, Taguchi); statistical process control (Shewhart/CUSUM/EWMA charts, capability indices, acceptance sampling); product/process comparisons (hypothesis tests and confidence intervals for 1/2/3+ groups, ANOVA, multiple comparisons); and reliability (lifetime & repair-rate models, accelerated testing, reliability growth). Scope limits: this is applied frequentist statistics for measurement and quality — it does NOT reproduce the per-distribution formula galleries, worked case studies, datasets, plot images, or Dataplot/R code of the original web Handbook (those are described, not copied); it is thin on modern machine learning, Bayesian methods beyond conjugate reliability priors, time-series/forecasting, and Bayesian experimental design.
日本語の概要は準備中です。原文の説明を表示しています。
jgsystemsconsulting/jgs-se-knowledge-packs☆ 82026年10月9日 更新
Knowledge base from NASA-STD-8729.1A — NASA's R&M (Reliability and Maintainability) standard covering spaceflight and support systems (Revision A, 2017-06-13). Use for NASA's objectives-driven R&M framework — the R&M objectives hierarchy (one Top Objective decomposed into four sub-objectives, paired with tailorable strategies), how R&M requirements are established inside the SMA Plan required by NPR 7120.5, the SMA Technical Authority concurrence/independent-evaluation governance, milestone vs. readiness review gates, the reliability/maintainability/availability vocabulary (Ai vs. Ao, the failure causal chain, risk as a triplet), and the R&M Evidentiary Methods catalogue (FMEA/FMECA, FTA, RBDA, RCM, LORA, plus the space-environment and parts-pedigree analyses). Scoped to NASA spaceflight and support systems and to assurance planning: it is an objectives-and-strategies standard, so it deliberately does NOT prescribe step-by-step analysis procedures (those live in the referenced NASA Preferred Reliability Practices) and is thin on facility R&M, detailed math/derivations, and non-NASA or commercial life cycles.
日本語の概要は準備中です。原文の説明を表示しています。
jgsystemsconsulting/jgs-se-knowledge-packs☆ 82026年10月9日 更新
Use when working with Blameless — blameless incident management, SLO tracking, retrospectives, and reliability insights. Covers incident lifecycle, blameless retrospective facilitation, follow-up tracking, reliability scorecards, and incident type categorization. Use when managing active incidents, conducting retrospectives, tracking SLO compliance, or analyzing reliability trends in Blameless.
日本語の概要は準備中です。原文の説明を表示しています。
cloudthinker-ai/CloudSkills☆ 62026年4月5日 更新
Google SRE (Site Reliability Engineering) practices distilled from the SRE book series and real Google Brain/Meta/ByteDance production experience. Covers SLI/SLO/Error Budget design, incident response, capacity planning, toil automation, monitoring golden signals, and chaos engineering principles. USE WHEN: designing reliability targets, setting up monitoring/alerting, planning capacity, reducing operational load, building incident response runbooks, implementing chaos engineering, defining error budgets, or establishing reliability culture that balances velocity with stability.
日本語の概要は準備中です。原文の説明を表示しています。
aAAaqwq/openclaw-team☆ 22026年6月18日 更新
Implements reliability patterns including circuit breakers, retries, fallbacks, bulkheads, and SLO definitions. Provides failure mode analysis and incident response plans. Use for "SRE", "reliability", "resilience", or "failure handling".
日本語の概要は準備中です。原文の説明を表示しています。
sathishssj3/Stereix-Engine☆ 22026年10月4日 更新
Complete observability & reliability engineering system. Use when designing monitoring, implementing structured logging, setting up distributed tracing, building alerting systems, creating SLO/SLI frameworks, running incident response, conducting post-mortems, or auditing system reliability. Covers all three pillars (logs/metrics/traces), alert design, dashboard architecture, on-call operations, chaos engineering, and cost optimization.
日本語の概要は準備中です。原文の説明を表示しています。
johnalbertini14-glitch/openclaw-skills☆ 22026年2月26日 更新
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Use when defining, reviewing, or operating SLOs/SLIs/error budgets. Triggers on "define an SLO", "what should our SLO be", "error budget", "burn rate", "SLI", "service level objective", "Google SRE workbook", "multi-window burn-rate alert", or any reliability-target question. Ships SLO designer, error-budget calculator with multi-window burn-rate thresholds, and SLO reviewer that catches the common bugs (target too aggressive, window too short, conflicting SLOs, no SLI definition). 4 references on SLO principles + SLI design + error budget math + composition with feature-flags-architect/chaos-engineering/kubernetes-operator. NOT a generic observability skill — specifically the SLO discipline.
日本語の概要は準備中です。原文の説明を表示しています。
alirezarezvani/claude-skills☆ 2.8万2026年8月30日 更新
Orchestrates comprehensive production readiness reviews and assessments for GKE clusters and workloads across scalability, security, reliability, observability, backup/DR, and cost optimization. Use when asked to productionize, prepare, assess, audit, or review a GKE cluster or workload before going live to production. Don't use for deep-dive single-domain implementation (use specific domain skills like gke-workload-scaling, gke-platform-security, gke-workload-security, gke-service-networking, gke-reliability instead).
日本語の概要は準備中です。原文の説明を表示しています。
google/skills☆ 2.1万2026年10月10日 更新
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for standard infrastructure monitoring unrelated to AI agents, or when the agent is not instrumented with OpenTelemetry (for Reliability, Cost, Safety, Security alerts). NOTE: Reliability, Cost, Safety, and Security alerts use generic OTel metrics and work across runtimes (such as Cloud Run, Vertex AI). Quality alerts rely on Vertex AI Online Monitors and are strictly bound to Vertex AI deployments.
日本語の概要は準備中です。原文の説明を表示しています。
google/skills☆ 2.1万2026年10月10日 更新
Makes systems debuggable and reliably operable — instrumentation, alerting that is worth waking for, service objectives, and learning from failure. Use this to instrument a service, fix alerting that is ignored, set error budgets or reliability targets, prepare for on-call, or run a blameless post-incident review.
日本語の概要は準備中です。原文の説明を表示しています。
cbrock84/headcount☆ 2,0312026年9月18日 更新
Assesses whether study results are trustworthy by auditing design integrity, sample structure, statistical handling, bias control, validation chain, and claim discipline. It identifies where results are robust, fragile, overfit, under-validated, or overclaimed. Always separate reported findings from reliability judgment. Never fabricate references, PMIDs, DOIs, trial identifiers, study features, or validation claims.
日本語の概要は準備中です。原文の説明を表示しています。
aipoch/medical-research-skills☆ 1,9392026年9月17日 更新
You are an SLO (Service Level Objective) expert specializing in implementing reliability standards and error budget-based practices. Design SLO frameworks, define SLIs, and build monitoring that balances reliability with delivery velocity.
日本語の概要は準備中です。原文の説明を表示しています。
rmyndharis/antigravity-skills☆ 1,7312026年10月1日 更新
Expert database administrator specializing in modern cloud databases, automation, and reliability engineering. Masters AWS/Azure/GCP database services, Infrastructure as Code, high availability, disaster recovery, performance optimization, and compliance. Handles multi-cloud strategies, container databases, and cost optimization. Use PROACTIVELY for database architecture, operations, or reliability engineering.
日本語の概要は準備中です。原文の説明を表示しています。
rmyndharis/antigravity-skills☆ 1,7312026年10月1日 更新
Make an AI agent or automation reliable enough to trust — the tests, checks, and guardrails that catch its failures before they reach anything real. Use when asked how do I test my AI agent, make my automation reliable, my agent works sometimes, or how do I trust an AI workflow in production. Produces a map of where the agent can fail (bad input, hallucination, wrong tool call, edge cases, silent errors), the checks that catch each (validation, evals on real cases, human-in-the-loop gates, monitoring), a right-sized reliability plan scaled to the stakes, and a rollout that earns trust incrementally — so an agent that works in a demo becomes one that works in reality. For builders putting AI agents into real workflows.
日本語の概要は準備中です。原文の説明を表示しています。
mohitagw15856/pm-claude-skills☆ 1,4362026年10月10日 更新
Design and operationalize SRE dashboards that surface reliability, latency, error, saturation, and capacity signals across services. Use when building observability views for SLOs, incident response, and executive reliability reporting.
日本語の概要は準備中です。原文の説明を表示しています。
BagelHole/DevOps-Security-Agent-Skills☆ 1,1542026年5月22日 更新
Audit a proposed assessment for construct validity, reliability, and alignment to learning objectives. Use when reviewing or quality-assuring assessments before deployment.
日本語の概要は準備中です。原文の説明を表示しています。
GarethManning/education-agent-skills☆ 8452026年8月29日 更新
Expert knowledge for Azure Service Health development including troubleshooting, decision making, limits & quotas, security, configuration, and integrations & coding patterns. Use when handling Service Health APIs, Resource Graph queries, webhooks, VM Resource Health, or retirement alerts, and other Azure Service Health related development tasks. Not for Azure Monitor (use azure-monitor), Azure Reliability (use azure-reliability), Azure Resiliency (use azure-resiliency).
日本語の概要は準備中です。原文の説明を表示しています。
MicrosoftDocs/Agent-Skills☆ 7772026年10月11日 更新
Expert knowledge for Azure Resiliency development including security, configuration, and deployment. Use when configuring resiliency drills, backup/replication policies, Recovery Plans RBAC, or Infra Resiliency Manager coverage, and other Azure Resiliency related development tasks. Not for Azure Reliability (use azure-reliability), Azure Site Recovery (use azure-site-recovery), Azure Backup (use azure-backup), Azure Monitor (use azure-monitor).
日本語の概要は準備中です。原文の説明を表示しています。
MicrosoftDocs/Agent-Skills☆ 7772026年10月11日 更新
Expert knowledge for Azure Sre Agent development including troubleshooting, best practices, decision making, architecture & design patterns, security, configuration, integrations & coding patterns, and deployment. Use when wiring SRE Agent to DevOps/GitHub, KQL telemetry, AKS Java apps, IaC deployments, or DR architectures, and other Azure Sre Agent related development tasks. Not for Azure Monitor (use azure-monitor), Azure Reliability (use azure-reliability), Azure Resiliency (use azure-resiliency), Azure Service Health (use azure-service-health).
日本語の概要は準備中です。原文の説明を表示しています。
MicrosoftDocs/Agent-Skills☆ 7772026年10月11日 更新
Expert knowledge for Chaos Studio development including troubleshooting, best practices, decision making, limits & quotas, security, configuration, integrations & coding patterns, and deployment. Use when automating Chaos Studio via CLI/ARM, configuring faults/targets, securing networks/identity, or choosing Workspaces vs classic, and other Chaos Studio related development tasks. Not for Azure Resiliency (use azure-resiliency), Azure Reliability (use azure-reliability), Azure Monitor (use azure-monitor), Azure Site Recovery (use azure-site-recovery).
日本語の概要は準備中です。原文の説明を表示しています。
MicrosoftDocs/Agent-Skills☆ 7772026年10月11日 更新
Check whether an x402 payment endpoint is reliably delivering value before spending USDC on it. Returns paid delivery rate, active incidents, latency, and a clear proceed/warn/block recommendation.
日本語の概要は準備中です。原文の説明を表示しています。
aeonfun/aeon☆ 7702026年10月10日 更新
Knowledge base from "Solving a Million-Step LLM Task with Zero Errors" (Meyerson, Paolo, Dailey, Shahrzad, Francon, Hayes, Qiu, Hodjat, Miikkulainen — arXiv:2511.09030). Use when designing long-horizon, multi-step agentic LLM pipelines that need very high reliability, applying Maximal Agentic Decomposition (MAD), first-to-ahead-by-k voting, or red-flagging error correction, estimating cost/reliability scaling laws for multi-agent systems, or studying Massively Decomposed Agentic Processes (MDAPs).
日本語の概要は準備中です。原文の説明を表示しています。
coco-research/coco☆ 5362026年10月11日 更新