本文へ移動
cccskills

「reliability」の検索結果

343 件 ・ 関連度順

概要と使いどころ

DevOps 与站点可靠性工程 (SRE) — 平台 / 基础设施 / 可靠性工程师的认知操作系统, 覆盖软件交付 + 运维全生命周期 (CI/CD 与发布工程 trunk-based + 渐进式发布 canary/blue-green/feature flag + GitOps Argo CD/Flux / 基础设施即代码 Terraform/OpenTofu/Pulumi/Ansible + policy-as-code OPA / 容器与编排 Docker/Kubernetes + Helm/Kustomize + service mesh Istio/Linkerd / 可观测性 Prometheus + Loki + OpenTelemetry + Honeycomb + eBPF + RED/USE / SLO-SLI-error budget 与可靠性工程 Google SRE 学科 + 容量规划 + 优雅降级 / 事件管理与 on-call 事件指挥 + PagerDuty + runbook + 无指责复盘 + MTTR / 云平台与 FinOps AWS/GCP/Azure + 成本优化 + 弹性伸缩 / 平台工程与开发者体验 IDP + Backstage + golden path + Team Topologies / DevSecOps 与供应链安全 shift-left + SBOM + SLSA + sigstore + Vault / 韧性与混沌工程 fault injection + game day + 安全科学 / DORA 指标与工程效能 部署频率 + 变更前置时间 + 变更失败率 + Accelerate 研究 / 数据库与有状态运维 schema 迁移 + 备份容灾) — 不含 通用应用开发 / 纯云销售认证速成 / 'DevOps = 跑 Jenkins 的岗位' 窄化误解 / ITIL 工单文化传统运维 (旧范式仅做边界) / 把手工运维 ClickOps 当稳态 (是 toil, 本 skill 核心反模式) (DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, Helm / Kustomize, service mesh Istio / Linkerd), (d) observability (the three pillars + beyond — metrics Prometheus / VictoriaMetrics, logs Loki / ELK, traces OpenTelemetry / Jaeger / Tempo, high-cardinality observability Honeycomb, eBPF, RED / USE methods, SLO-based alerting), (e) SLO / SLI / error budgets & reliability engineering (Google SRE discipline — service level objectives, error budget policy, toil budgets, capacity planning, load shedding, graceful degradation), (f) incident management & on-call (incident command, PagerDuty / Opsgenie, runbooks, blameless postmortems, MTTR / MTTD, error budget burn), (g) cloud platforms & FinOps (AWS / GCP / Azure well-architected, multi-region, cost optimization, autoscaling), (h) platform engineering & developer experience (internal developer platforms, Backstage, golden paths, self-service, Team Topologies), (i) DevSecOps & supply-chain security (shift-left, SAST / DAST, SBOM, SLSA, sigstore / cosign, secrets management Vault), (j) resilience & chaos engineering (chaos experiments, fault injection, game days, resilience engineering / safety science), (k) DORA metrics & engineering effectiveness (deployment frequency, lead time, change failure rate, MTTR, the Accelerate research), (l) databases & stateful operations (schema migrations, backups / DR, replication); NOT generic software development / app feature coding (是 平行学科, DevOps/SRE 关注 delivery + operability 不是 product feature), NOT pure cloud sales / certification cram without operational depth, NOT 'DevOps = a job title that runs Jenkins' 的窄化误解 (DevOps 是 文化 + 实践, SRE 是 Google 对 reliability 的工程化具体实现), NOT ITIL-heavy 传统运维 工单文化 (是 被 DevOps 取代的旧范式, 仅做边界标注), NOT manual ops / ClickOps as a steady state (是 toil, 本 skill 的核心反模式).) Master OS — automated mastery of DevOps & Site Reliability Engineering — the cognitive operating system of platform / infrastructure / reliability practitioners who own the full software delivery + operational lifecycle, covering (a) CI/CD & release engineering (build pipelines, trunk-based development, progressive delivery — canary / blue-green / feature flags, GitOps with Argo CD / Flux), (b) Infrastructure as Code (Terraform / OpenTofu, Pulumi, CloudFormation, Ansible, Crossplane — module design, state management, drift, policy-as-code OPA / Sentinel / Checkov), (c) containers & orchestration (Docker / OCI, Kubernetes — scheduling, networking CNI, storage CSI, operators / CRDs, He

日本語の概要は準備中です。原文の説明を表示しています。

swaylq/master-skill1492026年9月6日 更新

Knowledge base from MIL-HDBK-338B (Electronic Reliability Design Handbook). Use for electronic reliability engineering: R/M/A theory, reliability specification/allocation/prediction, parts management and derating, reliable circuit and fault-tolerant design, environmental and human performance reliability, FMEA/FMECA/FTA/sneak circuit analysis, design reviews and testability, FRACAS, reliability demonstration and growth testing, and systems-level R&M parameters. Covers selected Part 2 design-guidance topics only; skips large annex/part-stress tables. Guidance handbook — not a contractual requirement text. Does not replace MIL-HDBK-217 prediction libraries, service-specific reliability regs, or safety standards (see mil-std-882).

日本語の概要は準備中です。原文の説明を表示しています。

jgsystemsconsulting/jgs-se-knowledge-packs82026年10月9日 更新

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1002026年10月10日 更新

Assess and improve the reliability posture of PaaS Applications (Azure Functions and Azure App Service). Scans deployed resources for zone redundancy, ZRS storage, health probes, and multi-region failover. Presents a feature-pivoted checklist, then drives staged remediation (CLI or IaC patches) end-to-end with user confirmation. WHEN: "assess reliability", "check reliability", "zone redundant", "multi-region failover", "high availability", "disaster recovery", "single points of failure", "reliability posture", "resiliency".

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/azure-skills1,5552026年10月10日 更新

[omh] Postmortem for an outage or SLO miss: postmortems, SLOs, error budgets, incident follow-ups, and service reliability evidence. Use when the user says: reliability-review, reliability review, incident review, incident postmortem, postmortem, post-mortem, slo review, slo.

日本語の概要は準備中です。原文の説明を表示しています。

rlaope/oh-my-hermes3,2592026年10月11日 更新

Analyze and design idempotency, bounded polling, per-service circuits, durable queues, DLQ/replay, degradation, and reconciliation for Adobe workloads. Use when the task requires adobe async reliability controls. Trigger with "Adobe reliability", "Firefly retry design", or "PDF job recovery".

日本語の概要は準備中です。原文の説明を表示しています。

jeremylongshore/tons-of-skills-marketplace2,8312026年10月11日 更新

Use this skill when fitting accelerated life models, computing mean time to failure (MTTF), B10 life and similar quantities, fitting life distributions like Weibull for reliability analysis, comparing distribution fits, computing confidence intervals on reliability quantities with bootstrapping, performing lifetime prediction, computing failure rates, plotting Kaplan-Meier curves, computing survival functions, or doing reliability analysis.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

Provides Site Reliability Engineering best practices for SLOs, SLIs, SLAs, error budgets, toil reduction, reliability reviews, and capacity planning. Use when defining service objectives, measuring reliability, reducing toil, planning capacity, or when user mentions 'SRE', 'SLO', 'SLI', 'SLA', 'error budget', 'toil', 'reliability', 'on-call', 'capacity planning'.

日本語の概要は準備中です。原文の説明を表示しています。

Tibsfox/gsd-skill-creator702026年7月20日 更新

faa-rma

無料

Knowledge base from FAA-HDBK-006D (2020), the FAA's System Reliability, Maintainability, and Availability (RMA) Handbook for National Airspace System (NAS) hardware and software. Use for RMA figures of merit (MTBF, MTBO, MTTR, the availability variants), the bathtub curve and probability distributions, the four-stage RMA lifecycle mapped to the FAA Acquisition Management System (AMS), the six-task RMA technical-management/acquisition process, service-thread criticality and the NAS-RD severity-to-target table, the RMA analysis toolbox (RBD, FMEA/FMECA, FTA, Ishikawa, Monte Carlo, Bayesian, FRACAS, ALT, reliability-growth/recovery tests), software reliability (early prediction + reliability growth), and tailorable example RMA requirements for a Program Requirements Document. Scope limits: it is FAA/NAS-specific guidance (not a requirement, by its own statement), centred on hardware-plus-software availability — it brackets out human/service/facility/communications availability, names but does not reproduce the math derivations or external standards (MIL-STDs, IEEE 1633, RTCA DO-178C/DO-278A, SAE JA1011), is thin on detailed worked numerical examples and on safety-case method, and predates any current revision beyond 006D.

日本語の概要は準備中です。原文の説明を表示しています。

jgsystemsconsulting/jgs-se-knowledge-packs82026年10月9日 更新

eval-harness

無料日本語概要

AIによる開発の成功条件を実装前に定義し、コード・ルール・モデル・人の評価で新機能と既存機能を確認しながら、複数回の試行から信頼性を記録するスキル。

  • AI実装前に成功条件を決めたいとき
  • プロンプトやエージェント変更後の回帰確認
  • 複数回の試行で信頼性を測りたいとき
affaan-m/ECC27.7万2026年10月10日 更新

Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in the Google Cloud Well-Architected Framework. Use when the user asks to evaluate, design, or improve the reliability, resilience, availability, or disaster recovery capabilities of Google Cloud workloads.

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for generating K8s YAML manifests (use gke-manifest-generation) or disaster recovery and cluster backups (use gke-backup-dr).

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Implement reliability patterns for Claude API: circuit breakers, graceful degradation, idempotency, and fallback strategies. Trigger with phrases like "anthropic reliability", "claude circuit breaker", "claude fallback", "anthropic fault tolerance".

日本語の概要は準備中です。原文の説明を表示しています。

jeremylongshore/tons-of-skills-marketplace2,8312026年10月11日 更新

Performance, load, and reliability (SLO/SRE) planning for production readiness. Use when defining service level objectives, load characterization, capacity, latency budgets, stress/soak/spike test plans, false-positive baselines, and reliability targets. USE FOR: SLO/SLA definition, load testing plan, performance budget, capacity planning, reliability/SRE backlog, latency targets, error-budget policy. DO NOT USE FOR: executing load tests (use Azure Load Testing tooling), security threat modeling, RAI assessment, privacy/compliance planning, or authoring/restating PRD requirements (cite the PRD's existing NFR/FR ids instead).

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/hve-core1,5182026年10月11日 更新

Compare this period's reliability against the prior period using Agent Monitor data — error rate (APIError/total) and tool-failure rate (PreToolUse→PostToolUse gap) — flag any regression where reliability got worse, and optionally wire a persistent alert rule so the dashboard catches the next regression automatically. Use when checking whether reliability degraded.

日本語の概要は準備中です。原文の説明を表示しています。

hoangsonww/Claude-Code-Agent-Monitor1,0612026年10月10日 更新

Generate a CCAM reliability report from session outcomes, hook events, alerts, tool failures, and data freshness. Use for health reviews, incident follow-up, hook-delivery audits, or reliability trend summaries.

日本語の概要は準備中です。原文の説明を表示しています。

hoangsonww/Claude-Code-Agent-Monitor1,0612026年10月10日 更新

Expert knowledge for Azure Reliability development including best practices, decision making, architecture & design patterns, limits & quotas, and deployment. Use when choosing Azure regions, availability zones, geo-paired deployments, Queue Storage limits, or Web PubSub apps, and other Azure Reliability related development tasks. Not for Azure Monitor (use azure-monitor), Azure Resiliency (use azure-resiliency), Azure Service Health (use azure-service-health), Azure Sre Agent (use azure-sre-agent).

日本語の概要は準備中です。原文の説明を表示しています。

MicrosoftDocs/Agent-Skills7772026年10月11日 更新

Plan, implement, or review google cloud waf reliability work in an existing codebase with compatibility, security, and verification controls. Use when the user explicitly requests google cloud waf reliability work.

日本語の概要は準備中です。原文の説明を表示しています。

sandbaseai/sandbase-skills2032026年9月26日 更新

Resilience testing specialist for failure injection, game day planning, and building confidence in system reliabilityUse when "chaos engineering, resilience testing, failure injection, game day, fault tolerance, chaos experiment, disaster recovery, reliability testing, chaos-engineering, resilience, failure-injection, game-day, fault-tolerance, reliability, testing, litmus, chaos-monkey, ml-memory" mentioned.

日本語の概要は準備中です。原文の説明を表示しています。

omer-metin/skills-for-antigravity1642026年1月22日 更新

Autonomous agents are AI systems that can independently decompose goals, plan actions, execute tools, and self-correct without constant human guidance. The challenge isn't making them capable - it's making them reliable. Every extra decision multiplies failure probability. This skill covers agent loops (ReAct, Plan-Execute), goal decomposition, reflection patterns, and production reliability. Key insight: compounding error rates kill autonomous agents. A 95% success rate per step drops to 60% by step 10. Build for reliability first, autonomy second. 2025 lesson: The winners are constrained, domain-specific agents with clear boundaries, not "autonomous everything." Treat AI outputs as proposals, not truth. Use when "autonomous agent, autogpt, babyagi, self-prompting, goal decomposition, react pattern, agent loop, self-correcting agent, reflection agent, langgraph, agentic ai, agent planning, autonomous, agents, langgraph, react, planning, reflection, guardrails, reliability, checkpointing" mentioned.

日本語の概要は準備中です。原文の説明を表示しています。

omer-metin/skills-for-antigravity1642026年1月22日 更新

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarksUse when "agent testing, agent evaluation, benchmark agents, agent reliability, test agent, testing, evaluation, benchmark, agents, reliability, quality" mentioned.

日本語の概要は準備中です。原文の説明を表示しています。

omer-metin/skills-for-antigravity1642026年1月22日 更新

Establish Service Level Objectives (SLO), Service Level Indicators (SLI), and Service Level Agreements (SLA) with error budget tracking, burn rate alerts, and automated reporting using Prometheus and tools like Sloth or Pyrra. Use when defining reliability targets for customer-facing services, balancing feature velocity against system reliability through error budgets, migrating from arbitrary uptime goals to data-driven metrics, or implementing Site Reliability Engineering practices.

日本語の概要は準備中です。原文の説明を表示しています。

pjt222/agent-almanac372026年10月10日 更新

Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in the Google Cloud Well-Architected Framework. Use when the user asks to evaluate, design, or improve the reliability, resilience, availability, or disaster recovery capabilities of Google Cloud workloads.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for disaster recovery setup or full cluster backups (use gke-backup-dr instead).

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新