本文へ移動
cccskills

「fault tolerance」の検索結果

41 件 ・ 関連度順

概要と使いどころ

Knowledge base from the NASA Fault Management Handbook (NASA-HDBK-1002, Draft 2, April 2012). Use for the systems-engineering discipline of designing what a system does when it fails — the fault/failure/anomaly vocabulary, the five FM strategies (two prevention strategies — design-time fault avoidance and operational failure avoidance — plus three tolerance strategies — failure masking, failure recovery, and goal change), the six FM process activities across NASA mission phases Pre-A to E, response-latency-versus-time-to-criticality (TTC) design trades, fault/failure containment regions and FEPPs, the four redundancy approaches, top-down FM requirements development, FM verification vs. validation, the seven dedicated FM milestone reviews (FMCR/FMARR/FMPDR/FMCDR/FMTRR/FMLRR/FMCERR), FM organizational structure, and mined NASA Lessons Learned. Space-mission oriented (flight/ground/operations); spacecraft examples throughout. IMPORTANT: built from an UNAPPROVED DRAFT — several sections (detailed architecture building blocks, the Assessment & Analysis guidance, and Operations & Maintenance) are placeholders in the source and are thin here. Thin on detailed FDIR algorithms, terrestrial/IT fault management, and any approved-standard terminology that may have changed since Draft 2.

日本語の概要は準備中です。原文の説明を表示しています。

jgsystemsconsulting/jgs-se-knowledge-packs82026年10月9日 更新

ray-train

無料

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

ray-train

無料

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault tolerance, elastic scaling. Use when training massive models across multiple machines or running distributed hyperparameter sweeps.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Resilience testing specialist for failure injection, game day planning, and building confidence in system reliabilityUse when "chaos engineering, resilience testing, failure injection, game day, fault tolerance, chaos experiment, disaster recovery, reliability testing, chaos-engineering, resilience, failure-injection, game-day, fault-tolerance, reliability, testing, litmus, chaos-monkey, ml-memory" mentioned.

日本語の概要は準備中です。原文の説明を表示しています。

omer-metin/skills-for-antigravity1642026年1月22日 更新

Foundational distributed computing algorithms for consensus, coordination, and fault tolerance

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/windags-skills132026年10月1日 更新

Foundational distributed computing algorithms for consensus, coordination, and fault tolerance

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Implement circuit breaker logic for agentic tool calls — tracking tool health, transitioning between closed/open/half-open states, reducing task scope when tools fail, routing to alternatives via capability maps, and enforcing failure budgets to prevent error accumulation. Separates orchestration (deciding what to attempt) from execution (calling tools), following the expeditor pattern. Use when building agents that depend on multiple tools with varying reliability, designing fault-tolerant agentic workflows, recovering gracefully from tool outages mid-task, or hardening existing agents against cascading tool failures.

日本語の概要は準備中です。原文の説明を表示しています。

pjt222/agent-almanac372026年10月10日 更新

Foundational impossibility result proving consensus cannot be guaranteed in asynchronous systems with even one faulty process

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/windags-skills132026年10月1日 更新

Chaos Engineering discipline — proactive experimentation on distributed systems to build confidence in their resilience. Covers Netflix-inspired chaos principles, experiment design, blast radius control, LitmusChaos and Chaos Mesh usage, game day planning, and production safety. USE WHEN: validating system resilience before production incidents, designing chaos experiments, planning game days, implementing CI/CD chaos pipeline, or building confidence in fault tolerance. Triggers on "chaos engineering", "chaos experiment", "game day", "resilience test", "fault injection", "Litmus", "Chaos Mesh", "break things on purpose".

日本語の概要は準備中です。原文の説明を表示しています。

aAAaqwq/openclaw-team22026年6月18日 更新

Foundational impossibility result proving consensus cannot be guaranteed in asynchronous systems with even one faulty process

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新

Master defensive Bash programming techniques for production-grade scripts. Use when writing robust shell scripts, CI/CD pipelines, or system utilities requiring fault tolerance and safety.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

Generates guidance for reliability, resilience, availability, redundancy, fault-tolerance, and disaster recovery (DR) for Google Cloud workloads based on the design principles and recommendations in the Google Cloud Well-Architected Framework. Use when the user asks to evaluate, design, or improve the reliability, resilience, availability, or disaster recovery capabilities of Google Cloud workloads.

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Field-tested methodology and concrete recipes for training and operating large-scale LLM/VLM/multi-modal models end to end - choosing and benchmarking accelerators, storage and network; SLURM/Kubernetes orchestration; maximizing training throughput and fitting models in memory; diagnosing and surviving training instabilities, NaN/Inf, and hardware/job failures; checkpointing and fault tolerance; inference performance and memory; debugging multi-node/ multi-GPU hangs; and writing/running tests. Use when the user is training or fine-tuning large models, hits low TFLOPS/MFU, OOM, slow dataloading, a loss spike/divergence, a NCCL/InfiniBand or multi-node hang, node/GPU failures, checkpoint or preemption problems, storage/network bottlenecks, or needs to pick GPUs/cloud/file-systems or size inference latency/throughput. Distilled from "Machine Learning Engineering", the latest version of which can be found at https://github.com/stas00/ml-engineering The latest SKILL.md version can be found at https://github.com/stas00/ml-engineering/blob/master/skills/ml-engineering/SKILL.md

日本語の概要は準備中です。原文の説明を表示しています。

stas00/ml-engineering1.9万2026年10月8日 更新

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5592026年10月10日 更新

Implement reliability patterns for Claude API: circuit breakers, graceful degradation, idempotency, and fallback strategies. Trigger with phrases like "anthropic reliability", "claude circuit breaker", "claude fallback", "anthropic fault tolerance".

日本語の概要は準備中です。原文の説明を表示しています。

jeremylongshore/tons-of-skills-marketplace2,8312026年10月11日 更新

Multi-protocol consensus for agent swarms supporting Raft leader election, Byzantine fault tolerance, Gossip state propagation, and CRDT conflict-free merging.

日本語の概要は準備中です。原文の説明を表示しています。

a5c-ai/babysitter1,8402026年9月17日 更新

Use when oTP behaviors including gen_server for stateful processes, gen_statem for state machines, supervisors for fault tolerance, gen_event for event handling, and building robust, production-ready Erlang applications with proven patterns.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Write idiomatic Elixir code with OTP patterns, supervision trees, and Phoenix LiveView. Masters concurrency, fault tolerance, and distributed systems. Use PROACTIVELY for Elixir refactoring, OTP design, or complex BEAM optimizations.

日本語の概要は準備中です。原文の説明を表示しています。

rmyndharis/antigravity-skills1,7302026年10月1日 更新

Master defensive Bash programming techniques for production-grade scripts. Use when writing robust shell scripts, CI/CD pipelines, or system utilities requiring fault tolerance and safety.

日本語の概要は準備中です。原文の説明を表示しています。

rmyndharis/antigravity-skills1,7302026年10月1日 更新

Multi-agent orchestration: lane contracts, worker-model doctrine, fault tolerance. Triggers: orchestrate, coordinate, agents, parallel, spawn, contract, tier 2, tier 3.

日本語の概要は準備中です。原文の説明を表示しています。

hashgraph-online/awesome-codex-plugins1,2752026年10月11日 更新

canary

無料

Post-deploy monitoring. Watches live app via headless browser for console errors, performance regressions, page failures, visual anomalies. Alert on deltas not absolutes, transient tolerance (2+ consecutive checks). Default 10-minute monitoring window. Baseline capture mode.

日本語の概要は準備中です。原文の説明を表示しています。

paperclipai/companies9192026年3月24日 更新

Master defensive Bash programming techniques for production-grade scripts. Use when writing robust shell scripts, CI/CD pipelines, or system utilities requiring fault tolerance and safety.

日本語の概要は準備中です。原文の説明を表示しています。

Microck/ordinary-claude-skills4042026年9月7日 更新

Provides fault tolerance patterns for Spring Boot 3.x using Resilience4j. Use when implementing circuit breakers, handling service failures, adding retry logic with exponential backoff, configuring rate limiters, or protecting services from cascading failures. Generates circuit breaker, retry, rate limiter, bulkhead, time limiter, and fallback implementations. Validates resilience configurations through Actuator endpoints.

日本語の概要は準備中です。原文の説明を表示しています。

giuseppe-trisciuoglio/developer-kit3572026年9月10日 更新

Use when building batch jobs, ETL pipelines, scheduled imports/exports, or any chunk-oriented bulk processing with Spring Batch. Covers the Spring Batch 6 / Boot 4 builder API, resourceless vs JDBC job repositories, restartability and idempotent job parameters, reader/writer thread-safety, fault tolerance, and chunk transaction boundaries.

日本語の概要は準備中です。原文の説明を表示しています。

rrezartprebreza/spring-boot-skills3012026年9月21日 更新