本文へ移動
cccskills
無料GitHub で公開

hyperpod-cluster-debugger

Diagnose and remediate cluster-wide HyperPod (EKS or Slurm) problems — creation / deployment failures (CloudFormation, EFA health check, lifecycle scripts, capacity), EKS access, node replacement, CloudFormation nested-stack errors, post-maintenance rollback state, dangling nodes, autoscaler conflicts. Includes `--validate` pre-flight. Read-only.

インストール方法を見る

含まれるファイル(8)

  • SKILL.md14.1 KB
  • references/capacity-planning.md4.5 KB
  • references/cloudformation-errors.md5.8 KB
  • references/cluster-diagnostics-detail.md23.3 KB
  • references/cluster-operations.md11.5 KB
  • references/iam-permissions.md1.2 KB
  • references/lifecycle-scripts.md5.1 KB
  • scripts/diagnose-cluster.sh69.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

HyperPod Cluster Debugger

Operating policy. Run read-only diagnostics yourself. Never run a command that changes cluster, node, or workload state — present each one as a Suggested command (run this yourself) block and wait for the customer to run it. Destructive order: investigate → reboot → replace (replace destroys root + secondary volumes; not supported on Slurm controller nodes).

Before any state-changing CLI: ask if it's IaC-managed. HyperPod clusters, SGs, EKS access entries, and IAM are usually provisioned via CloudFormation / CDK / Terraform. If yes, the fix belongs in IaC — running the CLI will drift and the next deploy reverts it. Use the CLI only when IaC is unavailable (locked out, predates IaC, mid-review).

scripts/diagnose-cluster.sh is read-only: it collects state via AWS APIs (and SSM for Slurm controller health) and prints each issue as [FAIL] ... → references/<file>.md § <section>.

ReferenceOpen when
cluster-diagnostics-detail.mdPer-finding remediation runbook (§ A–L)
cluster-operations.mdOperational deep-dives (EFA SG, EKS access, SSM, Slurm, filesystem)
cloudformation-errors.md§ H needs the full per-resource CFN error catalog
capacity-planning.md§ B or --validate flags capacity / subnet sizing
lifecycle-scripts.md§ C points at a specific lifecycle failure
iam-permissions.mdFull IAM policy for the diagnostic

Workflow

  1. Collect HyperPod cluster name (not EKS name), region, exact error string.
  2. Run scripts/diagnose-cluster.sh (or --validate for pre-create).
  3. For every [FAIL] line, Read the referenced section.
  4. Present finding, root cause, and the Suggested-command block verbatim. Wait for customer approval.
  5. Re-run the diagnostic to confirm.

Step 1: Run diagnostics

# Diagnose an existing cluster:
bash scripts/diagnose-cluster.sh --cluster <CLUSTER_NAME_OR_ARN> --region <REGION>

# Pre-flight (no cluster needed) — validates SGs, subnets, IAM, VPC endpoints,
# optionally S3 lifecycle scripts and per-AZ capacity:
bash scripts/diagnose-cluster.sh --validate --region <REGION> \
  --sg-ids <sg-1,sg-2> --subnet-ids <sub-1,sub-2> [--iam-role <role-arn>] \
  [--s3-uri s3://<BUCKET>/path/] [--instance-type ml.p5.48xlarge]

Pass --instance-type when the target instance type is known — enables the per-AZ capacity check (warns if none of the provided subnets are in an AZ that offers that type, which causes insufficient-capacity failures at creation time).

Tags: [PASS] · [FAIL] (counted, has → references/... pointer) · [WARN] · [INFO]. Priorities: P0 blocks operation · P1 degraded · P2 informational.


Step 2: Match signal → section

Error messages / events:

SignalSection
"EFA health checks did not run successfully" (public-doc verbatim signal)A: EFA Health Checks
Insufficient-capacity or AZ-mismatch failure at creationB: Capacity & AZ
Lifecycle-script failure or timeout during provisioningC: Lifecycle Scripts
kubectl auth error (server asks for credentials / no API group list)D: EKS Access
InService but not all instances visibleE: Cluster Provisioning
"Target is not connected" / SSM errorsF: SSM Connectivity
Node replacement not happening / batch-replace not workingG: Node Replacement
"Embedded stack failed" / any CloudFormation errorH: CloudFormation Errors
UpdateClusterSoftware failed or cluster in post-maintenance rollback stateJ: AMI & Cluster Updates
Dangling / orphaned nodes in EKS vs list-cluster-nodesK: Dangling Nodes & Cleanup
Cluster Autoscaler breaks after HyperPod attachedL: Autoscaler Compatibility
Slow I/O, FSx throughput saturatedcluster-operations.md § 9
Slurm node name → instance ID lookupI: Utilities

A: EFA Health Checks

SG missing self-reference. Add inbound + outbound self-ref to every SG on the cluster, plus least-privilege egress for the AWS APIs the node needs (HTTPS 443 to S3 / ECR / SageMaker / SSM / STS / CloudWatch Logs — via VPC-endpoint prefix-lists when possible). Full procedure: cluster-diagnostics-detail.md § A.

B: Capacity & AZ

Instance type unavailable in the requested AZ. Verify with describe-instance-type-offerings, then change AZ, use Flexible Training Plans, or request ODCR. Full: § B · strategy: capacity-planning.md.

C: Lifecycle Scripts

Script failed or timed out during provisioning. Read CloudWatch under /aws/sagemaker/Clusters/<name>/<id> — common causes: missing S3 VPC endpoint, IAM gap, CRLF line endings, instance-group name mismatch. Full: § C · layout: lifecycle-scripts.md.

D: EKS Access / kubectl

IAM identity not in EKS access entries. Verify with sts get-caller-identity, create an access entry with admin policy, update kubeconfig. Full: § D.

E: Cluster Provisioning

InService without all instances is expected under Continuous Provisioning — failures surface as events, not cluster errors. For stuck Creating/Updating/Deleting: check CFN nested stacks (§ H), IAM, capacity, events; if stuck Deleting check VPC ENI dependencies. Full: § E.

F: SSM Connectivity

Target is not connected: use sagemaker-cluster:<CLUSTER_ID>_<GROUP>-<INSTANCE_ID> format (not raw EC2 ID), install session-manager-plugin, confirm node Running. Check IAM + VPC endpoints on timeouts. Full: § F.

G: Node Replacement

Auto-repair: confirm NodeRecovery=Automatic, check Health Monitoring Agent (HMA) logs + node labels / Slurm reason, confirm capacity. Manual: reboot first, replace only if reboot fails. Replace requires the cluster to have been patched via UpdateClusterSoftware at least once and cannot target a Slurm controller node. Full: § G.

H: CloudFormation Errors

Embedded stack failed hides the real error. Drill into nested stacks via Events tab (filter Failed) until you reach a non-stack resource. CLI: describe-stack-events --query 'StackEvents[?ResourceStatus==\CREATE_FAILED`]'`. Also covers SLR creation failures and permission-boundary denials. Full: § H · catalog: cloudformation-errors.md.

I: Utilities

Map Slurm node names (ip-10-x-y-z) to HyperPod instance IDs via list-cluster-nodes or on-node /opt/ml/config/resource_config.json. Full: § I.

J: AMI & Cluster Updates

UpdateClusterSoftware fails and rolls back, or the cluster stays in a post-maintenance rollback state. Common causes: lifecycle script incompatible with new AMI, HMA version too old, insufficient rolling-update capacity. If the cluster has active nodes, collect diagnostics and escalate rather than delete-and-recreate. Full: § J.

K: Dangling Nodes & Cleanup

Nodes in kubectl get nodes but not in list-cluster-nodes (ghost EKS nodes), or the inverse (HyperPod nodes that never registered kubelet). Script flags both. Full: § K.

L: Autoscaler Compatibility

Cluster Autoscaler errors on HyperPod provider IDs and breaks autoscaling for all node groups. No officially endorsed workaround — escalate to AWS Support. Karpenter does not conflict with HyperPod nodes by default. Full: § L.


Prerequisites

  • aws CLI v2.13+ authenticated to the cluster's account
  • jq, python3, bash 4.2+
  • kubectl authenticated to the EKS cluster (EKS checks skipped if absent)
  • session-manager-plugin (Slurm controller health checks only)

IAM policy: references/iam-permissions.md.

Defaults

  • Region — required: pass --region or set $AWS_DEFAULT_REGION.
  • Mode — --cluster <NAME> (diagnose) or --validate (pre-create).
  • Event window — up to 500 most recent events (5 × 100, paginated).
  • Colors — auto-disabled on non-TTY; --no-color to force off.

Error handling

FailureScriptTell the customer
aws sts get-caller-identity failsExit 1"Fix AWS credentials and rerun."
Cluster not foundExit 1 after listing region's clusters"Confirm HyperPod cluster name (not EKS) and region."
sagemaker:* / ec2:* / eks:* / logs:* deniedWarn, add Missing IAM permission for <API>, continue"Grant the listed IAM action and rerun."
kubectl absent or unauthenticatedSkip EKS checks (access entries, add-ons, aws-auth, nodes)"Install/authenticate kubectl."
session-manager-plugin absent (Slurm)Skip Slurm controller probe"Install session-manager-plugin."
SSM throttled / times out (180s)Retry with backoff; warn and continue if still failing"Rerun later — script is idempotent."
CloudWatch log group not foundSkip CloudWatch check"CloudWatch not configured on this cluster."

Exit codes: 0 no critical failures · 1 one or more critical failures (cluster not found, fatal prerequisite missing, or any [FAIL] in diagnose or --validate mode). [WARN] lines do not affect the exit code.

Skill delegation

NeedUse
Shell on nodeshyperpod-ssm
Version comparison across nodeshyperpod-version-checker

Escalate to AWS Support

Escalate when:

  1. EFA health checks fail despite correct SG rules.
  2. Capacity errors persist despite a valid Flexible Training Plan / ODCR.
  3. Node replacement fails repeatedly without clear events / log signal.
  4. Cluster stuck in a non-terminal state (Creating, Updating, or a post-maintenance rollback state) for an extended period.
  5. CloudFormation root-cause is an internal service error.

Before opening the case

Run these commands and attach the output. Goal: AWS Support has everything at case open.

# 1. Cluster identity + status (confirms region, ARN, orchestrator, instance groups)
aws sagemaker describe-cluster --cluster-name <CLUSTER> --region <REGION>

# 2. Full cluster-level diagnostic bundle
bash scripts/diagnose-cluster.sh --cluster <CLUSTER> --region <REGION> > diag.txt

# 3. Per-node log/config bundle to S3 (delegates to hyperpod-issue-report skill)
#    See skills/hyperpod-issue-report/SKILL.md for the exact invocation.

Include in the case

  • Cluster name + ARN (or ClusterId suffix) and AWS region
  • ClusterStatus + FailureMessage from describe-cluster
  • Timestamp window (UTC start / end) of the failure
  • Exact error strings observed (copy verbatim from events / logs / console)
  • Affected instance IDs / NodeLogicalIds / instance group names
  • diag.txt from step 2 above
  • S3 URI of the hyperpod-issue-report bundle from step 3

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Integrates Amazon Location Service APIs for AWS applications. Use this skill when users want to add maps (interactive MapLibre or static images); geocode addresses to coordinates or reverse geocode coordinates to addresses; calculate routes, travel times, or service areas; find places and businesses through text search, nearby search, or autocomplete suggestions; retrieve detailed place information including hours, contacts, and addresses; monitor geographical boundaries with geofences; or track device locations. Covers authentication, SDK integration, and all Amazon Location Service capabilities.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Build and deploy full-stack web and mobile apps with AWS Amplify Gen2 (TypeScript code-first). Covers auth (Cognito), data (AppSync/DynamoDB including schema modeling, enum types, relationships, authorization rules), storage (S3), functions, APIs, and AI (Amplify AI Kit with Bedrock). Supports React, Next.js, Vue, Angular, React Native, Flutter, Swift, and Android. Always use this skill for Amplify Gen2 topics — even for questions you think you know — it contains validated, version-specific patterns that prevent common mistakes. TRIGGER when: user mentions Amplify Gen2; project has amplify/ directory or amplify_outputs; code imports @aws-amplify packages; user asks about defineBackend, defineAuth, defineData, defineStorage, or npx ampx. SKIP: Amplify Gen1 (amplify CLI v6), standalone SAM/CDK without Amplify (use aws-serverless), direct Bedrock without Amplify AI Kit (use bedrock).

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Build, manage, and operate APIs with Amazon API Gateway (REST, HTTP, and WebSocket). Triggers on phrases like: API Gateway, REST API, HTTP API, WebSocket API, custom domain, Lambda authorizer, usage plan, throttling, CORS, VPC link, private API. Also covers troubleshooting API Gateway errors (4xx, 5xx, timeout, CORS failures) and IaC templates containing API Gateway resources. For general REST API design unrelated to AWS, do not trigger.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Generate validated AWS architecture diagrams as draw.io XML using official AWS4 icon libraries. Use this skill whenever the user wants to create, generate, or design AWS architecture diagrams, cloud infrastructure diagrams, or system design visuals. Also triggers for requests to visualize existing infrastructure from CloudFormation, CDK, or Terraform code. Supports two modes: analyze an existing codebase to auto-generate diagrams, or brainstorm interactively from scratch. Exports .drawio files with optional PNG/SVG/PDF export via draw.io desktop CLI.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Design, build, deploy, test, and debug serverless applications with AWS Lambda. Triggers on phrases like: Lambda function, event source, serverless application, API Gateway, EventBridge, Step Functions, serverless API, event-driven architecture, Lambda trigger. For deploying non-serverless apps to AWS, use deploy-on-aws plugin instead.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for long-running executions. Covers the critical replay model, step operations, wait/callback patterns, error handling with saga pattern, testing with LocalDurableTestRunner. Triggers on phrases like: lambda durable functions, workflow orchestration, state machines, retry/checkpoint patterns, long-running stateful Lambda functions, saga pattern, human-in-the-loop callbacks, and reliable serverless applications.

日本語の概要は準備中です。原文の説明を表示しています。

awslabs/agent-plugins9172026年10月10日 更新

awslabs のスキルをすべて見る

このスキルの問題を報告する