本文へ移動
cccskills
無料GitHub で公開

ops-inspector

AIOps-style CloudBase inspection skill (v3). Use when users need health checks, log diagnosis, alarm interpretation (CPU alert normal?, peak QPS), metrics via queryEnv(action=metrics), or fault playbooks for 429 / function 404 / ACCESS_TOKEN_INVALID / zero invocations. Triggers on 巡检, 诊断, 告警, 峰值 QPS, 限频, 调用量为 0, troubleshooting.

インストール方法を見る

含まれるファイル(4)

  • SKILL.md13.9 KB
  • LICENSE.md1.0 KB
  • references/alarm-interpretation.md5.0 KB
  • references/fault-playbooks.md7.2 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Sibling skills (local only)

Sibling CloudBase skills ship beside this skill. Use local relative paths such as ../auth-tool-cloudbase/SKILL.md.

If a referenced sibling skill file is missing from this environment, ask the user to install the full CloudBase plugin (or the missing skill). Do not HTTP-fetch remote skill or protocol markdown into the agent context.

Activation Contract

Use this first when

  • The user wants to check the health or status of CloudBase resources (cloud functions, CloudRun, databases, storage, etc.).
  • The user reports errors, failures, or abnormal behavior and wants a quick diagnosis.
  • The user asks for an "inspection", "health check", "巡检", "诊断", or "troubleshooting" of their CloudBase environment.
  • The user wants to review recent error logs across services.
  • The user asks 告警解读 questions: whether a CPU 告警 is normal, what 峰值 QPS was, or whether throttle/error metrics look healthy.
  • The symptom matches a v3 fault playbook: 429 / 限频, 云函数 404, ACCESS_TOKEN_INVALID, or 调用量为 0.

Read before writing code if

  • The inspection reveals code-level issues in cloud functions or CloudRun services — then read the relevant implementation skill before suggesting fixes.
  • The user wants to fix a problem found during inspection rather than just diagnose it.

Then also read

  • Alarm interpretation baselines -> references/alarm-interpretation.md
  • Fault playbooks (429 / 404 / token / zero calls) -> references/fault-playbooks.md
  • Cloud function issues -> ../cloud-functions/SKILL.md
  • CloudRun issues -> ../cloudrun-development/SKILL.md
  • Database issues -> ../postgresql-development-cloudbase/SKILL.md for CloudBase PG / PostgreSQL, ../relational-database-mcp-cloudbase/SKILL.md for MySQL, or ../cloudbase-document-database-web-sdk/SKILL.md for NoSQL
  • Auth readiness (token failures) -> ../auth-tool-cloudbase/SKILL.md
  • Platform overview -> ../cloudbase-platform/SKILL.md

Do NOT use for

  • Deploying new resources or writing application code. This skill is read-only and diagnostic.
  • Replacing proper monitoring/alerting infrastructure. It provides point-in-time inspection, not continuous monitoring.
  • Directly fixing problems — it diagnoses and recommends; actual fixes should use the appropriate implementation skill.
  • Fetching metrics by guessing cloud API Actions. Never use callCloudApi for monitor curves — always use queryEnv(action="metrics").

Common mistakes / gotchas

  • Running a full inspection without first confirming the environment is bound (auth tool must show logged-in and env-bound state).
  • Ignoring CLS log service status — if CLS is not enabled, queryLogs will fail; always check first with queryLogs(action="checkLogService").
  • Searching logs without a time range — this can return excessive or irrelevant results. Always scope searches to a relevant time window.
  • Treating a single error log as the root cause without correlating across resources. A function error may stem from a database or config issue.
  • Answering "峰值 QPS" / "CPU 告警是否正常" from screenshots or memory instead of queryEnv(action="metrics").
  • Calling callCloudApi with invented GetMonitorData / DescribeCurveData parameters — the metrics branch already wraps Manager SDK.

Minimal checklist

  • Environment is bound and accessible (queryEnv(action="info"))
  • Metrics pulled with queryEnv(action="metrics") when the question involves QPS / CPU / throttle / invocation volume
  • CLS log service is enabled (queryLogs(action="checkLogService")) when log diagnosis is needed
  • Matching fault playbook selected when symptoms match 429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0
  • Time range is specified for any log or metrics searches
  • Findings are summarized with severity levels, 告警解读, and actionable recommendations

How to use this skill (for a coding agent)

Ops Inspector v3 additions

v3 adds two mandatory capabilities on top of log/resource inspection:

  1. 告警解读 — pull metrics, compare to baselines in references/alarm-interpretation.md, answer CPU-alert / peak-QPS style questions in plain language.
  2. 故障剧本 — when symptoms match, follow references/fault-playbooks.md instead of ad-hoc tool fishing.

Inspection Modes

ModeWhen to useScope
Full inspectionUser asks for a general health check / 巡检 / 全面检查All resource types + core metrics
Targeted inspectionUser reports a specific error or asks about a specific resourceOne resource type or playbook
Alarm interpretationUser asks CPU 告警是否正常 / 峰值 QPS / 是否限流Metrics-first, then logs
Fault playbook429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0Playbook steps only

Full Inspection Workflow

Follow these steps in order for a comprehensive environment health check:

Step 1 — Environment Check

queryEnv(action="info")

Confirm the environment is accessible. Record the envId for console link generation.

Step 2 — Metrics snapshot (v3)

queryEnv(action="metrics", envId="<EnvId>", metricName="GatewayTraceEnvQPS")
queryEnv(action="metrics", envId="<EnvId>", metricName="FunctionInvocation")
queryEnv(action="metrics", envId="<EnvId>", metricName="MysqlCpuUsageRate")

Use returned Summary.max / avg / allZero / peakTimestamp. Add FunctionError, FunctionThrottle, or CloudRun Tke* metrics when those resources exist. Read references/alarm-interpretation.md before concluding.

Step 3 — Log Service Status

queryLogs(action="checkLogService")

If CLS is not enabled, note this as a warning — log-based diagnosis will be unavailable. Recommend enabling CLS in the console: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log

Step 4 — Cloud Functions Inspection

queryFunctions(action="listFunctions")

For each function, check:

  • Status: Is the function in an active/deployed state?
  • Recent errors: queryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<recent>")
  • Common issues:
    • Timeout errors (execution exceeded limit)
    • Memory limit exceeded
    • Runtime errors (unhandled exceptions)
    • Cold start frequency
    • Zero invocations while traffic is expected → Playbook 4

Step 5 — CloudRun Services Inspection

queryCloudRun(action="list")

For each service, check:

  • Status: Is the service running?
  • Detail: queryCloudRun(action="detail", detailServerName="<name>")
  • Metrics: queryEnv(action="metrics", metricName="TkeQPSService", resourceID="<serviceName>") (resourceID required)
  • Common issues:
    • Service not running (scaled to zero or crashed)
    • Image pull failures
    • OOMKilled events
    • Health check failures

Step 6 — Error Log Aggregation (if CLS is enabled)

queryLogs(action="searchLogs", queryString="ERROR", service="tcb", startTime="<24h-ago>", limit=50)
queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", startTime="<24h-ago>", limit=50)

Look for patterns:

  • Repeated error messages (same error many times)
  • Cascading failures (errors in multiple services around the same time)
  • Timeout / 429 / 404 / ACCESS_TOKEN_INVALID patterns → jump to the matching playbook

Step 7 — Summary Report

Generate a structured report:

# CloudBase Resource Inspection Report

**Environment**: ${envId}
**Inspection Time**: ${timestamp}

## Overall Health: ✅ Healthy / ⚠️ Warnings Found / ❌ Issues Found

## 告警解读
| 问题 | 指标 | 窗口峰值 | 基线 | 结论 |
|------|------|----------|------|------|
| 峰值 QPS | GatewayTraceEnvQPS | ... | package default 500 unless known | ... |
| CPU 告警是否正常 | MysqlCpuUsageRate | ... | warn≥80 / crit≥90 | ... |

### Cloud Functions
| Function | Status | Recent Errors | Invocations | Severity |
|----------|--------|---------------|-------------|----------|
| ... | ... | ... | ... | ... |

### CloudRun Services
| Service | Status | Issues | Severity |
|---------|--------|--------|----------|
| ... | ... | ... | ... |

### Error Log Summary
- Total errors in last 24h: N
- Top error patterns: ...

## Recommendations
1. ...
2. ...

## Console Links
- Cloud Functions: https://tcb.cloud.tencent.com/dev?envId=${envId}#/scf
- CloudRun: https://tcb.cloud.tencent.com/dev?envId=${envId}#/platform-run
- Logs: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops/log
- Monitor: https://tcb.cloud.tencent.com/dev?envId=${envId}#/devops

Targeted Inspection Workflow

When the user specifies a resource type or a specific resource:

  1. Cloud function errors: queryFunctions(action="listFunctionLogs", functionName="<name>") then queryLogs(action="searchLogs", queryString="* AND functionName:<name> AND level:ERROR", ...)
  2. CloudRun errors: queryCloudRun(action="detail", detailServerName="<name>") then queryLogs(action="searchLogs", queryString="ERROR", service="tcbr", ...)
    • If logs show DB / Redis connection failures (ECONNREFUSED, timeout, "could not connect"): check whether VpcConf is set and matches the database VPC. See cloudrun-development/references/vpc-and-database.md.
  3. Database issues: Check queryPgDatabase(action="context"|"metadata"|"objects") for CloudBase PG, queryMysqlDatabase for MySQL, or readNoSqlDatabaseStructure for NoSQL depending on type; for CPU/disk alerts also pull MysqlCpuUsageRate / MysqlStorageUsage metrics
  4. General error search: queryLogs(action="searchLogs", queryString="<error-keyword>", ...)
  5. Alarm / QPS questions: follow references/alarm-interpretation.md
  6. 429 / function 404 / ACCESS_TOKEN_INVALID / 调用量为 0: follow references/fault-playbooks.md

AIOps Methodology

This skill follows AIOps principles for intelligent inspection:

  1. Data Collection: Gather metrics (queryEnv metrics), logs, and resource states via MCP tools — never via ad-hoc callCloudApi
  2. Pattern Recognition: Identify recurring errors, anomaly patterns, and correlations across services
  3. Baseline Comparison: Compare metric Summary values to skill baselines (告警解读)
  4. Root Cause Hypothesis: Based on error patterns + metrics, suggest likely root causes
  5. Actionable Recommendations: Provide specific, prioritized remediation steps with links to relevant skills and console pages

Severity Levels

LevelIconMeaning
Critical❌Service is down or data is at risk; requires immediate action
Warning⚠️Errors detected but service is still partially functional; investigate soon
Infoℹ️No errors found; informational status only
Healthy✅Resource is operating normally

Preferred Tool Map

OperationMCP Tool Call
Check environmentqueryEnv(action="info")
Query metrics (QPS/CPU/invocations)queryEnv(action="metrics", envId, metricName="...")
Check CLS statusqueryLogs(action="checkLogService")
List cloud functionsqueryFunctions(action="listFunctions")
Get function detailqueryFunctions(action="getFunctionDetail", functionName="<name>")
Get function logsqueryFunctions(action="listFunctionLogs", functionName="<name>", startTime="<time>", endTime="<time>")
Get function log detailqueryFunctions(action="getFunctionLogDetail", requestId="<id>")
List CloudRun servicesqueryCloudRun(action="list")
Get CloudRun detailqueryCloudRun(action="detail", detailServerName="<name>")
Search CLS logsqueryLogs(action="searchLogs", queryString="<query>", service="tcb|tcbr", startTime="<time>", endTime="<time>")
Check NoSQL structurereadNoSqlDatabaseStructure(action="listCollections")
Check PostgreSQL contextqueryPgDatabase(action="context")
Check PostgreSQL metadataqueryPgDatabase(action="metadata", limit=20)
Check MySQL statusqueryMysqlDatabase(action="getContext")
Auth provider readinessqueryAppAuth / auth-tool skill (for ACCESS_TOKEN_INVALID)

Common CLS Query Patterns

ScenarioqueryString
All errorsERROR
Function timeouttimeout OR 超时
Function OOMOOM OR out of memory OR 内存超限
CloudRun crashcrash OR OOMKilled OR Error
Specific function errorsfunctionName:<name> AND level:ERROR
5xx HTTP errorsstatusCode:>499
429 / throttle429 OR throttle OR 限流 OR FREQUENCY
Function 404404 OR FUNCTION_NOT_FOUND
Token invalidACCESS_TOKEN_INVALID OR token invalid
Cold start issuescoldStart OR 冷启动

Time Range Guidance

  • Quick check: Last 1 hour (startTime = 1 hour ago)
  • Standard inspection: Last 24 hours
  • Trend analysis: Last 7 days
  • Specific incident: Narrow to the reported time window

Always use ISO-like YYYY-MM-DD HH:mm:ss for metrics startTime/endTime, e.g., "2026-08-17 00:00:00".

Related Skills

  • cloud-functions — Cloud function development, deployment, and debugging
  • cloudrun-development — CloudRun backend deployment and management
  • cloudbase-platform — General platform knowledge and console navigation
  • postgresql-development-cloudbase — CloudBase PostgreSQL / PG diagnostics and schema/RLS checks
  • relational-database-mcp-cloudbase — MySQL database management and diagnostics
  • auth-tool-cloudbase — Auth provider readiness for token failures

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill for Node.js backend AI via @cloudbase/node-sdk (>=3.16.0) — cloud functions, CloudRun, Express/Koa/NestJS, serverless APIs, scheduled jobs, LLM proxies, agent orchestration. The only SDK supporting image generation (ai.createImageModel + generateImage). Text via ai.createModel with groups cloudbase, hunyuan-exp, or custom-*; model ids (e.g. deepseek-v4-flash, glm-5, kimi-k2.6) go in the `model` field of generateText/streamText. MUST run two-step preflight before code — see body. NOT for browser/Web (use ai-model-web) or Mini Program (use ai-model-wechat).

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

Use this skill when a browser/Web app (React, Vue, Next, Nuxt, static sites, SPAs, dashboards, AI chat UI, 页面, 前端, 网页) needs AI models via @cloudbase/js-sdk. Default routing for Web/frontend AI — call directly from the browser, do NOT propose a Node.js proxy. Covers generateText and streamText; models via ai.createModel with groups cloudbase, hunyuan-exp, or custom-*, model id in the `model` field. MUST run two-step preflight before code — see body. NOT for Node.js backend (use ai-model-nodejs), Mini Program (use ai-model-wechat), or image generation (Node SDK only).

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

Use this skill for WeChat Mini Program AI via wx.cloud.extend.AI (小程序, wx.cloud apps). Covers generateText and streamText with callbacks (onText, onEvent, onFinish); streamText needs a data wrapper, generateText returns the raw response. Models via wx.cloud.extend.AI.createModel with groups hunyuan-exp (小程序成长计划), cloudbase (main managed), or custom-*; model id goes in the data wrapper `model` field. MUST run two-step preflight before code — see body. NOT for browser/Web (use ai-model-web), Node.js backend (use ai-model-nodejs), or image generation (use ai-model-nodejs).

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

Use when auditing CloudBase cloud API wrappers, MCP tools, generated action metadata, or related docs for outdated or incorrect action names, parameters, casing, request shapes, or missing contract tests, especially during periodic quality review or before preparing corrective PRs.

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

CloudBase Node SDK auth guide for server-side identity, user lookup, and custom login tickets. This skill should be used when Node.js code must read caller identity, inspect end users, or bridge an existing user system into CloudBase; not when configuring providers or building client login UI.

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

CloudBase auth provider configuration and login-readiness guide. This skill should be used when users need to inspect, enable, disable, or configure auth providers, publishable-key prerequisites, login methods, SMS/email sender setup, or other provider-side readiness before implementing a client or backend auth flow.

日本語の概要は準備中です。原文の説明を表示しています。

TencentCloudBase/CloudBase-AI-Toolkit1,1362026年10月10日 更新

TencentCloudBase のスキルをすべて見る

このスキルの問題を報告する