本文へ移動
cccskills
無料GitHub で公開日本語紹介

openclaw-qa-testing

OpenClawのQAシナリオを模擬環境や実サービスで実行し、失敗原因の修正、結果レポートの確認、複数モデルの応答スタイル比較まで支援するスキル。

原文Run, watch, debug, extend, or explain OpenClaw qa-lab and qa-channel scenarios, artifacts, and live lanes.

インストール方法を見る

こんなときに便利

  • 模擬環境や実サービスでQAを実行
  • 失敗シナリオの原因調査と再検証
  • 複数モデルのキャラクター比較
  • ローカルで処理追跡データを検証

日本語での紹介

できること

OpenClawのqa-labとqa-channelを使い、品質確認のシナリオを実行・調査・追加します。模擬応答を使う環境と実サービスでの検証を選び、成功・失敗の件数やレポート、保存された検証資料を確認します。失敗時は製品側やテスト実行の仕組みを調べ、修正後に対象のテスト一式を再実行します。複数モデルの話し方やキャラクターを比較し、順位や会話全文を含む評価結果も得られます。

こんなときに便利

OpenClawの動作を一連の利用シナリオで確認したい開発者や、モデルごとの応答スタイルを比較したい人に向いています。OpenTelemetryの処理追跡データを、外部サービスの認証情報なしで検証する用途にも使えます。

使い方の例

  • 「模擬環境でQA一式を実行し、失敗原因を調べて」
  • 「このキャラクター設定を複数モデルで比較して」
  • 「ローカルで追跡データの出力を確認して」

注意点

ソースを取得した開発環境向けです。通常のビルドにはQA機能が含まれません。実サービスの検証には対象に応じた認証情報が必要で、WhatsAppでは専用のテスト用アカウントを使います。認証情報はログや画像に残さない運用が求められます。

この紹介文は、公開されている SKILL.md をもとに AI(Claude Haiku)が作成しました。正確な仕様は下の原文を確認してください。

含まれるファイル(2)

  • SKILL.md11.3 KB
  • agents/openai.yaml500 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

OpenClaw QA Testing

Use this skill for qa-lab / qa-channel work. Repo-local QA only.

Read first

  • docs/concepts/qa-e2e-automation.md
  • docs/help/testing.md
  • docs/channels/qa-channel.md
  • qa/README.md
  • qa/scenarios/index.yaml
  • extensions/qa-lab/src/suite.ts
  • extensions/qa-lab/src/character-eval.ts

Model policy

  • Normal live suite runs rely on QA Lab source- and auth-aware defaults.
  • Do not pass --model, --alt-model, or --fast by default. Omitted --fast does not mean fast is disabled; fast behavior is source-owned.
  • For scenario-specific runs, the complete execution.summary is authoritative and overrides generic default guidance, including when it requires other flags. Add explicit provider/model pins only when execution.config.requiredProvider or requiredModel requires them.

Default workflow

  1. Read the scenario pack and current suite implementation.
  2. Decide lane:
    • mock/dev: mock-openai
    • real validation: live-frontier
  3. For a normal live suite, use:
pnpm openclaw qa suite \
  --provider-mode live-frontier \
  --output-dir .artifacts/qa-e2e/run-all-live-frontier-<tag>
  1. Watch outputs:
    • summary: .artifacts/qa-e2e/run-all-live-frontier-<tag>/qa-suite-summary.json
    • report: .artifacts/qa-e2e/run-all-live-frontier-<tag>/qa-suite-report.md
  2. If the user wants to watch the live UI, find the current openclaw-qa listen port and report http://127.0.0.1:<port>.
  3. If a scenario fails, fix the product or harness root cause, then rerun the full lane.

OTEL smoke

For local QA-lab OpenTelemetry validation, use:

pnpm qa:otel:smoke

This starts a local OTLP/HTTP trace receiver, runs the otel-trace-smoke scenario through qa-channel, decodes the emitted protobuf spans, and verifies the exported trace names and privacy contract. It does not require Opik, Langfuse, or external collector credentials.

QA credentials and 1Password

  • Use op only inside tmux for QA secret lookup in this repo.
  • Quick auth check inside tmux:
op account list
  • Direct Telegram npm live test secrets currently live in 1Password item:
    • vault: OpenClaw
    • item: Telegram E2E
  • That item is the first place to look for:
    • OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN
    • OPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN
    • OPENCLAW_QA_PROVIDER_MODE
    • OPENCLAW_NPM_TELEGRAM_PACKAGE_SPEC
  • Convex QA secrets currently live in 1Password items:
    • vault: OpenClaw
    • item: OPENCLAW_QA_CONVEX_SITE_URL
    • item: OPENCLAW_QA_CONVEX_SECRET_MAINTAINER
    • item: OPENCLAW_QA_CONVEX_SECRET_CI
  • Additional related notes/login items seen during QA credential work:
    • vault: Private
    • items: OPENCLAW QA, Convex, Telegram
  • If a required value is missing from those notes:
    • do not guess
    • ask the maintainer/operator for the current value or the current 1Password item name
    • for Telegram direct runs, OPENCLAW_QA_TELEGRAM_GROUP_ID may be stored separately from Telegram E2E
    • for Convex runs, the leased Telegram credential should provide the Telegram group id and bot tokens together; do not require a separate OPENCLAW_QA_TELEGRAM_GROUP_ID
    • for Convex runs, prefer OpenClaw/OPENCLAW_QA_CONVEX_SITE_URL; if that is stale or unclear, ask for the active pool URL before running
  • Prefer direct Telegram envs for the npm Telegram Docker lane when available:
OPENCLAW_QA_TELEGRAM_GROUP_ID="..." \
OPENCLAW_QA_TELEGRAM_DRIVER_BOT_TOKEN="..." \
OPENCLAW_QA_TELEGRAM_SUT_BOT_TOKEN="..." \
OPENCLAW_QA_PROVIDER_MODE="mock-openai" \
OPENCLAW_NPM_TELEGRAM_PACKAGE_SPEC="openclaw@beta" \
pnpm test:docker:npm-telegram-live
  • Prefer Convex mode when the goal is stable shared QA infra:
    • round-robin credential leasing
    • thinner wrapper for channel-specific setup
    • CLI/admin flows around the pooled credentials
  • Live npm Telegram Docker lane note:
    • scripts/e2e/npm-telegram-live-runner.ts reads OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE
    • do not assume OPENCLAW_QA_PROVIDER_MODE is consumed by that wrapper
    • if a 1Password note only gives OPENCLAW_QA_PROVIDER_MODE, map it explicitly to OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE before running the Docker lane
  • Verified live shape:
    • Convex mode can pass the real Docker lane without direct Telegram env vars
    • leased Telegram payload includes the group id coupled to the driver/SUT tokens
    • a real run of pnpm test:docker:npm-telegram-live passed with:
      • OPENCLAW_QA_CREDENTIAL_SOURCE=convex
      • OPENCLAW_QA_CREDENTIAL_ROLE=maintainer
      • OPENCLAW_QA_CONVEX_SITE_URL
      • OPENCLAW_QA_CONVEX_SECRET_MAINTAINER
      • OPENCLAW_NPM_TELEGRAM_PROVIDER_MODE=mock-openai
  • If direct Telegram env is missing locally and op signin blocks, prefer dispatching the manual GitHub lane because the qa-live-shared environment already has Convex CI credentials:
gh workflow run "NPM Telegram Beta E2E" --repo openclaw/openclaw --ref main \
  -f package_spec=openclaw@YYYY.M.D-beta.N \
  -f package_label=openclaw@YYYY.M.D-beta.N \
  -f provider_mode=mock-openai
  • Poll the exact run id from the dispatch URL. gh run view --json artifacts is not supported; list artifacts with:
gh api repos/openclaw/openclaw/actions/runs/<run-id>/artifacts

WhatsApp live credentials

Use this when setting up or replacing Convex kind=whatsapp credentials.

  • Treat WhatsApp QA credentials as operator-owned live accounts, not generated fixtures.
  • Use two dedicated WhatsApp-capable test numbers: one driver account and one SUT account. Do not use personal numbers or personal OpenClaw WhatsApp accounts in the shared pool.
  • Register and link each account manually with WhatsApp or WhatsApp Business, storing Web auth only in isolated local auth dirs outside the repo.
  • For group coverage, create a dedicated test group that includes both QA accounts and store its JID as groupJid; otherwise the group mention-gating scenario should be skipped by default and fail when explicitly requested.
  • Package the two Baileys auth dirs into base64 .tgz payload fields and add a new active Convex credential row. Prefer adding a fresh row and disabling stale/broken rows over overwriting credentials in place.
  • Expected payload fields: driverPhoneE164, sutPhoneE164, driverAuthArchiveBase64, sutAuthArchiveBase64, and optional groupJid.
  • Keep credential material out of the repo, logs, PRs, and screenshots. Redact phone numbers unless the operator explicitly asks for local debugging.
  • Validate with pnpm openclaw qa whatsapp --credential-source convex --credential-role maintainer --provider-mode mock-openai and preserve artifact paths plus redacted pass/fail summaries.
  • If WhatsApp expires or invalidates a linked Web session, relink locally, package fresh auth archives, add a new Convex row, then disable the stale row.

Character evals

Use qa character-eval for style/persona/vibe checks across multiple live models.

pnpm openclaw qa character-eval \
  --output-dir .artifacts/qa-e2e/character-eval-<tag>
  • Runs local QA gateway child processes, not Docker.
  • Packaged pnpm build omits QA Lab + qa-channel by design (source-checkout only). To exercise openclaw qa/qa-channel from a built dist, build with OPENCLAW_BUILD_PRIVATE_QA=1 pnpm build (emits dist/plugin-sdk/qa-lab.js, qa-runtime.js, dist/extensions/{qa-lab,qa-channel}) or run via pnpm dev.
  • With no model flags, character eval uses its current source-defined candidate, judge, thinking, and fast defaults.
  • Repeat --model provider/model,thinking=<level>[,fast|,no-fast|,fast=<bool>] or --judge-model ... only to replace the corresponding inventory explicitly.
  • Do not add new examples with separate --model-thinking; keep that flag as legacy compatibility only.
  • Report includes judge ranking, run stats, durations, and full transcripts; do not include raw judge replies. Duration is benchmark context, not a grading signal.
  • Candidate and judge concurrency default to 16. Use --concurrency <n> and --judge-concurrency <n> to override when local gateways or provider limits need a gentler lane.
  • Scenario source is YAML-only under qa/scenarios/: use index.yaml and per-scenario *.yaml files with top-level title, scenario, and optional flow. Never add fenced qa-scenario / qa-flow Markdown files.
  • For isolated character/persona evals, write the persona into SOUL.md and blank IDENTITY.md in the scenario flow. Use SOUL.md + IDENTITY.md only when intentionally testing how the normal OpenClaw identity combines with the character.
  • Keep prompts natural and task-shaped. The candidate model should receive character setup through SOUL.md, then normal user turns such as chat, workspace help, and small file tasks; do not ask "how would you react?" or tell the model it is in an eval.
  • Prefer at least one real task, such as creating or editing a tiny workspace artifact, so the transcript captures character under normal tool use instead of pure roleplay.

Codex CLI model lane

Use model refs shaped like codex-cli/<codex-model> whenever QA should exercise Codex as a model backend.

Examples:

pnpm openclaw qa suite \
  --provider-mode live-frontier \
  --model codex-cli/<codex-model> \
  --alt-model codex-cli/<codex-model> \
  --scenario <scenario-id> \
  --output-dir .artifacts/qa-e2e/codex-<tag>
pnpm openclaw qa manual \
  --model codex-cli/<codex-model> \
  --message "Reply exactly: CODEX_OK"
  • Treat the concrete Codex model name as user/config input; do not hardcode it in source, docs examples, or scenarios.
  • Live QA preserves CODEX_HOME so Codex CLI auth/config works while keeping HOME and OPENCLAW_HOME sandboxed.
  • Mock QA should scrub CODEX_HOME.
  • If Codex returns fallback/auth text every turn, first check CODEX_HOME, relevant secret-backed auth, and gateway child logs before changing scenario assertions.
  • For model comparison, include codex-cli/<codex-model> as another candidate in qa character-eval; the report should label it as an opaque model name.

Repo facts

  • Seed scenarios live in qa/scenarios/index.yaml and qa/scenarios/<theme>/*.yaml.
  • Main live runner: extensions/qa-lab/src/suite.ts
  • QA lab server: extensions/qa-lab/src/lab-server.ts
  • Child gateway harness: extensions/qa-lab/src/gateway-child.ts
  • Synthetic channel: extensions/qa-channel/

What “done” looks like

  • Full suite green for the requested lane.
  • User gets:
    • watch URL if applicable
    • pass/fail counts
    • artifact paths
    • concise note on what was fixed

Common failure patterns

  • Live timeout too short:
    • widen live waits in extensions/qa-lab/src/suite.ts
  • Discovery cannot find repo files:
    • point prompts at repo/... inside seeded workspace
  • Subagent proof too brittle:
    • prefer stable final reply evidence over transient child-session listing
  • Harness “rebuild” delay:
    • dirty tree can trigger a pre-run build; expect that before ports appear

When adding scenarios

  • Add or update scenario YAML under qa/scenarios/; do not add .md scenario files or fenced YAML blocks.
  • Keep kickoff expectations in qa/scenarios/index.yaml aligned
  • Add executable coverage in extensions/qa-lab/src/suite.ts
  • Prefer end-to-end assertions over mock-only checks
  • Save outputs under .artifacts/qa-e2e/

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

1password

無料日本語概要

1Password CLIの導入と認証を確認し、保存したパスワードやAPIキーをコマンドや設定へ渡します。デスクトップ連携やサービスアカウントにも対応します。

  • 1Password CLIを導入したいとき
  • APIキーをコマンドに渡したいとき
  • CIでサービスアカウント認証を使う
openclaw/openclaw39.2万2026年10月10日 更新

acp-router

無料日本語概要

OpenClawへの自然な言葉の依頼をClaude Codeなどの外部コーディングエージェントへ振り分け、作業の開始や継続、スレッド内の会話をつなぐスキルです。

  • Claude Codeをスレッドで開始
  • 外部エージェントの作業を続けたいとき
  • acpxから直接指示を渡したいとき
openclaw/openclaw39.2万2026年10月10日 更新

Add and live-prove a model provider with non-interactive config one-liners, without exposing credentials.

日本語の概要は準備中です。原文の説明を表示しています。

openclaw/openclaw39.2万2026年10月10日 更新

Requested GitHub PR/issue agent transcripts: redact, trim, preview, and insert safely.

日本語の概要は準備中です。原文の説明を表示しています。

openclaw/openclaw39.2万2026年10月10日 更新

apple-notes

無料日本語概要

macOSのApple Notesをエージェントから作成・検索・編集・削除し、フォルダ間の移動やHTML・Markdownへの書き出しを行うスキル。

  • タイトルを付けてメモを作りたいとき
  • フォルダ指定やあいまい検索でメモ探し
  • メモの編集とフォルダ整理
openclaw/openclaw39.2万2026年10月10日 更新

apple-reminders

無料日本語概要

Apple Remindersの予定付きToDoをMacから確認・追加・編集するスキル。リストの管理や完了・削除にも対応し、iPhoneやiPadで見るタスクを整理できます。

  • 今日のタスクや期限超過を確認したいとき
  • 期限付きの個人ToDoを追加したいとき
  • iPhoneやiPadのタスクを整理
openclaw/openclaw39.2万2026年10月10日 更新

openclaw のスキルをすべて見る

このスキルの問題を報告する