本文へ移動
cccskills
無料GitHub で公開

senpi-qa

Manual QA harness for the senpi coding agent itself. MUST USE after changing packages/ai, packages/agent, packages/coding-agent, or packages/tui — a green typecheck and `npm test` are NOT QA. Drives the real CLI from source in an isolated sandbox (never touches ~/.senpi or real credentials) across four channels: remote RPC (--mode rpc JSONL stdio), TUI smoke (node-pty on Windows, tmux on POSIX), mock loop (a local fake model server for deterministic, zero-token agent-loop runs), and CLI smoke (--help/--print/--list-models). Every helper ships a --self-test. Use whenever someone says qa senpi, test the agent, verify my change, rpc qa, tui qa, mock-loop qa, smoke the cli, or needs evidence that an agent-loop, tool, keybinding, or provider change works end to end. Capture evidence to local-ignore/qa-evidence/.

インストール方法を見る

含まれるファイル(154)

  • SKILL.md10.7 KB
  • .gitignore14 B
  • AGENTS.md3.6 KB
  • evals/evals.json1.2 KB
  • package-lock.json910 B
  • package.json472 B
  • references/credential-injection.md2.2 KB
  • references/env-vars.md2.6 KB
  • references/mock-loop.md9.1 KB
  • references/provider-error-tui.md2.3 KB
  • references/rpc-protocol.md2.7 KB
  • references/tui-driving.md2.5 KB
  • scripts/anthropic-oauth-callback-bind-fallback-probe.mjs2.5 KB
  • scripts/anthropic-subscription-auth-spike.mjs4.2 KB
  • scripts/anthropic-subscription-autocompact-settings-probe.mjs1.6 KB
  • scripts/anthropic-subscription-autocompact-spike.mjs10.6 KB
  • scripts/anthropic-subscription-fullstack-probe.mjs12.4 KB
  • scripts/anthropic-subscription-headless-restart-probe.mjs7.0 KB
  • scripts/anthropic-subscription-native-inline-spike.mjs12.0 KB
  • scripts/anthropic-subscription-persistent-query-spike.mjs10.9 KB
  • scripts/anthropic-subscription-reattach-spike.mjs13.2 KB
  • scripts/anthropic-subscription-registry-probe.mjs14.6 KB
  • scripts/anthropic-subscription-stream-stall-retry-probe.mjs7.1 KB
  • scripts/anthropic-subscription-stream-start-knob-probe.mjs6.8 KB
  • scripts/anthropic-subscription-sysprompt-spike.mjs3.6 KB
  • scripts/anthropic-subscription-toolless-compact-probe.mjs10.4 KB
  • scripts/astra-ultrafast-mock-loop.mjs4.5 KB
  • scripts/cli-smoke.mjs3.0 KB
  • scripts/compaction-remote-qa.mjs12.2 KB
  • scripts/eval-hard-limit-rpc-qa.mjs6.1 KB
  • scripts/footer-abbrev-qa.mjs4.1 KB
  • scripts/glm-5.3-preset-mock-loop.mjs2.9 KB
  • scripts/gpt-6-1-sol-preset-mock-loop.mjs6.6 KB
  • scripts/gpt-6-astra-preset-mock-loop.mjs4.6 KB
  • scripts/gpt-6-family-preset-mock-loop.mjs5.5 KB
  • scripts/gpt-test-decision-preset-mock-loop.mjs5.9 KB
  • scripts/grok-neo-drive.mjs20.0 KB
  • scripts/lib/anthropic-policy-refusal-server.mjs4.0 KB
  • scripts/lib/anthropic-subscription-fullstack-harness.mjs11.2 KB
  • scripts/lib/anthropic-subscription-fullstack-support.mjs7.7 KB
  • scripts/lib/anthropic-subscription-hermetic-env.mjs3.9 KB
  • scripts/lib/anthropic-subscription-matrix-assert.mjs3.7 KB
  • scripts/lib/anthropic-subscription-matrix-constants.mjs260 B
  • scripts/lib/anthropic-subscription-matrix-phases.mjs5.7 KB
  • scripts/lib/anthropic-subscription-matrix-run.mjs8.9 KB
  • scripts/lib/anthropic-subscription-matrix.mjs5.0 KB
  • scripts/lib/anthropic-subscription-reattach-worker.mjs8.1 KB
  • scripts/lib/anthropic-subscription-spike-credentials.mjs2.5 KB
  • scripts/lib/anthropic-subscription-spike-support.mjs9.7 KB
  • scripts/lib/anthropic-subscription-toolless-compact-server.mjs5.1 KB
  • scripts/lib/cache-warm-ready-rpc.mjs2.9 KB
  • scripts/lib/cache-warm-ready-scenario.mjs9.2 KB
  • scripts/lib/codex-websocket-mock.mjs6.3 KB
  • scripts/lib/common.mjs14.8 KB
  • scripts/lib/common.test.mjs539 B
  • scripts/lib/fake-model-server.mjs19.9 KB
  • scripts/lib/fallback-abort-server.mjs2.9 KB
  • scripts/lib/hint-429-server.mjs5.0 KB
  • scripts/lib/mock-loop-cli.mjs1.3 KB
  • scripts/lib/mock-loop-hint-429.mjs19.8 KB
  • scripts/lib/mock-loop-kimi-thinking-recovery.mjs3.5 KB
  • scripts/lib/mock-loop-policy-refusal.mjs4.9 KB
  • scripts/lib/mock-loop-retry.mjs7.5 KB
  • scripts/lib/mock-loop-support.mjs9.7 KB
  • scripts/lib/mock-loop-text-leak.mjs5.7 KB
  • scripts/lib/mock-loop-ttsr.mjs14.4 KB
  • scripts/lib/output-safety.mjs373 B
  • scripts/lib/rpc-client.mjs3.0 KB
  • scripts/lib/rpc-qa-client.mjs3.0 KB
  • scripts/lib/target-rpc-client.mjs2.7 KB
  • scripts/lib/terminal-screenshot.mjs4.8 KB
  • scripts/lib/tmux-tui-driver.mjs2.7 KB
  • scripts/lib/tui-resume-args.mjs5.7 KB
  • scripts/lib/tui-resume-args.test.mjs2.4 KB
  • scripts/lib/tui-resume-evidence.mjs3.7 KB
  • scripts/lib/tui-resume-pty.mjs4.5 KB
  • scripts/lib/tui-resume-pty.test.mjs3.4 KB
  • scripts/lib/tui-resume-signal.mjs1.6 KB
  • scripts/lib/tui-resume-signal.test.mjs2.3 KB
  • scripts/lib/tui-resume-teardown.mjs3.5 KB
  • scripts/lib/with-timeout.mjs672 B
  • scripts/look-at-qa.mjs3.6 KB
  • scripts/loop-guard-hard-escalation-qa.mjs5.9 KB
  • scripts/loop-guard-qa.mjs6.9 KB
  • scripts/mcred-slot-preservation-qa.mjs3.1 KB
  • scripts/mock-loop-anthropic-toolsearch.mjs3.6 KB
  • scripts/mock-loop-codex-websocket-liveness.mjs8.7 KB
  • scripts/mock-loop-credits-fallback.mjs7.2 KB
  • scripts/mock-loop-forbidden-fallback-return.mjs8.6 KB
  • scripts/mock-loop-long-retry-after-2446.mjs6.9 KB
  • scripts/mock-loop-stall-fallback.mjs6.1 KB
  • scripts/mock-loop-stream-retry.mjs4.8 KB
  • scripts/mock-loop-stream-start-timeout-steering.mjs12.6 KB
  • scripts/mock-loop-transport-timeout-recovery.mjs7.3 KB
  • scripts/mock-loop.mjs29.2 KB
  • scripts/moved-path-guard-qa.mjs14.2 KB
  • scripts/probes/cursor-cli/permissions-config-probe.mjs22.7 KB
  • scripts/probes/cursor-cli/prompt-ceiling-probe.mjs35.4 KB
  • scripts/probes/cursor-cli/resume-model-swap-canary.mjs15.5 KB
  • scripts/pty-drive.mjs4.9 KB
  • scripts/qa-1524-resume.mjs7.3 KB
  • scripts/qa-2480-overflow-ladder.mjs9.3 KB
  • scripts/rpc-drive.mjs11.3 KB
  • scripts/scenarios/anthropic-mid-output-fallback-qa.mjs5.7 KB
  • scripts/scenarios/cache-warm-ready-rpc-tui.mjs281 B
  • scripts/scenarios/compaction-abort-standdown-qa.mjs11.6 KB
  • scripts/scenarios/compaction-absolute-cap-qa.mjs8.7 KB
  • scripts/scenarios/compaction-body-too-large-qa.mjs17.2 KB
  • scripts/scenarios/compaction-retry-cycle-repro.mjs9.4 KB
  • scripts/scenarios/compaction-shared-retry-qa.mjs14.9 KB
  • scripts/scenarios/compaction-wave1.mjs5.9 KB
  • scripts/scenarios/core-route-warm-handoff-qa.mjs8.7 KB
  • scripts/scenarios/cursor-cli-oauth-auto-bootstrap-qa.mjs9.6 KB
  • scripts/scenarios/cursor-cli-oauth-qa.mjs31.7 KB
  • scripts/scenarios/cursor-composition-schema-qa.mjs6.5 KB
  • scripts/scenarios/cursor-exec-lifecycle-qa.mjs12.0 KB
  • scripts/scenarios/cursor-exec-lifecycle/cli-turn.mjs2.7 KB
  • scripts/scenarios/cursor-exec-lifecycle/frame-log.mjs2.7 KB
  • scripts/scenarios/cursor-exec-lifecycle/run-server.mjs4.3 KB
  • scripts/scenarios/cursor-exec-lifecycle/scenarios.mjs2.5 KB
  • scripts/scenarios/cursor-exec-lifecycle/wire.mjs4.1 KB
  • scripts/scenarios/cursor-oauth-catalog-refresh-qa.mjs8.1 KB
  • scripts/scenarios/cursor-oauth-catalog-refresh/setup.ts9.6 KB
  • scripts/scenarios/dollar-invocation-qa.mjs4.1 KB
  • scripts/scenarios/dollar-skill-invocation-qa.mjs9.3 KB
  • scripts/scenarios/durable-session-id-qa.mjs5.5 KB
  • scripts/scenarios/eval-execution-event-qa.mjs10.7 KB
  • scripts/scenarios/eval-live-headline-qa.mjs12.8 KB
  • scripts/scenarios/eval-throughput-badge-qa.mjs12.2 KB
  • scripts/scenarios/fallback-chains-kimi-k3-qa.mjs5.4 KB
  • scripts/scenarios/fallback-selector-nav-repro.mjs10.6 KB
  • scripts/scenarios/footer-monitor-elapsed-rpc.mjs4.1 KB
  • scripts/scenarios/fork-only-catalog-skip-qa.mjs3.0 KB
  • scripts/scenarios/goal-policy-rejection-fixture.ts3.1 KB
  • scripts/scenarios/goal-policy-rejection-qa.mjs4.4 KB
  • scripts/scenarios/goal-provider-auth-qa.mjs7.1 KB
  • scripts/scenarios/goal-reload-reengagement-qa.mjs13.6 KB
  • scripts/scenarios/grok-model-spec-qa.mjs6.9 KB
  • scripts/scenarios/idle-compaction-bracket-repro.mjs7.3 KB
  • scripts/scenarios/issue-1329-goal-continuation-compaction-qa.mjs8.9 KB
  • scripts/scenarios/issue-1329-qa.mjs10.2 KB
  • scripts/scenarios/per-model-thinking-memory-qa.mjs23.1 KB
  • scripts/scenarios/provider-error-tui-preload.mjs3.8 KB
  • scripts/scenarios/provider-error-tui-qa.mjs2.8 KB
  • scripts/scenarios/reload-stale-warmup-crash-qa.mjs11.9 KB
  • scripts/scenarios/rpc-fast-mode-surface-qa.mjs11.7 KB
  • scripts/scenarios/rpc-input-hardening-qa.mjs5.2 KB
  • scripts/scenarios/sdk-lane-delegated-admission-qa.mjs12.3 KB
  • scripts/scenarios/tui-long-session-profile.mjs7.0 KB
  • scripts/scenarios/warm-summary-anchor-qa.mjs8.4 KB
  • scripts/tui-fallback-abort-history.mjs6.8 KB
  • scripts/tui-resume.mjs4.1 KB
  • scripts/tui-scenario.mjs13.7 KB
  • scripts/tui-smoke.mjs8.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

senpi QA

QA the senpi coding agent (packages/{ai,agent,coding-agent,tui}) by driving the REAL CLI — not by reading code or trusting unit tests. Each channel runs the agent from source via tsx in an isolated sandbox and asserts observable behavior, so a passing run is evidence the user-facing surface actually works.

Every helper script ships a --self-test (or --self-check) that asserts its scenario against this machine. The scripts are therefore both the QA tools and their own regression checks.

Golden rules (read before running anything)

  • Isolation is mandatory. Everything spawns the CLI with SENPI_CODING_AGENT_DIR / SENPI_CODING_AGENT_SESSION_DIR pointed at a temp sandbox and PI_OFFLINE=1. QA must never write into the real ~/.senpi. scripts/lib/common.mjs does this for you — use it.
  • Never read or modify the real credentials. ~/.senpi/agent/auth.json is the user's real key store. Every script snapshots its sha256 and asserts it is unchanged at the end. If you script a run by hand, do the same (guardRealAuth() in common.mjs).
  • Deterministic loop = mock loop. To exercise the agent loop without real tokens, use Channel 3 (a local fake model server). Real-provider runs are for final smoke only and must use the user's existing auth, never a new key.
  • No src/ edits from this skill. It verifies; it does not fix. If QA finds a bug, report it with the captured evidence and let a follow-up change fix it.
  • The captured artifact IS the evidence. Write it under local-ignore/qa-evidence/<YYYYMMDD>-<slug>/. No artifact == the QA did not happen. local-ignore/ is gitignored — never commit evidence.

Setup (once)

node scripts/devenv-setup.mjs        # installs skill deps (node-pty), wires .env.local + .claude/skills
node .agents/skills/senpi-qa/scripts/lib/common.mjs --self-check   # confirm the harness

common.mjs --self-check confirms the repo resolves, a sandbox is created and auto-removed, a free port is allocatable, and the real auth file is untouched.

Router: match QA to your change

You changed…Run this channelReference
Agent loop, tools, sessions, provider/model resolution, RPCChannel 1 (RPC) — and Channel 3 for a deterministic loopreferences/rpc-protocol.md
Interactive TUI, keybindings, rendering, composerChannel 2 (TUI smoke)references/tui-driving.md
Anything where you want a full agent turn with ZERO tokensChannel 3 (mock loop)references/mock-loop.md
CLI flags, --help, --print, model listingChannel 4 (CLI smoke)—
Persistent terminal tools / packages/pty / PTY sessionsChannel 5 (pty-drive)—
Added a provider / auth pathChannel 3 + 4, and update references/env-vars.mdreferences/credential-injection.md

When in doubt, run the channel closest to your change AND Channel 3 (mock loop): the mock loop is the cheapest end-to-end proof that the agent still completes a turn.

Channels

All commands are run from the repo root.

Channel 1 — Remote RPC (scripts/rpc-drive.mjs)

Drives --mode rpc (JSON lines over stdio). get_state round-trips with no API call; --prompt drives a real turn and captures the event stream.

node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --self-test
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --state
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --prompt "say PONG" --provider mock --model mock-model --evidence rpc-pong

Channel 2 — TUI smoke (scripts/tui-smoke.mjs)

Boots the interactive TUI in a real pseudo-terminal, confirms it renders and a keystroke reaches the composer, then tears it down. Uses node-pty (ConPTY on Windows — no WSL) and falls back to tmux on POSIX.

node .agents/skills/senpi-qa/scripts/tui-smoke.mjs --self-test
node .agents/skills/senpi-qa/scripts/tui-smoke.mjs --self-test --driver tmux --evidence tui

TUI smoke proves boot/render/input, not fine-grained output. For behavioral assertions use Channel 1 or 3.

Channel 3 — Mock loop (scripts/mock-loop.mjs)

Starts a local fake model server, registers it via a baseUrl override in an isolated models.json, and drives a REAL turn — deterministic, zero tokens. Covers all three wire formats senpi uses, so baseUrl override is QA'd for both OpenAI and Anthropic (pick with --api; default openai-completions):

--apiprovider overriddenpath / auth
openai-completionsmock/v1/chat/completions · Bearer
anthropic-messagesanthropic/v1/messages · x-api-key
openai-responsesopenai/v1/responses · Bearer

--self-test (no --api) round-trips all three and exercises the three error/retry scenarios below. --with-tool proves the full loop (model → bash tool → final text). --with-mcp-tool registers a sandbox extension that proxies mcp_fx_tool_<n> to the local MCP stdio fixture, then asserts the fixture call log exists and the model's second request contains the fixture result. The loop is hermetic: provider key env vars are stripped so only the inline mock key is ever used.

node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test --api anthropic-messages
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-tool --api openai-responses
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-mcp-tool mcp_fx_tool_1 --tool-args '{"value":"ok"}'
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --with-eval-hard-limit --evidence eval-hard-limit
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario transient-recover
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario budget-exhaust
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --scenario long-retry-after
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --run "summarize this repo" --evidence mock-summary

Channel 4 — CLI smoke (scripts/cli-smoke.mjs)

Fast, no model: --help, --version, offline --list-models, unknown-flag handling.

node .agents/skills/senpi-qa/scripts/cli-smoke.mjs --self-test

Channel 5 — Persistent terminal / PTY (scripts/pty-drive.mjs)

Drives the real @earendil-works/pi-pty runtime that backs the built-in terminal tools (bash / bash_output / bash_input / bash_resize / kill_bash) through the canonical scenarios: background command + monitor-style line watch and peek, stdin steering, screen snapshot + resize reflow, and registry teardown with no orphans. Uses native PTY when a host prebuild is present, else the pipe fallback (--force-pipe forces it).

node .agents/skills/senpi-qa/scripts/pty-drive.mjs --self-test --evidence terminal
node .agents/skills/senpi-qa/scripts/pty-drive.mjs --self-test --force-pipe

Scripts index (each is its own regression test)

Script--self-test / --self-check asserts
scripts/lib/common.mjs --self-checkrepo + tsx resolve; sandbox created and auto-removed; free port; real auth.json unchanged
scripts/lib/fake-model-server.mjs --self-testOpenAI SSE contract: scripted text streams back, [DONE] sent, request recorded
scripts/rpc-drive.mjs --self-testget_state returns the documented RpcSessionState, no API call, auth unchanged
scripts/mock-loop.mjs --self-testscripted marker returns through the real loop via the mock provider; retry error injection proves same-model recovery, retry-budget fallback, and long-retry-after fallback; zero real calls; auth unchanged
scripts/mock-loop.mjs --with-toolfull loop: two model turns served, bash tool ran, final text returned
scripts/mock-loop.mjs --with-mcp-tool <tool>full loop with a registered sandbox MCP stdio fixture proxy; fails if the requested mcp_fx_tool_<n> is not registered, invoked, and fed back to the model
scripts/mock-loop.mjs --with-eval-hard-limitfull loop where a never-returning eval cell is killed by its wall-clock hard limit and the kill is reported back to the model
scripts/eval-hard-limit-rpc-qa.mjs --self-testRPC channel: a DETACHED eval cell killed by the hard limit injects its <system-reminder> kill notice into the model's next request (print mode cannot show this)
scripts/tui-smoke.mjs --self-testTUI boots, renders, accepts a keystroke, tears down; auth unchanged
scripts/cli-smoke.mjs --self-test--help/--version/--list-models work offline; unknown flag reported; auth unchanged
scripts/pty-drive.mjs --self-testPTY runtime backing the terminal tools: background line watch + peek, stdin steering, screen snapshot + resize, registry teardown (no orphans); auth unchanged

Run the whole suite:

for s in lib/common.mjs:--self-check lib/fake-model-server.mjs:--self-test \
         rpc-drive.mjs:--self-test mock-loop.mjs:--self-test \
         tui-smoke.mjs:--self-test cli-smoke.mjs:--self-test; do
  node ".agents/skills/senpi-qa/scripts/${s%%:*}" "${s##*:}" || echo "FAILED: $s"
done

Capturing evidence

ev="local-ignore/qa-evidence/$(date +%Y%m%d)-senpi-qa-<slug>"; mkdir -p "$ev"
node .agents/skills/senpi-qa/scripts/mock-loop.mjs --self-test | tee "$ev/mock-loop.txt"
node .agents/skills/senpi-qa/scripts/rpc-drive.mjs --prompt "say PONG" \
  --provider mock --model mock-model --evidence senpi-qa-<slug>

Most channels accept --evidence <slug> and write artifacts to local-ignore/qa-evidence/<date>-<slug>/ themselves.

References

  • references/rpc-protocol.md — RPC command/response catalog, turn completion, examples
  • references/tui-driving.md — node-pty vs tmux, keybindings files, fragility, isolation
  • references/mock-loop.md — fake server, custom-provider models.json shape, in-process faux alternative
  • references/credential-injection.md — per-harness credential injection + masking
  • references/env-vars.md — provider keys + isolation env vars

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

bun-1-4

無料

Read before a js eval cell that installs a package, spawns a server or PTY, or starts a long run, and before any bun -e, script, server, CLI, test, bundle, or package-management work: this session's eval js kernel runs Bun 1.4+ (this skill is present only when it does). Bun 1.4 ships builtins that replace 15+ npm deps — check here BEFORE installing sharp, puppeteer/playwright (scraping), marked, node-cron, node-pty, concurrently, serve-static, tar, json5, fast-xml-parser, string-width. Triggers: eval js, bun, Bun.serve, bun test, bun build, bun install, bun run, image resize, headless browser, markdown render, cron, PTY.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

Implements a single feature in the pi-mono `todotools` builtin extension work. Use for refactoring, continuation runtime, config resolver, prompt builder, test authoring, golden snapshots, CHANGELOG entries, and harness helpers. Does NOT do manual tmux QA.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

Example skill loaded from resources_discover

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

MUST read before generating images. Prompt-crafting guide for gpt-image-2.5 covering tool routing (native image_generation server tool vs the generate_image tool), model and quality selection, prompt structure, exact text rendering, reference-image editing, transparent assets, output formats, and multi-turn refinement.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

Sync a fork branch with an upstream remote using a history-preserving merge. Use this whenever the user says /merge-upstream, merge upstream, sync upstream, sync fork, or wants upstream changes integrated without rebasing or force-pushing.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

Release, publish, CalVer tag push, npm publish, 배포, 릴리즈 for senpi. Use for the canonical release flow, GitHub tag/release publication, and npm publishing.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/senpi4742026年10月11日 更新

code-yeongyu のスキルをすべて見る

このスキルの問題を報告する