本文へ移動
cccskills
無料GitHub で公開

optimize-agent-prompt

Builds and improves Browserbase Agent API demos through an Autobrowse-style outer loop: run a fixed task, collect Agent messages and session logs, score the result, revise one system-prompt heuristic, and confirm convergence. Use when creating a Browserbase Agents demo or POC, optimizing an Agent system prompt, diagnosing flaky Agent runs, or applying auto-research/autobrowse to the Browserbase Agents API.

インストール方法を見る

含まれるファイル(7)

  • SKILL.md5.5 KB
  • agents/openai.yaml232 B
  • evals/evals.json2.2 KB
  • LICENSE.txt1.0 KB
  • references/api.md3.0 KB
  • references/evaluation.md2.3 KB
  • scripts/optimize_agent_prompt.mjs11.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Optimize Agent Prompt

Optimize a Browserbase Agent's systemPrompt while holding its task, result schema, variables, and evaluation criteria fixed. Treat the outer agent as the teacher and each Browserbase Agent run as an inner-agent rollout.

Use Node.js 18 or later and set BROWSERBASE_API_KEY. The harness uses only Node.js built-in modules.

Set up the experiment

Choose a short experiment name and create an isolated workspace inside the demo or POC repository:

node <skill-dir>/scripts/optimize_agent_prompt.mjs init \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --name <experiment-name>

Edit the generated files:

  • task.json: keep task, resultSchema, variables, browser settings, and evaluation oracle stable across iterations.
  • prompts/iteration-001.md: write the minimal baseline system prompt. Include irreversible-action guardrails when applicable.

Use concrete success criteria. Prefer a strict JSON Schema with required fields and null for unavailable facts. Add known-field regexes and factuality-warning regexes under evaluation when a truth oracle exists. Read references/evaluation.md when designing the task or score.

Run the baseline

node <skill-dir>/scripts/optimize_agent_prompt.mjs run \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --prompt prompts/iteration-001.md \
  --label iteration-001

The harness creates one reusable Browserbase Agent, updates its systemPrompt on later iterations, starts the run, polls messages and status, and writes:

runs/<label>/
├── system-prompt.md
├── created-run.json
├── run.json
├── messages.json
├── session-logs.json
└── summary.json

It stops a run after the configured message budget instead of paying for an unproductive spiral. Use --max-messages, --timeout-ms, --proxies, or --verified only when the task needs different values from task.json.

Diagnose from observable evidence

Start with the compact trajectory:

node <skill-dir>/scripts/optimize_agent_prompt.mjs inspect \
  --workspace ./agent-prompt-optimization/<experiment-name> \
  --label iteration-001

Then read summary.json and drill into messages.json at the first wrong or wasted turn. Agent messages expose ordered tool calls, tool results, errors, and final output. A reasoning part may contain no readable text; never require hidden chain-of-thought for the teacher loop.

Read session-logs.json only when browser-level evidence can distinguish the cause—for example, a redirect, 403, failed request, console error, or hidden endpoint. Empty session logs can mean the Agent completed with search/fetch tools and never drove its browser.

See references/api.md for endpoint shapes, pagination, result normalization, and trace caveats.

Improve one heuristic

Find the earliest consequential failure and state one counterfactual:

If the system prompt had instructed X, the Agent would have avoided Y, as shown by tool result Z.

Copy the current prompt to prompts/iteration-NNN.md and make one attributable change. Typical improvements are:

  • cap retries after a repeated block or identical error;
  • distinguish public identifiers from private/internal IDs;
  • prefer search/fetch before launching a browser when interaction is unnecessary;
  • separate current snapshots from dated historical events;
  • define when a qualified fallback counts as completed;
  • require null instead of guessed values;
  • add a tool-call or evidence budget.

Keep wins. If the new run regresses, restore the previous prompt and test a different hypothesis rather than stacking more rules.

Judge and converge

Generate the comparison table after each run:

node <skill-dir>/scripts/optimize_agent_prompt.mjs report \
  --workspace ./agent-prompt-optimization/<experiment-name>

Judge more than field completeness. Require:

  • terminal status COMPLETED;
  • required fields populated or explicitly nullable;
  • known-fact checks passing when available;
  • no factuality-warning match;
  • provenance and safety constraints preserved;
  • fewer messages or lower duration without quality loss.

Once a prompt wins, run it again unchanged with a new label. Converge only after it passes at least two of the last three runs and one pass is an unchanged confirmation. Do not call a prompt globally optimal from one task; describe it as the best prompt for the tested task distribution.

Graduate into the demo

Use the confirmed prompt as the Agent's production systemPrompt. Keep the strict result schema and per-run variables. Preserve the experiment workspace or its report so reviewers can audit why each instruction exists.

In the final handoff, report:

  • baseline versus winning score, duration, and message count;
  • the first wrong turn each prompt change fixed;
  • whether session logs added evidence;
  • the winning prompt path;
  • confirmation-run results;
  • limitations and the next holdout matrix.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Analyze an existing web application, identify safe user-visible capabilities across routes, forms, server actions, handlers, and schemas, then implement first-party WebMCP tools and validate discovery and invocation with Stagehand. Use when the user asks to make a codebase agent-ready, expose website features as WebMCP tools, or add WebMCP directly to an app rather than generating a standalone injection script from a URL.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

Audit the developer experience of a product, SDK, docs site, or SKILL.md by dropping multiple Claude subagents at it with only a tiny task prompt and real tools (WebFetch, Bash, Write). Agents must discover the docs themselves, install deps, ask for credentials if needed, and attempt real execution. The skill captures each agent's trace — tool calls, retries, wall time, errors — and scores on Setup Friction, Speed, Efficiency, Error Recovery, and Doc Quality, then emits an HTML report with an A–F grade and concrete fixes. Use when the user asks to audit agent experience, test a skill, audit docs for agents, check if a SDK is agent-friendly, validate a SKILL.md, measure agent DX, or benchmark how painful onboarding is for an AI agent. Triggers: 'audit agent experience', 'test this skill', 'audit docs for agents', 'is my SDK agent-friendly', 'run a DX audit', 'agent experience test', 'test my docs', 'how do agents do with my product'.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

Self-improving browser automation via the auto-research loop. Iteratively runs a browsing task, reads the trace, and improves the navigation skill (strategy.md) until it reliably passes. Supports parallel runs across multiple tasks using sub-agents. Use when you want to build or improve browser automation skills for specific website tasks.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

browser

無料

Automate web browser interactions using natural language via CLI commands. Use when the user asks to browse websites, navigate web pages, extract data from websites, take screenshots, fill forms, click buttons, or interact with web applications. Supports remote Browserbase sessions with Browserbase Identity, Verified browsers, automatic CAPTCHA solving, and residential proxies — ideal for protected websites and JavaScript-heavy pages.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

Turn a website's observable HTTP traffic into a best-effort OpenAPI 3.1 spec by analyzing a `browser-trace` capture. Use when the user wants to discover/extract API endpoints from a browser session, build an OpenAPI doc from network traffic, or document a third-party site's XHR/fetch surface for client integration.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

Capture a full DevTools-protocol trace of any browser automation — CDP firehose, screenshots, and DOM dumps — then bisect the stream into per-page searchable buckets. Use when the user wants to debug a failed run, audit network/console/DOM activity, attach a trace to an in-progress session, or feed structured per-page summaries back into an agent loop so its next iteration learns from the last one.

日本語の概要は準備中です。原文の説明を表示しています。

browserbase/skills3,7312026年9月30日 更新

browserbase のスキルをすべて見る

このスキルの問題を報告する