本文へ移動
cccskills
無料GitHub で公開

browser-bench-task

Author or expand browser-agent benchmark tasks in data/real_world_bench.json (Odysseys schema, per-rubric graded). Use when creating, adding, or expanding real-world browser tasks, writing rubrics, or picking bot-friendly sites — every task MUST be verified runnable through ego-browser before it is written.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md5.9 KB
  • references/authoring.md7.0 KB
  • references/ego-browser-testing.md10.4 KB
  • references/schema.md4.5 KB
  • references/site-selection.md2.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Browser benchmark task author

Why this skill exists

You are building tasks for data/real_world_bench.json — a browser-agent benchmark where each task is graded checkpoint-by-checkpoint (per-rubric). A good task is real-world, completable, and free of any risk-control block the autonomous run can't get past: an agent driving ego-browser must be able to actually finish it, and a grader must be able to verify each checkpoint from screenshots, trajectory, or final output.

The one rule that matters most, and the reason this skill exists:

Never write a task you have not run through ego-browser first. Web research and model priors lie about bot-friendliness — sites that look "low risk" on paper hard-block real automation (Product Hunt → Cloudflare "Just a moment"; Etsy → DataDome

  • forced login at checkout). The only proof is a live ego-browser probe. Then calibrate the rubrics to what you actually observed, not to what you imagined.

This operationalizes the repository's core authoring rule: calibrate every task from a real ego-browser run.

The loop

Do one task at a time. Read the referenced file the first time you need its detail.

  1. Load context. Read data/real_world_bench.json (tail for the live schema; scan all website + task_id values to avoid duplicating a site or play you already have). Skim references/schema.md for the schema and references/authoring.md for rubric discipline. Recall the bot-friendliness memory if present.

  2. Pick a candidate. Honor the user's constraints (US mainstream site, required level, cross-site wanted?, a mandated site such as reddit?). Avoid sites already in the dataset unless the play is genuinely new. To cover breadth fast you may fan out research with a Workflow (parallel web-search lanes → shortlist ranked by bot-risk), but research only prioritizes probe order — it never replaces the probe. → references/site-selection.md

  3. Probe with ego-browser — this is the gate. Drive the EXACT interactions the task will need: reach the start page, confirm the autonomous run won't be walled off, apply every filter/sort/form step, and confirm the target data is actually extractable. Capture the real trajectory and concrete values. An automated hard wall (Cloudflare / DataDome / PerimeterX that won't clear) or an unreachable key interaction → drop it and pick another, do not write the task. A human-verification / captcha is the one carve-out: pause and hand off to the user to solve it once — if it then clears and does not return on repeated probing, treat it as a one-time trust gate (not a per-request wall) and the site is usable; if the SAME site re-challenges after that solve, skip it — a per-entry captcha can't be driven by the agent at run time. Probes share one real browser, so run them sequentially, never in parallel. → references/ego-browser-testing.md

  4. Design + write confirmed_task. Fit a proven, gradeable pattern (aggregate-and-compute / form-fill-then-stop / calculator / cross-site or cross-app), with concrete filters, fields, computations, and stop-boundaries. Write it in natural user voice — first person, the motivation woven in, like the odysseys medium/hard tasks. No operator meta-instructions ("read-only", "don't log in", "must use old.reddit.com") — a real user would not say those; put any such rationale in the note field. → references/authoring.md

  5. Write rubrics (3–6). Each rubric = one independently verifiable checkpoint; verification is state-based (what a grader sees in a screenshot / final output, not an action sequence); operation rubrics must be "applied AND confirmed"; all of it calibrated to your real run. → references/authoring.md

  6. Append + validate. Append the object surgically (do not reformat or rewrite existing entries — keep the diff to an append). Then validate:

    uv run python -c "import json; json.load(open('data/real_world_bench.json')); print('json ok')"
    uv run python -c "from ego_bench.datasets import _ADAPTERS; ts=_ADAPTERS['real-world-bench']().load(); t=[x for x in ts if x.task_id==NEW_ID][0]; print(len(ts), 'rubrics:', len(t.metadata['rubrics']))"
    

    Missing/empty rubrics fail-loud in Stage B, so confirm the new entry carries them.

  7. Capture learnings. Update the bot-friendliness memory with any new pass/fail site and its evidence.

Definition of done

  • Ran end-to-end through ego-browser — no automated hard wall, and any human-verification cleared on a single handoff and did not recur; key interactions reached; data extracted.
  • confirmed_task reads like a real user request — first person, motivated, zero operator meta-instructions.
  • 3–6 rubrics, each one verifiable checkpoint, calibrated to the real run (not invented).
  • JSON valid, adapter loads, new entry has rubrics.
  • level / site_mode / site_count / categories / reference_length filled; note carries operator hints + the calibration record.

Reference files

  • references/schema.md — the Odysseys schema fields, semantics, and a worked JSON example.
  • references/ego-browser-testing.md — the gate: risk-control signals, copy-paste probe recipes, special cases (React inputs, cross-origin editors, captcha handoff), and the known bot-friendly / bot-hostile site table.
  • references/authoring.md — the proven task patterns, the natural-phrasing rules (with before/after), rubric discipline with examples, and the anti-patterns to avoid.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Analyze Codex provider rollout JSONL and ego-benchmark-harness run artifacts to identify slow or failed individual tasks, tool retries, token/cost anomalies, and differences in agent execution paths. Use when users ask to analyze a Codex session or rollout, explain why a Codex task was slow or failed, inspect a Codex provider execution trace, compare Codex runs, count function_call or tool output events, or investigate agent-home/.codex/sessions.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

Export Codex rollout/session JSONL as a self-contained HTML report that can be opened offline. Use when a user provides a complete session file path or Codex session ID, or asks to export a Codex session to HTML, generate a session report, or convert a rollout to HTML. Supports exact full-ID lookup in the CODEX_HOME sessions and archived_sessions directories.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

Verify a tested web agent's real behavior from its recorded Pi or Codex session and selected screenshots before scoring. The judge prompt provides absolute artifact paths; use targeted queries to compare actual outputs, errors, and the final answer with the recorded behavior, then return only the verdict JSON. Only for ego-bench judging—never to drive a browser or rerun the agent.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

Generate self-contained HTML/CSS benchmark and comparison charts in an ego-inspired Browser-Native Editorial style. Use for blue-white product metrics, model benchmarks, A/B comparisons, efficiency reports, difficulty breakdowns, and compact technical data panels.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

ego-browser (ego-lite) is a Chromium-based browser that gives AI agents a CLI-accessible Node.js runtime for driving a real browser. Use this skill whenever the user needs to interact with a website opening pages, filling forms, clicking buttons, taking screenshots, extracting page data, testing web apps, logging into sites, automating browser operations, or any other browser automation task. Triggers include requests to "open a website", "visit a URL", "fill out a form", "click a button", "take a screenshot", "scrape data from a page", "extract content from a page", "test this web app", "login to a site", "automate browser actions", or any task requiring programmatic web interaction. Also used for exploratory testing, dogfooding, QA, bug hunting, or reviewing app quality. Prefer ego-browser over any built-in browser automation, web fetch, or other web tools.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

keel

無料

Architecture design and governance protocol for system structure, boundaries, contracts, APIs, schemas, migrations, rewrites, and long-lived codebase health. Use when designing or reviewing architecture, planning structural refactors or migrations, governing tech debt, or containing rewrite risk. Keeps the load-bearing spine small, owned, checked, and deletable over time.

日本語の概要は準備中です。原文の説明を表示しています。

citrolabs/ego-browser-benchmark-framework92026年8月10日 更新

citrolabs のスキルをすべて見る

このスキルの問題を報告する