本文へ移動
cccskills
無料GitHub で公開

browser

Drives a real browser through the omowright library from the js eval kernel: sites the user is already signed into, forms and clicks, JS-rendered pages, screenshots, web QA, extension popups, a human handoff for login, CAPTCHA or OTP, and a browser you own for scraping, bot-scored targets, network capture and QA traces. Use for any interactive browser task; not for a plain search or an unblocked static fetch.

インストール方法を見る

含まれるファイル(18)

  • SKILL.md8.6 KB
  • agents/openai.yaml42 B
  • ATTRIBUTION.md1.8 KB
  • references/commands.md6.2 KB
  • references/install.md5.8 KB
  • references/owned-engine/frames-and-humans.md2.6 KB
  • references/owned-engine/ladder.md3.8 KB
  • references/owned-engine/network.md2.9 KB
  • references/owned-engine/README.md3.3 KB
  • references/recipes/1password.md2.6 KB
  • references/remote.md1.1 KB
  • runtime/omowright/index.js933.4 KB
  • runtime/omowright/manifest.json347 B
  • runtime/omowright/page-bundle.js45.2 KB
  • scripts/browser-doctor.mjs4.1 KB
  • scripts/browser-engine-guard.mjs20.1 KB
  • scripts/browser-install.mjs2.8 KB
  • scripts/omowright.mjs1.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Browser

One library, two engines. omowright ships inside this skill; choose the engine before you act:

You needEngineEntry point
A site the user is signed into, their open tabs, a form, a click-through, a screenshot, web QA, an extension popupattached — the user's own browser through BrowserSkillconnectBrowserSkill()
A throwaway profile, bot-scoring evasion, a CAPTCHA, network interception, a QA flight trace, coordinate control, headless runsowned — a browser your code launchesconnectPipe() / connectCloakProfile() — references/owned-engine/README.md
Text out of a URL, a 403 bypass, a platform that blocks fetchersneitherthe ultimate-browsing skill

Attached is the default, because it is the only engine carrying the user's logins and the only one where a human is a single call away. Never substitute one engine for the other silently: if the attached engine is not set up, run the onboarding script and tell the user its one remaining step.

Step 0 — which engine this session is allowed to use

When OMO_BROWSER_ENGINE is set (the OmO desktop app sets it for every session), it wins over the table above:

ValueWhat you do
connectedUse connectBrowserSkill() only. If the user's browser is not connected you get a "Connect your browser" error: relay it and stop. Never open another browser
builtinDo not call connectBrowserSkill(); use the app's in-app browser tools
noneDo not do browser work. Say that agent browser access is off for this project
unsetThe table above, as before (terminal use)

While any engine is set, loadOmowright() returns a guarded library. The owned engine (connectPipe, connectCloakProfile, connect) and every other export that acts on a browser is refused, so the table above does not apply: do not look for a way around it, and tell the user what the session allows. Under connected the app sees what the browser is doing, and before a click, Enter or script that sends, posts, pays, orders, subscribes, deletes or closes an account, and before Enter in a message box, it asks the user first. A "No" fails the action with BrowserActionDeclinedError: report that, never retry it or go around it (session.tool() lets only reads through; evaluate is guarded too). If the user presses Stop, the next call throws BrowserUserStoppedError: tell the user browser use was stopped and start no new session this turn.

The guard prevents mistakes by a cooperating agent. It is not a security boundary: code that imports the raw entry (resolveOmowrightEntry()) is not guarded, and a host without the omo_browser_bridge tool cannot show state or honor Stop, though questions are still asked.

Step 1 — load omowright and prove the stack

const { loadOmowright } = await import("<skill-root>/scripts/omowright.mjs")
const { omowright } = await loadOmowright()          // { connectBrowserSkill, bskSnapshot, connectPipe, ... }
node "<skill-root>/scripts/browser-doctor.mjs" --json
StateMeaningNext
readyCLI, daemon and a connected browserstart a session
no-cli / no-daemon / no-extensionsomething is missingnode "<skill-root>/scripts/browser-install.mjs" [--browser=<id>] prepares everything it can for the browser the user uses, then prints the single step only the user can do (relaunch that browser and click Enable); relay it verbatim, wait, re-run the doctor
choose-browserthe signals do not single out one browser (Safari/Firefox default, an idle default while another browser runs, several in use)nothing was installed; take the browser from memory or ask the user, then browser-install.mjs --browser=<id>
no-browser-supportno Chromium-family profile on this machinesay so and stop

Install into the browser the user actually uses, never into whatever happens to be on disk. Before installing, check your memory for the user's browser; otherwise read the doctor's browser (picked from the OS default browser, running apps and recent use — candidates shows the evidence). If memory and the doctor disagree, or the doctor says choose-browser, ask the user. Pass the answer as --browser=<id> and record it in memory. A Chrome that is merely installed is not their browser.

Never launch a headless browser because the attached one is missing. It has none of the user's sessions, so every login turns into a ladder you should not be climbing. Say which state you hit and ask.

The loop (attached)

const session = await omowright.connectBrowserSkill({ name: "<task>", focused: false })
try {
  await session.navigate("https://example.com/", { waitUntil: "load" })
  const { tree, refs, css } = await omowright.bskSnapshot(session, { interactive: true })  // OmOWright tree + refs, no trace in the page
  await session.click({ selector: css.e3 })                                                 // css[ref] is null inside shadow roots:
  const vom = await session.observe({ maxTokens: 4000 })                                    //   then read the daemon's own tree ...
  await session.click("@e7")                                                                //   ... and click its @eN ref
  await session.fill(css.e5, "hello")
  await session.press("Enter")
  await session.waitForNavigation({ waitUntil: "load" })
  const shot = await session.screenshot()                                                   // { buffer, width, height, captureId }
} finally {
  await session.stop()                                                                      // success AND failure; returns borrowed tabs
}
  1. Read before every action. bskSnapshot refs and observe @eN refs are reissued on each call; use a ref in the same cycle you read it.
  2. Navigation and large DOM changes stale every ref. Read again rather than reusing.
  3. Two identical failures mean change approach, not retry. A third identical attempt is a defect.
  4. Borrow a user tab explicitly (tabList({ scope: "user" }), tabBorrow(id), tabReturn(id)). Borrowing prompts the user; never invent tab ids and never repeat a denied borrow.
  5. Always stop() the session, on success and on failure.

Every method, its options, and the failure codes are in references/commands.md.

When a human is the only way through

Login, CAPTCHA, OTP, a payment confirmation, a consent dialog:

const outcome = await session.requestHelp({ prompt: "<what you need done>", targets: ["@e4"], timeoutMs: 300_000 })

Then read the page again. Respect a cancelled or timed_out outcome; do not work around it by changing the extension's automation settings.

Rules

  • Never read credentials through the page. No evaluate that extracts a password, token, cookie or recovery code. The value of the attached engine is that the browser is already signed in.
  • Never clear cookies, cache or site data. It is the user's real profile; clearing it logs them out everywhere. No flow here needs it.
  • focused: false by default. The browser belongs to someone who is probably using it.
  • One short, named session per task, always stopped.
  • Bot-scored or WAF targets go to the owned engine (not while OMO_BROWSER_ENGINE is set: then say the site needs a browser the session does not allow). The attached engine's daemon enables console capture on every tab it drives, which is a known automation signal; CloakBrowser through connectCloakProfile() is the stealth path.

Where the rest lives

TopicRead
Session methods, targets, options, error codesreferences/commands.md
Installing: CLI, daemon, extension, the one human step, blocklisted extensionreferences/install.md
Agent on one machine, browser on anotherreferences/remote.md
Owned engine: launch, snapshot ladder, network, frames, human handoffreferences/owned-engine/README.md
Reading a 1Password vault the user has unlockedreferences/recipes/1password.md

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

ast-grep

無料

Searches and rewrites code by AST shape across 25 languages. Use when the target is a syntax pattern (every call/class/import shaped like X, a codemod, a YAML rule) rather than literal text; for plain strings, comments, or filenames, use rg.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

Finds, reads, and reconstructs coding-agent sessions across Codex, Claude, OpenCode, OMO/Senpi, and other local agent logs. Use when asked to find or search past sessions, transcripts, or subagent runs, or to recover what an earlier session did.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

Use when Codex needs to understand or respond to automatic comment-checker feedback emitted after an edit-like PostToolUse hook.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

Processes and analyzes data with resident-kernel engines (DuckDB, Polars) and one-shot tools. Use for CSV/parquet/JSON analysis, group-by/join/aggregation, time series, distributions, cleaning, or plotting a dataset.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

debugging

無料

Runs a hypothesis-driven debugging loop across any language or binary, escalating to orthogonal oracle angles and locking the fix with a failing test. Use for crashes, silent failures, hangs, wrong responses, memory leaks, async misbehavior, or reverse engineering.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

frontend

無料

Builds, styles, and polishes web UI and UX. Use for any frontend, page, component, styling, layout, animation, or visual-quality task, or when asked to make an interface look or feel a certain way.

日本語の概要は準備中です。原文の説明を表示しています。

code-yeongyu/lazycodex3,7572026年10月10日 更新

code-yeongyu のスキルをすべて見る

このスキルの問題を報告する