Drive an installed LaRuche App through its declared actions, never its files.
日本語の概要は準備中です。原文の説明を表示しています。
Drive the machine itself, mouse, keyboard and screen, without needing vision
インストールする前に、エージェントに与えられる指示の中身を確認できます。
The computer tool drives the machine outside the browser: desktop applications, installers,
native dialogs, anything that is not a web page.
For a web page, use browser instead. It reads the DOM and clicks by element number,
so it is faster, cheaper in tokens and it does not miss.
This is the one that matters, and the only one that works if you cannot see images.
windows lists every open window, with its size, position, application and which
screen it is on. The focused one is marked *, and the screen your last screenshot
came from is marked <- captured. Minimised windows are listed separately, by title:
they are still open, and focus_window or restore_window brings one back.
On a multi-monitor setup this is the line that matters. A window on screen 2 cannot be
clicked with coordinates read off a capture of screen 1: those numbers are valid only
for the screen they came from, so you would click a real point on the wrong monitor.
Either capture that screen first, or act by ref, which needs no coordinates at all and
does not care which monitor the window sits on. A window straddling two monitors is
attributed to the one showing most of it, which is the one a human would name.
focus_window { "window": "Notepad" } brings one to the front, matching a piece of its
title.
read returns a numbered map of that window's controls, plus the static text on it:
ref_1 <document> Text editor = the current contents
ref_7 <tab> notes.txt
ref_13 <menuitem> File
ref_21 <button> Minimise
click, double_click, right_click and fill take a ref instead of coordinates.
focus gives a control focus without acting on it.
read again to see what changed.
Acting by ref goes through the OS accessibility API. It calls the control directly instead
of simulating someone aiming at it, so it is deterministic, it cannot land a pixel off, and
it works even when the window is not in front. The result says which pattern was used
(invoke, toggle, select, expand, setvalue, rangevalue, default action) or
tells you it had to fall back to a real click.
Refs come only from the latest read or find. Every read renumbers them, and changing
window throws them away. Read again after acting, exactly as with the browser tool.
read uses UI Automation, so it is wired for Windows today. Elsewhere, use the pixel path
below; windows still works everywhere.
read stops at 300 controls and 120 lines of static text, and says so when it truncates. On
a large Electron app or an office suite you will hit that, and you will pay for it every
time.
find { "text": "Save" } runs the same read and returns only the controls whose label
matches. Use it whenever you already know what you are looking for. It renumbers refs like
any other read, and the numbers it hands back are usable immediately.
Every one of these acts on the element, not on a guess:
Action with a ref | What happens |
|---|---|
left_click | its automation pattern, or a real click if it exposes none |
right_click, double_click, triple_click, middle_click | a real click at its centre |
mouse_move (or hover) | the pointer moves onto it and nothing is clicked |
scroll | it is scrolled into view through its own pattern |
left_click_drag | it is the grab point, to_x and to_y are the drop |
fill | text into a field, or a number into a slider or spinner |
focus | focus only, no action |
A slider or a spinner is set, not clicked. click on one is refused, and the refusal
names its minimum, maximum and current value. Use fill with the number you want. A click
would have landed at the centre of the control's rectangle, which is halfway along its
travel, and reported success.
scroll on a ref is the only way to reach an off-screen control without a screenshot.
read skips anything scrolled out of view, so a control below the fold does not exist as
far as the tree is concerned. Scroll a nearby ref into view, then read again.
Hand-drawn interfaces expose nothing: games, canvases, poorly tagged Electron apps. When
read comes back empty, fall back to looking.
screens lists the monitors: rank, id, size, position, scale, name.screenshot captures one (screen picks it by id, rank or name; default is the primary)
and returns it as an image, 1280px wide by default. That capture sets the coordinate
system.mouse_move, left_click, left_click_drag, scroll at x,y read off that image.Coordinates are always pixels of the last screenshot. Never desktop pixels. The tool converts, including on a display scaled to 150% and on a mixed setup where a 4K sits next to a 1080p. Pointing before any screenshot is refused on purpose: without a capture there is no shared coordinate system, and treating the numbers as desktop pixels would work on a simple setup and fail silently everywhere else.
max_width raises the resolution when you genuinely cannot read what you need. It costs
tokens in proportion.
move_window takes window plus any of x, y, width, height, in desktop pixels.
A maximised window is restored first, because Windows ignores a placement request on one
and would report a success that changes nothing.
minimize_window, maximize_window and restore_window do what they say.
close_window is the cross, not a kill: the application gets the request, may put up a
save prompt, and may refuse. Call windows again to see whether it actually went, and
read the screen if a dialog appeared. To really kill a process, that is shell_exec.
All of them take a piece of the title, like focus_window. An exact title wins over a
partial match, and an ambiguous piece is refused with the list of candidates rather than
acted on: closing the wrong window is not recoverable.
left_click_drag does the whole gesture in one call: press, move in steps, release. It
holds briefly at each end, because HTML5, Electron and most file managers treat a press
that moves immediately as a click and drop nothing. hold_ms changes that pause,
button makes it a right or middle drag, which is how CAD and 3D tools orbit and pan.
For anything the single call cannot express, mouse_down and mouse_up hold a button
across several calls, exactly as key_down and key_up do for the keyboard:
key_down Shift, mouse_down, mouse_move, mouse_up, key_uptype types into whatever has focus, up to 5000 characters. For more than that, or for
anything with awkward characters, put it on the clipboard with write_clipboard and press
Control+v: there is no length limit that way, and no per-character timing to go wrong.
key presses one key or a chord.
Both tell you which window they went to. Read that line. A keystroke goes to the front
window, not to the one you have in mind, and on a multi-screen setup the window you were
thinking of is often on the other screen with something else in front. If the reported focus
is not what you expected, nothing you typed went where you thought: focus_window first,
then type again. Typing into LaRuche's own window is refused outright.
Named keys: Enter, Tab, Escape, Backspace, Delete, Space, Home, End,
PageUp, PageDown, Up, Down, Left, Right, F1 to F12, or a single character.
Chord with Control+c, Shift+Tab, Alt+F4, Meta+r. repeat presses several times,
hold_ms holds each press.
key_down and key_up hold a key while other gestures happen, which is what a game needs
and what a drag with a modifier needs.
Anything left held is released for you after a minute, and release_all releases
everything on demand. Release deliberately anyway: a key still down changes what every
later click means, including the user's own, and the automatic release is a safety net for
an abandoned mission, not a substitute for finishing the gesture.
read_clipboard and write_clipboard move text in and out. This is the shortest route
into and out of a native application, and it is the only way to read what a Control+c
actually copied.
wait { "ms": 1500 } pauses while an installer advances, a menu animates open, or a
dialog appears. Use it instead of taking a screenshot you do not need: a capture costs
tokens and a wait costs nothing.
The screen being driven gets an amber frame, a floating panel that names each action as it happens, and a ring that follows the cursor and pulses on every click. The cursor glides to its target instead of teleporting, text appears character by character, and a screenshot triggers a brief flash once the capture is taken. It all fades a few seconds after LaRuche stops acting. The panel can be dragged anywhere and stays where the user puts it.
glow: false removes the decoration, animate: false acts instantly on a long sequence,
speed above 1 slows the motion down for a demonstration. Acting by ref shows the panel
line without moving the mouse at all, because nothing needs to be aimed at.
Two refusals are by design, not bugs to work around:
type is checked too, even though it aims at no point: a long burst of
typing is exactly what someone reaching for the mouse wants to stop.A window running as administrator cannot be driven at all. Windows blocks synthetic
input across that privilege boundary, silently: the click is sent, does nothing, and
returns no error, which is indistinguishable from aiming badly. focus_window and type
warn you when the target is elevated, and the window actions refuse outright. This cannot
be worked around and should not be. Say that the window has to be driven by hand, or that
LaRuche has to be restarted as administrator, and stop trying.
Ctrl+Alt+Shift+H stops everything, immediately. It is registered globally as soon as
the tool is first used, and it works whether or not the glow is on. Pressing it releases
every held key and button, refuses the call in flight, and interrupts a long burst of
typing part way through, reporting how many characters had already landed.
If a call comes back saying the user aborted, they took the machine back on purpose. Say so and stop. Do not retry the gesture, and do not look for another way to do the same thing: that is the one response the shortcut exists to prevent.
Approval comes in two classes. Approving one observing action (screens, screenshot,
cursor_position) does not approve a click, and the first acting call is asked
separately.
The whole tool can be switched off on a node with LARUCHE_COMPUTER=0. If every call comes
back saying GUI control is disabled, that is why, and it is the operator's decision: say so
rather than looking for another way in.
browser.shell_exec or file_*. Driving a file
manager by clicking is slower and fails in more ways.systeminfo prints is wasted work and wasted tokens.The right use is the case with no other door: a native application with no CLI, an installer, a dialog the OS put on top of everything.
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Drive an installed LaRuche App through its declared actions, never its files.
日本語の概要は準備中です。原文の説明を表示しています。
Find academic papers on arXiv, with citation counts and BibTeX.
日本語の概要は準備中です。原文の説明を表示しています。
Render text or an image as ASCII art for terminal-friendly output.
日本語の概要は準備中です。原文の説明を表示しています。
Track RSS/Atom feeds and blogs via blogwatcher-cli.
日本語の概要は準備中です。原文の説明を表示しています。
Drive a real web browser: navigate, read, find, click, fill, screenshot
日本語の概要は準備中です。原文の説明を表示しています。
Measure a codebase: lines of code, language mix, and symbol lookups.
日本語の概要は準備中です。原文の説明を表示しています。