How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
日本語の概要は準備中です。原文の説明を表示しています。
Detect all UI elements on screen using OmniParser
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Parse the screen with OmniParser v2 to find all visible UI elements. TestDriver returns structured data. This includes element types, text content, interactivity levels, and bounding box coordinates.
This method examines all of the screen and returns each element that it finds. Use it for these:
const result = await testdriver.parse()
None.
Promise<ParseResult> - Object containing detected UI elements
| Property | Type | Description |
|---|---|---|
elements | ParsedElement[] | Array of detected UI elements |
annotatedImageUrl | string | URL of the annotated screenshot with bounding boxes |
imageWidth | number | Width of the analyzed screenshot |
imageHeight | number | Height of the analyzed screenshot |
| Property | Type | Description |
|---|---|---|
index | number | Element index |
type | string | Element type (e.g. "text", "icon", "button") |
content | string | Text content or description of the element |
interactivity | string | Interactivity level (e.g. "clickable", "non-interactive") |
bbox | object | Bounding box in pixel coordinates {x0, y0, x1, y1} |
boundingBox | object | Bounding box as {left, top, width, height} |
const result = await testdriver.parse();
console.log(`Found ${result.elements.length} elements`);
result.elements.forEach((el, i) => {
console.log(`${i + 1}. [${el.type}] "${el.content}" (${el.interactivity})`);
});
const result = await testdriver.parse();
const clickable = result.elements.filter(e => e.interactivity === 'clickable');
console.log(`Found ${clickable.length} clickable elements`);
clickable.forEach(el => {
console.log(`- "${el.content}" at (${el.bbox.x0}, ${el.bbox.y0})`);
});
const result = await testdriver.parse();
// Find a "Submit" button
const submitBtn = result.elements.find(e =>
e.content.toLowerCase().includes('submit') && e.interactivity === 'clickable'
);
if (submitBtn) {
// Calculate center of the bounding box
const x = Math.round((submitBtn.bbox.x0 + submitBtn.bbox.x1) / 2);
const y = Math.round((submitBtn.bbox.y0 + submitBtn.bbox.y1) / 2);
await testdriver.click({ x, y });
}
const result = await testdriver.parse();
// Get all text elements
const textElements = result.elements.filter(e => e.type === 'text');
textElements.forEach(e => console.log(`Text: "${e.content}"`));
// Get all icons
const icons = result.elements.filter(e => e.type === 'icon');
console.log(`Found ${icons.length} icons`);
// Get all buttons
const buttons = result.elements.filter(e => e.type === 'button');
console.log(`Found ${buttons.length} buttons`);
import { describe, expect, it } from "vitest";
import { TestDriver } from "testdriverai/vitest/hooks";
describe("Login Page", () => {
it("should have expected form elements", async (context) => {
const testdriver = TestDriver(context);
await testdriver.provision.chrome({
url: 'https://myapp.com/login',
});
const result = await testdriver.parse();
// Assert expected elements exist
const textContent = result.elements.map(e => e.content.toLowerCase());
expect(textContent).toContain('email');
expect(textContent).toContain('password');
// Assert there are clickable elements
const clickable = result.elements.filter(e => e.interactivity === 'clickable');
expect(clickable.length).toBeGreaterThan(0);
});
});
const result = await testdriver.parse();
result.elements.forEach(el => {
// Pixel coordinates
console.log(`Element "${el.content}":`);
console.log(` bbox: (${el.bbox.x0}, ${el.bbox.y0}) to (${el.bbox.x1}, ${el.bbox.y1})`);
console.log(` size: ${el.boundingBox.width}x${el.boundingBox.height}`);
console.log(` position: left=${el.boundingBox.left}, top=${el.boundingBox.top}`);
});
const result = await testdriver.parse();
// The annotated image shows all detected elements with bounding boxes
console.log('Annotated screenshot:', result.annotatedImageUrl);
console.log(`Image dimensions: ${result.imageWidth}x${result.imageHeight}`);
```javascript
// Prefer this for clicking a specific element
await testdriver.find("Submit button").click();
// Use parse() for full UI analysis
const result = await testdriver.parse();
const allButtons = result.elements.filter(e => e.type === 'button');
```
</Accordion>
<Accordion title="Filter by interactivity">
Use the `interactivity` field to distinguish between clickable and non-interactive elements.
```javascript
const result = await testdriver.parse();
const interactive = result.elements.filter(e => e.interactivity === 'clickable');
const static_ = result.elements.filter(e => e.interactivity === 'non-interactive');
```
</Accordion>
<Accordion title="Wait for content to load">
If elements aren't being detected, the page may not be fully loaded. Add a wait first.
```javascript
// Wait for page to stabilize
await testdriver.wait(2000);
// Then parse
const result = await testdriver.parse();
```
</Accordion>
<Accordion title="Use the annotated image for debugging">
The `annotatedImageUrl` provides a visual overlay showing all detected elements with their bounding boxes — great for debugging.
```javascript
const result = await testdriver.parse();
console.log('View annotated screenshot:', result.annotatedImageUrl);
```
</Accordion>
</AccordionGroup>
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
日本語の概要は準備中です。原文の説明を表示しています。
Execute natural language tasks using AI
日本語の概要は準備中です。原文の説明を表示しています。
Make AI-powered assertions about screen state
日本語の概要は準備中です。原文の説明を表示しています。
Deploy TestDriver on your AWS infrastructure using CloudFormation
日本語の概要は準備中です。原文の説明を表示しています。
Speed up tests with screenshot-based caching
日本語の概要は準備中です。原文の説明を表示しています。
How TestDriver learns your app and caches what it discovers for instant, deterministic replays
日本語の概要は準備中です。原文の説明を表示しています。