How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
日本語の概要は準備中です。原文の説明を表示しています。
Build TestDriver tests iteratively using MCP tools with visual feedback
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Build automated tests by directly controlling a sandbox through MCP tools. Every action returns a screenshot AND the generated code to add to your test file.
Use this skill when:
session_start, find, click, etc.)Use MCP tools to:
session_start({ type: "chrome", url: "https://your-app.com" })
This provisions a sandbox with Chrome and navigates to your URL. You'll see a screenshot and the provision code:
Add to test file:
await testdriver.provision.chrome({ url: "https://your-app.com" });
For local development (pointing to a custom API endpoint):
session_start({
type: "chrome",
url: "https://your-app.com",
apiRoot: "https://your-ngrok-url.ngrok.io"
})
For self-hosted AWS instances (your own Windows EC2):
session_start({
type: "chrome",
url: "https://your-app.com",
os: "windows",
ip: "1.2.3.4" // IP from your AWS instance
})
See AWS Setup Guide to deploy your own infrastructure.
Find elements and interact with them. Each action returns a screenshot AND generated code:
find_and_click({ description: "Sign In button" })
→ Returns: screenshot with element highlighted
→ Add to test file: await testdriver.find("Sign In button").click();
type({ text: "user@example.com" })
→ Returns: screenshot showing typed text
→ Add to test file: await testdriver.type("user@example.com");
After performing actions, use check to verify they worked:
check({ task: "Was the text entered into the field?" })
→ Returns: AI analysis of whether the task completed, with screenshot
check({ task: "Did the button click navigate to a new page?" })
→ Returns: AI compares previous screenshot to current state
Use assert for boolean pass/fail conditions that get recorded in test files:
assert({ assertion: "the login form is visible" })
→ Returns: pass/fail with screenshot
→ Add to test file:
const assertResult = await testdriver.assert("the login form is visible");
expect(assertResult).toBeTruthy();
As you perform actions, append the generated code to your test file:
/**
* Login Flow test
*/
import { describe, expect, it } from "vitest";
import { TestDriver } from "testdriverai/lib/vitest/hooks.mjs";
describe("Login Flow", () => {
it("should complete login", async (context) => {
const testdriver = TestDriver(context);
// Append generated code here as you go:
await testdriver.provision.chrome({ url: "https://app.example.com" });
await testdriver.find("email input field").click();
await testdriver.type("user@example.com");
// ... more code as you perform actions
});
});
Run the test from scratch to validate it works:
verify({ testFile: "tests/login.test.mjs" })
| Tool | Description |
|---|---|
session_start | Start sandbox with browser/app, returns screenshot + provision code |
session_status | Check session health and time remaining |
session_extend | Add more time before session expires |
Each tool returns a screenshot AND the generated code to add to your test file.
| Tool | Description |
|---|---|
find | Locate element by description, returns ref for later use |
click | Click on element ref |
find_and_click | Find and click in one action |
type | Type text into focused field |
press_keys | Press keyboard shortcuts (e.g., ["ctrl", "a"]) |
scroll | Scroll page (up/down/left/right) |
| Tool | Description |
|---|---|
check | For AI to understand screen state. Analyzes current screen and tells you (the AI) whether a task/condition is met. Use this after actions to verify they worked. |
assert | AI-powered boolean assertion for test files (pass/fail for CI). Returns generated code. |
screenshot | For showing the user the screen. Captures and displays a screenshot. Does NOT return analysis to you (the AI). |
exec | Execute JavaScript, shell, or PowerShell in sandbox. Returns generated code. |
| Tool | Description |
|---|---|
verify | Run test file from scratch to validate it works |
Every tool returns a screenshot showing:
Don't try to build the entire test at once:
# Step 1: Get to login page
session_start({ url: "https://app.com" })
→ Add to test: await testdriver.provision.chrome({ url: "https://app.com" });
# Step 2: Verify you're on the right page
check({ task: "Is this the login page?" })
# Step 3: Fill in email
find_and_click({ description: "email input field" })
→ Add to test: await testdriver.find("email input field").click();
type({ text: "user@example.com" })
→ Add to test: await testdriver.type("user@example.com");
# Step 4: Check if email was entered
check({ task: "Was the email entered correctly?" })
# Step 5: Continue with password...
After each action, use check to verify it worked:
find_and_click({ description: "Submit button" })
check({ task: "Was the form submitted?" })
The check tool compares the previous screenshot (from before your action) with the current state, giving you AI analysis of what changed and whether the action succeeded.
For AI understanding: Use check to analyze the screen:
check({ task: "Did the form submit successfully?" })
→ Returns AI analysis you can read and understand
For user visibility: Use screenshot to show the user:
screenshot()
→ Displays to user, no analysis returned to you
Action tools (find, click, find_and_click) return screenshots automatically, which the user can see. But if you need to understand the state, use check.
If elements take time to appear, use find with timeout:
find({ description: "Loading complete indicator", timeout: 30000 })
Sessions expire after 5 minutes by default. Use session_status to check time remaining and session_extend to add more time:
session_status()
→ "Time remaining: 45s"
session_extend({ additionalMs: 60000 })
→ "New expiry: 105s"
After each successful action, append the generated code to your test file. This ensures you don't lose progress and makes the test easier to debug.
If find fails:
find({ description: "...", timeout: 10000 })scroll({ direction: "down" })If the session expires:
session_start again with the same URLverify to get back to last stateIf verify fails:
When creating a new test project, use these exact dependencies:
package.json:
{
"type": "module",
"devDependencies": {
"testdriverai": "canary",
"vitest": "^4.0.0"
},
"scripts": {
"test": "vitest"
}
}
Important: The package is testdriverai (NOT @testdriverai/sdk). Always install from the canary tag.
Create test files using this standard format. Append generated code inside the test function:
/**
* Login Flow test
*/
import { describe, expect, it } from "vitest";
import { TestDriver } from "testdriverai/lib/vitest/hooks.mjs";
describe("Login Flow", () => {
it("should complete Login Flow", async (context) => {
const testdriver = TestDriver(context);
// Append generated code from each action here:
await testdriver.provision.chrome({ url: "https://app.example.com" });
await testdriver.find("email input field").click();
await testdriver.type("user@example.com");
await testdriver.find("password field").click();
await testdriver.type("secret123");
await testdriver.find("Sign In button").click();
const assertResult = await testdriver.assert("dashboard is visible");
expect(assertResult).toBeTruthy();
});
});
You can use your own AWS-hosted Windows instances instead of TestDriver cloud. This gives you:
Deploy AWS infrastructure using CloudFormation
Spawn an instance:
AWS_REGION=us-east-2 \
AMI_ID=ami-0504bf50fad62f312 \
AWS_LAUNCH_TEMPLATE_ID=lt-xxx \
bash setup/aws/spawn-runner.sh
Output: PUBLIC_IP=1.2.3.4
Connect via session_start:
session_start({
type: "chrome",
url: "https://example.com",
os: "windows",
ip: "1.2.3.4"
})
Terminate when done:
aws ec2 terminate-instances --instance-ids i-xxx --region us-east-2
You can also set TD_IP environment variable in your MCP config instead of passing ip to each session:
{
"mcpServers": {
"testdriver": {
"env": {
"TD_API_KEY": "your-key",
"TD_IP": "1.2.3.4"
}
}
}
}
Every action returns generated code - Look for "Add to test file:" in responses and append that code
Use check to understand the screen - This is how you (the AI) see and analyze the current state
Use screenshot to show the user - This displays the screen to the user, but does NOT return analysis to you
Use check after every action - Verify your actions succeeded before moving on
Be specific with element descriptions - "the blue Sign In button in the header" is better than "button"
Use check for verification, assert for test files - check gives detailed AI analysis, assert gives boolean pass/fail for CI
Write code incrementally - Append generated code to your test file after each successful action
Extend session proactively - If you have complex workflows, extend before you run out of time
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
日本語の概要は準備中です。原文の説明を表示しています。
Execute natural language tasks using AI
日本語の概要は準備中です。原文の説明を表示しています。
Make AI-powered assertions about screen state
日本語の概要は準備中です。原文の説明を表示しています。
Deploy TestDriver on your AWS infrastructure using CloudFormation
日本語の概要は準備中です。原文の説明を表示しています。
Speed up tests with screenshot-based caching
日本語の概要は準備中です。原文の説明を表示しています。
How TestDriver learns your app and caches what it discovers for instant, deterministic replays
日本語の概要は準備中です。原文の説明を表示しています。