BLOG
Agent-Browser CLI: Run AI Browser Agents Remotely
The agent-browser CLI drives hosted Chromium for AI agents. Learn setup, schema, CDP wiring, and when to move from local to remote runtime.
# Agent-Browser CLI: Run AI Browser Agents Remotely
The agent-browser cli is a command-line interface that lets AI agents drive a real browser through a small, scriptable surface. Instead of embedding a full automation framework in your agent loop, you issue commands—navigate, click, type, extract—and the CLI translates them into browser actions. That design is convenient for local prototyping, but it raises a production question fast: where does the browser actually run? This guide covers what the agent-browser cli does, how its schema and skills work, how to wire it to a hosted Chromium runtime over CDP, and the criteria that decide when local execution stops being viable.
If you already know you need a remote runtime, you can skip ahead to the connection section. If you are still evaluating, read the trade-offs first—they explain why most teams eventually separate the CLI from the browser process.
What the agent-browser CLI actually is
The agent-browser CLI is a thin control layer. It exposes browser operations as commands an agent (or a shell script, or a CI job) can call in sequence. Typical operations map to the primitives every automation stack needs:
- Navigation — open a URL, wait for load state, go back or forward.
- Interaction — click, type, select, scroll, hover.
- Extraction — read text, attributes, or structured page data.
- Session control — start, stop, and inspect a browser session.
The value is the interface, not the browser. The CLI does not ship a rendering engine; it connects to one. That distinction matters because it means the same command surface can point at a local Chrome install during development and a hosted Chromium session in production, with no rewrite of your agent logic.
This is also why the CLI shows up in agent frameworks as a *skill* or *tool*: it gives the model a bounded, predictable set of verbs instead of raw DOM access.
The agent-browser schema and skills model
Most CLI-driven agent setups define a schema—a structured description of each command, its parameters, and its expected output. The schema is what the model reads to decide which action to take. A well-formed schema keeps the agent honest: it cannot invent a download_pdf command that does not exist.
Skills are the higher-level layer. A skill bundles several CLI calls into one intent, such as "log in and reach the dashboard." In practice:
- Schema = the low-level contract (one command, typed inputs, typed outputs).
- Skills = composed workflows built from schema-valid commands.
If you are evaluating an agent-browser setup, check the schema first. A vague schema produces vague agent behavior, and no amount of prompt engineering fixes an ambiguous tool definition.
Local CLI vs. hosted runtime: the real decision
The CLI runs fine on a laptop. The problems start when the workload does not fit on a laptop. Here is the comparison that matters for production planning.
| Dimension | Local CLI + local Chrome | Agent-browser CLI + hosted Chromium |
|---|---|---|
| Scaling | One machine, manual parallelism | Sessions provisioned per task via API |
| Session persistence | Tied to the local process | Persistent profiles survive restarts |
| Environment drift | Chrome version, OS, fonts vary per dev | Consistent hosted image |
| Network/IP | Your office or home IP | Configurable proxy settings |
| Debugging | Attach to local browser | Live viewer + CDP access |
| CI/CD fit | Fragile, needs a display server | Headless by default, no display needed |
| Cost model | Your hardware and time | Metered per browser-hour (see /pricing) |
The pattern is familiar from CI: local is great until you need reproducibility and concurrency. At that point the browser becomes infrastructure, and infrastructure belongs on a server.
Connecting the CLI to a remote browser over CDP
The Chrome DevTools Protocol (CDP) is the standard wire format for driving Chromium remotely. Playwright, Puppeteer, and Selenium all speak it, and a hosted runtime exposes a CDP endpoint you connect to instead of launching a local browser. The official Playwright CDP documentation is the authoritative reference for connectOverCDP.
Here is a minimal TypeScript example that connects to a hosted session and runs a task. The endpoint comes from your runtime provider; treat it as a secret.
import { chromium, Browser, Page } from 'playwright';
interface SessionInfo {
cdpUrl: string;
}
async function runTask(session: SessionInfo): Promise<string> {
const browser: Browser = await chromium.connectOverCDP(session.cdpUrl);
// Reuse the default context so persistent profile state is available.
const context = browser.contexts()[0] ?? (await browser.newContext());
const page: Page = context.pages()[0] ?? (await context.newPage());
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.getByRole('link', { name: /docs/i }).click();
await page.waitForLoadState('networkidle');
const heading = await page.locator('h1').first().innerText();
// Do not close the browser if the session is reused across tasks.
await browser.close();
return heading;
}
runTask({ cdpUrl: process.env.CDP_URL! }).catch((err) => {
console.error('task failed', err);
process.exit(1);
});Two details are easy to get wrong. First, connectOverCDP attaches to an *existing* browser—it does not launch one, so there is no launchOptions to tune. Second, if your runtime reuses sessions, closing the browser tears down state you may want to keep. Decide per workflow whether the CLI owns the session lifecycle or the runtime does.
For a deeper walkthrough of the connection model, see Remote Browser for AI agents.
Production criteria to evaluate before you commit
Not every CLI-plus-runtime combination is production-ready. These are the criteria that separate a demo from something you can run on a schedule.
Session isolation. Each task should get its own browser context so cookies and storage do not leak between jobs. Shared state is a correctness bug waiting to happen.
Persistent profiles. Some workflows need to stay logged in across runs. Persistent profiles let you resume a session without re-authenticating every time. This is a runtime feature, not a CLI feature—verify it exists before you depend on it.
Proxy and network controls. IP reputation affects success rates on protected sites. A hosted runtime should let you configure proxy settings per session rather than hardcoding them.
Live debugging. When an agent fails at step seven, you need to see the page. A live viewer plus CDP access means you can attach DevTools and inspect the actual DOM state, not a screenshot from three seconds earlier.
Usage controls. Metered browser-hours are the norm. Understand how sessions are billed and how to cap runaway jobs. Current rates and limits are on the pricing page.
Framework compatibility. If your stack is Playwright today and Puppeteer tomorrow, a CDP-based runtime keeps both options open. Avoid runtimes that only expose a proprietary SDK with no escape hatch.
Where the CLI fits in a larger agent stack
The CLI is one layer in a stack that usually looks like this:
- Model / planner — decides what to do next.
- Agent framework — manages the loop, memory, and tool calls.
- Agent-browser CLI — translates intent into browser commands.
- Runtime — hosts Chromium, manages sessions, profiles, and proxies.
- Target site — the actual web application.
Failures cluster at the boundaries. The model-to-CLI boundary fails when the schema is ambiguous. The CLI-to-runtime boundary fails when sessions drop or endpoints expire. The runtime-to-site boundary fails on bot detection or network issues. Debugging means knowing which boundary broke—which is another argument for a runtime that gives you visibility into the browser itself.
If you are running this in the cloud, Remote Browser online covers the provisioning model, and remote web browser explains the access patterns for teams that need shared sessions.
Common mistakes when adopting a CLI-first approach
Treating the CLI as the runtime. The CLI is a control surface. If you assume it also handles scaling, persistence, and isolation, you will rebuild those features badly.
Ignoring the schema. A loose schema lets the agent call commands with wrong arguments. Tighten types and validate inputs before they reach the browser.
Hardcoding endpoints. CDP URLs are credentials. Rotate them, store them in a secret manager, and never commit them.
Skipping the viewer. Without live inspection, every failure becomes a log-reading exercise. The viewer turns a 30-minute debug into a 2-minute one.
Assuming no concurrency limits. Every runtime has a maximum number of concurrent sessions. Check yours before you design a fan-out job that assumes many parallel sessions.
When to move off local
The signal is usually one of these:
- You need more than a handful of concurrent sessions.
- Your CI pipeline cannot reliably run a browser.
- Sessions need to survive process restarts.
- You are hitting bot detection from a single IP.
- You cannot reproduce a failure because the local environment differs from production.
Any one of these is enough. The migration itself is small if you built against CDP from the start—you change the endpoint, not the agent logic. Start with the documentation to see the connection flow, then wire your CLI to a hosted session and run one real task end to end before scaling.
The agent-browser CLI is a good interface. It is not a runtime. Treat those as two separate decisions and you will avoid most of the pain that shows up when a prototype meets production traffic.