BLOG
Vercel-Labs Agent-Browser: CLI, Schema, and Production Path
Vercel-Labs agent-browser is a CLI for AI browser agents. Here is what it does, its schema, and how to run it against hosted Chromium in production.
# Vercel-Labs Agent-Browser: CLI, Schema, and Production Path
Vercel-Labs agent-browser is a browser automation CLI for AI agents. It gives an agent a small, scriptable surface for driving a browser — navigate, click, type, extract — without the agent author writing raw Playwright or Puppeteer for every step. The interesting question for production teams is not whether the CLI works. It is where the browser actually runs, how sessions persist, and what happens when you need ten of them at once. This guide covers the CLI, its schema, and the path from a local demo to a hosted runtime.
What Vercel-Labs Agent-Browser Actually Is
The project lives at vercel-labs/agent-browser on GitHub and ships as an npm package. It is a command-line tool, not a hosted service. You install it, point it at a browser, and issue commands that map to browser actions. The design goal is to make browser control legible to an LLM: commands are discrete, outputs are structured, and the agent does not need to hold a full DOM in context.
That framing matters. A general-purpose automation library like Playwright is built for deterministic test code written by humans. An agent-browser CLI is built for a loop where a model decides the next action from the current state. The two overlap heavily — both ultimately speak CDP to a Chromium instance — but the ergonomics differ.
If you have read our agent-browser CLI setup guide, the short version is: the CLI is the control surface, and the browser is a separate concern you can host anywhere.
The CLI Surface
The CLI exposes a set of verbs an agent can call. Exact flags change between releases, so treat this as the shape rather than a frozen spec:
- Navigation — open a URL, go back, reload.
- Interaction — click an element, type text, press keys, select options.
- Inspection — read the accessibility tree, get element text, screenshot the viewport.
- Waiting — wait for a selector, wait for network idle, wait for a navigation.
- Session — start, list, and stop browser sessions.
The output is the part that matters for agents. Rather than returning a raw DOM dump, the CLI tends to return a trimmed representation — often accessibility-tree-shaped — so the model gets the interactive elements without thousands of tokens of markup. That is the same principle behind most modern browser-use tooling, and it is why these CLIs exist at all.
Why a CLI Instead of an SDK
A CLI is language-agnostic. An agent written in Python, TypeScript, Go, or a shell script can shell out to the same binary. That is a real advantage when your agent framework is not the same language as your automation library. The trade-off is process overhead and the awkwardness of parsing stdout for structured data. For high-frequency loops, an SDK or a direct CDP connection is usually cleaner.
The Agent-Browser Schema
The "schema" in agent-browser refers to the structured contract between the agent and the browser: what actions exist, what parameters they take, and what the result looks like. A well-defined schema is what lets a model plan reliably instead of guessing at free-form commands.
A typical action schema looks like this in spirit:
{
"action": "click",
"target": { "role": "button", "name": "Sign in" },
"result": { "ok": true, "url": "https://example.com/dashboard" }
}Two design choices are worth noting. First, targeting by role and accessible name is more robust than CSS selectors, because it survives markup changes that would break a brittle selector. Second, the result includes the post-action state, so the agent does not need a separate observation call after every step.
When you evaluate any agent-browser schema, ask:
- Is the action set closed? A fixed vocabulary is easier for a model to use correctly than an open-ended one.
- Are results deterministic? If the same action returns different shapes, your agent's parsing logic will drift.
- Does it expose enough state? Too little and the agent flails; too much and you burn context window.
Running the CLI Against a Hosted Browser
Here is the production question. The CLI needs a browser to drive. By default it will try to launch a local Chromium. That works on a laptop and falls apart in a container, a CI runner, or a serverless function — no display, no persistent disk, no stable IP.
The fix is to point the CLI at a remote browser over CDP. Remote Browser provides hosted Chromium sessions with a CDP endpoint, so the CLI connects instead of launching. The same pattern applies whether you use the CLI, Playwright, or Puppeteer.
import { chromium } from "playwright";
// Remote Browser exposes a CDP endpoint per session.
// Get the current endpoint from your session API or dashboard.
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP_URL!;
async function runAgentStep() {
const browser = await chromium.connectOverCDP(CDP_ENDPOINT);
// Reuse the existing context so cookies and storage persist
// across agent steps within the same session.
const context = browser.contexts()[0] ?? (await browser.newContext());
const page = context.pages()[0] ?? (await context.newPage());
await page.goto("https://example.com/login", {
waitUntil: "domcontentloaded",
});
// Target by accessible role, not brittle CSS.
await page.getByRole("textbox", { name: "Email" }).fill("agent@example.com");
await page.getByRole("button", { name: "Continue" }).click();
await page.waitForLoadState("networkidle");
const title = await page.title();
console.log("Landed on:", title);
// Do NOT call browser.close() here — that tears down the
// remote session. Disconnect instead.
await browser.close();
}
runAgentStep().catch(console.error);Two details trip people up. First, connectOverCDP attaches to an existing browser; it does not launch one. Second, calling browser.close() on a CDP connection can terminate the remote session depending on how the endpoint is managed — check your provider's semantics and prefer disconnecting when you intend to resume later. The Playwright CDP documentation covers the connection contract in detail.
Local CLI vs Hosted Runtime
The decision is not about which tool is better. It is about which environment the browser lives in.
| Criterion | Local Chromium + CLI | Hosted Chromium + CLI |
|---|---|---|
| Setup | Install browser, manage versions | Connect to a CDP endpoint |
| Scaling | One machine, manual parallelism | Session-per-task, managed |
| Persistence | Local profile dir, lost on rebuild | Persistent profiles across sessions |
| IP / geo | Your machine's IP | Configurable proxy settings |
| Debugging | Local DevTools | Live viewer + CDP |
| CI / serverless | Fragile, often blocked | Designed for headless infra |
| Cost model | Your compute | Metered per browser-hour |
For a demo, local is fine. For anything that runs on a schedule, in a container, or against sites that care about IP reputation, the hosted path removes a category of failure you would otherwise debug at 2 a.m.
If you want the longer version of that argument, see remote browser for AI agents.
Production Criteria Before You Commit
Before you standardize on any agent-browser CLI, check these against your workload:
Session lifecycle. Can you start a session, disconnect, and reconnect to the same browser state? Agents that run multi-step tasks across minutes or hours need this. A CLI that only supports launch-and-die is a demo tool.
Profile persistence. Login flows, MFA, and cart state all depend on cookies surviving between runs. Persistent profiles are the difference between an agent that logs in once and one that fights a login wall every task.
Concurrency controls. You need to know how many sessions you can run and how they are isolated. Session isolation prevents one agent's cookies from leaking into another's.
Observability. When an agent fails, you need to see what it saw. A live viewer and session recordings turn a mystery into a diff.
Proxy and browser settings. Sites treat datacenter IPs differently from residential ones. Configurable browser settings and proxy options let you match the environment to the target without rewriting your agent.
Usage controls. Metered browser-hours are predictable; runaway loops are not. Set limits before you scale.
Remote Browser exposes these as runtime primitives — hosted Chromium, CDP access, persistent profiles, a live viewer, and configurable browser and proxy settings. Current limits and pricing are on the pricing page; check there rather than assuming a number.
Where the CLI Fits in a Larger Stack
The CLI is one layer. A production agent stack usually looks like this:
- Agent framework — decides what to do next (your LLM loop).
- Control surface — the CLI or SDK that translates decisions into browser actions.
- Browser runtime — the actual Chromium instance, hosted or local.
- Session store — profiles, cookies, and state that persist across runs.
- Observability — viewer, logs, recordings.
The CLI handles layer two. It does not solve layers three through five, and no CLI can. That is why teams that start with a local CLI eventually separate the control surface from the runtime. The remote web browser guide walks through that separation in more detail.
Practical Migration Path
If you are already using the CLI locally, the migration is small:
- Keep the CLI. Do not rewrite your agent. The command surface stays the same.
- Swap the browser target. Replace local launch with a CDP endpoint from a hosted session.
- Move profile state. Point persistent profiles at the hosted runtime so logins survive.
- Add observability. Wire the live viewer into your debugging flow.
- Set usage limits. Cap browser-hours per task before you scale out.
Steps one and two are an afternoon. Steps three through five are where the reliability actually comes from.
Common Mistakes
Treating the CLI as the runtime. The CLI is a client. If you conflate the two, you will try to solve persistence and scaling problems in the wrong layer.
Closing the browser between steps. browser.close() on a CDP connection can end the session. Disconnect instead if you plan to resume.
Ignoring IP reputation. A perfectly written agent will still fail against a site that blocks your datacenter IP. Proxy configuration is not optional at scale.
Skipping the viewer. Debugging an agent without seeing the page is guesswork. Turn on the live viewer early.
Assuming the schema is stable. Pin your CLI version and test after upgrades. Action vocabularies evolve.
Summary
Vercel-Labs agent-browser gives AI agents a clean, scriptable CLI for browser control, with a schema designed around accessibility-tree targeting rather than brittle selectors. It is a good control surface. It is not a runtime. The production decision is where the Chromium instance lives and how sessions persist, scale, and stay observable. Point the CLI at hosted Chromium over CDP, keep profiles persistent, and set usage limits — and the same CLI that ran a demo on your laptop will hold up under real traffic. Start with the documentation to wire the connection, and check pricing for current session limits.