← Blog

BLOG

Browser Agent Vercel: Running Agent-Browser on Hosted Chromium

Browser agent Vercel setups need a real runtime. Learn how vercel-labs/agent-browser works, where it breaks in production, and how to connect it to hosted Chromium.

October 8, 202610 min readRemote Browser

# Browser Agent Vercel: Running Agent-Browser on Hosted Chromium

A browser agent on Vercel usually means one specific thing: vercel-labs/agent-browser, the CLI that Vercel Labs published for driving Chrome from an AI agent. It is a thin, well-scoped tool — a command surface over Chrome DevTools Protocol (CDP) that lets a model click, type, navigate, and extract without writing Playwright glue by hand. What it is not is a runtime. The moment you try to run it inside a Vercel Function, on a build machine, or across a fleet of serverless invocations, you hit the same wall every browser agent hits: there is no Chrome where your code is running, and there is no session that survives the next cold start. This post covers what agent-browser actually does, why Vercel's execution model fights it, and how to pair the CLI with a hosted Chromium runtime so the browser agent on Vercel has somewhere real to run.

What vercel-labs/agent-browser actually is

The repository describes itself plainly: a browser automation CLI for AI agents. That framing matters. It is not a framework, not an orchestration layer, and not a browser. It is a command-line interface that an agent — or a shell script, or a tool-calling loop — invokes to perform discrete browser actions.

The design follows the pattern that has become standard for agent tooling: expose a small set of verbs, return structured output the model can reason about, and keep the transport layer out of the way. Underneath, it speaks CDP to a Chrome instance. That is the same protocol Playwright and Puppeteer use, which is why agent-browser composes cleanly with existing automation stacks rather than replacing them.

Two properties follow from this design, and both matter for deployment:

  • It needs a browser endpoint. The CLI does not ship Chromium. It connects to one. Where that browser lives is entirely your problem.
  • It is stateless by default. Each invocation is a process. Session continuity — cookies, localStorage, an authenticated tab — depends on the browser you point it at, not on the CLI.

If you are still forming a mental model of what these tools are, What Is a Browser Agent? covers the runtime-versus-tool distinction in more depth. The short version: the CLI is the hands, the browser is the environment, and the environment is where production problems live.

Why Vercel's execution model complicates browser agents

Vercel is excellent at what it is designed for: stateless HTTP request handling, edge routing, and fast deploys. Browser agents are the opposite workload. They are long-running, stateful, and resource-hungry. Four specific frictions show up.

No persistent Chrome process. Serverless functions spin up, handle a request, and freeze or terminate. A browser agent task — log in, navigate three pages, fill a form, verify a result — takes tens of seconds to minutes. You cannot hold a Chrome process across that boundary.

Bundle size and cold starts. Chromium is roughly 150–300 MB depending on build. Even if you could ship it, you would pay that cost on every cold start, and most function size limits will not accommodate it.

Session state evaporates. Cookies, session tokens, and in-page state live in the browser profile. When the function instance is recycled, that profile is gone. For agents that authenticate, this means re-login on every invocation — which is both slow and a detection signal.

Concurrency is bounded by memory. Each headless Chrome instance consumes a substantial amount of memory. Running several in one function is not viable, and Vercel's model does not give you a long-lived pool to schedule against.

None of this is a Vercel defect. It is a mismatch. The fix is not to fight the platform; it is to move the browser out of the function and keep only the orchestration there.

The split architecture: agent on Vercel, browser elsewhere

The pattern that works is straightforward: your Vercel deployment handles routing, auth, model calls, and tool dispatch. The browser lives in a hosted runtime that exposes a CDP endpoint. The agent-browser CLI — or the Playwright code it wraps — connects over the network.

Vercel Function  ──►  agent-browser CLI  ──►  CDP over WebSocket  ──►  Hosted Chromium
   (stateless)          (tool layer)              (transport)            (stateful session)

This gives you three things the all-in-function approach cannot:

  1. Session persistence. The hosted browser keeps its profile between invocations. Your agent resumes an authenticated session instead of rebuilding it.
  2. Bounded function cost. The function does orchestration and returns. It does not pay for browser memory or CPU.
  3. Observability. A hosted runtime typically exposes a live viewer, so you can watch what the agent is doing rather than inferring it from logs.

Remote Browser is built for exactly this shape: hosted Chromium sessions with CDP access, Playwright/Puppeteer/Selenium compatibility, persistent profiles, configurable browser and proxy settings, session isolation, and a live viewer. The remote browser for AI agents post walks through the runtime layer in detail.

Connecting agent-browser to a hosted CDP endpoint

The mechanics are the same whether you drive the CLI or write the connection yourself. You need a WebSocket CDP endpoint and, ideally, a session identifier so you can reconnect to the same browser.

Here is the Playwright path, which is the most portable and the one most teams end up using in production. It connects over CDP to a remote Chromium instance, reuses a persistent context, and runs a task:

import { chromium, Browser, BrowserContext, Page } from 'playwright';

interface RemoteSession {
  cdpUrl: string;      // wss://... from your runtime provider
  sessionId: string;   // stable ID so reconnects hit the same browser
}

async function runAgentTask(
  session: RemoteSession,
  task: (page: Page) => Promise<void>
): Promise<void> {
  // connectOverCDP attaches to an already-running browser.
  // No local Chromium download, no launch flags to tune.
  const browser: Browser = await chromium.connectOverCDP(session.cdpUrl, {
    timeout: 30_000,
  });

  // A persistent remote profile usually exposes one default context.
  // Reuse it so cookies and localStorage survive across invocations.
  const context: BrowserContext =
    browser.contexts()[0] ?? (await browser.newContext());

  const page: Page = context.pages()[0] ?? (await context.newPage());

  try {
    await task(page);
  } finally {
    // Close the CDP connection, not the remote browser.
    // The session stays alive for the next function invocation.
    await browser.close();
  }
}

// Inside a Vercel route handler:
export async function POST(req: Request) {
  const { cdpUrl, sessionId } = await getSessionForUser(req);

  await runAgentTask({ cdpUrl, sessionId }, async (page) => {
    await page.goto('https://example.com/dashboard', {
      waitUntil: 'domcontentloaded',
    });
    await page.getByRole('button', { name: 'Export' }).click();
    await page.waitForEvent('download');
  });

  return Response.json({ ok: true });
}

Two details are easy to get wrong. First, browser.close() on a CDP connection detaches the client; it does not necessarily terminate the remote browser. That is what you want — the session persists. Second, browser.contexts()[0] assumes the remote runtime exposes a default context. If your provider returns a fresh context per connection, you will lose profile state and should instead pass a persistent context identifier at session creation time.

For the CLI itself, the same endpoint is what you configure. The CLI's job is to issue CDP commands; where those commands land is a connection-string question. If you want the full picture of how CDP wiring works across tools, Agent-Browser CDP covers the transport layer, and the Chrome DevTools Protocol documentation is the authoritative reference for the domains you will actually call.

agent-browser vs browser-use vs a hosted runtime

These three get compared constantly, and the comparison is usually confused because they operate at different layers. Here is the honest breakdown.

Layeragent-browser (Vercel Labs)browser-useHosted runtime (Remote Browser)
What it isCLI for driving Chrome via CDPAgent framework for LLM-driven browsingBrowser infrastructure: sessions, profiles, proxies
Ships a browser?NoNoYes — hosted Chromium
State modelStateless processIn-process agent loopPersistent remote sessions
Runs on Vercel?Only as a clientOnly as a clientN/A — it is the remote side
Session persistenceDepends on endpointDepends on endpointBuilt in
Live debuggingNoLimitedLive viewer
Best fitTool-calling agents that need a small verb setTeams building the agent loop itselfProduction workloads needing a stable browser

The practical takeaway: agent-browser and browser-use are not competitors to a hosted runtime. They are clients of one. You can run browser-use's agent loop against a Remote Browser session, and you can point agent-browser's CLI at the same CDP endpoint. The runtime question — where does the browser live, how does it persist, how do you observe it — is orthogonal and, in production, more consequential.

If you are weighing the framework side specifically, Agent Browser vs Browser Use goes deeper on where each fits. And if you are comparing hosted providers, Browserbase vs Browser Use is a useful reference for the infrastructure trade-offs.

Production criteria for a Vercel-hosted browser agent

Once the architecture is split, the remaining decisions are about the runtime. These are the criteria that actually determine whether an agent survives contact with real sites.

Session persistence and reconnection. Can you reconnect to the same browser after a function cold start? Does the profile survive? This is the single biggest determinant of agent reliability for authenticated workflows.

CDP compatibility. Does the runtime expose a standard CDP WebSocket endpoint, or a proprietary API? Standard CDP means Playwright, Puppeteer, Selenium, and agent-browser all work without adapters.

Proxy and network configuration. Agents hit rate limits and geo-restrictions. You want configurable proxy settings at the session level, not per-request hacks. Note that "configurable browser settings" is the accurate framing here — specific anti-detection behavior varies by provider and should be verified against their documentation rather than assumed.

Observability. A live viewer is not a nice-to-have. When an agent fails at step seven of a twelve-step task, logs alone will not tell you why. Being able to watch the session — or replay it — collapses debugging time.

Isolation. Sessions must not leak state into each other. If you run multiple agents, each needs its own browser context and profile.

Usage controls. Browser time is metered. You want per-session visibility into consumption so a runaway agent does not quietly burn budget. Current rates and limits are on the pricing page.

Concurrency model. Do not assume a fixed ceiling of parallel sessions. Plan for a bounded pool and queue work against it. The remote browser online guide covers what a managed session looks like in practice.

When to keep the browser local instead

Hosted Chromium is not always the right answer, and it is worth being explicit about when it is not.

  • Short, unauthenticated, one-shot tasks. If the agent navigates to one public page and extracts text, a local headless browser in a container is fine.
  • Development and debugging. Running Chrome locally while you iterate on agent logic is faster than round-tripping to a remote session.
  • Strict data residency requirements where the browser must not leave your network. In that case, self-host a CDP endpoint and point agent-browser at it — the architecture is identical, only the endpoint changes.

The split-architecture pattern is portable. What you are buying from a hosted runtime is operational relief: no Chromium builds to maintain, no profile storage to design, no viewer to build, no proxy pool to manage. Whether that trade is worth it depends on how many sessions you run and how much of your team's time goes to browser infrastructure instead of agent logic.

Putting it together

A browser agent on Vercel is a viable architecture, but not by running Chrome inside a function. The working shape is: agent-browser (or Playwright, or browser-use) as the tool layer, a Vercel deployment as the orchestration layer, and a hosted Chromium runtime as the stateful browser layer connected over CDP.

Start by getting a CDP endpoint and a persistent session. Point your existing agent code at it. Verify that reconnection works across invocations — that is the test that separates a demo from something you can ship. Then tune proxies, isolation, and usage controls against real traffic.

The documentation covers session creation, CDP connection strings, and profile handling. If you want to see the runtime in action before wiring it into a Vercel deployment, the remote web browser walkthrough is the fastest path from zero to a live session.