← Blog

BLOG

Browser Use Model: Choosing a Runtime for AI Agents

A browser use model defines how AI agents drive the web. Compare local vs hosted runtimes, CDP wiring, and production criteria for remote Chromium.

October 3, 20269 min readRemote Browser

# Browser Use Model: Choosing a Runtime for AI Agents

A browser use model is the combination of a language model and the browser runtime it drives. The model decides what to click; the runtime decides whether that click lands on a real page, in a real session, with a real profile. Most teams spend weeks tuning the first half and almost no time on the second, then wonder why their agent works in a demo and fails in production. This guide covers what the runtime half actually involves, how local and hosted browser use models differ, and the criteria that separate a prototype from something you can run on a schedule.

Remote Browser is a hosted Chromium runtime for AI agents and browser-use workflows. It exposes sessions over CDP, works with Playwright, Puppeteer, and Selenium, and adds a live viewer, persistent profiles, configurable browser settings, and session isolation. You can read the connection details in the /documentation and current usage terms at /pricing.

What "browser use model" actually refers to

The term gets used loosely. In practice it describes a two-layer system:

  • The reasoning layer — an LLM (or a planner plus a smaller executor model) that receives an observation, decides on an action, and emits something like click("#submit") or type("input[name=q]", "query").
  • The execution layer — a browser process that receives those actions, renders pages, manages cookies and storage, handles navigation, and returns observations (DOM snapshots, screenshots, accessibility trees) back to the model.

The reasoning layer is where the interesting research lives. The execution layer is where production breaks. A model that scores well on a benchmark can still fail on a real site because the browser it controls has no session state, gets blocked at the network edge, or dies mid-task when the host machine sleeps.

When people compare browser use models, they usually mean the reasoning layer. When they ship, the execution layer is what they end up debugging.

Local browser use vs hosted browser use

The first architectural decision is where Chromium runs.

DimensionLocal browser (your machine or CI box)Hosted browser runtime
SetupInstall Chromium, match versions, manage driversConnect over CDP with a URL
Session persistenceTied to process lifetimeProfiles survive across sessions
ScalingOne process per worker; memory-boundSessions provisioned per request
Network identityYour IP or a proxy you configureConfigurable proxy and browser settings
DebuggingLocal DevTools, screenshots to diskLive viewer plus CDP access
Failure modeHost sleeps, OOM, version driftSession-level isolation and teardown
Cost modelCompute you already pay forMetered browser time — see /pricing

Local is the right default for development. You get fast iteration, full DevTools, and no network hop. The problems start when you move to concurrency. Each headless Chromium instance is a few hundred megabytes of RSS before the page even loads, and agent workloads are long-lived — a single task can hold a browser open for minutes. Ten parallel agents on one box is often the practical ceiling before you are tuning the OOM killer instead of the agent.

Hosted runtimes invert that trade-off. You give up local DevTools convenience and pay per unit of browser time, but you get isolation, a stable network identity, and the ability to run sessions that outlive any single worker process. For agent workloads that run on a schedule or in response to events, that matters more than raw latency.

If you want the longer version of this argument, /blog/remote-browser-for-ai-agents covers the runtime layer in depth, and /blog/remote-browser-online walks through running Chromium without a local install.

How the model talks to the browser

Almost every modern browser use model speaks CDP (Chrome DevTools Protocol) under the hood, either directly or through a framework. Playwright and Puppeteer both expose a connection path that lets you attach to an already-running browser instead of launching one.

The key primitive is connectOverCDP. Instead of chromium.launch(), you call chromium.connectOverCDP(endpoint) and drive the remote instance exactly as you would a local one. The Playwright CDP documentation covers the full API surface; the practical shape looks like this:

import { chromium, Browser, Page } from "playwright";

type SessionInfo = {
  cdpUrl: string;      // wss://... from your runtime provider
  sessionId: string;
};

async function runAgentTask(session: SessionInfo, task: string) {
  const browser: Browser = await chromium.connectOverCDP(session.cdpUrl, {
    timeout: 30_000,
  });

  // Reuse the existing context so profile state (cookies, storage) persists.
  const context = browser.contexts()[0] ?? (await browser.newContext());
  const page: Page = context.pages()[0] ?? (await context.newPage());

  page.setDefaultTimeout(20_000);

  await page.goto("https://example.com/login", { waitUntil: "domcontentloaded" });

  // Your model decides these actions; the runtime executes them.
  await page.fill("input[name=email]", process.env.AGENT_EMAIL!);
  await page.fill("input[name=password]", process.env.AGENT_PASSWORD!);
  await page.click("button[type=submit]");

  await page.waitForLoadState("networkidle");

  // Observation returned to the reasoning layer.
  const snapshot = await page.accessibility.snapshot();
  const url = page.url();

  // Do NOT close the browser if the session should persist.
  // await browser.close();

  return { snapshot, url, task };
}

Three details matter here and are easy to get wrong:

  1. Do not call `browser.close()` if you want the session to survive. Closing the CDP connection is fine; closing the browser tears down the context and loses profile state.
  2. Reuse the existing context. Calling newContext() on every task throws away cookies and storage, which is the most common cause of "the agent logged in yesterday but not today."
  3. Set explicit timeouts. Agent-driven navigation is unpredictable; a 30-second default will hang your worker on a slow page.

If you are wiring this up for the first time, /blog/remote-control-browser covers the connection patterns and failure modes in more detail.

Production criteria for a browser use model

Once the agent works locally, the question shifts from "does it work" to "does it keep working." These are the criteria that actually predict that.

Session isolation

Two agents running the same task should not share cookies, storage, or a network identity. Isolation is what makes concurrent runs reproducible. In a hosted runtime this is usually a per-session guarantee; in a local setup it is something you build with separate user-data directories and hope you got right.

Persistent profiles

Many real tasks require authentication. A profile that survives across sessions means the agent logs in once and reuses the session, rather than re-authenticating on every run. This is also where you need to think about credential handling — profiles store cookies, so treat them as secrets.

Network identity

Sites rate-limit, geo-restrict, and block. A runtime that lets you configure proxy settings and browser-level options gives you a lever when a target site starts rejecting requests. Be careful with claims here: "configurable browser settings" is accurate; blanket promises about evading detection are not, and no responsible provider makes them.

Observability

When an agent fails at step 7 of 12, you need to see what the page looked like. A live viewer plus CDP access means you can attach DevTools to a running session, or replay a screenshot sequence. Without that, debugging is guesswork.

Usage controls

Agent workloads are bursty. You want per-session visibility into browser time and a way to cap runaway tasks. Metered billing is the norm; check /pricing for how Remote Browser meters sessions rather than assuming a flat rate.

Where the model choice and the runtime choice interact

The reasoning layer and execution layer are not independent. A few interactions are worth knowing about:

  • Observation format drives model cost. Sending full DOM snapshots to a large model is expensive. Accessibility trees or pruned DOM are cheaper and often more reliable, but they depend on the runtime exposing them cleanly.
  • Latency budget is shared. If your model takes 8 seconds to decide and the browser takes 2 seconds to navigate, the browser is not your bottleneck. Optimizing the runtime first is often wasted effort.
  • Retry semantics belong to the runtime. A model that emits a bad action should get a clean error, not a corrupted session. Session isolation makes retries safe.
  • Model swaps are cheap; runtime swaps are not. If you have wired your agent to CDP, you can change the reasoning model in an afternoon. Changing the execution layer means re-testing every task.

That last point is the practical argument for investing in the runtime early. The model landscape moves fast; the browser protocol has been stable for years.

Common failure patterns

A short list of things that break browser use models in production, roughly in order of frequency:

  • Session state lost between runs. Cause: newContext() per task, or closing the browser instead of the connection.
  • Version drift. Cause: local Chromium updated by the OS or a package manager, breaking a selector or a CDP method. Hosted runtimes pin versions.
  • Resource exhaustion. Cause: too many concurrent local browsers on one host. Symptoms look like flaky tests but are actually OOM.
  • Network blocks. Cause: a datacenter IP hitting a site that expects residential traffic. Fixable with proxy configuration, not with prompt engineering.
  • Silent timeouts. Cause: no explicit timeout on navigation or action. The agent hangs, the worker is stuck, and the task never reports failure.

None of these are model problems. All of them are runtime problems, and all of them are cheaper to fix at the runtime layer than by adding retries to the agent loop.

Choosing a runtime: a short checklist

Before committing to a browser use model stack, confirm:

  • [ ] The runtime exposes CDP so you are not locked into one framework.
  • [ ] Sessions are isolated by default.
  • [ ] Profiles can persist across sessions, with a clear story for credential storage.
  • [ ] You can configure proxy and browser settings per session.
  • [ ] There is a live viewer or equivalent for debugging failed runs.
  • [ ] Usage is metered and visible, with a way to cap spend.
  • [ ] The provider documents version pinning so upgrades do not break selectors.

Remote Browser satisfies these through hosted Chromium sessions, CDP access, Playwright/Puppeteer/Selenium compatibility, persistent profiles, a live viewer, and per-session usage controls. The connection details are in /documentation; if you want to see what a session looks like end to end, /blog/remote-web-browser is a good starting point.

The short version

A browser use model is a reasoning layer plus an execution layer, and the execution layer is where production reliability is won or lost. Local Chromium is fine for development and fine for a handful of concurrent tasks. Past that, the constraints — session persistence, isolation, network identity, observability — push you toward a hosted runtime that speaks CDP.

Pick the runtime first, wire your agent to it over connectOverCDP, and keep the reasoning layer swappable. The model will change. The protocol will not.