BLOG
Browser Use Production Browser Automation
Learn how browser use production browser automation works: hosted Chromium, CDP, persistent profiles, and the criteria that separate demos from reliable agents.
# Browser Use Production Browser Automation
Browser use production browser automation is the practice of running an AI agent's browser session as managed infrastructure rather than a Chrome window on someone's laptop. In a demo, that distinction barely matters. In production, it decides whether your agent finishes tasks at 3 a.m. when a container restarts, a login expires, or a page layout shifts. This guide covers what changes when browser use moves from a local script to a hosted runtime, the criteria that matter, and how to wire it up with Playwright and CDP.
If you already understand the basics of agent-driven browsing, the short version is this: the browser becomes a service with a connection URL, a lifecycle, and an isolation boundary. Your agent code stays the same. Everything around it gets stricter.
What "production" actually changes
A local browser-use loop and a production one run the same actions: navigate, read the DOM, click, type, extract. The difference is everything the local version silently assumes.
- The browser is always there. A local Chrome depends on a machine that can sleep, reboot, or run out of memory. A hosted session is created on demand and torn down on a schedule you control.
- State survives. Cookies, localStorage, and logged-in sessions need to persist across runs. That requires persistent profiles, not a fresh incognito context every time.
- Sessions are isolated. Two agents must not share cookies, tabs, or credentials. Isolation is a runtime property, not a coding convention.
- Failures are observable. When a task fails, you need the DOM at failure time, a screenshot, console logs, and network activity — not just a stack trace.
- Cost is metered. Browser time is a real resource. Production systems track it per task so you can attribute spend to a workflow.
None of these are exotic requirements. They are the baseline that separates an agent you can demo from an agent you can operate. The production runtime overview covers the architecture in more depth; this post focuses on the browser-use-specific decisions.
The runtime layer browser-use needs
Browser-use libraries give you the agent loop: a model decides on an action, the library executes it against a browser, and the result feeds back into the next step. That loop is model-agnostic and, in practice, runtime-agnostic too — as long as the runtime exposes a standard automation interface.
The runtime layer is responsible for:
- Session provisioning. Starting a Chromium instance with the right flags, viewport, locale, timezone, and proxy configuration.
- Connection exposure. Handing your agent a CDP endpoint or a WebSocket URL it can attach to.
- Profile persistence. Mounting a storage volume so cookies and site data survive across sessions.
- Isolation. Guaranteeing that one session's storage, network identity, and tabs never leak into another.
- Observability. Streaming a live view and recording artifacts for post-hoc debugging.
- Usage control. Metering browser time and enforcing concurrency limits per account.
When you run browser-use locally, you are implementing all six of these yourself, usually by accident. A hosted runtime makes them explicit and configurable. That is the entire value proposition — not magic, just fewer undocumented assumptions.
Local Chrome vs hosted Chromium for browser use
The comparison below is about operational characteristics, not raw capability. Both run real Chromium and speak CDP.
| Dimension | Local Chrome / self-managed | Hosted Chromium runtime |
|---|---|---|
| Session startup | Manual launch, machine-dependent | API call, deterministic |
| Profile persistence | You manage user-data-dir | Persistent profiles per session |
| Isolation | Process-level, easy to leak | Session-scoped by default |
| Scaling | Add machines, add ops | Request more sessions |
| Debugging | Local DevTools only | Live viewer + recorded artifacts |
| Network identity | Whatever your IP is | Configurable proxy settings |
| Cost model | Fixed infra + your time | Metered browser time |
| Failure recovery | Restart the script | Reconnect or recreate session |
The local column is not wrong — it is just expensive in a way that does not show up on an invoice. Every hour spent debugging a stale Chrome process is an hour not spent on the agent's actual task logic.
Connecting browser-use to a hosted session
The connection path is the same one Playwright and Puppeteer already use: Chrome DevTools Protocol over a WebSocket endpoint. If your browser-use stack is built on Playwright, you attach with connectOverCDP. The Playwright CDP documentation is the authoritative reference for the client side.
Here is a TypeScript example that connects to a hosted session, reuses a persistent context, and runs a task with basic failure capture:
import { chromium, Browser, Page } from 'playwright';
interface SessionInfo {
cdpUrl: string; // wss://... from your runtime's session API
sessionId: string;
}
async function runTask(session: SessionInfo, task: (page: Page) => Promise<void>) {
let browser: Browser | undefined;
try {
// Attach to the hosted Chromium instance over CDP.
browser = await chromium.connectOverCDP(session.cdpUrl, {
timeout: 30_000,
});
// Reuse the existing context so persistent profile state is available.
const context = browser.contexts()[0] ?? (await browser.newContext());
const page = context.pages()[0] ?? (await context.newPage());
page.setDefaultTimeout(20_000);
await task(page);
} catch (err) {
// Capture artifacts before the session is torn down.
const pages = browser?.contexts().flatMap((c) => c.pages()) ?? [];
for (const p of pages) {
await p.screenshot({ path: `failure-${session.sessionId}.png` }).catch(() => {});
}
throw err;
} finally {
// Disconnect without killing the remote browser; the runtime owns lifecycle.
await browser?.close();
}
}Two details matter here. First, browser.close() on a CDP connection disconnects the client — it does not necessarily terminate the remote browser. Session lifecycle belongs to the runtime, which is what you want: a crashed agent process should not orphan or destroy a session mid-task. Second, capturing artifacts in the catch block is not optional in production. Without a screenshot and the page state at failure, you are debugging blind.
For a broader look at how connection URLs and session APIs fit together, see the remote browser online guide.
Production criteria for a browser-use runtime
When you evaluate a runtime for browser-use workloads, score it against these criteria rather than feature lists.
Session lifecycle control
Can you create, inspect, and terminate sessions programmatically? Can you set a maximum lifetime so a runaway agent does not burn browser hours indefinitely? A runtime without explicit lifecycle control will eventually cost you money in a way you cannot explain.
Persistent profiles
Browser-use tasks frequently involve authenticated flows. Persistent profiles let a session resume with cookies and storage intact. Check how profiles are scoped — per user, per workflow, per tenant — and whether they can be reset cleanly. The session and profile API overview is the place to confirm what your runtime exposes.
Isolation guarantees
Ask directly: if two sessions run concurrently, can they observe each other's storage or network identity? The answer should be a structural no, not a "we recommend separate profiles." Isolation that depends on developer discipline is isolation you will eventually lose.
Network configuration
Many production tasks need specific egress behavior: a particular region, a stable IP, or configurable browser settings for sites that treat automation differently. Look for configurable proxy and browser settings rather than promises about detection evasion. What you can verify — region, IP stability, header behavior — is what you should build on.
Observability
A live viewer is useful during development. Recorded artifacts — screenshots, DOM snapshots, console and network logs — are what you need at 2 a.m. when a scheduled run fails. Confirm both exist before committing.
Cost transparency
Browser time is metered. You should be able to see per-session duration and attribute it to a task or tenant. Current rates and plan details live on the pricing page; the important production question is whether usage is visible at the granularity you need for chargeback.
Common failure modes and how to avoid them
Most browser-use production incidents fall into a small number of categories.
Stale sessions. An agent crashes, the session lingers, and the next run attaches to a browser in an unknown state. Fix: always create a fresh session per task, or explicitly reset state at the start. Never assume the previous run left things clean.
Expired authentication. A persistent profile outlives the site's session cookie. Fix: detect login walls early in the task and fail fast with a clear error rather than letting the model improvise.
Timeout cascades. One slow page load consumes the task budget, and every subsequent step fails. Fix: set per-action timeouts and a hard task deadline. The example above sets a 20-second default; tune it per site.
Silent selector drift. A layout change breaks a selector, and the agent clicks the wrong element. Fix: prefer role- and text-based locators, and assert on expected page state after navigation.
Unbounded concurrency. A queue backs up and the system spawns sessions faster than the runtime allows. Fix: treat concurrency as a configured limit and queue work behind it. Do not assume there is no ceiling on parallel sessions.
Missing artifacts. The task fails, the session is destroyed, and there is nothing to inspect. Fix: capture screenshots and logs in the failure path, as shown earlier.
Each of these is cheap to prevent and expensive to diagnose. The pattern is consistent: production browser-use is mostly about making failure states explicit.
Where hosted runtimes fit — and where they do not
A hosted Chromium runtime is the right choice when you need isolation, persistence, observability, and metered cost without building a browser fleet. It is a poor fit when your task is a one-off local scrape, when you need a browser extension loaded into a personal profile, or when your compliance model forbids third-party infrastructure entirely.
For the middle ground — teams running browser-use agents on a schedule, across tenants, with real users waiting — the runtime layer is usually the difference between an agent that works and an agent that works reliably. The remote web browser guide walks through the deployment patterns, and the remote control browser overview covers interactive debugging when you need to drive a session by hand.
A practical adoption path
If you are moving browser-use from local to production, sequence it like this:
- Instrument first. Add per-task timing and artifact capture to your local runs. You will need this data later.
- Externalize the connection. Replace local browser launch with a CDP connection URL. Keep the agent loop unchanged.
- Add profile persistence. Move from fresh contexts to persistent profiles for authenticated flows.
- Set limits. Define max session lifetime, per-action timeouts, and a concurrency ceiling.
- Meter. Track browser time per task and review it weekly before you scale.
- Harden the failure path. Screenshots, logs, and a clear error taxonomy so on-call has something to act on.
Each step is independently useful, and none require rewriting your agent. That is the point of treating the browser as a runtime: the agent logic stays yours, and the operational concerns move somewhere they can be managed.
Production browser automation with browser-use is not a different kind of automation. It is the same automation with the assumptions written down, the failure paths instrumented, and the browser treated as infrastructure. Start with the connection layer, measure everything, and let the runtime handle the parts that were never really your agent's job.