BLOG
Simple Remote Browser: Hosted Chromium for AI Agents
A simple remote browser gives AI agents hosted Chromium over CDP. Learn how to connect Playwright, manage sessions, and run browser-use in production.
# Simple Remote Browser: Hosted Chromium for AI Agents
A simple remote browser is a hosted Chromium instance you connect to over a WebSocket endpoint instead of launching locally. You get a browser process running in the cloud, a CDP URL, and a session you can drive with Playwright, Puppeteer, or Selenium. That's the whole idea. The complexity lives in everything around it: session lifecycle, profile persistence, proxy configuration, and knowing when a task actually finished.
This guide covers what a simple remote browser is, how it differs from running Chrome on your own machine, and how to wire one into an agent workflow without rebuilding infrastructure every time your concurrency needs change.
What "simple" actually means here
The word simple is doing real work in that phrase. A remote browser is simple when the connection step is one line and the operational surface is small. It is not simple when you have to provision a VM, install Chromium, manage a display server, patch the browser monthly, and babysit zombie processes.
The practical definition: you receive a connection URL, you pass it to your automation library, and you start issuing commands. Everything else — the container, the browser binary, the network egress, the cleanup — is someone else's problem.
That matters most for AI agents, because agent workloads are bursty and unpredictable. A browser-use task might take four seconds or four minutes. It might need one page or forty. If you're running your own Chromium fleet, you're paying for idle capacity between bursts and scrambling during them.
Remote browser vs. local Chrome: the real trade-offs
Local Chrome is the default for a reason. It's free, it's fast to start, and you can see what's happening. For a script you run once, it's the right call.
The problems show up at the second and third iteration.
| Dimension | Local Chrome | Simple remote browser |
|---|---|---|
| Startup | Instant, but cold on fresh CI runners | Connection URL, no install step |
| Environment parity | Depends on your OS and Chrome version | Consistent hosted Chromium |
| Concurrency | Bounded by your machine's RAM | Scales with the runtime, subject to plan limits |
| Session persistence | Lost when the process exits | Persistent profiles available |
| Debugging | You watch the window | Live viewer over the session |
| Network identity | Your IP | Configurable proxy settings |
| Ops burden | You patch and monitor it | Managed by the provider |
| Cost model | Your compute, always on | Metered browser time — see /pricing |
The honest trade-off: you give up some control and you pay per unit of browser time. In exchange you stop treating browser infrastructure as a side project. For teams running agents in production, that exchange usually pays for itself the first time a Chrome update breaks a selector at 2 a.m.
If you want the fuller argument for why agents specifically need this layer, /blog/remote-browser-for-ai-agents covers it in depth.
Connecting with Playwright over CDP
The Chrome DevTools Protocol is the wire format. Playwright's connectOverCDP is the standard entry point. Here's a minimal TypeScript example that connects to a hosted session, opens a page, and extracts structured data:
import { chromium, Browser, Page } from 'playwright';
interface SessionInfo {
cdpUrl: string; // e.g. wss://<host>/cdp/<session-id>
}
async function runTask(session: SessionInfo): Promise<string> {
let browser: Browser | null = null;
try {
browser = await chromium.connectOverCDP(session.cdpUrl);
// Reuse the existing context when the runtime provides one.
const context = browser.contexts()[0] ?? await browser.newContext();
const page: Page = await context.newPage();
await page.goto('https://example.com/dashboard', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
// Wait on a real signal, not a fixed sleep.
await page.waitForSelector('[data-testid="account-balance"]', {
state: 'visible',
timeout: 15_000,
});
const balance = await page
.locator('[data-testid="account-balance"]')
.innerText();
return balance.trim();
} finally {
// Close the page, not the browser, if the session is reused.
if (browser) await browser.close();
}
}Three details matter more than the rest of that snippet.
Wait on state, not time. waitForSelector with a state option is deterministic. page.waitForTimeout(3000) is a guess that will eventually be wrong on a slow network.
Respect the session lifecycle. If your runtime hands you a persistent session, closing the browser may terminate it. Close pages and contexts instead, and let the session manager handle teardown.
Treat the CDP URL as a secret. It grants full control of the browser. Keep it out of logs and client-side code.
The Playwright documentation on connecting to browsers over CDP is worth reading before you build retry logic around it.
What a hosted runtime gives you beyond the connection
A connection URL is table stakes. The differences between providers show up in the operational features around it.
Session isolation. Each agent task should get its own browser context at minimum, and its own browser process when the workload is sensitive. Shared state between concurrent tasks is a correctness bug waiting to happen — cookies, localStorage, and service workers all leak across contexts if you're careless.
Persistent profiles. Some workflows need to stay logged in across runs. A persistent profile stores cookies and storage state server-side so the agent doesn't re-authenticate on every invocation. This is the difference between a demo and something you can schedule hourly.
Live viewer. When an agent fails, you need to see the page it saw. A live viewer lets you attach to a running session and inspect the DOM, the network tab, and the console. Without it, debugging is guesswork.
Configurable browser settings. Hosted runtimes typically expose options for user agent, viewport, locale, timezone, and proxy routing. These are configuration knobs, not magic. Be skeptical of any claim that a setting makes you undetectable — the honest framing is that you can align browser signals with the identity you're presenting.
Usage controls. Browser time is metered. You want per-session visibility and hard caps so a runaway loop doesn't quietly consume your budget. Current rates and plan details live at /pricing.
Where the agent layer sits
There's a distinction worth keeping straight: the browser runtime and the agent framework are different things.
The runtime is the browser. It exposes CDP. It doesn't know or care whether a human, a Playwright script, or an LLM is issuing commands.
The agent framework — browser-use, a custom planner, a Claude tool loop — decides *what* to do. It reads the page, reasons about the next action, and emits commands.
You can mix and match. A browser-use agent pointed at a hosted CDP endpoint behaves the same as one pointed at local Chrome, except it survives your laptop closing. The CLI tools in that ecosystem (agent-browser and similar) are thin wrappers that take a connection URL and expose navigation, clicking, and extraction as commands.
The reason this separation matters: when your agent framework changes — and it will — your browser infrastructure doesn't have to. If you've built a clean CDP boundary, swapping the reasoning layer is a config change.
Production criteria to evaluate before you commit
If you're choosing a runtime for real workloads, these are the questions that separate options.
- Does it speak raw CDP? If the answer is "we have our own SDK and that's it," you're locked in. CDP is the escape hatch.
- Can you pin the browser version? Agent prompts and selectors are sensitive to rendering differences. Unpinned browser versions mean unpredictable regressions.
- What happens on session timeout? Does the session resume, or does the task fail? Know this before you build retries.
- How is browser time metered? Per session, per hour, per task? Understand the unit before you estimate cost.
- Is there a live viewer? Non-negotiable for anything you'll operate.
- What's the proxy story? Datacenter, residential, or bring-your-own. Match it to the sites you're targeting.
- What are the concurrency limits? Check the actual ceiling for your plan rather than trusting a marketing headline.
If you're comparing this against running Chromium yourself, /blog/remote-browser-online walks through the setup differences.
A realistic agent loop
Here's how the pieces fit in a task that runs unattended.
- Request a session. The runtime allocates a browser and returns a CDP URL plus a session ID.
- Connect. Playwright or Puppeteer attaches over CDP.
- Navigate and observe. The agent reads the accessibility tree or a DOM snapshot.
- Decide and act. The model picks an action: click, type, scroll, extract.
- Verify. After each action, confirm the page reached the expected state. Don't chain actions blindly.
- Extract or terminate. Pull the result, then close the session.
- Record. Log the session ID, the actions taken, and the outcome. You'll want this when something fails next week.
Steps 3 through 5 are where most agent reliability is won or lost. A remote browser makes step 1 and step 6 trivial, which is exactly the point — it removes the infrastructure questions so you can spend your effort on the reasoning loop.
Common mistakes when moving to a remote browser
Ignoring the session timeout. Sessions expire. If your agent takes longer than the timeout, it fails mid-task. Either keep tasks short or use a runtime that supports session extension.
Reusing one browser across unrelated tasks. It's tempting to save on startup cost. It also means one task's cookies and storage contaminate the next. Isolate unless you have a specific reason not to.
Skipping the viewer until something breaks. Attach to a session early in development. You'll catch selector drift and timing bugs before they hit production.
Assuming the network is fast. Hosted browsers have their own egress path. Latency to your target site may differ from your local machine. Measure it.
Hardcoding the CDP URL. Endpoints change between sessions. Always fetch a fresh one.
When a simple remote browser is the wrong choice
Be honest about the cases where local Chrome wins.
If you're writing a one-off script, launching Chrome locally is faster than provisioning anything. If you need to interact with a browser extension, hosted runtimes generally don't support that. If your workflow depends on a specific Chrome profile with local certificates, you'll fight the abstraction.
And if your task volume is genuinely tiny — a handful of runs a day — the operational overhead of any hosted service may not be worth it. The calculus flips when you need concurrency, persistence, or the ability to run from a server that has no display.
Getting started
The shortest path: get a connection URL, paste it into chromium.connectOverCDP(), and run your existing Playwright script unchanged. If it works locally, it should work remotely, because the protocol is the same.
From there, add the pieces you actually need — persistent profiles for logged-in workflows, a viewer for debugging, proxy configuration for geo-targeted tasks. Don't build all of it upfront.
The documentation covers session creation, CDP endpoints, and the available configuration options. If you want to see what a session looks like end to end before writing code, /blog/remote-web-browser has a walkthrough.
The goal isn't to make browser automation sound complicated. It's to make the infrastructure disappear so the only thing left to solve is the task itself.