BLOG
Hermes Browser CDP: Connect Agents to Hosted Chromium
Hermes Browser CDP explained: how to wire a Hermes agent browser to hosted Chromium over CDP, with Playwright code, trade-offs, and production checks.
# Hermes Browser CDP: Connect Agents to Hosted Chromium
Hermes Browser CDP is the connection layer that lets a Hermes agent drive a real Chromium instance over the Chrome DevTools Protocol instead of bundling a browser into the agent process. If you are wiring a Hermes agent browser to a remote runtime, the practical question is not "does CDP work" — it does — but which endpoint you connect to, how sessions and profiles persist, and what breaks when you move from a laptop to a fleet of workers. This guide covers the Hermes Browser CDP connection model, a working Playwright example, and the production criteria that separate a demo from a runtime you can leave running.
The short version: point your Hermes agent at a CDP WebSocket endpoint, attach with connectOverCDP, and let the hosted runtime own the browser lifecycle. Your agent owns the reasoning loop. The browser becomes infrastructure.
What "Hermes Browser CDP" actually refers to
The phrase gets used loosely, so it helps to separate the pieces:
- Hermes agent — the reasoning and tool-calling layer. It decides what to click, type, or read.
- Browser skill / tool — the interface the agent calls to act on a page (navigate, snapshot, click, extract).
- CDP — the protocol that carries those actions to a real browser. Chrome DevTools Protocol is a JSON-over-WebSocket protocol maintained by the Chromium project; it exposes domains like
Page,Runtime,Network, andTarget. - The browser runtime — the process that actually renders pages, holds cookies, and executes JavaScript.
When people search for "hermes browser cdp," they usually mean one of three things: how to connect a Hermes agent to a browser over CDP, why their local connection fails, or how to run it in the cloud without managing Chrome. All three reduce to the same architecture decision: local browser or hosted browser.
CDP is the right abstraction because it is browser-native. Playwright, Puppeteer, and Selenium all speak it (Puppeteer and Playwright natively; Selenium via ChromeDriver). That means your Hermes browser skill does not have to be rewritten when you change runtimes — only the connection URL changes.
Why local Chromium breaks down for Hermes agents
Running Chromium next to your agent works until it doesn't. The failure modes are predictable:
- Resource contention. Each Chromium instance is memory-hungry. Ten concurrent agents on one box will thrash. See Selenium Playwright headless browser resource usage for the numbers.
- State leakage. A shared local browser carries cookies and localStorage between tasks. Agents that log into different accounts on the same profile will cross-contaminate sessions.
- No isolation. One crashed tab can take down the whole browser process, and with it every agent sharing it.
- Deployment friction. Your CI runner, your serverless function, and your laptop all need a compatible Chromium build. Version drift is constant.
- No live visibility. When an agent gets stuck, you have no way to watch the page without SSH-ing in and hoping.
A hosted runtime addresses these by giving each session its own isolated Chromium, a stable CDP endpoint, and a live viewer. The agent code stays the same; the browser moves out of process.
The connection model: endpoint, session, profile
A Hermes browser connect flow has three moving parts.
1. The CDP endpoint. The runtime exposes a WebSocket URL, typically wss://.../devtools/browser/<id>. This is the browser-level endpoint. You can also connect to a page-level target, but browser-level is what connectOverCDP expects.
2. The session. A session is one isolated browser context with its own cookies, storage, and tabs. Sessions are the unit of concurrency and the unit of billing on most hosted runtimes. When a session ends, its ephemeral state goes with it.
3. The profile. A profile is persistent state you can reattach across sessions — logins, cookies, localStorage. Profiles are how an agent stays logged in between runs without re-authenticating every time. This is the difference between a stateless scraper and an agent that maintains a working identity.
The important design point: sessions and profiles are separate concerns. You can run many ephemeral sessions, or attach a persistent profile to a session when the task needs continuity. Mixing them up is the most common source of "why is my agent logged out" bugs.
A working Playwright + CDP example
Here is the minimal TypeScript path. It assumes the runtime gives you a CDP URL and, optionally, a profile ID.
import { chromium, Browser, BrowserContext, Page } from 'playwright';
interface HermesBrowserConfig {
cdpUrl: string; // wss://.../devtools/browser/<id>
profileId?: string; // attach persistent state if needed
}
async function connectHermesBrowser(
config: HermesBrowserConfig
): Promise<{ browser: Browser; context: BrowserContext; page: Page }> {
const browser = await chromium.connectOverCDP(config.cdpUrl, {
// Keep the connection alive across long agent loops.
timeout: 30_000,
});
// Reuse the default context the runtime created, or make a new one.
const context = browser.contexts()[0] ?? (await browser.newContext());
if (config.profileId) {
// Runtime-specific: attach a persistent profile before navigation.
// Consult your runtime's docs for the exact call.
}
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
return { browser, context, page };
}
async function runHermesTask(cdpUrl: string) {
const { browser, page } = await connectHermesBrowser({ cdpUrl });
try {
// Your agent loop drives the page here.
const title = await page.title();
console.log('Page title:', title);
} finally {
// Close the page, not the browser, if the runtime owns lifecycle.
await page.close();
await browser.close();
}
}Two details matter more than they look. First, connectOverCDP is Chromium-only in Playwright — it will not attach to Firefox or WebKit. If your Hermes agent needs cross-browser coverage, you need a different transport. Second, decide who owns the browser lifecycle. If the runtime created the session, you generally close the page and let the runtime reap the session, rather than calling browser.close() and tearing down a shared instance.
For the protocol-level view, the Chrome DevTools Protocol documentation lists every domain and method available once you are attached.
Local vs hosted: the trade-off table
| Criterion | Local Chromium | Hosted Chromium (CDP) |
|---|---|---|
| Setup | Install + version-pin per machine | Paste a CDP URL |
| Concurrency | Bounded by local RAM/CPU | Scales with sessions |
| Isolation | Manual (contexts, containers) | Per-session by default |
| Persistent profiles | You manage the user-data dir | Runtime-managed profiles |
| Live debugging | SSH + screenshots | Live viewer in browser |
| Proxy / network config | Manual flags | Configurable browser settings |
| Failure blast radius | Whole process | Single session |
| Cost model | Your infra | Metered per browser-hour |
The trade-off is real: hosted runtimes cost money per session-hour, and you give up some low-level control over the Chromium build. If you are running one agent on one task, local is fine. If you are running many agents, or agents that must survive restarts, hosted wins on operational cost even before you count engineering time.
Production criteria for a Hermes browser runtime
Before you commit, check these against any runtime you evaluate — including Remote Browser.
CDP compatibility. Does it expose a standard browser-level WebSocket endpoint that connectOverCDP and Puppeteer's connect accept? Anything proprietary forces a rewrite.
Session isolation. Are sessions truly isolated, or do they share a browser process? Shared processes leak state and amplify crashes.
Profile persistence. Can you attach a named profile to a session and reattach it later? This is non-negotiable for authenticated workflows.
Live viewer. Can you watch a session in real time when an agent misbehaves? Debugging blind is expensive.
Network controls. Proxy configuration and configurable browser settings matter for sites that behave differently by region or IP reputation. Ask vendors to document exactly which browser settings and network options are configurable, and how they are applied per session.
Usage controls. Per-session timeouts, concurrency caps, and clear metering. See /pricing for how Remote Browser meters browser-hours.
Framework compatibility. Playwright, Puppeteer, and Selenium should all work against the same endpoint. If only one does, you are locked in.
Where the Hermes browser skill fits
The Hermes browser skill is the agent-facing wrapper around CDP. It exposes high-level actions — navigate, snapshot, click, type, extract — that the agent calls, and translates them into CDP commands. The skill should not care whether the browser is local or remote; it should take a connection URL as configuration.
That separation is what makes the architecture portable. Your skill defines *what* the agent can do. The runtime defines *where* it happens. When you move from a local Chromium to a hosted session, you change one config value.
If you are building the skill from scratch, keep the CDP connection in a thin adapter layer. Do not let Playwright types leak into your agent's tool definitions — you will want to swap transports eventually.
Common Hermes browser connect failures
Most "hermes browser not working" reports trace to a handful of causes:
- Wrong endpoint type. Connecting to a page target when the library expects a browser target (or vice versa). Check the URL shape.
- Expired session. Hosted sessions have timeouts. A stale CDP URL will fail to connect. Re-request a session.
- Missing profile attach. The agent navigates before the profile is attached, so it lands on a logged-out page.
- Firefox/WebKit expectation.
connectOverCDPis Chromium-only. If you need Firefox, use a different path. - Network egress. Your worker cannot reach the CDP host — usually a firewall or VPC issue, not a browser issue.
Log the CDP URL, the session ID, and the profile ID on every connect attempt. That one habit resolves most of these in minutes.
When to move to a hosted runtime
Move when any of these is true:
- You run more than a handful of concurrent agents.
- Your agents need to stay logged in across runs.
- You need to debug sessions you cannot reproduce locally.
- Your deployment target (serverless, CI, containers) cannot host Chromium reliably.
- You are spending more time on browser infra than on agent logic.
Remote Browser is built for exactly this: hosted Chromium sessions with CDP access, Playwright/Puppeteer/Selenium compatibility, a live viewer, persistent profiles, and per-session isolation. You can read the connection details in the /documentation and see how sessions and profiles are modeled in Remote Browser for AI agents.
If you want to see the runtime without installing anything, start with Remote Browser online. If you are comparing connection models, Remote control browser walks through the code-and-agent split.
The bottom line
Hermes Browser CDP is not a product — it is an architecture. The agent reasons, the skill translates, CDP transports, and the runtime renders. Get the Hermes Browser CDP connection model right — browser-level endpoint, isolated sessions, attachable profiles — and your Hermes agent browser will run the same locally and in the cloud. Get it wrong, and you will spend your time debugging Chromium instead of shipping agents.
Start with a hosted session, connect over CDP, and keep the browser out of your agent process. That is the whole trick.