BLOG
Hermes Browser Connect: Wire Agents to Hosted Chromium
Hermes browser connect explained: how to link a Hermes agent to a cloud browser over CDP, what to configure, and where hosted Chromium fits.
# Hermes Browser Connect: Wire Agents to Hosted Chromium
Hermes browser connect is the step where a Hermes agent stops driving a local Chrome window and starts driving a browser that lives somewhere else. If you are searching for this term, you probably already have a Hermes agent that works on your laptop and now needs to run on a server, in CI, or across many concurrent tasks. The connection itself is not exotic: Hermes talks to a browser over the Chrome DevTools Protocol (CDP), and a hosted runtime gives you a CDP endpoint plus a session you can inspect. This guide covers what "connect" actually means, the two wiring patterns you will use, the production criteria that matter, and the failure modes that waste the most time.
What "Hermes Browser Connect" Actually Means
Hermes is an agent framework that exposes browser actions as tools. Underneath, those tools need a browser process. There are only two ways to give it one:
- Local launch. The agent spawns Chromium on the same machine. Fine for development, painful at scale.
- Remote connect. The agent receives a WebSocket CDP endpoint and attaches to a browser running elsewhere.
Hermes browser connect refers to the second pattern. You are not installing a browser; you are pointing the agent at one. The agent's tool calls become CDP commands over a socket, and the browser executes them on a machine you do not manage.
This matters because the local pattern breaks in predictable ways. A container without a display server, a CI runner that gets recycled mid-task, a laptop that sleeps during a long scrape — all of these kill the session. A remote endpoint decouples the agent's lifetime from the browser's lifetime, which is the whole point.
If you want the broader architectural context before wiring anything, Remote Browser for AI agents covers why the runtime layer exists separately from the agent layer.
The Two Connection Patterns
There are two ways to connect Hermes to a remote browser, and picking the wrong one causes most of the confusion in this space.
Pattern A: CDP endpoint (connect over WebSocket)
The runtime gives you a URL like wss://.../cdp or an HTTP endpoint that resolves to one. You pass it to Playwright's connectOverCDP, Puppeteer's connect, or a raw CDP client. The agent then has full protocol access: pages, targets, network events, and the ability to create new contexts.
This is the most flexible pattern. It is also the one that requires you to manage session lifecycle yourself — you decide when to close the browser, and you decide how to handle reconnects.
Pattern B: SDK or REST session creation
The runtime exposes an API that creates a session and returns a connection URL. You call it at task start, get back an endpoint, connect, do the work, and release the session. This is Pattern A with lifecycle management wrapped around it.
For Hermes specifically, Pattern B is usually the better fit because agent tasks are short-lived and bursty. You do not want a browser sitting idle between tasks, and you do not want to hand-roll cleanup logic that leaks sessions when a task throws.
| Criterion | Local launch | CDP endpoint (Pattern A) | SDK session (Pattern B) |
|---|---|---|---|
| Setup effort | Low | Medium | Low |
| Survives agent restart | No | Yes | Yes |
| Concurrent tasks | Limited by host | Limited by your pool | Scales with runtime |
| Session cleanup | Manual | Manual | Handled by API |
| Live debugging | Local DevTools | Viewer or DevTools | Viewer or DevTools |
| Best for | Prototyping | Custom orchestration | Production agents |
The practical rule: prototype locally, then move to Pattern B for anything that runs unattended.
Wiring Hermes to a Hosted Browser with Playwright
Most Hermes browser skills are built on Playwright or Puppeteer. The connection code is small, but the details matter. Here is a TypeScript example using Playwright's CDP connection against a hosted Chromium session.
import { chromium, Browser, BrowserContext, Page } from "playwright";
interface SessionInfo {
cdpUrl: string;
sessionId: string;
}
// Your runtime's session API returns a CDP endpoint.
// Replace this with your actual session-creation call.
async function createSession(): Promise<SessionInfo> {
const res = await fetch(`${process.env.BROWSER_API}/sessions`, {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.BROWSER_API_KEY}`,
},
body: JSON.stringify({
// Configurable browser settings: region, proxy, profile reuse.
profile: "hermes-default",
timeoutSeconds: 900,
}),
});
if (!res.ok) throw new Error(`session create failed: ${res.status}`);
return res.json();
}
async function runHermesTask(taskUrl: string) {
const session = await createSession();
let browser: Browser | undefined;
try {
// connectOverCDP attaches to the remote browser over WebSocket.
browser = await chromium.connectOverCDP(session.cdpUrl, {
timeout: 30_000,
});
// Reuse the default context so persistent profile state applies.
const context: BrowserContext = browser.contexts()[0]
?? await browser.newContext();
const page: Page = context.pages()[0] ?? await context.newPage();
await page.goto(taskUrl, { waitUntil: "domcontentloaded" });
// Hand the page to your Hermes agent loop here.
// The agent issues CDP-backed actions through this page object.
const title = await page.title();
return { ok: true, title };
} finally {
// Closing the connection does not always release the remote session.
// Call your runtime's release endpoint explicitly.
if (browser) await browser.close().catch(() => {});
await fetch(`${process.env.BROWSER_API}/sessions/${session.sessionId}`, {
method: "DELETE",
headers: { Authorization: `Bearer ${process.env.BROWSER_API_KEY}` },
}).catch(() => {});
}
}Three things in that snippet are worth calling out.
`connectOverCDP` is Chromium-only. Playwright's CDP connection targets Chromium-based browsers. If your Hermes skill assumes Firefox or WebKit, CDP is not the transport — you need a different integration path. The Playwright CDP documentation is the authoritative reference here.
`browser.close()` is not the same as releasing a session. Closing the CDP connection detaches your client. Whether the remote browser is torn down depends on the runtime. Always call an explicit release endpoint, or you will accumulate orphaned sessions and wonder why your usage graph looks wrong.
Reuse the existing context. When you connect over CDP, the remote browser already has a default context. Creating a new one is sometimes correct (isolation between tasks) and sometimes wrong (you lose the persistent profile you configured). Decide deliberately.
What to Configure Before You Connect
Connection is the easy part. The configuration around it determines whether your agent succeeds on real sites.
- Persistent profiles. If your Hermes agent logs in once and expects the session to survive, you need a profile that persists across sessions. Without it, every task starts from a logged-out state.
- Proxy and region. Sites behave differently by geography. If your agent needs a specific region, set it at session creation, not after.
- Configurable browser settings. Hosted runtimes expose browser-level settings you can tune. Be precise about what you need rather than assuming a magic anti-detection switch exists.
- Timeouts. Agent tasks can run long. Set a session timeout that exceeds your worst-case task duration, and handle the case where the session dies mid-task.
- Session isolation. If you run concurrent tasks, each needs its own session. Sharing a browser across agents creates race conditions that are miserable to debug.
The documentation covers the specific knobs available and how they map to session creation parameters.
Production Criteria for a Hermes Browser Runtime
Once you move past "does it connect," these are the questions that decide whether the setup survives contact with production.
Does the runtime give you a live viewer? When an agent fails at step 7 of 12, a screenshot from the failure point is worth more than any log. A live viewer lets you watch the session in real time and inspect the DOM at the moment of failure.
Can you reconnect to a running session? If your agent process restarts, can it reattach to the browser that is still running? This is the difference between losing one task and losing a batch.
How is usage metered? Browser time is the usual unit. Understand whether you are billed for wall-clock session duration or active browser time, because idle sessions add up. Current rates and limits are on the pricing page.
What happens on session failure? Networks drop, sites change, sessions time out. A runtime that returns a clear error and lets you retry cleanly is worth more than one that silently hangs.
Is the CDP surface complete? Some hosted browsers expose a subset of CDP. If your Hermes skill relies on network interception, request modification, or target management, verify those work before you commit.
Common Connection Failures and How to Read Them
Most "Hermes browser connect" problems fall into a small number of buckets.
Connection refused or 401. The endpoint is wrong or the auth token is missing/expired. Check that you are passing the token in the header the runtime expects, not as a query parameter.
Connected but no pages. You attached successfully but context.pages() is empty. Create a page explicitly rather than assuming one exists.
Actions time out after connect. The connection is fine but the browser cannot reach the target site — usually a proxy or DNS issue inside the runtime, not your code.
Session dies mid-task. Either the session timeout is shorter than your task, or the runtime reclaimed an idle session. Increase the timeout and make sure your agent is actually issuing commands.
Works locally, fails remotely. Almost always a profile or proxy difference. Your local Chrome has cookies and a residential IP; the remote session does not, unless you configured it.
If you are debugging a specific failure, Hermes remote browser not working walks through the diagnostic sequence in order.
Where Hosted Chromium Fits vs Alternatives
The "browserbase alternative" searches usually come from teams who hit a wall with one of three things: pricing predictability, CDP completeness, or session debugging. The honest framing is that hosted browser runtimes differ on a few axes, and the right choice depends on which axis you care about.
| Axis | What to check |
|---|---|
| Protocol access | Full CDP vs a restricted action API |
| Session lifecycle | Explicit release vs auto-teardown |
| Debugging | Live viewer, session replay, DOM snapshots |
| Profiles | Persistent vs ephemeral only |
| Pricing model | Per browser-hour vs per task vs per seat |
| Framework support | Playwright, Puppeteer, Selenium, raw CDP |
If your Hermes agent is built on Playwright and you need full CDP access, a runtime that exposes a real CDP endpoint is the right shape. If you only need a handful of actions and never touch the protocol, a higher-level action API may be simpler — at the cost of flexibility when a site does something unexpected.
For a comparison of the runtime layer against self-hosted Playwright infrastructure, Remote Browser online covers the operational trade-offs.
A Practical Migration Path
If you have a working Hermes agent on local Chrome, here is the sequence that causes the least disruption.
- Abstract the browser launch. Replace direct
chromium.launch()calls with a function that returns a connectedBrowserobject. Keep the local path working. - Add a remote path behind a flag. Implement
connectOverCDPagainst a hosted session. Run both paths against the same task and compare results. - Move profile state. Export cookies and storage from your local profile and seed the remote profile. Verify login-dependent tasks still pass.
- Add session cleanup. Wire the release call into your task's
finallyblock before you run anything at volume. - Instrument failures. Log the session ID with every error so you can pull up the viewer for the exact session that failed.
- Scale concurrency gradually. Start with a small number of parallel sessions and watch for rate limits or resource contention before you push higher.
Step 2 is where most teams learn something. Tasks that passed locally often fail remotely for reasons that have nothing to do with the connection — different IP reputation, missing cookies, different viewport size. Finding those early is cheaper than finding them at scale.
Summary
Hermes browser connect is a CDP attachment problem, not a framework problem. You create a session, get an endpoint, connect over WebSocket, and drive the browser through the same Playwright or Puppeteer APIs your agent already uses. The connection code is a dozen lines. The work is in profiles, proxies, timeouts, session cleanup, and having a way to see what happened when a task fails.
Get those right and the local-versus-remote distinction stops mattering to your agent. Get them wrong and you will spend your time debugging sessions instead of shipping tasks.