BLOG
Hermes Agent Browser: Connect a Cloud Runtime
Hermes agent browser setup explained: how to connect a Hermes agent to hosted Chromium over CDP, what the browser skill does, and when to move off local Chrome.
# Hermes Agent Browser: Connect a Cloud Runtime
A Hermes agent browser setup is the layer that lets a Hermes agent drive a real Chromium instance — navigating pages, filling forms, reading DOM state, and returning results — without that browser running on the same machine as the agent process. In practice, most teams get there by pointing the agent at a remote browser over the Chrome DevTools Protocol (CDP) instead of launching a local Chrome binary. This guide covers what the Hermes browser skill actually does, how to wire a hosted runtime into it, and the production criteria that decide whether local Chrome is still good enough.
If you already know you want a hosted runtime and just need the connection details, start with the documentation and come back for the trade-offs.
What "Hermes agent browser" actually refers to
The phrase gets used for three different things, and conflating them causes most of the setup confusion:
- The browser skill. A capability module that gives the agent a set of browser actions — navigate, click, type, extract, screenshot — usually exposed as tool calls the model can invoke.
- The browser runtime. The actual Chromium process the skill drives. This can be local, in a container next to the agent, or a hosted session reachable over a WebSocket endpoint.
- The connection layer. The protocol and credentials that join the two. For Chromium-family browsers this is almost always CDP, which Playwright, Puppeteer, and Selenium can all speak.
The first two are frequently bundled in tutorials, which is why people assume the browser has to live inside the agent's process. It doesn't. The skill is just an interface; the runtime is a separate concern you can host anywhere reachable over the network.
Why the runtime location matters more than the skill
If your Hermes agent runs on a laptop and browses three pages a day, local Chrome is fine. The problems start when any of these become true:
- The agent runs in a container or serverless function. There's no display, no persistent disk, and often no Chromium binary. You either bake a browser into the image or connect to one elsewhere.
- Sessions need to survive a restart. A local browser dies with the process. Persistent profiles on a hosted runtime keep cookies, localStorage, and login state across runs.
- You need to watch what the agent is doing. Debugging a failed task from logs alone is slow. A live viewer lets you see the actual page state at the moment of failure.
- You're running more than a handful of concurrent tasks. Each local Chromium instance costs memory and CPU on the same box as your agent. That ceiling arrives fast.
None of these are about the skill's quality. They're about where the browser lives.
Connecting a Hermes agent to a hosted browser over CDP
The connection pattern is the same one Playwright uses for any remote Chromium: get a CDP WebSocket endpoint, then call connectOverCDP instead of launch. The Playwright CDP documentation covers the API surface; the Hermes-specific part is where the endpoint comes from.
A typical flow:
- Request a session from the runtime's API. You get back a session ID and a CDP endpoint URL.
- Hand that endpoint to the browser skill, or connect to it directly from your own Playwright code.
- Run the task. The agent's tool calls execute against the remote Chromium.
- Close the session when done so the browser-hour meter stops.
Here's the connection in TypeScript, using Playwright against a hosted Chromium session:
import { chromium, Browser, Page } from 'playwright';
interface SessionInfo {
sessionId: string;
cdpEndpoint: string; // wss://... from the runtime API
}
async function runHermesTask(session: SessionInfo): Promise<void> {
let browser: Browser | undefined;
try {
// connectOverCDP attaches to an already-running Chromium.
// Do NOT call chromium.launch() here — that starts a local browser.
browser = await chromium.connectOverCDP(session.cdpEndpoint, {
timeout: 30_000,
});
// A hosted session usually exposes one existing context.
const contexts = browser.contexts();
const context = contexts.length > 0
? contexts[0]
: await browser.newContext();
const page: Page = context.pages()[0] ?? (await context.newPage());
await page.goto('https://example.com', {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
// Agent actions go here: click, fill, extract, screenshot.
const title = await page.title();
console.log('page title:', title);
} finally {
// close() disconnects the CDP client. Whether it also terminates the
// remote browser depends on the runtime — check your provider's docs.
await browser?.close();
}
}Two details that trip people up:
- `connectOverCDP` vs `launch`. If you call
launch, you've built a local browser and the whole point is lost. The endpoint must come from the runtime API. - Context ownership. With a remote session, the browser may already have a default context. Creating a second one is legal but can split your cookies and storage. Read
browser.contexts()first.
If you're using Puppeteer instead, the equivalent is puppeteer.connect({ browserWSEndpoint }). Selenium reaches the same runtime through a remote WebDriver or CDP bridge depending on your setup.
Hermes browser skill vs raw CDP: which to use
You don't have to choose permanently, but it helps to know what each gives you.
| Approach | What you get | Where it hurts |
|---|---|---|
| Hermes browser skill | Tool-call interface the model can invoke directly; less glue code | Less control over timing, retries, and page-level detail |
| Raw CDP via Playwright/Puppeteer | Full control over navigation, waits, selectors, network interception | You write and maintain the orchestration yourself |
| Skill + hosted runtime | Model-friendly actions on a browser that survives restarts and scales independently | Requires an endpoint-management step and session lifecycle handling |
| Local Chrome + skill | Zero network setup, works offline | Dies with the process, no live viewer, hard concurrency ceiling |
The common production pattern is the third row: keep the skill for the agent's decision-making, but point it at a hosted runtime so the browser isn't coupled to the agent's process lifetime.
Production criteria for a Hermes browser runtime
When you're evaluating where the browser should live, these are the questions that actually change outcomes:
Session isolation. Each agent task should get its own browser context so cookies and storage don't leak between runs. Shared contexts cause the worst kind of bug: intermittent, task-dependent, and nearly impossible to reproduce.
Persistent profiles. If your agent logs in once and reuses that session, you need profiles that outlive a single browser process. This is the difference between a task that takes 40 seconds and one that takes 4 minutes of re-authentication.
Live viewer. When a task fails at step 7 of 12, a screenshot of the final state tells you almost nothing. A live or recorded view of the session tells you what the page looked like when the selector stopped matching.
Proxy and browser settings. Sites behave differently by geography and IP reputation. Configurable browser settings — including proxy routing — let you test and adjust without rebuilding your agent.
Usage controls. Browser time is metered. You want per-session visibility into how long tasks run, so a runaway loop doesn't quietly burn budget. See pricing for how Remote Browser meters this.
CDP compatibility. If the runtime speaks standard CDP, your existing Playwright, Puppeteer, or Selenium code works with a one-line change. If it requires a proprietary SDK for basic operations, you've added a migration cost to every future decision.
Where Remote Browser fits
Remote Browser is a browser API and runtime built for AI agents and browser-use workflows. It provides hosted Chromium sessions with CDP access, so Playwright, Puppeteer, and Selenium clients connect without modification. Sessions are isolated, profiles persist across runs, and a live viewer shows what the agent sees.
For a Hermes agent, the integration is the connectOverCDP pattern above: request a session, pass the endpoint to the browser skill or your own code, run the task, close the session. The agent's logic stays where it is; only the browser's location changes.
Two related reads if you're mapping out the architecture:
- Remote browsers for AI agents covers the runtime layer in more depth.
- Remote browser online walks through running Chromium without local setup.
Common failure modes and how to avoid them
The agent launches a local browser anyway. Some skill implementations default to launch() when no endpoint is configured. Check the skill's config path and confirm it's reading your CDP URL, not falling back silently.
Sessions leak. If you don't close sessions explicitly, they keep running and keep metering. Wrap task execution in try/finally and close in the finally block, as in the code above.
Timeouts are too short for real pages. A 5-second navigation timeout works on a test page and fails on anything with third-party scripts. Start at 30–45 seconds and tune down once you have real latency data.
Selectors assume a fresh DOM. Persistent profiles mean the page may load in a logged-in state with different elements present. Write selectors that handle both states, or reset to a known state at task start.
Concurrency assumptions. Don't assume high parallel session counts without checking your plan. Verify your plan's limits before designing a fan-out architecture, and refer to pricing for current capacity details.
When local Chrome is still the right call
Hosted runtimes aren't automatically better. Local Chrome wins when:
- The agent runs on a developer machine and the task is exploratory.
- You need to inspect the browser with local devtools in real time.
- The site requires a browser extension or a specific local profile you can't replicate remotely.
- Task volume is low enough that setup overhead exceeds the benefit.
The crossover point is usually when the agent moves off the developer's machine — into CI, a container, or a server — or when more than one person needs to see what the agent did. At that point, the browser becomes infrastructure, and infrastructure belongs somewhere you can manage it.
Getting started
The shortest path from a working local Hermes agent to a hosted one:
- Read the documentation for session creation and the CDP endpoint format.
- Replace
chromium.launch()withchromium.connectOverCDP(endpoint)in your task runner. - Verify the agent still completes a known task end to end.
- Add session cleanup in a
finallyblock. - Turn on the live viewer and run a task you expect to fail, so you can see what failure looks like before it happens in production.
The browser skill doesn't change. The runtime does. That separation is the whole point — and it's what makes a Hermes agent browser setup portable across machines, environments, and scale.