BLOG
Hermes Agent Browser UI: Connect a Cloud Runtime
Hermes agent browser UI explained: how the Hermes browser skill drives a hosted Chromium session over CDP, and when to move off local Chrome.
# Hermes Agent Browser UI: Connect a Cloud Runtime
The Hermes agent browser UI is the surface where a Hermes agent observes and acts on a web page — a rendered viewport, a DOM snapshot, and a set of tool calls that click, type, and navigate. What most teams discover after the first demo is that the UI is the easy part. The hard part is the browser behind it: a real Chromium instance with a stable session, a persistent profile, and a CDP endpoint the agent can reach from wherever it runs. This guide covers how the Hermes browser skill actually drives a browser, where local Chrome breaks down, and how to point Hermes at a hosted runtime instead.
What the Hermes agent browser UI actually is
Hermes is an agent framework, not a browser. When people say "Hermes agent browser UI," they usually mean one of three things:
- The agent's observation layer — screenshots, accessibility trees, or extracted DOM that the model reasons over.
- The browser skill — the tool definitions Hermes exposes to the model so it can call
navigate,click,type,scroll, andextract. - The live viewport — a human-readable window into what the agent is doing, used for debugging and intervention.
All three depend on the same underlying resource: a browser process the agent can control. In a local setup, that's Chrome or Chromium launched on the same machine as the agent. In a production setup, it's a remote browser session reachable over the Chrome DevTools Protocol.
The distinction matters because the browser skill is stateless from the model's perspective. It issues commands. The browser holds the state — cookies, localStorage, the current page, the open tab. If that state lives on a laptop that sleeps, or in a container that gets recycled between agent turns, the UI will look fine while the task silently fails.
How the Hermes browser skill drives a browser
Most Hermes browser integrations follow the same pattern. The skill wraps a browser automation library — Playwright, Puppeteer, or a raw CDP client — and exposes a small set of tools to the model. Each tool call maps to a browser action.
The connection target is the variable. Playwright's connectOverCDP method accepts a WebSocket endpoint and returns a Browser object you can drive exactly like a locally launched one. That's the seam where a hosted runtime plugs in.
import { chromium, Browser, Page } from 'playwright';
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP_URL!;
async function connectHermesBrowser(): Promise<{ browser: Browser; page: Page }> {
const browser = await chromium.connectOverCDP(CDP_ENDPOINT, {
timeout: 30_000,
});
// Reuse the existing context so the profile (cookies, storage) persists
// across agent turns instead of starting from a blank slate.
const context = browser.contexts()[0] ?? (await browser.newContext());
const page = context.pages()[0] ?? (await context.newPage());
page.setDefaultTimeout(15_000);
return { browser, page };
}
async function runHermesStep(page: Page, action: string, selector: string) {
switch (action) {
case 'click':
await page.locator(selector).click();
break;
case 'type':
await page.locator(selector).fill('query');
break;
case 'navigate':
await page.goto(selector, { waitUntil: 'domcontentloaded' });
break;
default:
throw new Error(`Unsupported Hermes action: ${action}`);
}
}Two details in that snippet carry most of the production weight. First, browser.contexts()[0] — reusing the existing context rather than creating a new one is what makes a login survive between agent turns. Second, the timeout. Agents retry; a browser call that hangs for two minutes will burn model tokens waiting on a page that will never load.
If you're wiring this for the first time, the Hermes browser CDP guide walks through endpoint formats and the failure modes that show up when the endpoint is wrong.
Where local Chrome breaks the Hermes agent browser UI
Local Chrome works until it doesn't. The failure modes are predictable:
Session loss between turns. An agent that logs into a dashboard on turn one and queries it on turn five needs the same browser context throughout. If the process restarts, the profile is gone and the agent re-authenticates — or gets locked out by rate limiting.
No isolation between tasks. Two concurrent Hermes runs sharing one Chrome profile will collide on cookies, tabs, and storage. You can launch separate profiles, but now you're managing profile directories, ports, and cleanup.
Headless detection. Many sites serve different content to headless Chrome. A local headless setup that works in testing may return a challenge page in production.
No remote debugging surface. If the agent runs in a container and Chrome runs on your laptop, there's no CDP path between them without a tunnel you have to maintain.
Resource contention. Chromium is memory-hungry. Running several instances alongside the agent process on one box degrades both.
None of these are fatal in isolation. Together they're why teams move the browser off the agent host.
Hosted Chromium as the Hermes browser runtime
A hosted browser runtime gives the Hermes browser skill a stable CDP endpoint instead of a local process. Remote Browser provisions Chromium sessions with a WebSocket URL you pass to connectOverCDP, plus a live viewer for debugging and persistent profiles for session continuity.
The practical differences from local Chrome:
| Concern | Local Chrome | Hosted runtime |
|---|---|---|
| Session persistence | Tied to process lifetime | Profile survives across sessions |
| Concurrency | Manual profile/port management | Isolated sessions per task |
| Agent location | Must share host with browser | Any host with network access |
| Debugging | Local DevTools only | Live viewer + CDP |
| Browser settings | Whatever you launch with | Configurable per session |
| Scaling | Add machines | Add sessions |
The trade-off is latency and cost. A hosted session adds network round-trips to every CDP call, and you pay per browser-hour rather than per machine. For short, interactive tasks on a fast connection the difference is negligible. For high-frequency DOM polling it's measurable, and worth benchmarking against your own workload before committing.
Pricing and session limits change; check /pricing for current numbers rather than planning against a figure from a blog post.
Choosing between a Hermes browser skill and a browser extension
There's a persistent question about whether to drive the browser through an extension instead of a CDP connection. Extensions attach to a user's existing browser, which is useful when the task depends on a logged-in human session you can't reproduce programmatically.
For agent workloads, extensions have three problems:
- They require a browser the human is using. That's not a production resource.
- They can't be provisioned per task. One extension, one browser, one profile.
- They expose a narrower API. Extension messaging is not CDP; you lose network interception, precise input events, and page-level evaluation.
CDP is the better fit when the agent owns the browser. Extensions are the better fit when the agent borrows one. Most Hermes production deployments are the former.
Production criteria for a Hermes browser runtime
Before you commit to a runtime, check these against your workload:
CDP compatibility. Does it expose a standard WebSocket endpoint that Playwright, Puppeteer, and Selenium can attach to? Vendor-specific SDKs are fine as a convenience layer, but the CDP path should work without them.
Profile persistence. Can you reuse a browser context across sessions so logins and storage survive? This is the single biggest determinant of agent reliability on authenticated sites.
Session isolation. Are concurrent sessions genuinely separate — separate cookies, separate storage, separate network identity? Shared state between agent runs is a correctness bug, not a performance trade-off.
Observability. Can you watch a session live while it runs? When an agent fails at step seven, a screenshot from step seven is worth more than a log line.
Configurable browser settings. Proxy configuration, user agent, viewport, and locale should be settable per session. Ask vendors to document exactly which browser properties a session can change, and treat undocumented claims about identity obfuscation with caution.
Usage controls. Per-session timeouts, hard caps, and a clear billing model. An agent stuck in a retry loop should hit a wall, not your budget.
Data handling. If your agent touches authenticated data, you need to know retention policy, whether sessions are wiped, and what compliance posture the vendor holds.
The remote browser for AI agents piece goes deeper on the runtime layer these criteria describe.
Wiring Hermes to a hosted session: the practical steps
The setup is short. The debugging is where time goes.
- Provision a session through the runtime's API or dashboard and copy the CDP WebSocket URL.
- Store the endpoint as an environment variable. Never hardcode it; sessions are ephemeral.
- Connect with `connectOverCDP` and reuse the existing context, as in the snippet above.
- Register the browser tools with Hermes so the model can call them.
- Open the live viewer in a second window before you run the agent. Watching the first run in real time saves an hour of log reading.
- Set a session timeout that matches your longest expected task, plus margin.
Common failure modes and their causes:
- `connectOverCDP` times out. The endpoint is wrong, expired, or the session was reaped. Re-provision.
- Agent logs in every turn. You're creating a new context instead of reusing
browser.contexts()[0]. - Actions succeed but the page doesn't change. The selector matched a hidden element, or the page uses a framework that needs a network-idle wait rather than
domcontentloaded. - Sessions interfere with each other. You're reusing one session across concurrent tasks. Provision one per task.
If you're comparing runtimes, remote browser online covers what to look for when evaluating hosted options, and the documentation has the endpoint and SDK reference.
When a hosted runtime is the wrong choice
Hosted browsers are not universally better. Skip them if:
- Your agent runs entirely on one machine and never needs to survive a restart.
- You're prototyping and the cost of a session outweighs the cost of local setup friction.
- Your task requires a browser the user is actively using — that's an extension case, not a CDP case.
- You have strict data residency requirements that a hosted vendor can't meet.
For everything else — authenticated workflows, concurrent tasks, agents in containers, anything that needs to run unattended — the hosted runtime removes a category of failure rather than a single bug.
The short version
The Hermes agent browser UI is a thin layer over a browser process. The model sees tools; the browser holds state. Local Chrome couples that state to a machine, which is fine for demos and fragile in production. A hosted Chromium session with a CDP endpoint decouples them: the agent runs anywhere, the browser persists, and the live viewer gives you a debugging surface that local headless Chrome doesn't.
Start by getting connectOverCDP working against a hosted endpoint with context reuse. That one change — reusing browser.contexts()[0] instead of creating a fresh context — fixes more Hermes reliability problems than any other single line of code.
For the underlying protocol details, the Chrome DevTools Protocol documentation is the authoritative reference. For the runtime side, see Remote Browser pricing and the documentation.