← Blog

BLOG

Hermes Browser Skill: Run Agents on Hosted Chromium

Hermes browser skill explained: what it does, how it connects to a remote browser over CDP, and how to run it in production without local Chrome.

October 1, 20269 min readRemote Browser

# Hermes Browser Skill: Run Agents on Hosted Chromium

A hermes browser skill is a packaged capability that lets a Hermes-based agent open pages, click, type, and read the DOM through a real browser. The skill itself is the interface. The part that decides whether it works in production is what sits behind it: a local Chrome process, or a hosted Chromium session you connect to over CDP.

This guide covers what the skill actually does, where local execution breaks down, and how to wire it to a remote browser runtime so sessions survive restarts, scale past one machine, and stay debuggable.

What a Hermes browser skill actually is

Hermes-style agent frameworks expose tools to a model. A browser skill is one of those tools. It usually wraps a driver — Playwright, Puppeteer, or raw CDP — and presents a small set of actions to the agent:

  • Navigate to a URL
  • Click an element by selector or accessibility role
  • Type into a field
  • Extract text or structured data from the page
  • Take a screenshot
  • Wait for a network condition or selector

The skill is a thin layer. It does not contain a browser. It needs a browser process to attach to, and that is the decision that matters.

There are three common attachment models:

  1. Launch local Chrome. The skill spawns a Chromium binary on the same machine as the agent.
  2. Attach to an existing local browser. The skill connects to a Chrome instance already running with a debugging port open.
  3. Connect to a remote browser. The skill receives a WebSocket endpoint and drives a Chromium session running elsewhere.

Most tutorials stop at option one. Most production deployments end up at option three.

Why local Chrome stops working

Local Chrome is fine for a demo. It becomes a liability the moment you have more than one agent, more than one machine, or more than one hour of runtime.

Resource contention. Each Chromium instance holds a substantial amount of RAM — often several hundred megabytes. Ten concurrent agents on one box is already uncomfortable. Fifty is not happening without a scheduler and a lot of headroom.

State that disappears. A local browser dies with the process. Cookies, localStorage, and logged-in sessions vanish. If your agent needs to stay authenticated across runs, you are rebuilding that state every time.

No shared visibility. When an agent fails at step seven, you want to see what the page looked like. With local Chrome on a worker you cannot easily reach, you get a log line and a guess.

Environment drift. Chrome versions, fonts, and OS libraries differ between your laptop, CI, and production. A selector that works locally can fail in the container because a font rendered differently and shifted the layout.

Detection. Datacenter IPs and default automation flags get flagged. Local Chrome on a cloud VM has the same problem, plus you own the mitigation work.

None of these are skill bugs. They are runtime problems. The fix is to move the browser out of the agent process.

The remote browser model

A remote browser runtime runs Chromium somewhere you control, exposes a CDP endpoint, and lets your skill connect to it. The agent process stays light. The browser becomes infrastructure.

Remote Browser provides hosted Chromium sessions with:

  • CDP access over a WebSocket endpoint you can pass to Playwright, Puppeteer, or Selenium
  • Playwright, Puppeteer, and Selenium compatibility through standard connection methods
  • A live viewer so you can watch a session in real time while it runs
  • Persistent profiles so cookies and storage survive across sessions
  • Configurable browser settings, including proxy and stealth-related options
  • Session isolation so one agent's cookies never leak into another's
  • Usage controls so you can meter and cap browser time

The skill does not change. What changes is the connection string.

Connecting a Hermes browser skill over CDP

Playwright's connectOverCDP is the most direct path. You get a Browser object, and from there everything behaves like a locally launched browser.

import { chromium, Browser, BrowserContext, Page } from "playwright";

// The endpoint comes from your Remote Browser session.
// Treat it like a secret: it grants full control of the browser.
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP_URL!;

async function runHermesBrowserTask(taskUrl: string) {
  const browser: Browser = await chromium.connectOverCDP(CDP_ENDPOINT, {
    timeout: 30_000,
  });

  // Reuse the default context so a persistent profile is honored.
  const context: BrowserContext = browser.contexts()[0] ?? (await browser.newContext());
  const page: Page = context.pages()[0] ?? (await context.newPage());

  try {
    await page.goto(taskUrl, { waitUntil: "domcontentloaded", timeout: 45_000 });

    // Wait on a real signal, not a fixed sleep.
    await page.waitForSelector("[data-testid='results']", { timeout: 20_000 });

    const rows = await page.$$eval("[data-testid='results'] li", (nodes) =>
      nodes.map((n) => (n.textContent ?? "").trim())
    );

    return rows;
  } finally {
    // Close the connection, not the remote browser, unless you own its lifecycle.
    await browser.close();
  }
}

Two details matter here.

First, browser.close() on a CDP connection closes the connection. Whether it terminates the remote browser depends on the runtime. If you want the session to persist for a follow-up task, check your runtime's semantics before assuming.

Second, browser.contexts()[0] is how you reach a persistent profile. If you call newContext() every time, you get a clean slate and lose the login state you were trying to keep.

For Puppeteer, the equivalent is puppeteer.connect({ browserWSEndpoint }). For Selenium, it is a remote WebDriver or a CDP bridge, depending on your setup. The endpoint format differs; the model is the same.

If you want the broader picture of how this fits into an agent stack, see Remote browsers for AI agents.

Local vs remote: what actually differs

DimensionLocal ChromeRemote browser runtime
SetupInstall Chromium per machinePass a CDP endpoint
ScalingBound by host RAM/CPUSessions provisioned per task
Session stateLost on process exitPersistent profiles available
DebuggingScreenshots and logs onlyLive viewer plus CDP access
Environment consistencyVaries by hostConsistent image per session
IP and networkHost IPConfigurable proxy settings
IsolationShared unless you build itPer-session isolation
Cost modelYour computeMetered browser time — see /pricing
Failure recoveryRestart the whole agentReconnect to a new session

The trade-off is real. Remote browsers add a network hop and a dependency. If you are running one agent on your laptop against a public page, local Chrome is simpler and you should keep it. The remote model pays off when concurrency, persistence, or observability become requirements.

Production criteria to evaluate before you commit

Do not pick a runtime on marketing copy. Check these.

CDP fidelity. Does the runtime expose a standard CDP endpoint, or a proprietary API that only works with their SDK? Standard CDP means you can swap runtimes without rewriting the skill.

Profile persistence. Can you attach a named profile to a session and reuse it later? This is the difference between an agent that logs in once and one that logs in every run.

Session isolation. Are cookies and storage scoped per session? Shared state across agents is a correctness bug waiting to happen.

Observability. Can you watch a live session, or only read logs after the fact? Live viewing cuts debugging time dramatically.

Proxy and network controls. Can you route traffic through a specific region or IP type? For sites with geo or bot constraints, this is not optional.

Lifecycle semantics. What happens to the browser when your client disconnects? What is the idle timeout? Get this in writing before you build retry logic around it.

Usage controls. Can you cap browser time per session or per project? Runaway agents are expensive. See /pricing for how metering works.

Where the skill ends and the runtime begins

It helps to be precise about the boundary, because it determines who fixes what.

The skill owns:

  • Which actions the agent can take
  • How the model's output maps to those actions
  • Retry and error-handling logic inside a task
  • What gets extracted and returned

The runtime owns:

  • Browser process lifecycle
  • Session isolation and profiles
  • Network egress and proxy configuration
  • CDP endpoint availability
  • Resource limits and metering

When an agent fails, the first question is which side broke. A selector that no longer matches is a skill problem. A session that dies mid-task is a runtime problem. Separating them makes debugging tractable.

Practical patterns that hold up

One session per task, not per agent. Long-lived sessions accumulate state and drift. Provision a session, run the task, release it. Use persistent profiles when you need continuity, not long-lived sessions.

Reconnect rather than retry from scratch. If a session drops, reconnect to the same profile and resume. Restarting the whole task wastes the work already done.

Wait on signals, not timers. waitForSelector, waitForResponse, and waitForLoadState are reliable. sleep(3000) is not, and it will fail on slow pages.

Log the CDP endpoint with the task ID. When something breaks, you want to correlate a failure to a specific session and open the viewer.

Keep credentials out of the skill. The skill should receive a session, not a password. Let the profile hold the authenticated state.

For a walkthrough of running Chromium without managing a local install, see Remote browser online.

When a remote browser is the wrong choice

Be honest about this.

  • Single agent, single run, public page. Local Chrome is fine.
  • Offline or air-gapped environments. A remote runtime is not available.
  • Sub-100ms interaction requirements. The network hop to a remote browser adds latency. If your task is latency-critical, measure before committing.
  • Heavy local file interaction. If the agent needs to read and write local files alongside browsing, a local browser is simpler.

The remote model is for concurrency, persistence, consistency, and observability. If you do not need those, you are adding a dependency for nothing.

Getting the connection right

The most common failure with a Hermes browser skill is not the skill. It is the connection.

Check these in order when something breaks:

  1. Is the CDP endpoint reachable from the agent's network? Firewalls and private networking are frequent culprits.
  2. Is the endpoint still valid? Sessions expire. A stale URL produces a connection refused, not a helpful error.
  3. Are you reusing the right context? browser.contexts()[0] for persistent profiles, newContext() for clean runs.
  4. Is the browser closing before the task finishes? Check idle timeouts and whether browser.close() terminates the remote session.
  5. Is the page actually loaded? domcontentloaded fires before client-side rendering completes. Wait on a selector that only exists after hydration.

For the underlying protocol details, the Chrome DevTools Protocol documentation is the authoritative reference for what your driver can and cannot do.

Summary

A Hermes browser skill is a tool wrapper. It does not solve concurrency, persistence, or observability — the runtime does. Connecting the skill to a hosted Chromium session over CDP keeps the agent process light, makes sessions reproducible, and gives you a live view when something goes wrong.

Start with the connection. Get connectOverCDP working against a real endpoint, confirm you can reuse a profile, and verify you can watch a session live. Everything else in the skill is easier to debug once those three things are true.

Read the documentation for endpoint formats and session lifecycle details, and check /pricing for current metering.