← Blog

BLOG

Hermes Remote Browser Tutorial: Connect Agents to Hosted Chromium

A practical Hermes remote browser tutorial: wire a Hermes agent to a hosted Chromium runtime over CDP, run Playwright tasks, and debug live sessions.

October 1, 20269 min readRemote Browser

# Hermes Remote Browser Tutorial: Connect Agents to Hosted Chromium

This Hermes remote browser tutorial walks through the practical wiring: taking a Hermes agent that already drives a local browser and pointing it at a hosted Chromium session instead. The goal is not to sell you a runtime — it is to show you exactly which connection string, which CDP endpoint, and which session settings a Hermes agent needs, and where the local-to-remote swap usually breaks.

If you have read the remote browser overview for AI agents, you already know why the runtime layer matters. This post is the implementation.

What "Hermes" means in a browser-agent stack

Hermes is used in a few overlapping ways in the agent ecosystem, and the ambiguity causes most of the setup confusion:

  • Hermes as the agent/orchestration layer. The component that decides what to click, type, or read. It holds the task loop and the model calls.
  • Hermes as a browser skill or tool. A capability the agent invokes — browser.navigate, browser.click, browser.screenshot — often exposed as a function-calling tool or an MCP-style skill.
  • Hermes as a browser extension or in-app control surface. A UI that lets a human watch or take over the session.

In all three cases, the browser itself is a separate concern. Hermes decides *what* to do; the browser runtime executes it. That separation is the whole reason a remote runtime is a drop-in change rather than a rewrite.

The practical implication: if your Hermes agent talks to the browser over CDP (Chrome DevTools Protocol) or through a Playwright/Puppeteer driver, you can redirect it to a hosted session without touching the agent's decision logic.

Why point Hermes at a remote browser instead of local Chrome

Running Hermes against a local Chrome instance works fine on a laptop. It stops working when:

  • You need more than a handful of concurrent sessions.
  • Sessions must survive a process restart or a deploy.
  • You need a stable egress IP or configurable browser settings for sites that behave differently by region.
  • You want a human to watch or intervene in a live session without screen-sharing a dev machine.

A hosted runtime addresses those by moving the browser off your machine while keeping the same control protocol. The trade-off is real: you add a network hop, you depend on the provider's uptime, and you need to think about session lifecycle and cost. For single-task local debugging, local Chrome is still simpler.

Prerequisites

Before you start, confirm you have:

  1. A Hermes agent (or any agent loop) that can call a browser tool.
  2. A Remote Browser account and an API key. Current plans and limits are on /pricing.
  3. Node.js 18+ if you are using the Playwright/CDP path shown below.
  4. Playwright installed locally as a *client* only — you do not need local browser binaries when connecting over CDP.

That last point trips people up. playwright install downloads browser binaries you will not use when connecting to a remote endpoint. You can skip it, or install only the driver.

Step 1: Create a hosted session

Sessions are the unit of isolation. Each session gets its own browser context, its own cookies and storage, and its own lifecycle. Create one through the API and capture the connection URL and CDP endpoint it returns.

The shape of the response is what matters, not the exact field names:

  • A CDP WebSocket endpoint (wss://...) for direct protocol access.
  • A connection URL you can paste into a Playwright connectOverCDP call.
  • A session ID you use to fetch the live viewer URL and to terminate the session.

Keep the session ID. You will need it for debugging and for cleanup, and orphaned sessions are the most common source of surprise cost.

Step 2: Connect Playwright to the hosted session

This is the core of the tutorial. The code below connects a Playwright client to a hosted Chromium session over CDP, runs a small task, and disconnects cleanly. It is TypeScript, but the same pattern applies in Python and in Puppeteer.

import { chromium, Browser, Page } from 'playwright';

// Provided by the Remote Browser session API.
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP!; // wss://...
const SESSION_ID = process.env.REMOTE_BROWSER_SESSION_ID!;

async function runHermesTask(): Promise<void> {
  let browser: Browser | undefined;

  try {
    // connectOverCDP attaches to an existing browser instead of launching one.
    browser = await chromium.connectOverCDP(CDP_ENDPOINT);

    // A hosted session usually exposes one default context.
    const context = browser.contexts()[0] ?? (await browser.newContext());
    const page: Page = context.pages()[0] ?? (await context.newPage());

    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });

    // This is where your Hermes agent's decision loop would run.
    // For now, extract something concrete to prove the wiring works.
    const title = await page.title();
    console.log('Page title:', title);

    await page.screenshot({ path: 'hermes-check.png' });

    // Do NOT call browser.close() on a remote session unless you intend
    // to destroy it. Disconnecting leaves the session alive for reuse.
    await browser.close();
  } catch (err) {
    console.error('Hermes remote browser task failed:', err);
    throw err;
  }
}

runHermesTask().catch(() => process.exit(1));

Two details matter more than the rest:

  • `connectOverCDP` vs `launch`. launch starts a local browser. connectOverCDP attaches to one that already exists. If your Hermes browser skill currently calls launch, that is the single line you change.
  • `browser.close()` semantics. With a remote session, closing the browser connection may or may not terminate the underlying session depending on how the runtime is configured. Check your provider's behavior and terminate sessions explicitly through the API when you are done.

The Playwright documentation on `browserType.connectOverCDP` is the authoritative reference for the client side of this call.

Step 3: Wire the Hermes browser skill to the session

If your Hermes agent uses a browser skill rather than raw Playwright, the change is usually one of:

  • Swap the endpoint. Point the skill's browserEndpoint or cdpUrl config at the hosted session instead of http://localhost:9222.
  • Swap the driver. If the skill launches Chrome itself, replace the launch call with a connect call.
  • Pass the session ID through. Some skills need the session ID to fetch screenshots or attach the live viewer.

A minimal skill config often looks like this:

{
  "browser": {
    "mode": "remote",
    "cdpUrl": "wss://<host>/<session-id>",
    "sessionId": "<session-id>",
    "timeoutMs": 30000
  }
}

The exact keys depend on your Hermes implementation. The invariant is that the skill must stop owning the browser lifecycle and start treating the browser as an external resource.

Step 4: Watch and debug the session

The reason to use a hosted runtime rather than a headless container you manage yourself is the debugging surface. A live viewer lets you watch the agent's session in real time — useful when a task fails on step 7 and you need to see what the page actually looked like.

For deeper debugging, connect a second CDP client to the same session. You can inspect the DOM, read console logs, and evaluate expressions without disturbing the agent's page state. This is the same pattern described in the remote control browser guide, applied to a Hermes agent.

Practical debugging checklist:

  • Screenshot on failure. Have the agent capture a screenshot before it gives up. A failed task with no screenshot is nearly impossible to diagnose.
  • Log the URL at each step. Most agent failures are navigation failures in disguise.
  • Check the live viewer before assuming a code bug. Sometimes the site changed, not your agent.

Local Chrome vs hosted Chromium for Hermes

CriterionLocal ChromeHosted Chromium
Setup timeFast for one sessionOne-time API wiring
ConcurrencyLimited by your machineScales with your plan (see /pricing)
Session persistenceLost on restartConfigurable persistent profiles
Egress IP controlYour network onlyConfigurable proxy settings
Live debuggingScreen share or local devtoolsBuilt-in live viewer + CDP
Cost modelYour hardwareMetered browser time
Best forSingle-task developmentProduction agents, parallel tasks

The honest read: local Chrome wins for the first hour of development. Hosted Chromium wins the moment you need a second concurrent session, a stable IP, or a session that outlives your process.

Production criteria before you ship

Before you move a Hermes agent from a demo to production, verify:

  1. Session cleanup. Every session has an explicit termination path. Orphaned sessions are the top cost leak.
  2. Timeout handling. Set a hard ceiling on task duration. Agents that retry forever on a broken page burn budget silently.
  3. Profile strategy. Decide whether sessions share a persistent profile or start clean. Persistent profiles help with logged-in workflows; clean sessions help with isolation.
  4. Proxy configuration. If the target site behaves differently by region, configure egress before you debug the agent.
  5. Failure observability. Screenshots, console logs, and the final URL on every failure.
  6. Cost visibility. Know your per-session cost before you scale concurrency. See /pricing for current rates.

None of these are exotic. They are the same criteria you would apply to any external dependency, and skipping them is why agent demos fail in production.

Common failure modes

The agent connects but every action times out. Usually a CDP endpoint mismatch — you connected to the browser-level endpoint but the skill expects a page-level one, or vice versa.

Sessions leak. The agent crashes before cleanup runs. Wrap session creation in a try/finally and terminate in the finally.

The agent works locally but not remotely. Almost always a timing difference. Remote sessions have a network hop; waitUntil: 'domcontentloaded' may fire before the content your agent needs is present. Prefer explicit waits on selectors over fixed sleeps.

Login state disappears between runs. You are creating a fresh session each time. Use a persistent profile if the workflow needs it.

Where this fits in a larger stack

A Hermes agent is one consumer of a browser runtime. The same hosted session model serves Playwright test suites, Puppeteer scrapers, and other agent frameworks. If you are evaluating runtimes, the remote browser online guide covers the trade-offs between managed sessions and self-hosted infrastructure, and the remote web browser overview covers the general architecture.

The short version: pick a runtime that speaks CDP, exposes a live viewer, and lets you control session lifecycle explicitly. Everything else is implementation detail.

Next steps

  1. Create a session and capture the CDP endpoint.
  2. Run the TypeScript snippet above against a trivial page to confirm the wiring.
  3. Swap your Hermes browser skill's launch call for a connect call.
  4. Add screenshot-on-failure and explicit session termination.
  5. Read the documentation for session, profile, and proxy configuration.

If you get the connection working on a trivial page, the rest is agent logic — which is the part you actually wanted to work on.