← Blog

BLOG

Hermes Browser GitHub: Run Agents on Hosted Chromium

Hermes Browser GitHub projects give you the agent loop. Here's how to wire them to a hosted Chromium runtime over CDP for production browser automation.

October 9, 202610 min readRemote Browser

# Hermes Browser GitHub: Run Agents on Hosted Chromium

If you searched for Hermes browser GitHub, you are probably looking for one of two things: the source repository for a Hermes-branded browser agent, or a way to point an existing Hermes agent at a browser it does not have to run locally. This post answers both. It covers what the Hermes browser ecosystem on GitHub actually contains, why the repository is only half the runtime, and how to connect a Hermes-style agent to hosted Chromium over the Chrome DevTools Protocol (CDP) so it survives past your laptop.

The short version: GitHub gives you the agent logic. It does not give you a browser that stays alive across deploys, holds a logged-in profile, or scales to concurrent sessions. That is a runtime problem, and it is the part most Hermes browser setups get wrong.

What "Hermes Browser GitHub" Usually Refers To

Search results for this term tend to mix several distinct things. Before you clone anything, sort out which one you actually need:

  • A Hermes agent repository. A repo containing the reasoning loop, tool definitions, and prompts that let a model decide when to click, type, or navigate. The browser is a dependency, not the product.
  • A browser automation CLI. Tools like agent-browser (Vercel Labs) wrap Playwright or Puppeteer behind a command-line interface so an agent can call open, click, and snapshot as shell commands. These are useful glue, but they still need a browser to talk to.
  • A CDP client. Code that connects to a running Chromium instance over a WebSocket endpoint. This is the layer that matters most for production, because it decouples your agent from the machine running the browser.
  • A fork or wrapper. Many "Hermes browser" repos are thin wrappers around Playwright with a Hermes-specific prompt layer. Read the package.json and the connection code before assuming they solve hosting.

The pattern across all four: the repository is the *control plane*. The browser is the *data plane*. If both live in the same process on the same machine, you have a demo, not a deployment.

Why the GitHub Repo Is Only Half the Runtime

A Hermes browser agent running from a cloned repo typically does this:

  1. Launches a local Chromium via Playwright's chromium.launch().
  2. Drives it through the agent loop.
  3. Exits, taking the browser and its cookies with it.

That works until you hit any of the following, which are all normal in production:

  • Session persistence. A login flow that requires a one-time code cannot complete inside a single agent run. You need a profile that survives between runs.
  • Concurrency. Ten agents launching ten local Chromium instances on one box will exhaust memory and CPU long before they exhaust your task queue.
  • Environment drift. The Chromium version, fonts, and system libraries on your CI runner differ from your dev machine. Selectors that worked locally fail in the pipeline.
  • Observability. When an agent fails at step 7 of 12, you need a live view or a recording, not a stack trace.
  • IP and network posture. Some sites behave differently depending on where the request originates. Local egress is a single point of failure.

None of these are solved by a better agent loop. They are solved by moving the browser out of the repo and into a runtime you connect to.

Hosted Chromium vs. Local Chromium for Hermes Agents

Here is the trade-off in concrete terms. This table assumes a Hermes-style agent that already works locally.

CriterionLocal Chromium (from the repo)Hosted Chromium (remote runtime)
Setupnpx playwright install per machineOne connection URL
Session persistenceManual; lost on process exitPersistent profiles across runs
ConcurrencyBounded by host RAM/CPUBounded by your plan; see /pricing
Environment consistencyVaries per machine and CI imageSame image every session
DebuggingLocal headed mode or videoLive viewer plus session logs
Network postureYour machine's IPConfigurable browser and proxy settings
Cost modelYour compute, always onMetered browser time
Best forDevelopment, one-off scriptsAgents that run on a schedule or at scale

The honest read: local Chromium is faster to start and free if you already have the hardware. Hosted Chromium costs money and adds a network hop. You move when persistence, concurrency, or consistency start costing you more than the runtime does.

Connecting a Hermes Agent to Hosted Chromium over CDP

The connection mechanism is the same one Playwright documents for connectOverCDP. Your runtime exposes a WebSocket endpoint; your agent connects to it instead of launching a browser. The Playwright CDP documentation covers the client side; the runtime covers everything behind the endpoint.

Here is a minimal TypeScript example. It assumes your runtime hands you a CDP URL, which is the standard shape for hosted Chromium providers.

import { chromium, Browser, BrowserContext, Page } from 'playwright';

// Your runtime returns a CDP WebSocket endpoint per session.
// Treat this like a credential: never commit it, never log it in full.
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP_URL!;

async function runHermesTask(task: string): Promise<void> {
  let browser: Browser | null = null;

  try {
    // Connect to hosted Chromium instead of launching locally.
    browser = await chromium.connectOverCDP(CDP_ENDPOINT, {
      timeout: 30_000,
    });

    // Reuse the existing context so persistent profile state carries over.
    const context: BrowserContext = browser.contexts()[0]
      ?? await browser.newContext({
        viewport: { width: 1280, height: 800 },
        locale: 'en-US',
      });

    const page: Page = context.pages()[0] ?? await context.newPage();

    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });

    // Hand control to your Hermes agent loop here. The agent decides
    // the next action; Playwright executes it against the remote page.
    await page.getByRole('button', { name: /sign in/i }).click();
    await page.getByLabel('Email').fill('agent@example.com');

    // Snapshot for the model. Keep it small; full HTML blows up context.
    const snapshot = await page.accessibility.snapshot();
    console.log(JSON.stringify(snapshot, null, 2));

    await page.waitForLoadState('networkidle');
  } catch (err) {
    // Surface the session ID from your runtime so you can pull the recording.
    console.error('Hermes task failed:', err);
    throw err;
  } finally {
    // Disconnect the client. The hosted session lifecycle is managed
    // by the runtime, not by this process.
    await browser?.close();
  }
}

runHermesTask('sign in and open the billing page').catch(() => process.exit(1));

Three details matter more than the rest:

  • `connectOverCDP` vs `launch`. You are not starting a browser. You are attaching to one that already exists. Anything you assumed about launch flags, userDataDir, or executable paths no longer applies.
  • Context reuse. browser.contexts()[0] returns the context the runtime created. Creating a new one each run throws away the persistent profile, which defeats the point.
  • `browser.close()` semantics. With connectOverCDP, closing the client disconnects it. Whether the underlying session terminates depends on your runtime's session policy. Check that before you rely on it for cleanup.

If you are using Puppeteer instead, the equivalent is puppeteer.connect({ browserWSEndpoint }). Selenium connects through its remote WebDriver endpoint. The runtime should expose all three; if it only speaks one protocol, that is a constraint worth knowing before you commit.

What to Verify Before You Commit to a Runtime

The Hermes browser repo you cloned will run against almost any CDP endpoint. That is exactly why you should test the endpoint, not the agent. Production criteria, in rough order of how often they bite:

  1. Session isolation. Two concurrent agents must not share cookies, storage, or tabs. Ask how sessions are separated and confirm it in a test.
  2. Profile persistence. Can a session resume with the same logged-in state after a process restart? This is the single biggest gap between local and hosted.
  3. Live debugging. When a run fails, can you watch it or replay it? A live viewer plus session logs shortens debugging from hours to minutes.
  4. Protocol coverage. CDP is the baseline. Playwright, Puppeteer, and Selenium compatibility means you are not locked into one client library.
  5. Configurable browser settings. User agent, viewport, locale, timezone, and proxy configuration should be settable per session. Be skeptical of any claim beyond what the provider documents.
  6. Usage controls. You want per-session and per-account limits so a runaway agent cannot burn your budget overnight.
  7. Pricing transparency. Metered browser time is the norm. Confirm what counts as a browser-hour and where the meter starts. Current details live at /pricing.

The Remote Browser documentation covers how sessions, profiles, and CDP endpoints are exposed if you want to compare against your current setup.

Where Hermes Fits in a Production Stack

A Hermes browser agent is a decision loop. In production, that loop sits between three other layers:

  • Orchestration. A queue or scheduler that decides which tasks run, when, and with what retry policy. This is where you enforce concurrency limits.
  • The browser runtime. Hosted Chromium sessions with persistent profiles, isolated storage, and a CDP endpoint per session.
  • Observability. Session recordings, live viewer access, and structured logs keyed by session ID so a failed run is reproducible.

The GitHub repo owns the first layer's logic and the agent's tool definitions. It should not own the second. When the browser is a connection string instead of a subprocess, you can redeploy the agent without losing state, run it from a serverless function, and scale it horizontally without rewriting the loop.

This is also why the "is Browserbase free" question keeps coming up alongside Hermes browser searches. Developers evaluating hosted Chromium want to know the entry cost before they refactor. The answer varies by provider and changes often, so check the provider's own pricing page rather than a blog post. The more useful question is whether the runtime supports persistent profiles and CDP, because a free tier without those will not carry a Hermes agent past the demo stage.

Common Failure Modes and Fixes

These are the ones that show up repeatedly when teams move a Hermes agent from local to hosted.

The agent connects but every selector times out. Usually a viewport or user-agent mismatch. The hosted browser renders a different layout than your local one. Pin the viewport explicitly in newContext and re-check your selectors against the hosted rendering.

Login state does not survive. You are creating a new context per run instead of reusing the runtime's context. Fix the contexts()[0] pattern shown above.

Sessions leak and you get billed for idle browsers. Your finally block disconnects the client but does not end the session. Confirm your runtime's session timeout and set an explicit close call if the API exposes one.

Concurrent runs interfere. Session isolation is not configured, or you are reusing one CDP endpoint across parallel tasks. Each task needs its own session and its own endpoint.

The agent works locally and fails in CI. Environment drift. This is the case for hosted Chromium in one sentence: the browser is identical everywhere because it is not on your machine.

A Practical Migration Path

You do not have to rewrite the agent to move it. The sequence that works:

  1. Keep the Hermes repo as-is. Change only the browser acquisition step.
  2. Replace chromium.launch() with chromium.connectOverCDP(endpoint).
  3. Reuse the runtime's context instead of creating a new one.
  4. Add a session ID to your logs so failures are traceable.
  5. Run one task end-to-end and confirm the profile persists across two runs.
  6. Only then add concurrency.

Steps 1 through 5 take an afternoon. Step 6 is where the runtime choice actually matters, because that is when isolation, metering, and session limits stop being theoretical.

If you want the broader context on why this split exists, the post on remote browsers for AI agents covers the runtime layer in more depth, and remote browser online walks through running real Chromium without managing Chrome yourself.

Bottom Line

The Hermes browser GitHub ecosystem gives you a capable agent loop and a set of CLI and CDP clients to drive it. What it does not give you is a browser that persists, isolates, and scales. That is a runtime decision, and it is the one that determines whether your Hermes agent ships or stays in a notebook.

Connect over CDP, reuse the runtime's context, keep the profile, and meter the sessions. Everything else in the repo can stay exactly where it is.