BLOG
Hermes Browser GitHub: Run Agents on Hosted Chromium
Hermes Browser GitHub projects give you the agent loop. Here's how to wire them to a hosted Chromium runtime over CDP for production browser automation.
# Hermes Browser GitHub: Run Agents on Hosted Chromium
If you searched for Hermes browser GitHub, you are probably looking for one of two things: the source repository for a Hermes-branded browser agent, or a way to point an existing Hermes agent at a browser it does not have to run locally. This post answers both. It covers what the Hermes browser ecosystem on GitHub actually contains, why the repository is only half the runtime, and how to connect a Hermes-style agent to hosted Chromium over the Chrome DevTools Protocol (CDP) so it survives past your laptop.
The short version: GitHub gives you the agent logic. It does not give you a browser that stays alive across deploys, holds a logged-in profile, or scales to concurrent sessions. That is a runtime problem, and it is the part most Hermes browser setups get wrong.
What "Hermes Browser GitHub" Usually Refers To
Search results for this term tend to mix several distinct things. Before you clone anything, sort out which one you actually need:
- A Hermes agent repository. A repo containing the reasoning loop, tool definitions, and prompts that let a model decide when to click, type, or navigate. The browser is a dependency, not the product.
- A browser automation CLI. Tools like
agent-browser(Vercel Labs) wrap Playwright or Puppeteer behind a command-line interface so an agent can callopen,click, andsnapshotas shell commands. These are useful glue, but they still need a browser to talk to. - A CDP client. Code that connects to a running Chromium instance over a WebSocket endpoint. This is the layer that matters most for production, because it decouples your agent from the machine running the browser.
- A fork or wrapper. Many "Hermes browser" repos are thin wrappers around Playwright with a Hermes-specific prompt layer. Read the
package.jsonand the connection code before assuming they solve hosting.
The pattern across all four: the repository is the *control plane*. The browser is the *data plane*. If both live in the same process on the same machine, you have a demo, not a deployment.
Why the GitHub Repo Is Only Half the Runtime
A Hermes browser agent running from a cloned repo typically does this:
- Launches a local Chromium via Playwright's
chromium.launch(). - Drives it through the agent loop.
- Exits, taking the browser and its cookies with it.
That works until you hit any of the following, which are all normal in production:
- Session persistence. A login flow that requires a one-time code cannot complete inside a single agent run. You need a profile that survives between runs.
- Concurrency. Ten agents launching ten local Chromium instances on one box will exhaust memory and CPU long before they exhaust your task queue.
- Environment drift. The Chromium version, fonts, and system libraries on your CI runner differ from your dev machine. Selectors that worked locally fail in the pipeline.
- Observability. When an agent fails at step 7 of 12, you need a live view or a recording, not a stack trace.
- IP and network posture. Some sites behave differently depending on where the request originates. Local egress is a single point of failure.
None of these are solved by a better agent loop. They are solved by moving the browser out of the repo and into a runtime you connect to.
Hosted Chromium vs. Local Chromium for Hermes Agents
Here is the trade-off in concrete terms. This table assumes a Hermes-style agent that already works locally.
| Criterion | Local Chromium (from the repo) | Hosted Chromium (remote runtime) |
|---|---|---|
| Setup | npx playwright install per machine | One connection URL |
| Session persistence | Manual; lost on process exit | Persistent profiles across runs |
| Concurrency | Bounded by host RAM/CPU | Bounded by your plan; see /pricing |
| Environment consistency | Varies per machine and CI image | Same image every session |
| Debugging | Local headed mode or video | Live viewer plus session logs |
| Network posture | Your machine's IP | Configurable browser and proxy settings |
| Cost model | Your compute, always on | Metered browser time |
| Best for | Development, one-off scripts | Agents that run on a schedule or at scale |
The honest read: local Chromium is faster to start and free if you already have the hardware. Hosted Chromium costs money and adds a network hop. You move when persistence, concurrency, or consistency start costing you more than the runtime does.
Connecting a Hermes Agent to Hosted Chromium over CDP
The connection mechanism is the same one Playwright documents for connectOverCDP. Your runtime exposes a WebSocket endpoint; your agent connects to it instead of launching a browser. The Playwright CDP documentation covers the client side; the runtime covers everything behind the endpoint.
Here is a minimal TypeScript example. It assumes your runtime hands you a CDP URL, which is the standard shape for hosted Chromium providers.
import { chromium, Browser, BrowserContext, Page } from 'playwright';
// Your runtime returns a CDP WebSocket endpoint per session.
// Treat this like a credential: never commit it, never log it in full.
const CDP_ENDPOINT = process.env.REMOTE_BROWSER_CDP_URL!;
async function runHermesTask(task: string): Promise<void> {
let browser: Browser | null = null;
try {
// Connect to hosted Chromium instead of launching locally.
browser = await chromium.connectOverCDP(CDP_ENDPOINT, {
timeout: 30_000,
});
// Reuse the existing context so persistent profile state carries over.
const context: BrowserContext = browser.contexts()[0]
?? await browser.newContext({
viewport: { width: 1280, height: 800 },
locale: 'en-US',
});
const page: Page = context.pages()[0] ?? await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
// Hand control to your Hermes agent loop here. The agent decides
// the next action; Playwright executes it against the remote page.
await page.getByRole('button', { name: /sign in/i }).click();
await page.getByLabel('Email').fill('agent@example.com');
// Snapshot for the model. Keep it small; full HTML blows up context.
const snapshot = await page.accessibility.snapshot();
console.log(JSON.stringify(snapshot, null, 2));
await page.waitForLoadState('networkidle');
} catch (err) {
// Surface the session ID from your runtime so you can pull the recording.
console.error('Hermes task failed:', err);
throw err;
} finally {
// Disconnect the client. The hosted session lifecycle is managed
// by the runtime, not by this process.
await browser?.close();
}
}
runHermesTask('sign in and open the billing page').catch(() => process.exit(1));Three details matter more than the rest:
- `connectOverCDP` vs `launch`. You are not starting a browser. You are attaching to one that already exists. Anything you assumed about launch flags,
userDataDir, or executable paths no longer applies. - Context reuse.
browser.contexts()[0]returns the context the runtime created. Creating a new one each run throws away the persistent profile, which defeats the point. - `browser.close()` semantics. With
connectOverCDP, closing the client disconnects it. Whether the underlying session terminates depends on your runtime's session policy. Check that before you rely on it for cleanup.
If you are using Puppeteer instead, the equivalent is puppeteer.connect({ browserWSEndpoint }). Selenium connects through its remote WebDriver endpoint. The runtime should expose all three; if it only speaks one protocol, that is a constraint worth knowing before you commit.
What to Verify Before You Commit to a Runtime
The Hermes browser repo you cloned will run against almost any CDP endpoint. That is exactly why you should test the endpoint, not the agent. Production criteria, in rough order of how often they bite:
- Session isolation. Two concurrent agents must not share cookies, storage, or tabs. Ask how sessions are separated and confirm it in a test.
- Profile persistence. Can a session resume with the same logged-in state after a process restart? This is the single biggest gap between local and hosted.
- Live debugging. When a run fails, can you watch it or replay it? A live viewer plus session logs shortens debugging from hours to minutes.
- Protocol coverage. CDP is the baseline. Playwright, Puppeteer, and Selenium compatibility means you are not locked into one client library.
- Configurable browser settings. User agent, viewport, locale, timezone, and proxy configuration should be settable per session. Be skeptical of any claim beyond what the provider documents.
- Usage controls. You want per-session and per-account limits so a runaway agent cannot burn your budget overnight.
- Pricing transparency. Metered browser time is the norm. Confirm what counts as a browser-hour and where the meter starts. Current details live at /pricing.
The Remote Browser documentation covers how sessions, profiles, and CDP endpoints are exposed if you want to compare against your current setup.
Where Hermes Fits in a Production Stack
A Hermes browser agent is a decision loop. In production, that loop sits between three other layers:
- Orchestration. A queue or scheduler that decides which tasks run, when, and with what retry policy. This is where you enforce concurrency limits.
- The browser runtime. Hosted Chromium sessions with persistent profiles, isolated storage, and a CDP endpoint per session.
- Observability. Session recordings, live viewer access, and structured logs keyed by session ID so a failed run is reproducible.
The GitHub repo owns the first layer's logic and the agent's tool definitions. It should not own the second. When the browser is a connection string instead of a subprocess, you can redeploy the agent without losing state, run it from a serverless function, and scale it horizontally without rewriting the loop.
This is also why the "is Browserbase free" question keeps coming up alongside Hermes browser searches. Developers evaluating hosted Chromium want to know the entry cost before they refactor. The answer varies by provider and changes often, so check the provider's own pricing page rather than a blog post. The more useful question is whether the runtime supports persistent profiles and CDP, because a free tier without those will not carry a Hermes agent past the demo stage.
Common Failure Modes and Fixes
These are the ones that show up repeatedly when teams move a Hermes agent from local to hosted.
The agent connects but every selector times out. Usually a viewport or user-agent mismatch. The hosted browser renders a different layout than your local one. Pin the viewport explicitly in newContext and re-check your selectors against the hosted rendering.
Login state does not survive. You are creating a new context per run instead of reusing the runtime's context. Fix the contexts()[0] pattern shown above.
Sessions leak and you get billed for idle browsers. Your finally block disconnects the client but does not end the session. Confirm your runtime's session timeout and set an explicit close call if the API exposes one.
Concurrent runs interfere. Session isolation is not configured, or you are reusing one CDP endpoint across parallel tasks. Each task needs its own session and its own endpoint.
The agent works locally and fails in CI. Environment drift. This is the case for hosted Chromium in one sentence: the browser is identical everywhere because it is not on your machine.
A Practical Migration Path
You do not have to rewrite the agent to move it. The sequence that works:
- Keep the Hermes repo as-is. Change only the browser acquisition step.
- Replace
chromium.launch()withchromium.connectOverCDP(endpoint). - Reuse the runtime's context instead of creating a new one.
- Add a session ID to your logs so failures are traceable.
- Run one task end-to-end and confirm the profile persists across two runs.
- Only then add concurrency.
Steps 1 through 5 take an afternoon. Step 6 is where the runtime choice actually matters, because that is when isolation, metering, and session limits stop being theoretical.
If you want the broader context on why this split exists, the post on remote browsers for AI agents covers the runtime layer in more depth, and remote browser online walks through running real Chromium without managing Chrome yourself.
Bottom Line
The Hermes browser GitHub ecosystem gives you a capable agent loop and a set of CLI and CDP clients to drive it. What it does not give you is a browser that persists, isolates, and scales. That is a runtime decision, and it is the one that determines whether your Hermes agent ships or stays in a notebook.
Connect over CDP, reuse the runtime's context, keep the profile, and meter the sessions. Everything else in the repo can stay exactly where it is.