← Blog

BLOG

Hosted Chromium GitHub: What to Check Before You Build

Hosted Chromium GitHub repos compared: what open-source runtimes give you, what they don't, and how to connect to a managed CDP endpoint.

October 7, 20269 min readRemote Browser

# Hosted Chromium GitHub: What to Check Before You Build

Searching for "hosted chromium github" usually means one of two things. Either you want to self-host Chromium and are looking for a repo to clone, or you want a managed Chromium endpoint and are checking whether the vendor's SDK is open source before you commit. Both are reasonable instincts, and both lead to the same question: what does the GitHub repository actually give you, and what does it leave for you to run?

This guide covers the open-source landscape, the specific things to verify in any repo, and how a hosted Chromium runtime fits when you'd rather not operate the browser fleet yourself. If you already know you want a managed endpoint, the remote browser for AI agents overview is a faster starting point.

What "hosted Chromium on GitHub" actually refers to

The phrase gets used for three distinct things, and conflating them causes most of the confusion.

1. Self-hosted browser infrastructure repos. Projects like Browserless, Steel, and various Playwright-in-Docker setups ship a server you run yourself. The GitHub repo is the product. You clone it, deploy it to your own cloud, and manage the Chromium fleet, scaling, and upgrades.

2. Client SDKs for a managed service. Some hosted browser vendors publish their SDK, CLI, or examples on GitHub while the actual browser runtime stays on their infrastructure. The repo is a thin client; the Chromium lives elsewhere.

3. Agent frameworks that assume a browser endpoint. Browser-use, Stagehand, and similar libraries are open source and browser-agnostic. They need a CDP endpoint — local or remote — and the repo tells you how to point them at one.

Knowing which category you're looking at determines what you're actually evaluating. A self-hosted server repo is an infrastructure commitment. A client SDK is an integration surface. An agent framework is a consumer of whichever runtime you choose.

What to verify in any hosted Chromium repo

Before you clone anything, check these properties. They separate a demo from something you can run in production.

  • CDP compatibility. Does it expose a real Chrome DevTools Protocol endpoint, or a proprietary API that only its own SDK speaks? CDP compatibility means Playwright, Puppeteer, and Selenium can all connect without vendor-specific glue.
  • Session isolation model. Are browser contexts isolated per session, or does everything share one browser process? Shared processes leak cookies, storage, and sometimes memory between tenants.
  • Profile persistence. Can a session keep cookies, localStorage, and login state across runs, or does every session start cold?
  • Proxy and network configuration. Can you route traffic through your own proxy, or set per-session network settings? This matters for geo-restricted content and for keeping your own egress IPs clean.
  • Resource limits and cleanup. What happens when a session is abandoned? Does the container get reaped, or does it leak until the host runs out of memory?
  • License. AGPL, SSPL, and permissive licenses have very different implications if you plan to embed the software in a commercial product.

That last point trips up more teams than any technical issue. A repo that looks perfect on the README can be unusable if its license doesn't fit your distribution model.

Open-source options and where they stop

The self-hosted category is mature. Here's a rough comparison of what you get from the common approaches.

ApproachWhat the repo gives youWhat you still operateBest fit
Browserless (self-hosted)Chromium server, REST + CDP endpoints, Docker imageHost, scaling, session reaping, upgradesTeams with existing container infra
Steel (self-hosted)Browser API, session management, CDPHost, storage, proxy config, monitoringTeams wanting a lighter server
Playwright in DockerA container running Chromium with CDPEverything: orchestration, isolation, cleanupPrototypes and CI jobs
Agent framework (browser-use, Stagehand)Agent logic, tool definitionsThe entire browser runtimeAny team, once a runtime exists
Managed hosted ChromiumA CDP URL and a session APINothing on the browser sideTeams shipping agents, not infra

The pattern is consistent: the open-source repos give you the browser, and you give them your operational attention. That's a fair trade if browser infrastructure is a core competency or if your volume is low enough that a single VM handles it. It stops being a fair trade when you're running many concurrent sessions, when sessions need to survive worker restarts, or when a stuck Chromium process takes down your whole automation pipeline at 3 a.m.

There's also a subtler cost. Self-hosted Chromium is a moving target. Chrome ships a major version roughly every four weeks, and each release changes fingerprinting surfaces, deprecates APIs, and occasionally breaks CDP methods that automation depends on. If you self-host, you own that upgrade treadmill. If you don't upgrade, you drift toward detection and compatibility problems.

Connecting to a hosted Chromium endpoint

The integration story is the same whether you self-host or use a managed runtime, because both speak CDP. Playwright's connectOverCDP method takes a WebSocket endpoint and returns a browser object you drive exactly like a locally launched one.

import { chromium, Browser, BrowserContext, Page } from 'playwright';

interface SessionInfo {
  cdpUrl: string;
  sessionId: string;
}

// Your runtime returns a CDP WebSocket URL for a fresh, isolated session.
async function createSession(): Promise<SessionInfo> {
  const res = await fetch('https://api.remote-browser.dev/v1/sessions', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.REMOTE_BROWSER_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      // Configurable browser settings: region, proxy, viewport.
      proxy: { country: 'us' },
      viewport: { width: 1440, height: 900 },
    }),
  });

  if (!res.ok) throw new Error(`Session create failed: ${res.status}`);
  return res.json() as Promise<SessionInfo>;
}

async function runTask(): Promise<void> {
  const { cdpUrl, sessionId } = await createSession();
  let browser: Browser | undefined;

  try {
    browser = await chromium.connectOverCDP(cdpUrl);
    const context: BrowserContext = browser.contexts()[0];
    const page: Page = context.pages()[0] ?? (await context.newPage());

    await page.goto('https://example.com/dashboard', {
      waitUntil: 'domcontentloaded',
    });
    await page.getByRole('button', { name: 'Export' }).click();
    await page.waitForEvent('download');
  } finally {
    // Closing the connection does not necessarily end the remote session.
    // Release it explicitly so the runtime can reap the container.
    await browser?.close();
    await fetch(`https://api.remote-browser.dev/v1/sessions/${sessionId}`, {
      method: 'DELETE',
      headers: { 'Authorization': `Bearer ${process.env.REMOTE_BROWSER_KEY}` },
    });
  }
}

runTask().catch((err) => {
  console.error(err);
  process.exit(1);
});

Two details matter here and are easy to get wrong. First, browser.close() on a CDP connection disconnects your client; it does not always terminate the remote browser. You need an explicit session-release call, or you'll accumulate orphaned sessions and pay for idle browser time. Second, browser.contexts()[0] assumes the runtime hands you a session with a context already created. Some runtimes return a bare browser with no contexts, in which case you call newContext() yourself. Check the runtime's docs rather than assuming.

The same endpoint works from Puppeteer via puppeteer.connect({ browserWSEndpoint }) and from Selenium through its CDP bridge. That portability is the main argument for insisting on real CDP rather than a vendor SDK.

Production criteria for choosing between them

Once you've confirmed CDP compatibility, the decision comes down to operational fit. These are the questions worth answering before you commit.

How many concurrent sessions do you actually need? A single 4-vCPU VM can handle a handful of Chromium instances before memory pressure starts killing them. If your peak is ten sessions, self-hosting is fine. If it's two hundred, you're now running a scheduling problem, and that's a different job than the one you were hired to do.

Do sessions need to outlive your workers? Serverless functions and short-lived containers lose their local browser state on every invocation. A hosted runtime decouples session lifetime from compute lifetime, which is what makes persistent profiles and resumable agent runs possible. See remote browser online for how that changes agent architecture.

How much do you care about IP reputation? Datacenter IPs from your own cloud get blocked by a predictable set of sites. Managed runtimes typically offer proxy configuration, sometimes with residential options. Self-hosting means sourcing and rotating proxies yourself.

What's your tolerance for the upgrade treadmill? Every Chrome release is a potential breaking change. A managed runtime absorbs that; you just reconnect to a new session.

What does your compliance posture require? If you need SOC 2, a signed DPA, or zero data retention, check whether the vendor offers it. Self-hosting gives you control but also puts the entire compliance burden on your team.

Pricing is the last variable, not the first. Browser time is usually metered per session-hour, and the difference between vendors is often smaller than the difference between running sessions efficiently and running them wastefully. Check /pricing for current rates rather than planning around numbers from a blog post — including this one.

When self-hosting is the right call

It would be dishonest to frame managed runtimes as universally better. Self-hosting wins in specific situations.

  • Your volume is low and predictable. A single VM running a few sessions a day costs less than any metered service.
  • Your data cannot leave your network. Regulated environments sometimes prohibit sending page content to a third party, and no amount of vendor certification changes that.
  • You need deep customization. Patching Chromium, injecting custom extensions, or running a modified build is easier when you control the binary.
  • Browser infrastructure is your product. If you're building a browser automation platform, you should own the runtime.

For everyone else — teams whose product is an agent, a scraper, or a QA pipeline — the browser is a dependency, not a differentiator. Treating it like one is usually the right call.

Practical migration path

If you're moving from a self-hosted repo to a managed endpoint, the change is smaller than it looks because CDP is CDP.

  1. Abstract session creation. Put your browser acquisition behind a single function that returns a CDP URL. Swapping runtimes then becomes a config change.
  2. Stop launching Chromium locally. Replace chromium.launch() with chromium.connectOverCDP(). Keep the rest of your Playwright code identical.
  3. Add explicit session release. Wrap your work in try/finally and delete the session on the way out. This is the single biggest source of wasted browser-hours.
  4. Move profile state to the runtime. If you were persisting cookies to disk, switch to the runtime's profile feature so state survives across workers.
  5. Instrument session duration. You can't optimize browser-hour spend without knowing which tasks hold sessions open longest. Most waste is a task that hangs on a selector that no longer exists.

The remote control browser guide covers the debugging side of this — attaching a live viewer to a running session so you can see what the agent sees when a task stalls.

The short version

A hosted Chromium GitHub repo is worth evaluating on CDP compatibility, session isolation, profile persistence, and license before anything else. Self-hosted options are genuinely good and genuinely yours to operate. Managed runtimes trade that operational load for metered billing and someone else's upgrade treadmill.

The deciding question isn't which is more powerful. It's whether browser infrastructure is the thing your team should be spending its engineering hours on. If it isn't, connect to an endpoint and get back to the agent. The documentation covers the connection details, and remote web browser walks through the runtime model in more depth.