← Blog

BLOG

Browser AI Agent GitHub: Open-Source Runtimes Compared

Browser AI agent GitHub projects compared: browser-use, agent-browser, and hosted runtimes. Learn what to self-host and when to connect a remote browser.

October 2, 20269 min readRemote Browser

# Browser AI Agent GitHub: Open-Source Runtimes Compared

Searching for a browser ai agent github repo usually means one of two things: you want to run an agent locally with full control, or you want to understand what the open-source ecosystem actually provides before committing to infrastructure. Most GitHub projects in this space solve the agent logic problem — planning, tool calling, DOM extraction — and leave the browser runtime problem to you. That split matters more than any benchmark table.

This guide maps the real open-source options, explains what each one gives you, and shows where a hosted runtime fits when local Chromium stops being practical.

What "Browser AI Agent" Repos Actually Ship

Open-source browser agent projects cluster into three layers. Knowing which layer a repo occupies tells you what you still have to build.

Agent frameworks — browser-use, AgentGPT-style planners, and similar projects. These ship the reasoning loop: take a task, observe the page, decide an action, execute it. They depend on a browser driver underneath.

Browser drivers and CLIs — Playwright, Puppeteer, and newer agent-specific CLIs like agent-browser. These ship the control surface: navigation, clicking, typing, screenshotting, and CDP access. They do not ship the agent.

Runtimes — the layer that actually hosts Chromium. This is where most GitHub projects are thin. A repo that says "just run npx" is assuming you have a machine, a display server (or headless config), a stable IP, and enough memory to hold N browser processes.

The gap is deliberate. Agent frameworks are easy to open-source because the interesting logic is portable. Runtime infrastructure is expensive to maintain, so it usually becomes a hosted product or stays as a Dockerfile you have to operate yourself.

The Main Open-Source Options

browser-use

The most-referenced project when people search for browser AI agents on GitHub. It provides a Python agent that drives a browser via Playwright, with an LLM deciding actions from an accessibility-tree-style representation of the page.

What you get: a working agent loop, model-agnostic design, and a large community of forks and examples.

What you still own: the browser. By default it launches local Chromium. In production you either run it on a VM with a persistent browser process or point it at a remote CDP endpoint. The second path is where hosted runtimes enter — see Remote Browser for AI agents for how that connection works.

agent-browser (Vercel Labs)

A CLI-oriented project that wraps browser automation for agents. It is closer to the driver layer than the agent layer: you get commands to drive a browser, and your agent (or your shell script) issues them.

This is useful when you want the agent logic in your own codebase and only need a reliable way to execute browser actions. It still needs a browser to talk to.

Playwright and Puppeteer

Not agent projects, but the substrate nearly every agent project sits on. Both support connecting to an existing browser over CDP rather than launching one locally. That single capability is what makes remote runtimes possible.

If you are evaluating any GitHub agent repo, check whether it exposes a way to pass a CDP endpoint or a browserWSEndpoint. If it does, you can run the agent anywhere and the browser somewhere else.

Self-Hosted vs Hosted Runtime: The Real Trade-off

The GitHub-first instinct is to self-host everything. That works until it doesn't. Here is the honest comparison.

DimensionSelf-hosted ChromiumHosted runtime (Remote Browser)
SetupDocker, deps, display/headless configConnect a CDP URL
ScalingYou manage process pools and memorySessions provisioned per request
Session persistenceManual profile dirs, volume mountsPersistent profiles built in
IP / proxy handlingYou configure and pay for proxiesConfigurable browser settings
DebuggingVNC or screenshots you wire upLive viewer per session
Cost modelFixed VM cost, idle or notUsage-based, see /pricing
MaintenanceOS patches, Chromium updates, crashesHandled by the runtime

The self-hosted column is not a strawman. For a single developer running a handful of tasks a day, a local Chromium is genuinely fine. The calculus changes when you need concurrent sessions, stable identities across runs, or an agent that runs while your laptop is closed.

When self-hosting breaks down

Three failure modes show up repeatedly:

  1. Memory pressure. Each Chromium instance is heavy. Ten concurrent sessions on one VM will swap. You end up building a scheduler, which is infrastructure work, not agent work.
  2. Session loss. Agents that log in and resume later need persistent profiles. Local profile directories work until you scale past one machine.
  3. Detection and IP quality. Datacenter IPs from a cloud VM get blocked more often than residential ones. Solving this means proxy management, which is its own project.

None of these are unsolvable. They are just not the problem you set out to solve when you started building an agent.

Connecting an Agent to a Remote Browser

The connection pattern is the same whether you self-host or use a managed runtime: get a CDP endpoint, connect over it, drive the browser. Here is the Playwright TypeScript version.

import { chromium, Browser, Page } from 'playwright';

// The endpoint comes from your runtime. For a hosted runtime,
// create a session via the API and read the CDP URL from the response.
const CDP_ENDPOINT = process.env.BROWSER_CDP_URL!;

async function runAgentTask(task: string) {
  const browser: Browser = await chromium.connectOverCDP(CDP_ENDPOINT);

  // Reuse the existing context so persistent profile state is preserved.
  const context = browser.contexts()[0] ?? await browser.newContext();
  const page: Page = context.pages()[0] ?? await context.newPage();

  await page.goto('https://example.com/login', { waitUntil: 'domcontentloaded' });

  // Your agent loop goes here: observe, decide, act.
  await page.fill('#email', process.env.AGENT_EMAIL!);
  await page.fill('#password', process.env.AGENT_PASSWORD!);
  await page.click('button[type="submit"]');

  await page.waitForLoadState('networkidle');

  // Do not close the browser if you want the session to persist.
  // Disconnect instead; the runtime keeps the session alive.
  await browser.close();
}

runAgentTask('log in and check the dashboard');

Two details matter here. First, connectOverCDP attaches to an existing browser rather than launching one — this is the mechanism that lets agent code run in one place and the browser in another. Second, calling browser.close() on a CDP connection disconnects your client; whether the underlying session survives depends on the runtime. With a hosted runtime, sessions are managed independently of your process, so a crashed agent does not necessarily kill the browser. The Playwright CDP documentation covers the connection semantics in detail.

If you are wiring this up for the first time, the documentation walks through session creation and endpoint retrieval.

What to Check Before You Commit to a Repo

When you find a browser AI agent repo on GitHub, evaluate it against these criteria rather than star count.

Does it separate agent logic from browser control? If the agent code calls chromium.launch() directly, you will be patching it to run remotely. If it accepts a CDP endpoint or a browser factory, you can swap runtimes without touching the agent.

Does it handle session state? Agents that need to stay logged in require persistent profiles. Check whether the project stores cookies and local storage, and where.

Does it expose the CDP connection? Some frameworks hide the browser object entirely. That is convenient until you need to debug a selector or inspect network traffic.

What is the failure behavior? Does a navigation timeout crash the whole run, or does the agent retry? Production agents need retry logic and a way to inspect what went wrong. A live viewer is worth more than a log line here — see Remote control browser for how that works in practice.

Is the license compatible with your use? Most of these projects are permissive, but check before you build a product on top.

Where Hosted Runtimes Fit

A hosted runtime is not a replacement for the agent framework. It replaces the part of the stack you would otherwise operate: the Chromium processes, the profile storage, the proxy configuration, and the debugging surface.

Remote Browser provides hosted Chromium sessions with CDP access, so any framework that can connect over CDP works without modification. Playwright, Puppeteer, and Selenium clients all connect the same way. Sessions are isolated, profiles persist across runs, and browser settings are configurable rather than hardcoded. You can watch a session live while it runs, which turns "the agent failed somewhere" into "the agent failed on this selector."

The practical split looks like this:

  • Keep on GitHub / local: agent reasoning, tool definitions, prompt logic, evaluation harnesses.
  • Move to a runtime: browser hosting, session persistence, proxy and IP handling, concurrency, observability.

This is the same separation that made databases managed and CI hosted. The logic is yours; the substrate is rented.

For teams comparing options, the relevant question is not "open source or not" but "which layer do I want to own." If you want to own the runtime, the GitHub projects above give you a starting point and a Dockerfile. If you want to own only the agent, connect to a runtime and skip the infrastructure work. Current usage details are on the pricing page.

A Practical Migration Path

If you are running a local agent today and hitting limits, the migration is smaller than it looks.

  1. Abstract the browser connection. Replace direct chromium.launch() calls with a function that returns a connected browser. Read the endpoint from an environment variable.
  2. Test against a remote session. Create a session, connect over CDP, run your existing task. Most agent code works unchanged.
  3. Move profile state. If you relied on a local user data directory, switch to the runtime's persistent profiles so logins survive across sessions.
  4. Add observability. Use the live viewer during development. It shortens the debug loop more than any logging change.
  5. Scale by session, not by VM. Once the connection is abstracted, concurrency becomes a matter of how many sessions you request, not how many machines you provision.

The remote browser online guide covers the connection mechanics if you want the shorter version.

Summary

The browser AI agent GitHub ecosystem gives you strong agent frameworks and solid browser drivers. What it does not give you, in most cases, is a runtime — the hosted Chromium, persistent profiles, and session management that production agents need. You can build that yourself, and for small workloads you should. For anything with concurrency, persistence, or uptime requirements, connecting an existing open-source agent to a hosted runtime over CDP is the shorter path. The agent logic stays yours; the browser becomes infrastructure.