← Blog

BLOG

Cloud Browser for AI Agents GitHub: A Practical Guide

A cloud browser for AI agents GitHub workflow: connect hosted Chromium over CDP, run Playwright or Puppeteer, and skip local Chrome setup.

October 1, 20269 min readRemote Browser

# Cloud Browser for AI Agents GitHub: A Practical Guide

If you searched for a cloud browser for AI agents GitHub, you probably want one of two things: a repository you can clone and run, or a hosted runtime your existing agent code can connect to without shipping a browser. This guide covers both, but it focuses on the second because that is where most production teams end up. You keep your agent logic in GitHub, and you point it at a remote Chromium session instead of bundling Chrome into your container.

Remote Browser is a browser API and runtime for AI agents and browser-use workflows. It provides hosted Chromium sessions, CDP access, Playwright/Puppeteer/Selenium compatibility, a live viewer, persistent profiles, configurable browser settings, session isolation, and usage controls. You can read the full surface in the documentation.

What "cloud browser for AI agents GitHub" actually means

The phrase mixes three different things, and the confusion causes a lot of wasted setup time:

  1. A GitHub repo that runs a browser agent locally. You clone it, npm install, and it launches Chrome on your machine or CI runner.
  2. A GitHub repo that is a client for a hosted browser. The repo contains agent logic and a connection layer; the browser lives elsewhere.
  3. A hosted browser service with a GitHub presence. The service exposes an API or CDP endpoint, and the GitHub repo is the SDK or example code.

Most teams start with option 1, hit reliability problems, and migrate to option 2 or 3. The migration is usually small: you replace a local chromium.launch() call with a connectOverCDP() call against a remote endpoint. The agent code, prompts, and tool definitions stay in your repository.

If you want the conceptual background first, Remote Browser for AI agents explains why the runtime layer matters more than the agent framework.

Why local Chrome breaks agent workflows

Running Chromium next to your agent process works fine on a laptop. It fails in predictable ways once you deploy:

  • Cold starts. Every worker that needs a browser pays the launch cost. In serverless or autoscaled environments, that is repeated constantly.
  • Memory pressure. A headless Chromium instance is not small. Running several per node competes with your model calls and application code.
  • Version drift. Your container image pins one Chromium build; a teammate's machine has another. Selectors and behaviors diverge.
  • No shared state. Cookies, logins, and session storage vanish when the process restarts, so agents re-authenticate on every run.
  • Debugging blind spots. When a run fails in CI, you get a stack trace and maybe a screenshot. You cannot watch the page or attach a devtools session.
  • Proxy and network setup. Routing agent traffic through specific egress is a separate infrastructure problem you now own.

None of these are agent-logic problems. They are runtime problems, and they are the reason hosted browser services exist.

What a hosted cloud browser provides

A hosted runtime moves the browser out of your process and behind a connection URL. The practical differences:

ConcernLocal Chrome in your repoHosted cloud browser
Browser installYou install and pin ChromiumManaged by the provider
Cold startPer worker, per launchSession created on demand
ScalingYou size nodes for peak browser loadSessions requested per task
Session stateLost on restart unless you persist itPersistent profiles available
DebuggingLogs and screenshots onlyLive viewer plus CDP access
Proxy/egressYou configure and maintain itConfigurable browser settings
Framework supportWhatever you installPlaywright, Puppeteer, Selenium, raw CDP
Cost modelCompute you provisionUsage-based; see pricing

The trade-off is real: you give up some control over the exact browser build and you add a network hop. In exchange you stop operating browser infrastructure. For most agent workloads, that is the right trade.

Connecting your GitHub-hosted agent to a remote browser

The connection pattern is the same across frameworks because they all speak CDP underneath. Playwright exposes chromium.connectOverCDP(); Puppeteer uses puppeteer.connect({ browserWSEndpoint }); Selenium uses a remote WebDriver endpoint. The Chrome DevTools Protocol is the common substrate.

Here is a TypeScript example using Playwright against a hosted Chromium session. It assumes your agent code lives in a GitHub repo and reads the connection URL from an environment variable, which is the pattern you want in CI and production:

import { chromium, Browser, Page } from 'playwright';

const CDP_URL = process.env.REMOTE_BROWSER_CDP_URL;

if (!CDP_URL) {
  throw new Error('REMOTE_BROWSER_CDP_URL is not set');
}

async function runAgentTask(task: string): Promise<string> {
  let browser: Browser | undefined;

  try {
    // Connect to the hosted Chromium session over CDP.
    browser = await chromium.connectOverCDP(CDP_URL);

    // Reuse the existing context so persistent profile state applies.
    const context = browser.contexts()[0] ?? await browser.newContext();
    const page: Page = context.pages()[0] ?? await context.newPage();

    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });

    // Your agent loop goes here: observe, decide, act.
    const title = await page.title();
    console.log(`Task "${task}" landed on: ${title}`);

    return title;
  } finally {
    // Disconnect without killing the remote session if you want to resume later.
    await browser?.close();
  }
}

runAgentTask('check pricing page').catch((err) => {
  console.error(err);
  process.exit(1);
});

Three details matter here:

  • `connectOverCDP` does not launch a browser. It attaches to one that already exists. If the endpoint is wrong, you get a connection error, not a new Chrome window.
  • `browser.close()` semantics differ from `launch()`. With a remote connection, closing may terminate the session or just disconnect, depending on how the runtime is configured. Check your provider's behavior before relying on session reuse.
  • Contexts and pages may already exist. A persistent profile can hand you a context with cookies and storage already populated. Reusing contexts()[0] preserves that; calling newContext() throws it away.

For the CDP-specific details, including how the endpoint is structured, see Agent-Browser CDP.

Choosing between a repo and a runtime

If you are evaluating options from GitHub, sort them by what they actually give you:

Self-hosted browser automation repos. These give you full control and no per-session cost. You pay in operational work: container images, scaling, proxy management, session cleanup, and debugging infrastructure. Reasonable if you already run browser infrastructure at scale.

Agent framework repos. These contain the reasoning loop, tool definitions, and prompt logic. They are framework-agnostic about where the browser lives, which is exactly why they pair well with a hosted runtime.

Hosted browser SDKs and examples. These are thin clients. The value is the runtime behind them, not the repo contents.

A common production setup combines all three: an agent framework from GitHub, your own orchestration code in your repo, and a hosted browser for the actual page interaction. That keeps your repository focused on agent behavior instead of browser plumbing.

Production criteria to check before you commit

Before you wire a hosted browser into your agent, verify these:

  • CDP compatibility. Does it expose a standard CDP endpoint, or a proprietary protocol? Standard CDP means your existing Playwright and Puppeteer code works unchanged.
  • Session lifecycle. Can you create, reuse, and explicitly terminate sessions? Can you resume a session after a worker restart?
  • Persistent profiles. Do cookies and storage survive across sessions? This is the difference between an agent that logs in once and one that logs in every run.
  • Isolation. Are sessions isolated from each other? Shared state between concurrent agent runs is a correctness bug waiting to happen.
  • Observability. Is there a live viewer? Can you attach devtools mid-run? Debugging a failed agent run without visual access is slow.
  • Network controls. Can you configure egress, proxies, and browser settings per session? Many agent tasks fail on network reputation, not on logic.
  • Usage controls. Are there limits, quotas, or spend caps you can set? See pricing for how Remote Browser meters usage.
  • Framework coverage. Playwright, Puppeteer, and Selenium all need to work, because your team will not standardize on one.

If a provider cannot answer these concretely, treat it as a prototype tool rather than production infrastructure.

Common mistakes when moving agents to a cloud browser

Launching instead of connecting. The most frequent error is leaving chromium.launch() in the code and expecting it to use the remote browser. It will not. You need connectOverCDP() or the equivalent.

Ignoring session cleanup. Remote sessions cost money and consume capacity while they exist. If your agent crashes without closing the session, you accumulate orphaned browsers. Wrap session creation in try/finally and terminate explicitly.

Assuming a fresh context. Persistent profiles are a feature, but they surprise people who expect a blank slate. Decide per task whether you want continuity or isolation, and select the context accordingly.

Hardcoding the endpoint. Connection URLs often include session identifiers and expire. Read them from environment variables or fetch them per run from an API call.

Skipping the live viewer during development. Watching the agent drive the page catches selector and timing bugs that logs never show. Use the viewer while building, then rely on logs and recordings in production.

Treating the browser as stateless. Agents that need to log in, navigate multi-step flows, or maintain a cart depend on session state. Design for it explicitly rather than hoping the browser remembers.

Where this fits in a real agent stack

A production agent stack typically looks like this:

  • Orchestration layer — your code, in your GitHub repo. Decides what to do next.
  • Model layer — the LLM that produces actions or tool calls.
  • Tool layer — the functions the agent can call, including browser actions.
  • Browser runtime — hosted Chromium, connected over CDP.
  • State layer — persistent profiles, session storage, and any external memory.

The browser runtime is the layer most teams under-invest in, because it looks like a solved problem until it is not. Hosted Chromium turns it into a connection string instead of a maintenance burden.

If you want to see what running a browser without local setup looks like end to end, Remote Browser online walks through the flow. For the broader picture of driving a browser from anywhere, remote control browser covers the control-plane side.

Getting started

The shortest path from a GitHub repo to a working cloud browser agent:

  1. Keep your agent logic where it is. Do not rewrite it.
  2. Replace the local browser launch with a CDP connection to a hosted session.
  3. Move the connection URL into an environment variable.
  4. Add explicit session termination in a finally block.
  5. Test with the live viewer open so you can see what the agent sees.
  6. Check pricing to understand how session time is metered before you scale.

The repository stays yours. The browser becomes infrastructure. That separation is what makes agent code portable across environments, and it is the main reason teams move off local Chrome in the first place.