BLOG
Puppeteer BrowserWSEndpoint Remote Browser Setup
Puppeteer browserWSEndpoint remote browser setup: get the WebSocket URL, connect over CDP, and harden the session for production.
# Puppeteer BrowserWSEndpoint Remote Browser Setup
A Puppeteer browserWSEndpoint is the WebSocket URL that puppeteer.connect() uses to attach to an already-running Chromium instance instead of launching one locally. In a remote browser setup, that endpoint points at a hosted Chromium session, so your script never downloads a browser binary, never manages a display server, and never leaves a zombie process behind. This guide covers how to obtain the endpoint, wire it into Puppeteer, and what to check before you run it in production.
The short version: you get a ws:// or wss:// URL from your browser provider, pass it to puppeteer.connect({ browserWSEndpoint }), and drive the remote session with the same Page and ElementHandle APIs you already use. The interesting parts are everything around that call — session lifecycle, reconnection, target management, and the failure modes that only show up under load.
What browserWSEndpoint actually is
Puppeteer talks to Chromium over the Chrome DevTools Protocol (CDP). When you call puppeteer.launch(), Puppeteer spawns a local Chromium process, reads the DevTools WebSocket URL from its stderr, and connects to it. When you call puppeteer.connect(), you skip the spawn step and supply that URL yourself.
That URL is the browserWSEndpoint. It looks like this:
ws://127.0.0.1:9222/devtools/browser/6b1f3c2a-...The /devtools/browser/<id> path is the browser-level target. Connecting there gives you a Browser object, from which you create pages, attach to targets, and manage contexts. This is distinct from a page-level WebSocket URL (/devtools/page/<id>), which only controls a single tab. For automation you almost always want the browser-level endpoint.
The Chrome DevTools Protocol is the underlying spec. Puppeteer is a client library over it, which is why any CDP-speaking runtime — including a hosted one — works with the same API surface.
Why point browserWSEndpoint at a remote browser
Running Chromium locally works until it doesn't. The failure modes are predictable:
- Binary drift. The Chromium version Puppeteer downloads changes between installs, and a page that rendered fine last month breaks after a
npm ci. - Resource contention. Headless Chromium is memory-hungry. Ten concurrent sessions on a 2 GB CI runner will OOM.
- No persistence. Close the process and your cookies, localStorage, and logged-in state are gone.
- No observability. When a job fails at 3 a.m., you have a stack trace and nothing else.
A hosted runtime moves the browser off your machine. You keep the Puppeteer API; the provider handles the process, the version pinning, and the network egress. The trade-off is a network hop between your code and the browser, which adds latency to every CDP command. For most automation that's tens of milliseconds and irrelevant. For tight interaction loops — many page.evaluate calls per second — it matters, and you should measure it rather than assume.
If you're evaluating whether this trade is worth it, the remote browser overview covers the runtime-layer argument in more depth.
Getting the endpoint from a hosted session
Every provider exposes the endpoint differently, but the shape is consistent: you create a session via an API call, and the response includes a WebSocket URL plus a session ID you'll need for cleanup.
With Remote Browser, a session create call returns a connection URL you can pass straight to Puppeteer. The important fields are the WebSocket URL, the session identifier, and any expiry. Treat the URL as a secret — anyone holding it can drive the browser.
A minimal flow:
POSTto create a session, with options for viewport, proxy, and profile.- Read
wsEndpoint(or equivalent) from the response. puppeteer.connect({ browserWSEndpoint }).- Do work.
- Close the browser connection and explicitly terminate the session.
Step 5 is the one people skip. Closing the Puppeteer connection does not necessarily end the remote session — depending on the provider, the session may keep running and keep billing until it times out. Always call the terminate endpoint in a finally block.
Connecting Puppeteer to a remote browser
Here's a working TypeScript example using Playwright's CDP path, which is the more common choice when you want cross-browser support and built-in tracing. Puppeteer's connect is nearly identical in shape.
import { chromium } from 'playwright';
const API_BASE = process.env.BROWSER_API_BASE!;
const API_KEY = process.env.BROWSER_API_KEY!;
async function withRemoteBrowser<T>(
fn: (page: import('playwright').Page) => Promise<T>
): Promise<T> {
// 1. Create a hosted session.
const createRes = await fetch(`${API_BASE}/sessions`, {
method: 'POST',
headers: {
'Authorization': `Bearer ${API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
viewport: { width: 1280, height: 800 },
// Provider-specific: proxy, profile id, timeouts.
}),
});
if (!createRes.ok) {
throw new Error(`session create failed: ${createRes.status}`);
}
const session = await createRes.json() as {
id: string;
wsEndpoint: string;
};
// 2. Attach over CDP using the browser WebSocket endpoint.
const browser = await chromium.connectOverCDP(session.wsEndpoint, {
timeout: 30_000,
});
try {
const context = browser.contexts()[0] ?? await browser.newContext();
const page = context.pages()[0] ?? await context.newPage();
return await fn(page);
} finally {
// 3. Always release both the connection and the session.
await browser.close().catch(() => {});
await fetch(`${API_BASE}/sessions/${session.id}`, {
method: 'DELETE',
headers: { 'Authorization': `Bearer ${API_KEY}` },
}).catch(() => {});
}
}
// Usage
await withRemoteBrowser(async (page) => {
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.title());
});The Puppeteer equivalent swaps two lines:
import puppeteer from 'puppeteer';
const browser = await puppeteer.connect({
browserWSEndpoint: session.wsEndpoint,
defaultViewport: null, // respect the remote session's viewport
});Two details worth noting. defaultViewport: null prevents Puppeteer from overriding the remote viewport, which matters if you configured a specific size at session creation. And connectOverCDP in Playwright only supports Chromium-family browsers — that's a protocol constraint, not a provider one. See the Playwright CDP documentation for the exact semantics.
Puppeteer connect vs launch: what changes
| Concern | puppeteer.launch() (local) | puppeteer.connect() (remote) |
|---|---|---|
| Browser binary | Downloaded per install | Managed by provider |
| Version pinning | Your package.json + lockfile | Provider's image |
| Startup latency | 300 ms–2 s | Network round trip + session boot |
| State persistence | None by default | Profiles, if configured |
| Concurrency ceiling | Your machine's RAM | Provider quota — check pricing |
| Debugging | Local DevTools | Live viewer, if the provider offers one |
| Failure isolation | One crash kills the host | Session-scoped |
| Cost model | Compute you already pay for | Per browser-hour |
The cost row is the one that surprises teams. Local Chromium feels free because the compute is already provisioned. Remote sessions are metered, so a leaked session is a real line item. Budget for it by enforcing session timeouts on your side, not just the provider's.
Production criteria for a remote browserWSEndpoint
Before you commit to a provider, check these against your workload:
Endpoint stability during a session. Some providers rotate the WebSocket URL on reconnect. If yours does, your reconnect logic needs to re-fetch the endpoint rather than retry the stale one.
Target and context semantics. Puppeteer's browser.targets() and browser.contexts() behave differently depending on whether the remote instance is a fresh browser or a shared one. Confirm you get an isolated browser context per session, not a shared one — otherwise cookies leak between jobs.
Profile persistence. If you need logged-in state across runs, the provider must support persistent profiles keyed by an ID you control. Without this, every run starts cold and you'll burn time on re-authentication.
Proxy and network controls. Residential or datacenter egress, geo-selection, and whether the proxy applies to the browser or the whole session. These are configurable browser settings, not magic — verify what's actually applied rather than assuming.
Observability. A live viewer is the difference between "the job failed" and "the job failed because the selector matched a cookie banner." If you can't watch a session in real time, debugging gets expensive.
Cleanup guarantees. What happens when your worker dies mid-session? The provider should reap orphaned sessions on a timeout. Ask what that timeout is.
The documentation covers the session lifecycle and the specific configuration knobs available.
Common failure modes
`Protocol error (Target.setAutoAttach): Target closed` — usually means the session expired or was terminated server-side while your script was still running. Check session TTL against your longest job.
Connection hangs on `connect()` — the endpoint is reachable but the browser isn't ready. Add an explicit timeout and retry with backoff; don't let it hang indefinitely.
Pages appear in an unexpected context — you attached to a shared browser rather than an isolated one. Verify isolation at session creation.
Slow commands, not slow pages — if page.evaluate round trips are the bottleneck, you're paying the network hop per call. Batch DOM reads into a single evaluate that returns an object.
Sessions that never die — your finally block isn't running, typically because the process was killed. Use a session TTL as a backstop.
When a remote endpoint is the wrong choice
If your automation is a one-off script, a local puppeteer.launch() is simpler and free. If you need sub-10 ms command latency, a network hop will hurt. If your workload is a single long-lived browser you keep open for days, a persistent VM may be cheaper than metered sessions.
The remote endpoint earns its place when you need concurrency, isolation, persistence, or observability — the things that are painful to build yourself. For a broader comparison of the runtime options, see remote web browser and remote control browser.
Wiring it up
The mechanics are small: get a WebSocket URL, pass it to puppeteer.connect(), do work, terminate the session. The engineering is in the lifecycle — timeouts, cleanup, isolation, and knowing what your provider actually guarantees.
Start by running one session end to end and watching it in the viewer. Then run ten concurrently and watch your error rate. The gap between those two numbers tells you more about a provider than any benchmark.