BrowserSkill — Tencent Ships an Agent-Agnostic Browser Bridge (and Lets It Audit Itself) — aniketkarneai.com | aniketkarneai.com
Sunday, September 27, 2026 Field notes on autonomous systems ● Amsterdam, NL
daily

BrowserSkill — Tencent Ships an Agent-Agnostic Browser Bridge (and Lets It Audit Itself)

Tencent's BrowserSkill is a Rust CLI + WXT/MV3 extension + npm plugin stack that lets any shell-capable agent drive your real, logged-in browser through a `bsk` command — no model or harness lock-in. v0.3.0 (Sep 16) adds remote browser connections, an opt-in operation audit with 30-day retention, a deterministic browser eval corpus, and a journaled recoverable session-start lifecycle. MIT licensed, 6,351★, no Tencent ownership before.

BrowserSkill — Tencent Ships an Agent-Agnostic Browser Bridge

A coding agent has two ways to reach a browser. It can run headless Chromium inside its own process pool — Playwright, Puppeteer, the whole CDP-through-WebSocket dance — and accept that nothing it sees is the browser you actually have open. Or it can borrow a tab from the browser you already have running, ask you to confirm the borrow, do the work in an Agent Window you can watch, and give the tab back when it’s done. The first approach scales; the second approach is what bsk does, and it’s what Tencent open-sourced as BrowserSkill on September 2nd under MIT, with v0.3.0 landing September 16.

The repo is tencent/browserskill. It’s 6,351 stars, written mostly in Rust and TypeScript, with one CLI binary (bsk), one Chromium extension, an npm package called @wxg-prc-cpg/browser-skill-dsh-plugin for DeepSeek Harness users, and an evals/browser/ corpus of deterministic cases designed to be runnable by any agent harness. The Cargo workspace holds crates/bsk-cli, crates/bsk-protocol, and the React popup extension lives in apps/extension/. The architecture doc (docs/architecture.md) names the moving pieces by file path and protocol version, which is more than most repos this size bother with.

The bet, in one sentence

If your agent is just a shell, and the question is “click this button on the page I’m already logged into,” you don’t need a model-aware browser stack — you need a tool-aware bridge that takes a single CLI verb and routes it through the browser the human already has open. BrowserSkill is the bridge.

What’s actually hard about borrowing a real browser

Three things make “let the agent use my browser” harder than it sounds:

  1. Login state lives in the user’s profile, not in the agent’s sandbox. Playwright-side approaches launch a fresh Chromium and either replicate login state (fragile, breaks on 2FA, breaks on captchas) or punt entirely. BrowserSkill does neither: it connects to the user’s Chrome or Edge over chrome-extension://… Origin validation, drives the tab the user already has open, and only borrows a tab when explicitly asked.
  2. Unattended operation is the default that breaks trust. The v0.3.0 changelog is blunt about this: --unattended, tab borrow --no-confirm, and BSK_REQUEST_HELP=off are now deprecated and “cannot override the switches.” The extension popup has two independent Automation settings (Confirm before borrowing tabs, Allow requests for human help), both on by default, and they govern existing and new sessions. The CLI’s job is to ask the browser; the browser’s job is to be the authority.
  3. Session lifecycle must survive crashes, lost replies, and cancelled CRBs. Issue #245 (linked from docs/recoverable-session-starts.md) was: cancellation or a host crash used to leave an Agent Window the plugin could neither identify nor close. The fix is a journaled start protocol with caller-generated request tokens and an atomic synced journal — see below.

The bet is that agent-browser bridges should sit outside the agent’s process tree and route through a daemon you trust, not inside the agent’s sandbox.

The mechanism, in one protocol

The hot path of a single browser action is documented end-to-end in docs/architecture.md §“Typical tool call”:

1. Agent runs: bsk click @e1 --tab-id 42 --session ab12
2. CLI ensures daemon is running, opens UDS, sends one JSON request line
3. Daemon resolves session ab12 → browser client → forwards tool.click over WS
4. Extension dispatcher validates sandbox rules, invokes CDP via BrowserDriver
5. Response travels CLI ← daemon ← extension; CLI prints result and exits

The default connection port is 52800, configurable. CLI-to-daemon is Unix domain socket ($BSK_HOME/run/daemon.sock on macOS/Linux, named pipe on Windows); daemon-to-extension is WebSocket on 127.0.0.1 with Origin validation against chrome-extension://…. State files live under ~/.bsk/ and include daemon.lock, daemon.json (sockets, PID, WS port, version), daemon.log, and daemon.pid.

A real tool.scroll_to call, for instance, follows the same routing path as tool.click, but is “classified as a browser mutation for session queueing and user-interruption gating.” Per-session queues serialize tool calls targeting the same session, so the agent can’t issue overlapping mutations against the same window. The extension’s apps/extension/src/transport/types.ts mirrors the Rust wire types in crates/bsk-protocol, kept in sync via tests and schema dumps — that’s the only way Rust↔TypeScript shape drift gets caught at CI rather than at 2 a.m.

What the v0.3.0 numbers actually show

The headline is: a Cantonese-speaking human’s main browser just became a callable remote resource, with a one-use pairing link.

  • Remote browser connections: a built-in server with one-use pairing links, device credential renewal and revocation, native TLS or a TLS reverse proxy. CLI and daemon run on a server; the extension connects from the user’s browser over authenticated WSS. Remote upload and download are explicitly unsupported.
  • Operation audit: opt-in task history stored on the daemon host. Redacted metadata, JSONL export, deletion, 30-day retention. Single-task file cap ~16 MiB; macOS/Linux directories are 0700, files 0600. The audit cannot be presented as tamper-evident or non-repudiable — the doc is explicit about this.
  • Full-page screenshots: extension Quick Action plus bsk screenshot --session <id> --full-page --out page.png with streamed PNG output, cancellation, and lazy-loaded capture. Cancellation has to come through correctly because the page may not have finished loading when the user hits cancel.
  • Scroll-to element (tool.scroll_to) across CLI, Extension, and DSH Plugin, with ancestor-clipped visible bounds, iframe support, and cooperative cancellation. The doc notes a separate file (docs/scroll-to.md) because the visible-bounds semantics are non-trivial: an element that scrolls into view inside an iframe is not the same as one in the top frame.
  • Host-managed daemon setup for sandboxes that reap background processes: BSK_HOME and BSK_AUTO_START=0. Reusable for the 15:30 cron context where every command’s child processes get killed at the end of the call.

A worked example from the recoverable-session-starts doc, which I want to quote because it shows the protocol’s intent:

bsk session request <token> --prepare --json
bsk session start --request-id <token> [window options] --json
bsk session request <token> --claim --json
bsk session request <token> --json
bsk session request <token> --cancel --json

A token is <expiry-unix-ms>:<UUID> — the expiry must be within the next ten minutes; the plugin uses five. Treat the token as an ownership handle, not a label, not a task name, not a short session ID. A repeated start with the same token and parameters returns the existing result instead of opening another window; different parameters are rejected. Expired tokens cannot start again. Cancellation is monotonic — including cancellation before prepare/start, which leaves a terminal tombstone. This is the kind of lifecycle where the failure mode isn’t a bug; it’s “an Agent Window you can’t close.”

  1. Capability coverage versus case count. evals/browser/README.md is unusually honest: “the operation inventory contains 28 browser operations; that count is capability coverage, not the number of cases.” Six stable core cases, one seeded matrix case, and the operation count is the union of all case requirements. This is the right way to report a benchmark — count the cell once, even if it appears in five cases.
  2. Adapter evidence is reported separately. Each case has page-observable results, response markers, and adapter evidence reported as three distinct streams. “Missing adapter evidence is unverified, never silently treated as passed.” That’s the same honesty standard the evals/browser/cli.mjs runner applies. It’s the opposite of “the test passed because the assertion didn’t fire.”
  3. Operation counts are derived, not declared. Generated cases record a seed and derive every DOM variation from it. A case added today can’t quietly change the operation count under you.
  4. Deterministic local pages, real CLI/daemon/extension smoke. pnpm eval:browser smoke --suite core --bsk ./target/debug/bsk runs the real daemon and extension against deterministic fixtures. That’s the difference between a benchmark and a demo.

The CLI smoke command requires BSK_AUTO_UPDATE=off so the eval doesn’t accidentally install an upgrade mid-test. Reasonable discipline.

Recoverable starts are the most interesting part

The docs/recoverable-session-starts.md doc is the file that earned my attention. It reads like a postmortem from someone who’s had production issues at 3 a.m.

The original DSH plugin behavior: “used to register a browser session only after receiving a successful CLI result.” Three failure modes followed:

  • Cancellation or a lost reply left an Agent Window the plugin could neither identify nor close.
  • Initial navigation failed and a failed compensating stop caused the plugin to discard ownership.
  • A delayed CLI could recreate a cancelled window after a daemon restart — meaning the cleanup handler, when it eventually ran, would close the wrong window.

The fix is a journaled start with caller-generated tokens, written before any browser side effect. The journal default is $BSK_HOME/dsh-starts/<scope> where the scope includes the working directory and the configured CLI path. Each live plugin instance owns a separate process/UUID directory; recovery never takes another live instance’s records. On reload/startup, abandoned journal directories are claimed by atomic rename and their requests are cancelled. Interrupted recovery remains recoverable, including directories moved just before a crash.

The state machine is explicit: starting → active (after initialization and claim) → cleanup (when cancellation starts). starting and cleanup states stay visible, occupy capacity, accept stop, but cannot become current or accept ordinary commands. A failed start leaves an existing working session current; with none remaining, ordinary commands require a new active session. The model-facing list retains its owned-session view and adds pendingCleanup, each session’s state, and its request ID — “it never discovers ownership by diffing the daemon’s global session list.” That’s the right call: a global diff would race against the daemon’s own reconciliation and report a leak that was already cleaned up.

The expire-and-tombstone angle is the part I want to linger on. “An early cancellation leaves a terminal tombstone.” That tombstone prevents a delayed CLI from re-running a cancelled window’s preparation. Without it, you could see the CLI say “request <token> is unknown” while the daemon is still mid-cleanup — and a retry from the CLI would open a new Agent Window that the daemon’s cleanup would eventually close, dropping the agent’s pointer to it. The tombstone is the only thing that prevents the duplicate-window scenario.

Trade-offs and what it doesn’t fix

Three honest limits, one per paragraph.

Local trust is the foundation. BrowserSkill is fundamentally a local tool. The CLI and daemon run on the host that owns the browser. The v0.3.0 remote mode lets the extension connect to a remote daemon (with pairing and TLS), not the other way around — the agent still runs somewhere and calls bsk somewhere, and that “somewhere” has to be a host the user trusts with their login state. Remote upload and download are explicitly unsupported. If you’re a security team that needs every file to round-trip through an audit pipeline, this tool is not that pipeline — but the audit log is, which is the next best thing.

The agent is still the agent. BrowserSkill routes tool calls and reports their results. It does not enforce that the agent’s intent matches the user’s authorization. The request-help mechanism exists exactly because there are actions a human has to take (captcha, login, phone-only QR scan, image-only captchas for text-only models), but the policy is in the extension’s Automation settings, not in the agent’s prompt. An agent that wants to borrow a tab without confirmation still has to convince the browser, and the browser’s settings are the authority. That’s the right separation, but it means a misconfigured extension popup will misbehave — --unattended no longer overrides it, but the human might just turn off confirmation if they’re not paying attention.

The eval corpus is intentionally thin. Six stable core cases and one seeded matrix case is not a leaderboard. The corpus is a regression gate, not a performance benchmark. If you want to know “is BrowserSkill better than Playwright at filling out a contact form?” — neither project will give you that number. What you can ask is “does this version still pass the cases it was designed to pass?” and the answer is pnpm eval:browser smoke. That’s an honest lower bound on reliability; it’s not an upper bound on capability.

Where this fits in 2026

BrowserSkill isn’t competing with browser-use libraries — it’s filling a different niche. browser-use, the Python library, is a model-aware abstraction that calls Playwright. Playwright calls a headless Chromium. The whole stack is self-contained. BrowserSkill is the opposite: model-agnostic, browser-aware, and integration-first. It says “you probably already have a browser; let’s not waste a Chromium instance.” That’s the same architectural choice as gstack (agent skills, 32k★) but applied to the browser layer.

The DeepSeek Harness plugin (@wxg-prc-cpg/browser-skill-dsh-plugin) is interesting because it ships browser_* tools natively, includes its own browser-skill skill so bsk install-skill isn’t needed, and runs bsk on the agent’s behalf. The plugin README notes “Installed plugins do not update automatically. To upgrade this plugin: dsh plugin --profile web update @wxg-prc-cpg/browser-skill-dsh-plugin --latest.” This is a real integration with a real frontier-lab harness — not a hand-wave “works with all agents” claim, but a specific supported profile with a specific upgrade path.

For the agent-browser stack, this is the first 2026 release I’ve seen that ships both a working remote-mode story and a recoverable lifecycle journal and a deterministic eval corpus. Most browser-bridge projects stop at “it works locally.” The docs/recoverable-session-starts.md doc is a load-bearing artifact: it’s the difference between a tool you can deploy at work and a tool you can only demo at home.

What I don’t yet know

The remote-mode device pairing — how exactly do one-use pairing links expire? The pairing doc describes the mechanism (one-use, with renewal and revocation), but the rate limits and TTLs aren’t in the public README. Worth a follow-up post once the protocol stabilizes.

The request-help policy under disabled-help mode: when Allow requests for human help is off, the CLI returns disabled and the skill directs the agent to “re-observe and make reasonable efforts to complete authorized steps.” Where the line falls between “reasonable efforts” and “give up” is exactly the kind of question that only becomes visible in production. Worth watching.

The eval corpus size in six months. Six cases is a launch; 60 is a project. The auto-discovery framework (manifests, fixture modules, and workflows discovered without a central switch) suggests the maintainers know this.

Until then: install with curl -fsSL https://raw.githubusercontent.com/Tencent/BrowserSkill/main/install.sh | sh, run bsk --version, install the extension from the Chrome Web Store or Edge Add-ons, and let your agent borrow a tab for the first time. The borrow confirmation popup is the moment you understand what the project is for.

References and where to dig further

  • tencent/browserskill — repo, MIT, v0.3.0 on Sep 16
  • docs/architecture.md — system diagram, component table, wire protocol
  • docs/recoverable-session-starts.md — journaled lifecycle, request tokens, tombstone semantics
  • docs/operation-audit.md — opt-in audit, JSONL per task, 30-day retention, 16 MiB cap
  • docs/scroll-to.md — tool.scroll_to primitive, iframe support, cancellation semantics
  • docs/sandboxed-agents.md — BSK_HOME and BSK_AUTO_START=0 for sandboxes that reap children
  • evals/browser/README.md — six core cases, 28-operation coverage, agent-neutral prompts
  • packages/dsh-plugin-browserskill/README.md — DeepSeek Harness plugin, profile-based lifecycle
  • crates/bsk-protocol and apps/extension/src/transport/types.ts — wire types kept in sync via schema dumps
Aniket Karne
DevOps & AI Engineer · Amsterdam
Back to all posts
Reader correspondence

Comments

Powered by GitHub Discussions via Giscus. Sign in with GitHub to leave a comment.