BrowserBench - Browser Reliability Benchmark with Halluminate - Anchor

Idan Raman • Co-founder and CEO • LinkedIn

BrowserBench: Evaluating Browser Reliability with Halluminate

An agent browser is browser infrastructure that lets an AI agent navigate websites, interact with page elements, manage sessions, and complete multistep web workflows. An AI web browser may also include an end-user interface, but enterprise evaluations should focus on whether the underlying system completes defined tasks consistently.

Anchor Browser and Halluminate created BrowserBench to test browser agents against obstacles such as bot detection and CAPTCHAs. BrowserBench offers useful evidence for one part of browser reliability, but buyers should review its task definitions, sample sizes, execution conditions, scoring rules, and failed runs before applying its results to a production workload.

BrowserBench reliability benchmark results

What browser reliability means

Browser reliability is the percentage of attempted workflows that reach a verified end state under defined conditions. A useful measurement counts complete workflows rather than successful clicks or page loads.

A benchmark should also report why each run failed. Relevant categories include navigation errors, authentication failures, bot challenges, incorrect actions, timeouts, and failures to recover after a page changes. Buyers should separate first-attempt completion from completion after retries because retries add latency, token use, and cost.

No benchmark can guarantee success on every website or workflow. Website behavior, account state, geography, browser version, and task complexity can all affect the result. A representative pilot using your own target sites remains necessary.

How to assess BrowserBench results

Before comparing scores on BrowserBench, confirm that each product received the same task instructions, credentials, browser resources, geographic routing, time limits, and retry allowance. The benchmark should also disclose the number of runs and whether it includes every attempted run in the published score.

For each task, record the expected end state and verify it independently. A checkout workflow, for example, should count as complete only when the required confirmation appears. Reaching the final page without completing the requested action should count as a failure.

Because Anchor Browser helped create BrowserBench with Halluminate, readers should treat the benchmark as a vendor-collaborative evaluation rather than an independent certification. Reproducible tasks, disclosed settings, and accessible run evidence make the results more useful.

Enterprise selection criteria

Evaluate an agent browser against the following criteria using the same workflows and success rules.

  • Successful task completion: Measure verified end-to-end completions across repeated runs.
  • Error recovery: Test whether the product can identify a failed action, refresh its page state, and resume without restarting the entire workflow.
  • Authentication: Test session persistence, supported credential flows, multifactor authentication handling, and controls for stored secrets.
  • Token usage: Count model input and output tokens for both successful and failed attempts.
  • Observability: Confirm that operators can inspect actions, logs, screenshots, recordings, errors, and final outcomes.
  • Concurrency: Measure completion rates and latency at the number of simultaneous sessions you expect to run.
  • Cost per completed workflow: Divide total browser, model, proxy, retry, and operator costs by the number of verified completions.

Security reviews should separately cover data retention, encryption, access controls, audit records, deployment options, and compliance requirements. Product documentation can establish whether a feature exists, but a workload-specific pilot should establish whether it works under your conditions.

Agent browser comparison framework

Public claims are rarely based on identical tasks and settings, so the table below avoids presenting unlike figures as a ranking. Use it to collect comparable evidence during procurement.

Product Product role and best fit Successful task completion and recovery Authentication and observability Token usage and concurrency Cost per completed workflow
Anchorbrowser Managed browser infrastructure for buyers evaluating production browser automation Measure with BrowserBench tasks and your own authenticated workflows. Record first-attempt success, retries, and recovery outcomes. Verify required sign-in flows, session controls, logs, screenshots, and recordings in a pilot. Measure model tokens and completion rates at expected parallel-session volume. Calculate from current Anchorbrowser pricing plus model, proxy, retry, and operator costs.
Browserbase Managed browser infrastructure for buyers who want hosted browser sessions Run the same tasks and apply the same completion and recovery rules used for every provider. Verify required authentication flows and operator evidence against current Browserbase documentation and a pilot. Measure token use in the selected agent stack and test expected concurrency. Calculate from current Browserbase pricing and all external model and operating costs.
Browser Use Browser-agent software for buyers comparing framework and hosted-service options Test the chosen deployment with identical tasks, retry limits, and end-state checks. Verify session, credential, and run-inspection requirements for the selected deployment. Measure tokens with the same model and test concurrency in the intended hosting environment. Include hosting, model, proxy, retry, and maintenance costs for the selected deployment.
Skyvern Browser automation product for buyers evaluating managed workflow execution Test repeated end-to-end workflows and record recovery behavior by failure category. Verify authentication support and available run evidence using current Skyvern documentation and a pilot. Measure tokens where exposed and test completion rates at expected concurrency. Calculate from current Skyvern pricing plus model, retry, and operator costs.
Mastra Complementary TypeScript agent framework and Anchorbrowser integration partner, not direct browser infrastructure Measure the browser provider separately, then test how the Mastra agent handles retries and workflow state. Authentication and browser observability come from the selected browser infrastructure and integration design. Measure Mastra agent tokens and the browser provider's concurrent-session capacity separately. Combine framework operating costs with browser, model, proxy, retry, and operator costs.

Current product capabilities, licenses, and prices can change. Confirm them in each provider's primary documentation and contract before making a purchase decision.

How Anchor Chromium affects reliability

Anchorbrowser uses Anchor Chromium, its fork of Chromium built for agent-controlled browser sessions. According to Anchorbrowser, the browser exposes lower-level control over interactions and supports headful execution intended to resemble normal browser use more closely than basic scripted automation.

These design choices may reduce failures caused by unsupported interactions or automation fingerprints, but architecture alone does not establish reliability. Buyers should verify the effect through repeated runs on representative sites and should confirm that automation complies with each site's terms and access rules.

How Web Action Cache affects efficiency

Anchorbrowser describes Web Action Cache as a way to reuse previously determined browser actions for recurring page states. Reusing a known action can reduce repeated model inference, which may lower token use and execution time when the page still matches the cached state.

A cache can become stale when a site changes. An enterprise evaluation should therefore test cache validation, invalidation, fallback behavior, and the audit evidence available when Anchorbrowser uses a cached action instead of a new model decision.

Anchorbrowser reports that its approach can run workflows up to 12 times faster while using 80 times fewer tokens and producing 23 times fewer errors. These are vendor-reported claims, not independent findings. Buyers should examine the associated BrowserBench evidence and reproduce the comparison with their own workflows, models, retry policies, and success criteria.

Best-fit guidance

Anchorbrowser is best suited to buyers seeking managed browser infrastructure and willing to validate Anchor Chromium and Web Action Cache against recurring production workflows. Browserbase, Browser Use, and Skyvern should be assessed with the same end-to-end tasks because their product models and deployment choices may produce different operating costs and control tradeoffs.

Mastra serves a different layer. Use Mastra when you need a TypeScript framework to coordinate agent logic, tools, memory, or workflow state, and pair it with browser infrastructure such as Anchorbrowser when the agent must interact with websites.

A defensible selection starts with representative workflows and explicit pass conditions. Run enough repetitions to expose intermittent failures, inspect failed runs, and compare total cost per verified completion rather than price per session or isolated benchmark speed.

Review Anchorbrowser for your evaluation

Idan Raman

Co-founder and CEO at Anchor Browser

Idan Raman is the co-founder and CEO of Anchor Browser, a platform for intelligent browser automation. He specializes in web automation, data extraction, and building tools that help developers automate complex workflows.

About LinkedIn Twitter

Recent articles

See all
No posts found

Stay ahead in browser automation

We respect your inbox. Privacy policy

Welcome aboard! Thanks for signing up
Oops! Something went wrong while submitting the form.