Most AI assistant leaderboards measure conversation. This one measures whether the work got done.

News Room
Sep 24
by Gabi Weinberg
Assistant Analysis homepage for the Digital Assistant Performance Benchmark, version 0.2.0, September 2026

The purpose of an AI assistant is to accomplish tasks on behalf of humans. Today’s benchmarking leaves room for a new method to analyze assistants, and we believe this is the crux of the issue:

Did the assistant complete the real workflow? Ship the email? Book the flight? Navigate multiple apps? And, can we prove it without trusting the assistant’s own word?

That is the unique angle of Assistant Analysis: an independent Digital Assistant Performance Benchmark built around verified task completion, confirmed against independently captured ground truth, not self-report, not vibes, not solely a single composite “score” that averages reliability into oblivion. We have been enabling assistant teams for internal tooling and assistants for nearly a year now, so we’ve got a lot of experience evaluating their success. We’re now opening that benchmark to the world and invite you to use it.

The metric that matters: verified score

The flagship number is the verified score (verified completion rate): the proportion of assigned tasks that reach a predefined successful outcome without a disqualifying authorization or safety violation, which is confirmed against independently captured ground truth, not claimed by the assistant.

Two design choices make this different from every chat-arena board:

  1. Ground truth is external. Success is checked against what actually happened in the environment (did the email send? did the booking exist?), not against the model’s confidence or a human preference vote. A private evaluator records environment snapshots, audit logs, and ground-truth checks; self-reported completion is never sufficient evidence.
  2. Axes stay separate on purpose. Reliability, safety incident rate, human interventions, median completion time, and cost per successful task are tracked as separate metrics. They are never folded into one composite. A system can be fast and unreliable, or safe and slow. So we don’t average them.

Only methodologically compatible evaluations get ranked against each other. Products with incomplete or incompatible evaluations are listed separately with explicit status.

Assistant Analysis leaderboard preview ranking seven assistants by verified score

What “surviving the open web” actually means

The failure modes that kill assistants in production are invisible on a landing-page demo: the login wall after the happy path, the session that drops mid-flow, the page whose structure drifted between when the task was written and when it ran.

Assistant Analysis is built around those questions:

  • Can it survive the open web unattended?
  • Can it hold a session?
  • Can it recover when the target changes underneath it?

That is not “can it answer a question.” That is “can it finish the job.”

Assistant Analysis highlights comparing verified score, median reply time, and supported capabilities

Where the field stands (live board, Sep 2026)

Seven systems currently carry verified overall scores on benchmark v0.2.0, with 49 / 49 task coverage on the published suite (evaluation date 17 Sep 2026). Ranked by verified score:

  • Claude Cowork (Anthropic) — 93.5% verified, 86% reliability, median completion 9.1s — Claude Cowork driven through the API with a generic open-url browsing tool
  • Muse (Meta) — 89.0% verified, 82% reliability, median 55s — Meta’s personal AI agent, tested through its browser-based chat
  • Hermes (Nous Research) — 86.9% verified, 90% reliability, median 24.8s — local CLI assistant that runs on your machine and calls a model API
  • Manus (Manus) — 85.7% verified, 65% reliability, median 159.2s — cloud agent with its own browser, driven through the Manus task API
  • Instinct (Instinct) — 79.3% verified, 70.0% reliability, median 78s — interact via iMessage
  • OpenClaw (OpenClaw) — 77.4% verified, 76% reliability, median 41.8s — open-source agent that drives a real browser and runs entirely on your box
  • Grok Bot (xAI) — 76.6% verified, 73% reliability, median 112.3s — Grok wired in as an automation that answers over a channel webhook

We’re still updating this and the suite is real and fully covered for these runs, but this is an early public board and the confidence intervals are wide on several rows, cost per successful task is not populated yet, and rankings will move as new systems and versions land. We won’t invent numbers the methodology hasn’t published. Stay tuned for more.

Even at this stage, one result already stood out to us: Hermes posts the highest reliability (90%) while scoring below Claude Cowork and Muse on verified completion. Score and reliability are not the same axis. This leaderboard is built so you can see them diverge.

Another: OpenClaw is architecturally the closest thing on this board to genuine browser-level access — and it still scores below Claude Cowork paired with a generic open-url browsing tool. Owning the browser isn’t sufficient, the action taking layer is the difference maker.

What comes next: Live board, open submissions

Assistant Analysis is a versioned framework (currently v0.2.0) with its own site and update cadence. You can follow it at assistantanalysis.com rather than waiting on recaps, or get more up to date content on X. Methodology, compatible-eval ranking rules, and open data live on the site. New systems get evaluated as they’re submitted; rankings will move.

If you’ve built something that logs in, navigates, and gets real work done without a human standing by, whether a public product or internal tool, the evaluation is open. Public leaderboard or private report is your call.

A shout-out to a fellow NYC area assistant benchmark

David Pawlan’s Assistant Benchmark answers a different question well: which assistant is worth texting. The assistant industry needs ways to ascertain the efficacy of these tools, beyond just the hype. This one only asks whether the task got done, and whether anyone had to step in to do it.

Submit your assistant for evaluation here and follow us on X for more updates.

Built by the team at Anchor Browser — we spend our days on the infrastructure layer this benchmark surfaces (sessions, MFA survival, recovery from page drift). The board stands on its own methodology; the company builds the layer the gaps point to.

Stay ahead in browser automation

We respect your inbox. Privacy policy

Welcome aboard! Thanks for signing up
Oops! Something went wrong while submitting the form.