“Benchmarks require humans to confirm that the thing you're claiming your browser use agent can do, it can in fact do. We chose to use Careerflow. They were absolutely amazing at this.”
Ricardo Spagni
Co-founder - Blink



































A benchmark for browser-use agents from the OSU NLP Group and UC Berkeley, with 300 real tasks across 136 live websites. Introduced in the COLM 2025 paper "An Illusion of Progress? Assessing the Current State of Web Agents."
WebJudge reaches ~85% agreement with humans and is great for iteration, but it still runs 6 to 10% off the true success rate. Human eval is the reference standard for a published result.
Careerflow, the benchmark's official human evaluation partner.
A per-task pass or fail, the evidence behind each judgment, and error analysis mapped to the paper's failure categories.
Confidential, with anonymized workflows and enterprise-grade security.
Yes. Careerflow's Human Data team handles human evaluation and labeling across agents, LLMs, and a wide range of domains.
