Evaluating AI Agents Live at the Grounded Reasoning Cup
Databricks hosted the Grounded Reasoning Cup, a competition evaluating AI agents on enterprise document reasoning tasks using the OfficeQA Pro V2 benchmark. Stanford won with a system achieving 63.3% accuracy, significantly outperforming baseline frontier agents. The event highlights current capabilities and limitations in grounded reasoning, noting that nearly 19% of questions remained unsolved by all participating teams.
