Databricks Releases OfficeQA Pro V2 Benchmark for Enterprise Grounded Reasoning
Databricks has introduced OfficeQA Pro V2, a new benchmark designed to evaluate AI agent performance on grounded reasoning tasks using enterprise-style data. This benchmark is built on a large corpus of U.S. Treasury PDFs and leverages synthetic data generation for scalability. While out-of-the-box agents show limited accuracy, Databricks' Genie demonstrates significant performance improvements, indicating that robust agent harnesses are crucial for unlocking AI capabilities in complex data scenarios. The release aims to drive progress in document retrieval, parsing, and analytical reasoning for enterprise applications.
- →New OfficeQA Pro V2 benchmark for enterprise grounded reasoning
- →Databricks Genie shows significant performance gains
- →Synthetic data generation scales benchmark creation
- →Benchmark challenges frontier agents and highlights need for better harnesses
- →OfficeQA Pro V2 builds on previous benchmark for broader evaluation
Features (2) ›
- New OfficeQA Pro V2 benchmark for enterprise grounded reasoning
Databricks released OfficeQA Pro V2, a new benchmark focused on evaluating AI agents' ability to perform grounded reasoning over enterprise-style data. It uses a corpus of approximately 1,400 U.S. Treasury PDFs, spanning 233 years and 120,000 pages, to test generalization capabilities.
- Databricks Genie shows significant performance gains
Databricks' agent harness, Genie, demonstrated up to a 92% relative improvement in accuracy on OfficeQA Pro V2 compared to out-of-the-box frontier agents. Genie achieved up to 60% accuracy, showcasing the impact of optimized agent harnesses on model performance.
Enhancements (1) ›
- Synthetic data generation scales benchmark creation
The benchmark was built using an internal synthetic data generation pipeline, enabling the creation of diverse and verified questions and answers. This approach allows for rapid development of new benchmarks to address customer needs like grounded reasoning.
Notes (2) ›
- Benchmark challenges frontier agents and highlights need for better harnesses
Out-of-the-box frontier agents averaged 26.0% accuracy on the benchmark, with some specialized agents reaching over 63.3%. This suggests that while AI models are advancing, significant headroom remains in grounded reasoning, particularly in handling complex document parsing, temporal reconciliation, and entity scope.
- OfficeQA Pro V2 builds on previous benchmark for broader evaluation
OfficeQA Pro V2 extends the original OfficeQA benchmark to test generalization beyond a single corpus. It was initially developed for the Databricks Grounded Reasoning Cup, involving academic teams and major AI labs, and is now being released broadly to AI practitioners.
https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning
Related releases
- Databricks SDK for Go v0.172.0 Enhances IAMv2 and Job Cluster Management Databricks Go SDK Releases ·
- Databricks Java SDK v0.146.0 Adds IAM V2 API, Updates Job Cluster Field Databricks Java SDK Releases ·
- Databricks SDK for Python v0.128.0 Adds Account and Workspace IAM V2 API Methods Databricks Python SDK Releases ·
- Amtrak Builds Unified Data Backbone with Databricks for Rail Network Transformation Databricks Blog ·
- Databricks Re-architects Serverless Network Config Delivery for 97.5% Latency Cut Databricks Blog ·
- How a Major Freight Railroad Scaled Pipeline Creation with Databricks Genie Code Databricks Blog ·