databricks Databricks Blog ·

Databricks Genie Code Agent Outperforms General Agents in Accuracy and Cost

blogaidatabricksengineer
announcement feature

Databricks evaluated its data agent, Genie Code, against three general coding agents on over 400 real-world tasks. Genie Code proved to be the most accurate and cost-efficient, delivering correct answers at less than half the cost of other agents. This performance stems from its deep semantic understanding of enterprise context and specialized capabilities, allowing it to navigate complex data workspaces more effectively than general-purpose agents. The evaluation included a wide spectrum of tasks such as data discovery, code creation, debugging, and data lookups, highlighting Genie Code's advantages for full-spectrum data work.

  • Databricks Genie Code Agent Demonstrates Superior Accuracy and Cost Efficiency
  • Genie Code Features for Enhanced Data Workspace Navigation
  • Evaluation Methodology and Task Distribution
  • Performance Metrics and Cost Analysis
  • Efficiency Driven by Data-Specific Capabilities
Features (1)
  • Databricks Genie Code Agent Demonstrates Superior Accuracy and Cost Efficiency

    Databricks' data agent, Genie Code, was evaluated against three leading general coding agents on over 400 real internal tasks. Results show Genie Code achieved higher accuracy and was more cost-effective, costing less than half of other agents for correct answers. This is attributed to its specialized capabilities for data-centric environments.

Enhancements (2)
  • Genie Code Features for Enhanced Data Workspace Navigation

    Genie Code is designed to address unique data agent challenges, such as discovering assets in dynamic workspaces and determining sources of truth from potentially contradictory metadata. It utilizes semantic search, persistent memory of user-reliant data, and deep enterprise context understanding to overcome limitations of general coding agents.

  • Efficiency Driven by Data-Specific Capabilities

    Genie Code averages fewer tool calls per task (8.3) compared to other agents, indicating greater efficiency in finding needed data. This reduced number of turns and tokens directly translates to lower costs without sacrificing accuracy, especially for complex discovery tasks.

Notes (3)
  • Evaluation Methodology and Task Distribution

    The evaluation compiled 401 self-contained tasks from real Databricks usage, covering discovery, code creation/modification, debugging, and data lookups. Tasks were run with a 20-minute budget, and answers were graded for correctness and usefulness by an independent judge.

  • Performance Metrics and Cost Analysis

    In the benchmark, Genie Code achieved 76.6% accuracy at $0.55 per task, significantly outperforming general agents like Agent X (72.1% accuracy, $1.09/task). General agents struggled with discovery tasks, leading to higher costs and lower accuracy due to inefficient workspace exploration.

  • Future Improvements and Broader Implications

    The Genie Ontology is expected to further strengthen Genie Code's advantage, though it was disabled for this evaluation. The findings emphasize that generic leaderboards may not reflect real-world user experience, advocating for evaluations based on actual usage patterns.

Read the original announcement →

https://www.databricks.com/blog/why-frontier-data-agent-outperforms-general-coding-agents-quality-and-cost

Related releases