This page compiles text-to-SQL accuracy data from the major academic benchmarks (Spider, BIRD, Spider 2.0, BIRD-Interact) and enterprise-deployment analysis (BlazeSQL, Datost), plus arXiv research. This is the most benchmark-rich topic in this coding set.
The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.
Why a single “text-to-SQL accuracy” number is misleading: accuracy is meaningless without naming the benchmark. Spider (clean, academic) and BIRD (dirty, realistic) and Spider 2.0 (enterprise, multi-step) measure increasingly hard versions of the same task, and scores fall from ~85% to ~33% across them — the same model, very different results. We report the hierarchy rather than a headline figure, because the gap between a benchmark question and a real one is the whole story.
Methodology notes: all accuracy figures use Execution Accuracy (EX) — whether the query returns the correct result — and come from public leaderboards that move with model and scaffold changes. Cross-benchmark numbers are not directly comparable (different databases, metrics, and setups). The context-dependent ranges (70–85% clean, 86–95% with semantic layer, 50–70% messy) are representative industry estimates, not a single controlled study. The “silent semantic error” risk is qualitative but widely documented.
Sources used on this page:
- Datost — How Accurate Is Text-to-SQL, Really? 2026 (benchmark hierarchy)
- BlazeSQL — Natural Language to SQL: 2026 Guide (enterprise cliff; context ranges)
- BIRD Benchmark — Realistic Text-to-SQL (EX/VES metrics; human bar)
- arXiv — Spider 2.0: Enterprise Text-to-SQL 2026 (enterprise difficulty)
Corrections or suggestions: research@aibehaviorindex.org