This page compiles data on AI for code review and debugging, drawing on GitHub Octoverse and Stack Overflow adoption analysis, the independent Macroscope bug-detection benchmark, the Martian Code Review Bench (~300,000 PRs, the largest independent benchmark), and CodeRabbit’s State of AI vs Human Code Generation report.
The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.
Why we lead with independent benchmarks here: AI code-review vendors publish their own favorable benchmarks, so we prioritize independent ones — Macroscope for bug-detection accuracy and Martian (built by researchers from DeepMind, Anthropic, and Meta) for precision based on real developer behavior across ~300k PRs. Where a figure is vendor-sourced (e.g. CodeRabbit’s own PR-quality study), we attribute it clearly rather than presenting it as neutral.
Methodology notes: adoption figures (40–50%) are estimates synthesized from GitHub and Stack Overflow data and vary by definition of “AI review.” Bug-detection accuracy is from the Macroscope benchmark; precision from Martian’s real-behavior benchmark — both independent. The 1.7x-more-issues figure is from CodeRabbit’s analysis of 470 PRs (vendor-conducted but methodologically transparent). Benchmark scores depend on test set and model version and move quickly. Review-time-savings figures are team-reported.
Sources used on this page:
- DEV — State of AI Code Review 2026 (adoption; market)
- BuildMVPfast / Macroscope — AI Code Review Tools 2026 (bug-detection accuracy)
- Martian Code Review Bench (via CodeRabbit) 2026 (precision; ~300k PRs)
- CodeRabbit — State of AI vs Human Code Generation 2026 (defect rates)
Corrections or suggestions: research@aibehaviorindex.org