This page examines AI for greenfield code generation, drawing on public coding benchmarks (HumanEval, SWE-bench Verified, the emerging FeatureBench), DX’s productivity data, and developer-workflow reporting. We are explicit about a real gap: greenfield work is less cleanly measured than bug-fixing.
The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.
Why “greenfield” accuracy is hard to pin down: the most-cited benchmarks don’t actually test it. HumanEval measures isolated function generation (near-saturated, ~90%+); SWE-bench measures repairing real repositories (~80–88%), not building net-new. Genuinely greenfield, multi-file feature development is under-benchmarked — which is why we present the function/repo accuracy figures as measured, but mark the “how devs use greenfield AI” split as an illustrative model, not a survey.
Methodology notes: benchmark figures are from public leaderboards and vary by agent scaffold, model version, and date — treat them as a moving snapshot. HumanEval near-saturation partly reflects benchmark age/contamination, not unlimited capability. The productivity figure (~3.6 hrs/week) is from DX’s large sample; the prototype-time gain and the use-case split are directional. “Greenfield” itself lacks a standard benchmark, which is the honest headline of this page.
Sources used on this page:
- DemandSphere — SWE-bench Verified Tracker 2026 (repo-scale accuracy)
- CodeSOTA — Code Generation Benchmarks 2026 (HumanEval; benchmark types)
- arXiv — FeatureBench: Agentic Coding for Feature Development 2026 (greenfield gap)
- Panto / DX — AI Coding Statistics 2026 (productivity)
Corrections or suggestions: research@aibehaviorindex.org