Home/By Use Case/AI for Code Generation (Greenfield)
By Use Case

AI for Code Generation (Greenfield)

“Greenfield” — building new features and apps from scratch — is where AI code generation feels most magical and, paradoxically, where it’s least cleanly measured. On the narrow task of writing a self-contained function from a description, frontier models are effectively saturated, scoring above 90% on HumanEval. On harder, repository-scale work they reach the low-to-high 80s and climbing. But here’s the catch the headlines miss: the most-cited coding benchmarks measure fixing existing code, not building net-new — SWE-bench is largely a bug-fixing benchmark — so true greenfield ability is under-benchmarked, and new feature-development benchmarks are only now emerging. In practice, developers lean on AI most for scaffolding, prototypes, and first-draft features, reserving novel architecture for humans. The speed gains are real (prototypes that took days now take hours), but net-new AI code carries the same logic and security caveats as all AI code. This page maps what’s solved, what’s still hard, how developers actually use greenfield generation, and the verify-everything rule underneath the prototype boom.

4 visualizations 4 sources Last updated June 2026 Free to embed
Loading chart...
Loading chart...
CHART 1 · FUNCTION-LEVEL IS NEAR-SOLVED

AI accuracy by code-generation task

AI code accuracy by task: function vs repository

Loading chart...
Source: HumanEval & SWE-bench leaderboards 2026: on isolated function generation (HumanEval), frontier models score ~90%+ — effectively saturated. On real repository work (SWE-bench Verified), top models reach ~80–88%, up from ~65% in early 2025. Self-contained generation is largely solved; whole-system work is not.
CHART 2 · BUT "GREENFIELD" IS UNDER-BENCHMARKED

What the benchmarks actually measure

What coding benchmarks actually measure (greenfield gap)

Loading chart...
Source: FeatureBench / SWE-bench analysis 2026: the dominant benchmark (SWE-bench) measures fixing existing repos, not building net-new — it “mainly focuses on bug fixing, with limited coverage of feature development.” New benchmarks (FeatureBench) are emerging precisely because greenfield/feature work is poorly measured. Genuine gap, flagged.
CHART 3 · HOW DEVS USE IT

Greenfield AI generation in practice

Where devs lean on AI for new work (illustrative)

Loading chart...
Source: AI Behavior Index synthesis of developer-workflow reporting 2026. Illustrative: for new work, developers lean on AI most for scaffolding/boilerplate, prototypes and “vibe-coded” MVPs, and first-draft features they then refine. Net-new architecture and novel logic remain human-led. No single survey cleanly measures this split — directional model.
CHART 4 · THE PROTOTYPE BOOM

Speed up, but verify

Generate fast — but verify hard

Loading chart...
Source: DX / industry data 2026: AI shines for greenfield speed — ~3.6 hrs/week saved per developer, and prototypes that took days now take hours. But generated net-new code carries the same quality caveats as all AI code (logic and security gaps), so “generate fast, verify hard” is the rule. Productivity figure is measured; the prototype-time gain is directional.

About this data

This page examines AI for greenfield code generation, drawing on public coding benchmarks (HumanEval, SWE-bench Verified, the emerging FeatureBench), DX’s productivity data, and developer-workflow reporting. We are explicit about a real gap: greenfield work is less cleanly measured than bug-fixing.

The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.

Why “greenfield” accuracy is hard to pin down: the most-cited benchmarks don’t actually test it. HumanEval measures isolated function generation (near-saturated, ~90%+); SWE-bench measures repairing real repositories (~80–88%), not building net-new. Genuinely greenfield, multi-file feature development is under-benchmarked — which is why we present the function/repo accuracy figures as measured, but mark the “how devs use greenfield AI” split as an illustrative model, not a survey.

Methodology notes: benchmark figures are from public leaderboards and vary by agent scaffold, model version, and date — treat them as a moving snapshot. HumanEval near-saturation partly reflects benchmark age/contamination, not unlimited capability. The productivity figure (~3.6 hrs/week) is from DX’s large sample; the prototype-time gain and the use-case split are directional. “Greenfield” itself lacks a standard benchmark, which is the honest headline of this page.

Sources used on this page:

Corrections or suggestions: research@aibehaviorindex.org

For Journalists & Researchers

Use this data in your work.

Every statistic, chart, and graphic in this index is free to use and cite, with full source attribution. Can’t easily find what you need? Use our search bar to search by keyword, topic, or category.

✉️
Talk to our research team
Need a specific cut of data, an interview, or a quote? Email us — we typically respond within one business day.
research@aibehaviorindex.org →

How the data works

Every statistic shown is sourced from a publicly available study, survey, or report. We aggregate, organize, and contextualize this data — but the underlying research is conducted by the cited sources. Click any source link to access the original methodology. If you run into any issues or have a study to suggest, contact us at research@aibehaviorindex.org.