Home/By Use Case/AI Writing Quality: Human vs Machine Detection
By Use Case

AI Writing Quality: Human vs Machine Detection

Can you tell whether a piece of text was written by a person or a machine? The research is remarkably consistent: since GPT-3, most people can’t. Across many peer-reviewed studies, human accuracy at distinguishing AI- from human-written text sits close to random guessing, and even trained experts — academics, linguists, ESL teachers — do only a little better, and inconsistently. The clearest exception is heavy AI users, who have learned to recognize the telltale patterns, and people can improve somewhat with immediate feedback. Automated AI detectors do outperform humans by reading statistical signals invisible to the eye — but they carry a serious cost: they falsely flag genuine human writing as AI a meaningful share of the time, which is why researchers conclude the tools are not reliable enough to police academic integrity on their own. This page lays out how poorly humans detect AI text, how little expertise helps, who actually can tell, and the real-world limits of detection tools.

4 visualizations 4 sources Last updated June 2026 Free to embed
Loading chart...
Loading chart...
Chart 1 · Humans are near chance

Can readers tell AI from human writing?

Human detection accuracy hovers near the 50% chance line

Loading chart...
Source: Multiple peer-reviewed studies (arXiv survey 2024–26): since GPT-3, humans distinguish AI- from human-written text at close to random-guessing (50%). Across studies, accuracy clusters near chance, with annotators often biased toward calling everything “human-written.” The machines now write past the human eye.
Chart 2 · Experts barely better

Detection accuracy by reader type

Expertise helps a little — and inconsistently

Loading chart...
Source: Nature / Scientific Reports (Mar 2026) & Wang et al.: on research abstracts, academics ranged 44–76% (avg ~72% but highly variable); ESL teachers hit ~61% (up to 67% with training); applied linguists fared no better than chance. Expertise helps a little — and inconsistently.
CHART 3 · WHO ACTUALLY CAN

Heavy AI users detect best

Heavy users do best; "which LLM?" falls below chance

Loading chart...
Source: The Scientist 2026 / feedback-training study 2025: the people who detect AI text best are heavy AI users — they recognize the patterns — and humans can improve with immediate feedback. But on “which LLM wrote this?” even NLP PhD students scored ~21%, below the 25% random-guess baseline.
CHART 4 · TOOLS BEAT HUMANS — BUT MISFIRE

AI detectors and the false-positive problem

Tools beat humans — but still falsely flag real writing

Loading chart...
Source: Detection-tool analysis 2026 / Weber-Wulff et al.: automated detectors outperform humans by reading statistical patterns people can’t see — but they falsely flag human writing as AI up to ~15% of the time (humans, up to ~20%), and researchers conclude current tools are insufficient to police academic integrity alone. Accuracy with a real cost.

About this data

This page synthesizes peer-reviewed research on detecting AI-generated text, including a March 2026 Nature/Scientific Reports study, multiple arXiv surveys and studies (Liu et al., Wang et al., Sarvazyan et al.), The Scientist’s 2026 coverage, and Weber-Wulff et al.’s detection-tool evaluation. This is the most academically-grounded topic in this set.

The AI Behavior Index is the research arm of OneChat AI, an integrated multi-model AI platform. We compile and analyze data from primary research sources to make AI adoption and market trends more accessible to journalists, researchers, and decision-makers.

Why detection-accuracy figures range so widely (44% to 94%): they depend heavily on text type, reader, and model. Scientific abstracts are hard even for experts (~60–84%); some Reddit-style text is easier (up to ~94%); poetry and student essays are near chance. “Accuracy” also differs from “above random” — a 60% score on a 50/50 task is only modestly better than guessing. We report ranges and the central finding (near-chance for most readers and texts) rather than a single headline number.

Methodology notes: figures come from peer-reviewed studies with varied designs (different text domains, reader expertise, LLM versions, and sample sizes — some small, e.g. six academics). “Near chance” means accuracy not reliably above 50% on a binary task. The false-positive figures (~15–20%) are from detection-tool analyses and matter most because flagging real human work as AI causes direct harm. As models improve, detectability generally falls, so older studies may understate the current difficulty.

Sources used on this page:

Corrections or suggestions: research@aibehaviorindex.org

For Journalists & Researchers

Use this data in your work.

Every statistic, chart, and graphic in this index is free to use and cite, with full source attribution. Can’t easily find what you need? Use our search bar to search by keyword, topic, or category.

✉️
Talk to our research team
Need a specific cut of data, an interview, or a quote? Email us — we typically respond within one business day.
research@aibehaviorindex.org →

How the data works

Every statistic shown is sourced from a publicly available study, survey, or report. We aggregate, organize, and contextualize this data — but the underlying research is conducted by the cited sources. Click any source link to access the original methodology. If you run into any issues or have a study to suggest, contact us at research@aibehaviorindex.org.