---
title: "Widely used AI-text detectors misclassify non-native English writing as AI-generated at high rates, undermining attempts to re-verify human effort"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/abs/2304.02819"
source_title: "GPT detectors are biased against non-native English writers"
source_author: "Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou"
source_date: "2023-04-18 (arXiv:2304.02819v2); published in Patterns, 2023"
source_quote: "these detectors consistently misclassify non-native English writing samples as AI-generated, whereas native writing samples are accurately identified... misclassified over half of the TOEFL essays as \"AI-generated\" (average false positive rate: 61.22%)"
source_tier: 1
audit_status: "capture-verified — the 2026-07-16 batch capture fetched the arXiv PDF directly and recorded this quote and the false-positive figure; promotion did not independently re-fetch (headless run). | 2026-07-19 cross-model audit (writer claude-sonnet-5, auditor claude-fable-5): independently re-fetched arXiv:2304.02819 (v3 PDF, sha256 9019ad9a…) — abstract and Results quotes verbatim; 61.22% avg FPR, seven detectors, 91 TOEFL / 88 ASAP 8th-grade essays, 18/91 (19.78%) unanimous misflags, and the prompt-bypass finding all confirmed against the primary. Upgraded: verified-verbatim."
provenance: "Promotion from 10-inbox/raw/2026-07-16-does-when-costliness-becomes-forgeable-explain-generative-ais.md, 2026-07-18"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-16-does-when-costliness-becomes-forgeable-explain-generative-ais.md"
date_created: "2026-07-18T00:00:00.000Z"
tags: ["generative-ai","ai-text-detection","credentials","higher-education","technical-mechanism","value-theory"]
audits: ["2026-07-19 claude-fable-5"]
---


Liang, Yuksekgonul, Mao, Wu, and Zou tested seven widely used GPT-text detectors against 91 human-authored TOEFL essays (written by non-native English speakers) and 88 US 8th-grade essays (written by native speakers). The detectors performed near-perfectly on the native-writer essays but "misclassified over half of the TOEFL essays as 'AI-generated' (average false positive rate: 61.22%)," and all seven detectors unanimously misflagged nearly 20% of genuinely human-written TOEFL essays. Simple prompting strategies that added linguistic variety to genuinely AI-generated text also evaded detection, cutting effectiveness on the other side of the ledger too.

This is a specific technical-mechanism finding — how detection tools actually perform, not merely that they exist — about why institutions cannot cheaply re-impose the cost of proving writing is human-produced once generative AI drove that production cost near zero. It bears directly on credentials, admissions essays, and any writing used as proof of a person's own effort: see [[claim-generative-ai-availability-compresses-university-grade-distributions]] for the university-grading side of the same problem, and [[claim-llm-collapse-of-costly-writing-signal-cuts-meritocratic-hiring]] for the hiring-market side.

In the vocabulary of [[claim-szabo-bit-gold-grounds-value-in-unforgeable-cost-of-production]], detection tools are an attempted furnace-repair: an effort to re-manufacture unforgeable costliness after the original cost floor collapsed. This finding shows that repair itself is leaky, and leaky in a way that penalizes exactly the writers — non-native speakers — least able to absorb a false accusation of dishonesty.

> [!note] Seek's commentary:
> The detail worth sitting with is not the 61% false-positive rate alone, it's who it lands on. A detector that fails randomly is a broken tool; a detector that fails specifically on non-native writers is a broken tool wearing the shape of a bias. The 2023 vintage of this study is itself a live question — a 2025 Booth working paper reportedly finds detectors have since diverged sharply in reliability, and I haven't read it. I'm not routing that as a formal open question here, because nothing this note asserts depends on today's detectors matching 2023's; it's a fact about a specific study at a specific time, and it stays true regardless of what Pangram or Turnitin do next. But it's the kind of thing I'd want fresher data on before ever citing a detector's *current* accuracy anywhere in this vault. — Seek
