Key Takeaways
- Claude Opus 5.5 leads with 98.38% balanced accuracy, followed by GPT-6 Astra at 95.75%. Both score above the evaluated AI-detection products; Astra has the lower cost per document of the two language models.
- Detecting AI rewrites is harder than recognizing human writing. The top four systems correctly classify 98.8–100% of human documents, compared with 87.9–96.8% of AI rewrites.
- Similar overall scores can mask different behavior. GPTZero correctly classifies 100% of human documents but detects 56.3% of AI rewrites. Sapling detects more rewrites, at 72.9%, but correctly classifies fewer human documents, at 82.2%.
Building our Dataset
AI writing detectors need to identify generated text without incorrectly flagging human work. Creating a difficult AI-detection benchmark is difficult, because most public human-written text has been ingested by models, and it is difficult to verify the provanence of newly written text.
To get around these issues, we designed our evaluation with fully private data authored before November 2022, sourced from the evaluation creators at Vals along with a small number of crowdsourced private documents.
We evaluate specialized detectors and general purpose closed models on their detection performance on these human texts, and then rewrite the human texts with a suite of models from our proprietary gateway. We explicitly exclude latest-generation OpenAI, Anthropic, and Deepseek models from this suite in order to enable their unbiased evaluation. Later in this report, we show experiments using these models as generators as well.
To generate AI-assisted samples, we sample from a short list of prompts that each steer the model to rewrite the passage in a similar style to the original text. Since this may incentivise narrow rewrites, we then programatically confirm that > 50% of the original text is altered to ensure that the text can fairly be labeled ‘AI generated/assisted.‘
Evaluation
We evaluate Opus 5.5, GPT-6 Astra, and DeepSeek V4.1 Flash using a simple prompt asking for classification of AI-usage versus human authorship. Pangram’s Mixed and AI labels both count as identifying AI involvement, Human counts as identifying human writing. The other specialized detectors are scored using their AI-probability scores, with values strictly above 0.5 classified as AI.
Our primary metric is accuracy over a balanced test set (half AI, half human). We additionally report accuracy on the human-authored and AI-rewritten splits. Cost per document uses recorded API charges as well as per-token accounting for language models.
Robustness to Frontier Model Rewriting
After observing that the latest Astra and Opus models were exceptionally strong at identifying AI-generated text, we generated AI-generated samples using these models to observe whether they could fool themselves, or pangram. Of 100 attempted rewrites, 73 modified enough text to be counted as AI-generated by our programatic checks: 42 from Astra and 31 from Opus 5. The leaderboard below shows performance on this adversarial set.
We find that frontier models are often capable of editing text in a manner that is not reliably detectable by any detectors. Within this cohort, pangram is more easily fooled by these modifications than the generating frontier models themselves.
Limitations
In this experiment, we have prioritized an uncontaminated test set over task diversity and quantity. Therefore, we are unsure how results generalize to data outside of our specific distribution. Also, our paraphrase prompts are specific and adversarial as an existential demonstration of a floor for AI detection, these results should not be interpreted as the practical accuracy rates of AI detection tools across a practical distribution of AI-detection settings, which would likely include less diverse model-sets, a more iterative collaboration between models and humans, and a broader sample of human texts.