What happened.

Epoch AI tested Pangram, GPTZero, and Originality.ai on AI-generated passages designed to mimic an author’s style. The reported results found that up to 18% of the AI passages went undetected overall. In scientific writing, the reported share rose as high as 48%.

Why it matters.

The result is relevant to academic and editorial tooling because scientific writing is a genre where these detectors are heavily relied on. A detector that misses style-imitated passages may be unsuitable as the sole basis for an academic or editorial decision. The finding also shifts attention from a headline accuracy figure toward the conditions under which a tool may fail.

Limit and next receipt.

This is a single reported evaluation, and the supplied record does not provide independent corroboration or fuller test details. The next useful receipt would be independent testing of these detectors on style-imitated writing, including whether the reported miss rates hold across scientific and other genres.

What to watch

Watch for independent evaluations that reproduce or challenge the reported miss rates for style-imitated AI text.

Sources and limits

Upstream references

Digest dated 2026-07-20 · upstream model claude-sonnet-4-6. Source IDs are preserved for audit; the publishing host does not receive the upstream URL map.

  1. 1
    937e0f01409cc92a6ac3c9243c6d1479594baf4bReference from the upstream research server

This Research brief was generated by Terra from a dated upstream research digest. It has not received the source-by-source human review required for Reviewed analysis. Material limit: The evidence is limited to a single Epoch AI-reported evaluation, without independent corroboration in the supplied record.