Structural 'Shape' of AI-Written Blog Posts Detected with 98% Accuracy
The paper builds on prior work (StoryScope) showing AI-generated text has a distinctive structural fingerprint, applying this to commercial blog content: 2,250 pre-ChatGPT human posts versus 11,250 AI-generated mirrors from five frontier models. Using a 214-feature instrument (187 purely structural, e.g. information order, evidence use, voice) scored by an LLM and validated against human annotators (high inter-rater agreement), the authors report 98.0 macro-F1 detection accuracy on held-out companies — robust even when AI posts are reworded by their own generating model — and can attribute posts to their source model at 79.3% accuracy against a 16.7% chance baseline. They argue human writing occupies structurally rare configurations that AI text rarely reproduces, and that these effects replicate and amplify the earlier fiction-domain findings. The original tweet thread (from one of the authors) highlighted the headline stat of only 19/1,740 misclassifications as evidence that AI text has a detectable 'shape' beyond word-level slop, framing it as a way to catch AI content even after paraphrasing defeats standard detectors. Discussion volume was otherwise limited, with no substantive critical pushback visible in the sampled commentary.
Discussion: 1 tweets from 1 authors · @jochenmadler