Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, John Wieting
University of Maryland, College Park | Google DeepMind
AI writing detectors rely heavily on surface artifacts (e.g., overuse of words like "delve" or specific sentence patterns). However, these stylistic signals are fleeting, easily bypassed via fine-tuning or post-hoc humanizing edits[cite: 1].
To process over 60,000 long-form stories objectively without style bias, the framework utilizes an automated structural translation sequence[cite: 1]:
Converts raw story text into highly detailed JSON schemas based on NarraBench dimensions (e.g., event tracking, perspective, time setup) to strip out lexical choices and focus purely on structure[cite: 1].
Executes cross-source template comparisons across parallel prompts using advanced LLM auditing[cite: 1]. Features are clustered via embedded vectors to yield a final structured space of $d = 304$ features[cite: 1].
Annotates the entire corpus using an aspect-isolated pipeline[cite: 1]. The resulting semantic mapping yields an encoded vector representation $x \in \mathbb{R}^D$ used to train tree-based classifiers[cite: 1].
Narrative features alone carry immense structural signatures, retaining over 97% of the performance of detectors that rely fully on prose style artifacts[cite: 1].
Even when surface style constraints are completely edited out via artifact-removal frameworks like LAMP, the narrative models maintain a high detection accuracy (93.9% F1)[cite: 1].
| Feature Profile | Count | Human vs AI F1 | 6-Way Class F1 |
|---|---|---|---|
| Narrative Only | 257 | 93.2% | 68.4% |
| Core Only | 30 | 84.8% | 46.5% |
| Core + Fingerprint | 101 | 91.1% | 63.4% |
| Narrative + Style | 304 | 96.0% | 77.3% |
* Evaluated on an independent test set of 1,384 prompts[cite: 1].
When projected onto Linear Discriminant components, the 5 prominent AI engines cluster tightly together[cite: 1]. They share an underlying set of storytelling defaults despite hailing from different engineering pipelines[cite: 1].
Human text populates a much broader, highly dispersed territory[cite: 1]. The average distance within human-authored configurations is 22% more spread out than the collective AI grouping radius[cite: 1].
While AI text generally clusters tightly together, individual model variants exhibit localized unique habits[cite: 1]:
Restraint & Flat Escalation
Displays notably flat event intensity paths. Deeply mirrors established traditions, relies heavily on extended epilogues, and actively actively blocks dream sequence hooks[cite: 1].
Social Dynamics & Gossip
Indexes heavily on community rumors as key plot machinery (64%). Frames stories through historical recollection and designs high-density social networks[cite: 1].
Tidy Resolution vs Front-Loading
Gemini favors exceptionally clean setups set in bleak, oppressive atmospheres[cite: 1]. DeepSeek immediately front-loads backstories that other engines distribute down the line[cite: 1].
The paper defines narrative originality by calculating Euclidean feature proximity to its 25 nearest structural neighbors[cite: 1].
Human Mean Rarity: 71st Percentile[cite: 1]
AI Combined Mean Rarity: 49th Percentile[cite: 1]