Research Presentation

StoryScope: Investigating Idiosyncrasies
in AI Fiction

Jenna Russell, Rishanth Rajendhran, Chau Minh Pham, Mohit Iyyer, John Wieting

University of Maryland, College Park  |  Google DeepMind

Moving Beyond Surface Style

The Detection Problem

AI writing detectors rely heavily on surface artifacts (e.g., overuse of words like "delve" or specific sentence patterns). However, these stylistic signals are fleeting, easily bypassed via fine-tuning or post-hoc humanizing edits[cite: 1].

The Shift: StoryScope evaluates whether AI fiction can be distinguished purely through discourse-level narrative choices like character agency, structural subplots, and chronological continuity[cite: 1].

The Experimental Framework

  • Massive Parallel Corpus: 10,272 writing prompts derived from real literature, matching human text against 5 state-of-the-art LLMs[cite: 1].
  • Total Generation Scale: 61,608 unique stories spanning roughly 5,000 words each[cite: 1].
  • Fine-Grained Space: 304 automatically induced structural narrative features across 10 core aspects[cite: 1].

The StoryScope Pipeline

To process over 60,000 long-form stories objectively without style bias, the framework utilizes an automated structural translation sequence[cite: 1]:

1. Template Induction

Converts raw story text into highly detailed JSON schemas based on NarraBench dimensions (e.g., event tracking, perspective, time setup) to strip out lexical choices and focus purely on structure[cite: 1].

2. Discovery & Deduplication

Executes cross-source template comparisons across parallel prompts using advanced LLM auditing[cite: 1]. Features are clustered via embedded vectors to yield a final structured space of $d = 304$ features[cite: 1].

3. Automated Featurization

Annotates the entire corpus using an aspect-isolated pipeline[cite: 1]. The resulting semantic mapping yields an encoded vector representation $x \in \mathbb{R}^D$ used to train tree-based classifiers[cite: 1].

Classification Performance

Key Takeaway

Narrative features alone carry immense structural signatures, retaining over 97% of the performance of detectors that rely fully on prose style artifacts[cite: 1].

Even when surface style constraints are completely edited out via artifact-removal frameworks like LAMP, the narrative models maintain a high detection accuracy (93.9% F1)[cite: 1].

Performance Table

Feature Profile Count Human vs AI F1 6-Way Class F1
Narrative Only 257 93.2% 68.4%
Core Only 30 84.8% 46.5%
Core + Fingerprint 101 91.1% 63.4%
Narrative + Style 304 96.0% 77.3%

* Evaluated on an independent test set of 1,384 prompts[cite: 1].

The Topology of Narrative Space

AI Convergence

When projected onto Linear Discriminant components, the 5 prominent AI engines cluster tightly together[cite: 1]. They share an underlying set of storytelling defaults despite hailing from different engineering pipelines[cite: 1].


Human Variance

Human text populates a much broader, highly dispersed territory[cite: 1]. The average distance within human-authored configurations is 22% more spread out than the collective AI grouping radius[cite: 1].

Centroid Projection (LD1 vs LD2)

LD1 (Separation Axis) LD2 Human Centroid AI Cluster Zone

The Human vs AI Divergence Profile

AI Narrative Fingerprints

  • Thematic Over-Determination: Tells the moral explicitly. Narrators state themes outright 77% of the time vs 52% for humans[cite: 1].
  • Streamlined Causality: Focuses on clean, single-track paths with minimal side subplots (79% feature "no subplots")[cite: 1].
  • Sensory Body Focus: Overly prone to illustrating emotional arcs through intense physical sensations (81% vs 38%)[cite: 1].

Human Narrative Fingerprints

  • Temporal Complexity: High reliance on nonlinear order hooks, flashbacks, and time jumps to strategically obscure setups[cite: 1].
  • Moral Ambiguity: Characters encounter complex ethical gray zones (59% framing vs 38% across AI models)[cite: 1].
  • Real-World Trajectories: Introduces direct named citations, real places, and structural fourth-wall reader engagement[cite: 1].

Individual Model Idiosyncrasies

While AI text generally clusters tightly together, individual model variants exhibit localized unique habits[cite: 1]:

Claude

Restraint & Flat Escalation

Displays notably flat event intensity paths. Deeply mirrors established traditions, relies heavily on extended epilogues, and actively actively blocks dream sequence hooks[cite: 1].

GPT

Social Dynamics & Gossip

Indexes heavily on community rumors as key plot machinery (64%). Frames stories through historical recollection and designs high-density social networks[cite: 1].

Gemini & DeepSeek

Tidy Resolution vs Front-Loading

Gemini favors exceptionally clean setups set in bleak, oppressive atmospheres[cite: 1]. DeepSeek immediately front-loads backstories that other engines distribute down the line[cite: 1].

Quantifying Rarity and Originality

Rarity Percentile Metric

The paper defines narrative originality by calculating Euclidean feature proximity to its 25 nearest structural neighbors[cite: 1].

Human Mean Rarity: 71st Percentile[cite: 1]

AI Combined Mean Rarity: 49th Percentile[cite: 1]

The Rarest Tail Distribution

  • Extreme Overrepresentation: 24.7% of all human narratives populate the top 10% rarest tier corpus-wide[cite: 1].
  • AI Scarcity in Innovations: Only 7.1% of AI generations touch that same rare frontier[cite: 1].
  • Prompt Ranking: When evaluated side-by-side, the human version registers as structurally unique 57.8% of the time[cite: 1].