Decoding AI Search Citations: New Study Reveals How Formatting, Position, and Model Randomness Shape Web Traffic

By Tech & SEO Research Desk
Published September 2026

As artificial intelligence increasingly mediates how humans discover information on the web, understanding the mechanics behind AI search engine citations has become the holy grail for digital marketers, content creators, and SEO professionals. Traditional search engine optimization (SEO) relied heavily on static variables like backlinks, keyword density, and traditional ranking algorithms. However, generative AI search agents operate on entirely different principles, synthesizing answers dynamically from multiple sources.

A new, non-peer-reviewed preprint published to arXiv on September 14 by researchers Sriram Selvam and Anneswa Ghosh sheds unprecedented light on this opaque ecosystem. The study investigates a fundamental question: Does tweaking a single variable—such as source position or page formatting—actually change whether an AI search agent cites a source, keeping everything else equal?

The findings challenge long-held assumptions about raw position metrics, introduce critical warnings regarding AI attribution sensitivity, and expose a staggering degree of non-determinism in how modern language models allocate credit.


1. Main Facts: Dissecting the Experiment

The research paper, titled in preprint and built around rigorous offline experimentation, focuses on a specific configuration: a GPT-5.4 search agent utilizing Exa as its underlying search provider. Crucially, the researchers did not execute live webpage edits. Instead, they relied on offline replayed conversations, examining how the model processed information when specific factors were altered.

Key Takeaways from the Study:

  • The Raw Position Illusion: In raw data, a top-ranked search result was cited roughly twice as often as the fifth-ranked result (85.1% vs. 42.8%). However, when researchers experimentally reversed the order of matched sources, the actual causal impact of position shrank dramatically, registering near zero in follow-up tests.
  • The Power of Structure: Pages rewritten with clean headings and lists captured an average of 0.50 more citation markers per answer compared to identical text formatted as plain paragraphs.
  • The Chaos of Model Randomness: When the researchers re-ran identical inputs 120 times, the decision to cite or omit a target page flipped in 15% of the cases—roughly one out of every seven runs. Statistical modeling suggested that up to 45% of single-run variance is driven purely by model randomness.
  • An Attribution Warning, Not a Playbook: The authors explicitly caution that their findings should be viewed as an "attribution-sensitivity warning" rather than a blueprint for tactical SEO optimization.

2. Chronology: How the Experiment Was Conducted

To understand the robustness of Selvam and Ghosh’s findings, it is essential to look at the step-by-step methodology deployed in the study.

Phase 1: Prompting and Transcript Gathering

The researchers prompted the GPT-5.4 search agent to answer 130 common, real-world questions by letting it perform independent, unconstrained web searches. Every interaction, message, and search result from 129 successfully addressed questions was meticulously recorded and saved.

Phase 2: Pair Identification and Fact-Checking

From these transcripts, the researchers filtered for pairs of web pages that met three strict criteria:

  1. They appeared in the same search engine result outcomes.
  2. They were independently screened as supporting the exact same factual claim.
  3. They offered genuine substitutes, meaning either page could have been fairly and logically cited by the AI.

This filtering process left 113 valid page pairs. To ensure absolute objectivity, a subsequent blinded human evaluation verified 103 of these pairs as authentic, interchangeable matches. Because both pages confirmed the exact same underlying fact, any disparity in how the AI distributed credit could be traced directly to formatting or positioning, rather than factual superiority.

Phase 3: The Replay Matrix

The researchers replayed each saved conversation four distinct ways:

  • By placing one target page either above or below the other within the prompt context.
  • By presenting the page text either as plain, dense paragraphs or as structured layouts featuring headings, bulleted lists, or tables.

To eliminate human writing bias, the text versions were generated via AI rewrites—primarily using Grok 4.3, with GPT-5.4 serving as a fallback for a single pair. A separate Grok review verified that the factual integrity remained intact across versions. Because the wording varied slightly between formats, the authors transparently noted that the test compared AI rewrites rather than isolating pure visual formatting.


3. Supporting Data: Position Gaps vs. Swap Effects

One of the most eye-opening revelations of the preprint involves the stark disconnect between raw observational data and true experimental causation.

The Raw Position Gap

In the initial search transcripts, pages appearing in the top position of an Exa search call were cited 85.1% of the time. In contrast, pages occupying the fifth position were cited only 42.8% of the time. This creates a massive raw gap of 42.3 percentage points.

However, the researchers emphasize that "position" in this context refers purely to the order of the five Exa results returned in a single API search call—not a page’s organic rank on Google or its position on the live web. Because search providers naturally place higher-relevance sources at the top, the raw 42.3-point gap conflates source quality with raw placement.

Testing the Causal Effect of Position

When the researchers directly intervened by moving the exact same page higher within its designated pair, its baseline chance of being cited increased by 7.9 percentage points. However, after adjusting for multiple statistical tests, this finding was not deemed statistically significant.

Furthermore, in a separate testing subset consisting of 56 pairs where source order was strictly reversed, the estimated causal effect of position dropped to 0.0 points, with a 95% confidence interval stretching from -5.4 to +5.4.

The data confirms a nuanced reality: while search position does influence AI citations in isolated scenarios, relying on raw observational position data as a definitive performance metric is fundamentally misleading.

Structured Rewrites and Credit Concentration

When examining page architecture, the study found that pages formatted with headings and lists secured an average of 0.50 more citation markers per answer than their plain-paragraph counterparts (95% confidence interval: 0.20 to 0.84).

Interestingly, the total number of citations per answer did not inflate; instead, the AI shifted its allocation, concentrating more credit onto the structurally optimized page while pulling weight away from competing sources.

When evaluating whether structured text increased the baseline likelihood of being cited at all, the researchers observed a 4.5 percentage point increase (95% CI: -1.4 to +10.4). Because the study’s statistical power could reliably detect effects only at or above 8.5 points, the authors characterized this specific metric as inconclusive.


4. Official Responses and Industry Context

The findings arrive amid a broader reckoning within the digital marketing and SEO communities regarding the reliability of correlation studies published by software vendors.

In May, an Ahrefs report revealed that web pages cited by AI search tools were roughly three times more likely to incorporate JSON-LD schema markup. Yet, when marketers subsequently tested adding schema to live pages, it failed to trigger a predictable, direct increase in AI citations.

Similarly, a landmark study published by SparkToro in January demonstrated that leading AI platforms—including ChatGPT and Google’s AI Overviews—produced identical brand recommendations less than 1% of the time when fed the exact same prompt repeatedly.

Selvam and Ghosh’s findings reinforce these industry anxieties. Their rerun tests—where 120 responses were processed multiple times using identical inputs—revealed that the AI’s binary decision to cite or ignore a target page flipped in 15% of cases (roughly one in seven runs). With model randomness accounting for an estimated 45% of total variance, the authors argue that evaluating an AI citation strategy based on a single search query is statistically meaningless.


5. Implications for Content Creators, Marketers, and AI Developers

The publication of this preprint holds profound implications for how the tech industry approaches search engine optimization, content structuring, and AI evaluation.

1. The Death of Single-Query Benchmarking

For months, brand managers have panicked or celebrated based on whether a single prompt in ChatGPT or Claude recommended their product. Selvam and Ghosh’s data proves that stochastic variance (model randomness) makes single-query testing a fool’s errand. Because roughly 45% of variance is driven by internal model noise, any serious audit of AI visibility must aggregate data across dozens or hundreds of repeated runs.

2. Correlation Does Not Equal Causation in AI SEO

The massive gap between raw position metrics (a 42.3-point spread) and true swap effects (approaching 0.0 points) serves as a stark warning to SEO tool vendors. Marketers must be skeptical of vendor reports claiming that specific tactics—such as keyword placement, schema injections, or specific heading structures—guarantee AI citations. Unless studies isolate variables through controlled, experimental swaps, they are merely capturing correlation, not causation.

3. The Limits of Offline Rewriting

It is vital to note the methodological boundaries of the paper. Because the study relied on offline replayed conversations and AI-generated text rewrites, it cannot confirm whether reformatting a live website will boost citations in the wild. The experiment bypassed critical real-world hurdles such as web crawling, server latency, tokenization limits, and dynamic ranking retrieval algorithms.

4. A Call for Methodological Rigor

In their discussion section, the authors hammer home their core thesis:

"This is an attribution-sensitivity warning, not an optimization tactic."

Moving forward, the researchers urge the AI evaluation community to adopt rigorous testing standards. This includes running scenarios multiple times to measure consistency, testing across diverse search providers and foundational models, and evaluating both the frequency of citations and the binary probability of inclusion.

As generative engines continue to reshape the digital landscape, studies like Selvam and Ghosh’s provide a much-needed anchor of scientific skepticism, reminding us that beneath the polished veneer of AI search answers lies a deeply volatile, probabilistic engine.

Leave a Reply

Your email address will not be published. Required fields are marked *