Digital Marketing

New Preprint Examines How AI Search Agents Allocate Citations and the Surprising Influence of Document Formatting

As artificial intelligence increasingly mediates how users discover information online, digital marketers, content creators, and SEO professionals have grown intensely focused on understanding how AI search engines choose which sources to credit. A new, non-peer-reviewed preprint posted to arXiv on September 14 by researchers Sriram Selvam and Anneswa Ghosh sheds light on this opaque ecosystem. By testing whether tweaking isolated components of a source changes how an AI search agent distributes citations—while keeping all other variables constant—the study challenges traditional assumptions about raw search rankings, formatting optimization, and the reliability of single-run AI evaluations.

The research focuses specifically on a GPT-5.4 search agent utilizing Exa as its underlying search provider. By replaying offline conversations without live webpage edits, the authors sought to isolate the mechanics of AI attribution. Their findings suggest that while structural adjustments like headings and lists can concentrate credit onto a specific page, raw position gaps in search outcomes do not necessarily translate to direct causal boosts. Furthermore, the high degree of variance observed across repeated model runs underscores the volatility inherent in generative engine optimization (GEO).

Methodology and Experimental Design of the Study

To evaluate citation attribution systematically, Selvam and Ghosh prompted the GPT-5.4 search agent to address 130 common questions through independent web searches. The system successfully addressed 129 of these prompts, and the researchers recorded every message and search result generated during the interactions.

From these transcripts, the researchers filtered for pairs of pages that appeared together within the same search outcomes and were independently verified as supporting the exact same factual claim. This stringent screening process ensured that either page could be legitimately and fairly cited by the model. When genuine matches were identified, any disparities in citation credit could be attributed entirely to how the model divided recognition between two equally valid sources.

This filtration left 113 valid pairs, of which 103 were confirmed as genuine matches via a subsequent blinded human evaluation. To test how presentation affected outcomes, the researchers replayed each saved conversation across four distinct variations. They positioned one page above or below its counterpart and rendered its text either as plain paragraphs or formatted with headings, lists, or tables. Crucially, only the final answer was regenerated by the model during these replays.

The textual variations were generated primarily through AI rewrites using Grok 4.3, with GPT-5.4 serving as a fallback for a single pair. A separate review by Grok confirmed that the factual accuracy remained consistent across versions. Because the wording varied between the standard and reformatted versions, the authors noted that the test compared two distinct rewrites rather than cleanly isolating pure layout adjustments.

Raw Position Gaps Versus Causal Swap Effects

One of the most revealing aspects of the study involves the discrepancy between raw observational data and controlled causal testing. In the initial positions of the search calls, pages appearing at the top were cited in 85.1% of saved transcripts, compared to just 42.8% for pages appearing in the fifth position. This created a raw difference of 42.3 percentage points.

However, the authors emphasize that "position" in this context refers strictly to the order of the five Exa results returned within a single search query, rather than traditional Google rankings or live web placement. Search providers inherently populate the top tiers with what algorithms deem to be the most relevant and authoritative pages. Consequently, the raw gap conflates raw presentation order with underlying content quality.

When the researchers experimentally swapped the positions of matched pages—moving the same page higher within its pair—the likelihood of it being cited increased by a modest 7.9 percentage points. Yet, this finding was not deemed statistically significant after controlling for multiple statistical tests. In a secondary testing subset comprising 56 pairs where only the order was inverted, the estimated impact of swapping was calculated at precisely 0.0 points, with a 95% confidence interval ranging from -5.4 to +5.4.

These results demonstrate that while search engine positioning correlates heavily with citation frequency in observational data, raw position averages cannot be reliably interpreted as direct causal drivers of AI preference.

Impact of Structured Text Rewrites on Citation Credit

While changing a source’s rank yielded minimal causal impact, altering its document structure produced more measurable shifts in attribution. Pages that were rewritten to incorporate structured elements—such as headings and bulleted lists—received an average of 0.50 more citation markers per answer than identical information formatted as plain paragraphs. This metric carried a 95% confidence interval ranging from 0.20 to 0.84, within a testing environment characterized by heavy citation density (a median of 29 markers across six documents).

Significantly, the total number of citations per answer did not increase when structured text was introduced, nor did the citation count for the opposing page change substantially. The authors interpreted this dynamic as a redistribution of credit, indicating that structured formatting successfully concentrated the model’s attention and recognition onto the reformatted page.

When examining the primary pre-planned hypothesis—whether structured text increased the baseline likelihood of a page being cited at all—the researchers observed a 4.5 percentage point increase, with a 95% confidence interval spanning from -1.4 to +10.4. Because the study could reliably detect effects only at or above approximately 8.5 points, the authors characterized this finding as inconclusive.

A stricter test, in which the wording remained completely identical while the layout was adjusted to isolate sentence-per-row list formatting, initially boosted citation rates across all 113 pairs. However, when this test was repeated on a subset, the effect reversed entirely. Summarizing these nuanced findings in the paper’s discussion section, the authors issued a clear caveat for digital strategists: "This is an attribution-sensitivity warning, not an optimization tactic."

Volatility and the Challenge of AI Reproducibility

Beyond structural and positional variables, the study uncovered significant baseline volatility in how AI search agents behave across repeated executions. When the researchers re-ran 120 responses using identical inputs, the decision of whether or not to cite the target page changed in 15% of the cases—roughly one out of every seven runs.

While the average count effect remained relatively stable across these reruns, the authors estimated that approximately 45% of the variation observed in a single test run stems purely from model randomness and internal stochasticity. In light of this, the researchers strongly recommend that future studies on AI citations execute tests across multiple iterations and explicitly report the consistency of outcomes.

This finding aligns with broader industry observations regarding the unpredictable nature of generative artificial intelligence. For instance, an analysis published by SparkToro revealed that ChatGPT and Google AI Overviews produced identical brand lists less than 1% of the time when subjected to repeated identical prompts. Such high volatility complicates the efforts of marketers trying to benchmark their visibility within AI-driven discovery channels.

Broader Industry Implications and Limitations

The findings of Selvam and Ghosh intersect with a growing body of empirical research examining the mechanics of Generative Engine Optimization (GEO). A previous report published by Ahrefs noted that web pages cited by AI search tools were approximately three times more likely to incorporate JSON-LD schema markup. However, controlled testing indicated that actively adding schema did not reliably or directly increase citation rates.

These combined insights raise critical questions for digital marketers who rely on vendor correlations or proprietary SEO tracking tools. The research demonstrates that observing a correlation—such as the presence of schema markup or top-tier search placement—does not guarantee a causal relationship. Furthermore, because a single AI-generated answer exhibits high variance, relying on one test run to declare a citation strategy successful or failed is statistically unsound.

The study also outlines clear methodological boundaries. Because the experiments relied on offline replays of text already retrieved by the system, the research could not evaluate how reformatting a live webpage might influence upstream crawling, vector retrieval, and ranking algorithms used by search providers.

Future Outlook and Recommendations for Researchers

To build upon these findings, the authors of the preprint recommend that future academic and commercial research adopt rigorous multi-run testing methodologies. Investigators are encouraged to evaluate multiple AI search providers and underlying foundation models, while simultaneously measuring both binary citation rates (whether a page is cited at all) and continuous citation counts (how frequently a page is referenced within an answer).

As artificial intelligence continues to reshape the landscape of web traffic and information dissemination, empirical studies like this one provide essential guardrails against premature optimization tactics. By demonstrating that AI attribution is sensitive to document structure yet heavily influenced by intrinsic model randomness, the research underscores the need for methodological rigor in the nascent field of AI search visibility.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button
Jar Digital
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.