AI Citation Test Finds Source Order Matters Less Than It Looks

A new preprint examined whether or not tweaking only one a part of a supply adjustments AI search citations, with all the things else saved the identical. Within the uncooked numbers, the highest outcome bought cited about twice as usually because the fifth. When researchers Sriram Selvam and Anneswa Ghosh reversed the order of matched sources, the influence was a lot smaller and measured zero in a follow-up take a look at.

Posted to arXiv on September 14, the paper isn’t peer-reviewed. The examine covers one GPT-5.4 search agent that makes use of Exa as its search supplier, that includes offline replayed conversations and no dwell webpage edits.

How The Check Labored

The researchers prompted the GPT-5.4 agent to reply 130 frequent questions by having it carry out unbiased internet searches. They recorded each message and search outcome from the 129 questions it addressed. From these transcripts, they selected pairs of pages that appeared in the identical search outcomes and had been each screened as supporting the identical reality. This screening aimed to search out conditions the place both web page could possibly be pretty cited. When true matches had been recognized, any credit score variations had been as a result of how the mannequin apportioned recognition between the 2 sources, each confirming the identical reality.

That left 113 pairs. A later blinded human verify confirmed 103 of them as real matches. The researchers replayed every saved dialog 4 methods, putting one web page above or beneath the opposite and displaying its textual content both as plain paragraphs or rewritten with headings and lists or a desk. Solely the ultimate reply was generated once more.

Each variations of the textual content had been generated by AI rewrites of the unique web page. Grok 4.3 created almost all of them, with GPT-5.4 used as a fallback for one pair, and a separate Grok assessment checked that the details matched. The wording varies between the 2 variations, so the authors be aware that the take a look at compares two rewrites however doesn’t particularly isolate formatting variations.

Uncooked Place Hole Was Bigger Than Swap Results

Within the preliminary place of a search name, pages had been cited 85.1% of the time in saved transcripts, in comparison with 42.8% for pages within the fifth place. This creates a distinction of 42.3 share factors.

Right here, ‘place’ merely means the order of the 5 Exa outcomes returned in a single search, not the place a web page ranks on Google or its place on the dwell internet.

The examine highlights that search suppliers often put extra related pages on the high, so the uncooked distinction displays each the place and the standard of the pages. When the researchers moved the identical web page increased inside its pair, the possibility it was cited in any respect went up by 7.9 share factors. Nonetheless, this discovering wasn’t thought of statistically vital after accounting for a number of assessments.

One other testing set with 56 pairs, the place solely the order was switched, confirmed an estimate of 0.0 factors, with a 95% confidence interval from -5.4 to +5.4.

The study explains that the uncooked hole and the swap outcomes measure completely different points. General, it means that place did affect citations in some instances, however averages from uncooked place information aren’t dependable.

Structured Rewrites Acquired Extra Credit score, Not Clearer Entry

Pages that had been rewritten with headings and lists acquired a mean of 0.50 extra quotation markers per reply in comparison with the identical pages written as plain paragraphs, with a 95% confidence interval from 0.20 to 0.84. The solutions within the take a look at had been closely cited, with a median of 29 markers throughout six paperwork.

The entire variety of citations per reply didn’t rise, and the quantity on the opposite web page barely modified. The authors see this as credit score being targeted extra on the rewritten web page.

The primary take a look at the researchers carried out, which they deliberate earlier than beginning the experiment, was to see if the web page bought cited in any respect. They discovered that utilizing structured textual content elevated that chance by 4.5 share factors, with a 95% interval from -1.4 to +10.4. The paper factors out that this outcome isn’t conclusive and mentions that the examine may reliably detect solely results of about 8.5 factors or extra.

A extra strict comparability, the place each phrase stayed the identical however the format was adjusted to 1 sentence per checklist row, boosted quotation charges throughout all 113 pairs. After they repeated the take a look at with a subset, the impact reversed.

Within the dialogue part of the paper, the authors shared these insights:

“That is an attribution-sensitivity warning, not an optimization tactic.”

Reruns Modified Quotation Outcomes

The researchers examined 120 responses once more utilizing the identical inputs, and located that the choice to quote or not for the goal web page modified in 15% of those instances, roughly one in seven.

The typical rely impact remained constant throughout these reruns. They estimate that about 45% of the variation in a single run’s impact is because of mannequin randomness.

The authors advocate rerunning quotation assessments a number of occasions and sharing how constant the outcomes are throughout these runs.

Moreover, SparkToro reported in January that ChatGPT and Google’s AI Overviews every produced the identical model checklist lower than 1% of the time when given the identical immediate repeatedly.

Why This Issues

The uncooked place hole on this take a look at was a lot bigger than the typical impact noticed when researchers swapped supply order. An Ahrefs report from Could confirmed pages cited by AI had been about thrice extra more likely to embody JSON-LD schema, however including schema didn’t clearly enhance citations.

This raises questions on whether or not a correlation in a vendor report or your monitoring was ever examined by altering the variable, and a single reply is a weak foundation for labeling a quotation as gained or misplaced.

The examine can’t affirm if reformatting a dwell web page boosts citations, since rewrites solely utilized to textual content already retrieved, excluding crawling, retrieval, and rating processes.

Trying Forward

The researchers re-ran the saved searches on Grok 4.3, discovering that the structured rewrites leaned the identical approach. Nonetheless, lower than half of Grok’s first replies adopted the right quotation format.

The authors advocate extra analysis to check every situation a number of occasions, discover completely different search suppliers and fashions, and take note of each how usually citations happen and if a web page is cited in any respect.

Featured Picture: Accogliente Design/Shutterstock


#Quotation #Check #Finds #Supply #Order #Issues

Leave a Reply

Your email address will not be published. Required fields are marked *