Quoting someone beats claiming it yourself
Two of the nine tactics tested are almost the same instinct, pointed in opposite directions. One makes your own text more authoritative — the paper’s Authoritative, defined as modifying the style of the content “to be more persuasive and authoritative”. The other adds someone else’s words — Quotation Addition, which “adds relevant citations and quotations from credible sources”.
Setting those two rows beside each other is our comparison, not a finding the authors state. They report both methods in one table, from one run against one benchmark, which is what makes the two numbers comparable at all; the conclusion drawn from the pair is ours.
Sounding credible yourself scores 21.8. Borrowing someone else’s credibility scores 27.8. An untouched page scores 19.5.
The gap, stated carefully
Those are absolute scores on the paper’s scale, not percentage gains. Expressed as change against the 19.5 baseline, writing more authoritatively is +11.8%; adding a quotation is +42.6%. On the paper’s overall column the authors put quotation at roughly +40.9% and call it the best-performing method of the nine.
The pattern underneath
The three methods that gained most are all forms of external corroboration rather than self-assertion: quotation at +42.6%, statistics at +32.8%, citing sources at +27.7% — against +11.8% for changing how your own voice sounds.
The authors’ explanation is short: citations “provide a source of verification for the facts presented, thereby enhancing the credibility of the response”. An engine assembling an answer is looking for material it can stand behind, and a page carrying verification is easier to stand behind than a page that asserts.
Where it applies unevenly
Quotation Addition was strongest in “People & Society”, “Explanation” and “History” queries; Statistics Addition in “Law & Government” and opinion questions; Cite Sources was “particularly beneficial for factual questions”. The paper reports no travel domain.
What this does not mean
It is not a licence to invent quotations. What was tested is adding a quotation from a credible source that genuinely exists. A fabricated one is a different act with different consequences, and nothing here measures it.
“Authoritative” is not “expert”. That method changes style, not substance — more persuasive phrasing over the same content. This is not evidence that expertise is worth less than a quote; it is evidence that sounding authoritative is.
The engine is not today’s. Answers were generated with GPT-3.5-turbo over the top five Google results, in 2023–2024. The direction is informative; the magnitude is dated.
A language model graded one of the two metrics. Subjective Impression was scored by GPT-3.5 acting as judge. No human evaluation is reported.
None of the queries are about travel. GEO-bench is built from general web-search benchmarks — 10,000 queries from nine public datasets across 25 domains. Nothing in it measures a dive centre, a DMC or a hotel. That part is ours to measure, and it is a separate exercise with separate numbers.
We did not run this experiment. This is a reading of someone else’s. Our own scans measure something different: whether a platform names a business, counted only where a citation backs it.
Source
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande. GEO: Generative Engine Optimization. KDD 2024. arXiv:2311.09735v3, 28 June 2024, Table 1. Read 6 September 2026.
How everything above is counted — methodology.