Keyword stuffing does nothing to AI search. Here is the measurement.
Every agency still selling keyword density is selling something that has been measured and did not work.
Not measured by us. In GEO: Generative Engine Optimization, the authors built a benchmark of 10,000 queries, generated AI answers over real search results, and tested nine ways of editing a source page to see which made the engine cite it more. Keyword stuffing — their term, defined as modifying content to include more keywords from the query, “as expected in classical SEO optimization” — was the only method that came back below the untouched page.
The numbers
On the paper’s citation-prominence metric an unedited page scores 19.5. Stuffed with the query’s keywords, the same page scores 17.8.
Two methods that did work, for contrast: adding a relevant quotation scores 27.8, adding statistics 25.9.
Those four are absolute scores on the paper’s own scale — Table 1 is captioned “absolute impression metrics”. They are not percentage gains. Stated as changes against the 19.5 baseline, quotation is +42.6%, statistics +32.8%, and keyword stuffing −8.7%. The absolute and the relative figures describe one result and must never be swapped for each other: a “+27.8%” that nobody wrote is how a real finding turns into a fake one.
What the number counts
Position-adjusted word count: how much of the generated answer consists of sentences citing your source, weighted so an early citation counts for more than a late one. It measures prominence inside one answer. It is not traffic, not clicks and not bookings, and the paper does not claim otherwise.
What this does not mean
It is not evidence that keyword stuffing is harmful. The drop is 1.7 points and that table publishes no error bars. The authors’ own wording is that such methods “offer little to no improvement”, and we will not put a stronger claim on it than they did. “Does nothing” is supported. “Actively hurts you” is not.
The engine is not today’s. Answers were generated with GPT-3.5-turbo over the top five Google results, in 2023–2024. The direction is informative; the magnitude is dated.
A language model graded one of the two metrics. Subjective Impression was scored by GPT-3.5 acting as judge. No human evaluation is reported.
None of the queries are about travel. GEO-bench is built from general web-search benchmarks — 10,000 queries from nine public datasets across 25 domains. Nothing in it measures a dive centre, a DMC or a hotel. That part is ours to measure, and it is a separate exercise with separate numbers.
We did not run this experiment. This is a reading of someone else’s. Our own scans measure something different: whether a platform names a business, counted only where a citation backs it.
Source
Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande. GEO: Generative Engine Optimization. KDD 2024. arXiv:2311.09735v3, 28 June 2024, Table 1. Read 6 September 2026.
How everything above is counted — methodology.