Skip to content

Methodology

Four decisions determine what actually gets measured. Each was made after the previous version produced a wrong result on live data — what went wrong is written out below.

1. A question about the category, not about you

The first version asked the platform best travel companies similar to <name>. The platform would go looking for that name, find the company’s own site and cite it — which handed a mention to almost any business with a working website.

That measures whether a company has a website, not whether it gets found when someone asks about the category. Those results are not used. The question is now phrased the way a traveler would ask it — “best dive centres in Seychelles” — with no name inside it.

2. A mention only counts with a citation

A name matching somewhere in the text does not count. Models routinely repeat the company name from the question in their own preamble with no source behind it — counting that means measuring the echo of our own query.

A mention requires a real citation: a source marker attached to the phrase where the company is named. This bug reached working code three times — in the Claude, Perplexity and ChatGPT providers — and all three times it was caught only after a live answer was read by eye.

3. Two variances, not one ±

Within-day spread (platform non-determinism) and between-day spread (drift in the results themselves) are counted separately and never collapsed into one number.

This changes the conclusion, not the presentation. On the first day the spread reached ±37.7 and was called “model noise.” Across three days it became clear that between-day spread almost always exceeds within-day spread: for one company, 42.3 against 15.1. The main source of movement is drift in the results, not randomness in the models.

The practical consequence: a single-day reading cannot be presented as a measurement.

4. Comparable by engine count, not by date

The score is divided by the number of platforms that answered. When one API key stopped working on 24 August, every niche ran on two engines — and that day’s average score came out higher than the day before with no improvement in visibility whatsoever: a mention by one platform out of two scores 50 instead of 33.

So days are selected by the number of engines that answered, recorded on every row, rather than by the calendar. A day with an incomplete set is excluded entirely.

The query language decides whose market you can see

The language is chosen by measurement, not by a country’s official language. Tested on four countries both ways: in Thailand the English query returned twice as many real operators as the Thai one; in Brazil English returned zero and Portuguese returned nine.

What predicts it is not the country’s language but whether the industry faces outward. Koh Tao and Phuket sell to foreigners and publish in English; Patagonia and the Brazilian coast sell to locals. Had the rule been “translate everything non-English,” Thailand would have been measured noticeably worse.

What these measurements do not show

This section is mandatory and is never shortened.

  • The score is not a share of travelers. It is the share of platforms with a confirmed mention, not the share of people who will see you.
  • One phrasing is not a niche. The same company scored 38, 100 and 67 on one question across three different days, with nothing changed on its side.
  • The data shows absence, not its cause. Why a company is missing from an answer does not follow from these measurements.
  • Position is not weighted. A mention on the first line and on the seventeenth line of an answer count the same.
  • Gemini is excluded. The Gemini API terms for Grounding with Google Search prohibit analysing results and using links to build an index — which is exactly what this measurement is. Perplexity, Claude and ChatGPT are measured.

What has changed in the method itself

The counting rule changed on 2 September 2026: previously, in an enumeration naming several companies under a single citation, only the last one named was credited. The error was one-directional — it understated visibility — and it lived in two engines out of three.

Measurements before and after that date are not comparable with each other. We write this here because any chart crossing 2 September will show a step, and it is easy to mistake for a rise in visibility.