The Hotel Website May Not Be Where AI Decides Which Hotels Matter

A new AGR study of 148 luxury hotels found website AI-readiness signals had no detectable association with recommendation frequency, while Forbes and Michelin tracked far more of the variation.

A study of 148 luxury hotels finds Forbes Travel Guide ratings and Michelin Keys explain 55% of AI recommendation frequency, while website schema and llms.txt files show near-zero correlation.

Hotels have spent the last year making their websites easier for AI to read.

Add structured data. Publish an llms.txt file. Keep the crawlers open. Make every property fact machine-readable. Clean up the entity record.

None of that is irrational.

But the question may be one stage too late.

What if part of the hotel-selection process occurs before the observable web retrieval sequence even begins?

We believe AI may already have enough internal knowledge to form an answer before the traveler asks the question, with web retrieval then validating, updating, or refining that answer. We did not test that proposition here.

A new Americas Great Resorts study of luxury hotel AI recommendations examines 148 luxury hotels that ChatGPT, Google AI Mode, or Gemini had already recommended at least once across six U.S. markets.

Its question is narrower: among hotels already inside the recommendation set, what distinguishes the properties named repeatedly from those named only occasionally?

That boundary matters. The study does not tell us what gets a hotel into the consideration set in the first place.

What it found inside the set raises a different problem.

The website variables barely separated the hotels

Each of the 148 recommended hotels was evaluated for lodging-specific schema, structured-data completeness, llms.txt presence, crawler directives, and site accessibility.

The correlation between structured-data completeness and recommendation frequency was 0.02.

A model containing website infrastructure variables plus market accounted for 2.8% of the variation in log recommendation frequency. Market alone accounted for 2.4%.

Hotels with an llms.txt file averaged 5.1 recommendation slots. Hotels without one averaged 5.7. The difference was not statistically distinguishable from zero.

That does not mean schema is useless. It does not mean crawler access is irrelevant. It does not mean a hotel's website cannot matter to initial consideration.

The study cannot answer that last question because every hotel in this analysis had already been recommended.

What it does show is narrower: among hotels already recommended at least once, the website variables measured in this study did not detectably distinguish the properties named most often from those named less often.

A separate September observation did not test recommendation frequency, but it suggested a potentially relevant part of the public record to examine.

We watched ChatGPT search

The study discloses the research sequence because it matters.

In September, three ChatGPT sessions were observed with a browser tool that exposes the web searches and retrieved pages associated with an answer. Those sessions are what prompted Forbes Travel Guide and Michelin to be added as variables for the broader 148-hotel analysis.

The three sessions asked for the top luxury hotels in Miami, required ChatGPT to choose a single hotel in Miami, and asked for the top luxury hotels in Napa Valley. Each was run in a fresh chat with personalization disabled.

Before any retrieved page came back, the first exposed Miami search query already named all five hotels that later appeared in the final answer.

The single-hotel Miami query already named the hotel ChatGPT ultimately selected.

In Napa Valley, the first exposed query named four of the five hotels that later appeared in the final answer.

Then the pages started coming back.

For the Miami top-five question, the trace exposed 98 distinct pages. Every one was on Forbes Travel Guide or Michelin Guide.

For Napa Valley, 36 of 45 pages, or 80%, came from Forbes or Michelin.

A fourth session was then run logged out in a private window. It returned the same top three for Miami, while drawing from a wider set of sources for positions four and five.

That does not tell us what happened before the first exposed query.

The tracing tool records only the searches and pages ChatGPT exposes. It cannot see model parameters, training data, stored memory, internal ranking logic, or any process that precedes that query.

What the trace does establish is narrower: in the two Miami sessions, the final-answer hotels were already present in the first exposed query, and in Napa four of five were. The pages subsequently returned in the observable retrieval trace were therefore not needed to introduce those final-answer hotel names into the first exposed query.

The concentration of subsequent retrieval on Forbes and Michelin was consistent with those sources being used to support or evaluate hotels already present in the first exposed query, rather than those hotels being identified through the exposed retrieval sequence itself.

Consistent with. Not proof of.

One disclosure belongs here. AGR publishes luxury hotel market rankings. AGR material appeared among cited sources in two of the six markets measured in the AGR Luxury Hotel AI Visibility Index and in one of four September ChatGPT sessions. We are part of the public record we are measuring.

So we tested the registries

The September sessions told us where to look. The next question was whether that observation had any relationship to the much larger July recommendation dataset.

Forbes Travel Guide and Michelin were useful variables for another reason: their relevant ratings and Keys had been published before the July 29 recommendation capture.

Among the 148 already-recommended hotels, a model including Forbes Travel Guide rating, Michelin Key count, and market accounted for 54.7% of the variance in log recommendation frequency.

Forbes Five-Star hotels averaged 13.4 recommendation slots, compared with 2.6 for hotels with no Forbes rating.

Hotels with three Michelin Keys averaged 13.0 slots, compared with 3.8 for hotels with no Key.

The result did not depend on one specification. Resampling the 148 hotels put the Forbes/Michelin model between 46% and 66%. Dropping any one market and refitting moved it only between 52% and 58%. A zero-truncated count model left both Forbes and Michelin significant at p < 0.001, while schema and llms.txt remained null.

The composite credential index was positively associated with slot counts in all six markets. The Forbes/Michelin relationship was also positive across all three platforms, though weaker on Google AI Mode than on ChatGPT or Gemini.

This is still an association.

Forbes and Michelin may be sources the systems consult. They may instead be measuring hotel qualities the models recognize through many other sources. Both may be happening.

The study cannot separate those explanations.

It can say that the strongest measured public correlates in this sample were not website configuration variables. They were the independently maintained credential variables measured in the study, specifically Forbes Travel Guide ratings and Michelin Key designations.

They were not the only outside record that tracked frequency. Coverage across six travel publications correlated 0.53 with recommendation frequency on its own and lifted the Forbes/Michelin model from 54.7% to 59.4% once the registries were already included. That coverage was measured after the July capture, so it is reported as an association and nothing stronger.

But it points in the same direction, and it argues against reading this as a two-registry problem.

Three different problems are being called “AI visibility”

This is where the distinction becomes operational.

Making a hotel's facts machine-readable is one problem.

Getting the hotel into the set of properties a model is willing to consider is a second, larger problem.

Being named repeatedly once the hotel is already under consideration is a third.

The AGR study tests only the third problem, and only among hotels already named at least once.

Within that boundary, measured site variables did not distinguish recommendation frequency. Forbes Travel Guide rating and Michelin Key count were associated with much more of the variation.

That is a reason to separate the workstreams, not abandon technical hygiene.

A hotel still needs an accurate, accessible entity record. Schema should match the property. Crawlers should be able to reach what they need. Old names, rebrands, closures, amenities, and property facts should not contradict one another across the hotel's own surfaces.

But website hygiene is not the same thing as consideration-set formation.

And neither is necessarily the same thing as repeated selection once a property is already inside the set.

The harder part lives outside the website

Forbes and Michelin are not hotel-controlled sources.

A property cannot edit its rating. It cannot add a Michelin Key through schema. It cannot publish an llms.txt file and manufacture an independent inspection record.

Those credentials are not necessary and they are not sufficient. Among the 148 already-recommended hotels, properties with no credential from any of the three registries still took 14% of all recommendation slots. And in a separate 67-hotel control group of credentialed luxury hotels that the platforms never named at all, Crosby Street Hotel held three Michelin Keys and a Forbes Four-Star and appeared in none of the 180 answers.

So this is not a new optimization checklist.

The point is structural.

Hotel companies can change a website this afternoon. The independent public record is slower, distributed across other organizations, and only partly controllable.

That record includes the independent sources that name the property, place it in a market hierarchy, and either corroborate or contradict the hotel's own description of itself.

If hotel leaders treat schema deployment, public-record formation, consideration-set inclusion, and recommendation frequency as one technical problem, they risk measuring activity at one layer while the competitive difference is being created at another.

If the observable search trace begins with candidate hotels already in view, where is that candidate set being formed, and what independent record is shaping it?

View story source
Technology Travel Recommendations Agent Engine Optimization Schema Markup Direct Booking Forbes Travel Guide USA & Canada United States

Andrew Paul is Founder and Managing Director of Americas Great Resorts, a luxury hospitality demand infrastructure company operating since 1993. He works with independent luxury hotels, resorts, and cruise lines on demand origin strategy, upstream guest acquisition, AI-mediated discovery, and the structural conditions that determine whether marketing investment compounds or resets.

Americas Great Resorts is a luxury hospitality demand infrastructure company operating since 1993. We work with independent luxury hotels, resorts, and cruise lines in North America, Mexico, the Caribbean, and select international markets.

Comments

Comments for this content

0 comments available
Loading comments...