How should an AI hotel concierge prove it is accurate?
TL;DR
A demo is not evidence. Fixed question sets with unanswerable questions, refusal accuracy, isolation checks and deploy gates — plus what FluxPMS publishes.
An AI concierge should prove accuracy the way any measurement gets proven: with a fixed question set that includes questions the system cannot possibly answer, separate scores for finding the right source and for answering from it, a refusal score that measures how often it declines instead of inventing, and a published threshold that blocks a release when it's missed — all stamped with the date and build the numbers came from. A demo proves nothing, because a demo only ever shows the questions the vendor chose.
Why a demo is not evidence
Every AI concierge demo works. That is the point of a demo. The vendor asks about breakfast times and parking, the assistant answers fluently, and the room nods. Nothing in that exchange tells you what happens at 11pm when a guest asks about a spa your hotel doesn't have, or quotes a rate that was never published, or asks about a cancellation policy that doesn't exist.
Those are the questions that cost money. A confidently wrong answer about a policy becomes a dispute at the desk; a confidently wrong rate becomes a guest holding a screenshot. The vendor's reputation doesn't absorb that cost. Yours does.
Get hotel revenue tips + product updates — no spam
The four properties of a real accuracy test
A fixed question set, written against real hotel content. Fixed matters more than large. If the questions change between runs, you can't compare this month's number to last month's, and you can't tell an improvement from an easier test.
Questions with no correct answer, deliberately included. This is the part vendors skip. A test set made only of answerable questions measures fluency, not accuracy. Inventing questions the content cannot support — a nonexistent amenity, an unpublished rate — is the only way to measure whether the system knows what it doesn't know.
Retrieval scored separately from answering. These fail differently. If the system never retrieved the passage containing the answer, the answer was always going to be a guess, and fixing the phrasing won't help. Splitting the two tells you which half broke.
A gate that blocks the release. A number on a page is a report. A number that stops a deploy when it's missed is a control. The difference is whether the vendor has to act on a regression or merely disclose it.
Refusal accuracy is the metric that matters most
Of everything measured, the share of unanswerable questions the system correctly refuses is the one that predicts real-world damage. A system scoring well on answerable questions and badly on refusals is one that will confidently invent a policy on a Tuesday night, and the hotel finds out from the guest.
The mirror metric matters too: how often the system refused something it did have the answer to. Refusing everything is a trivially perfect refusal score and a useless product. Both numbers have to be read together, which is why an accuracy claim consisting of a single percentage is not a claim you can evaluate.
Isolation is an accuracy question too
If your assistant runs on a platform serving other hotels, one more thing needs proving: that another property's content can never answer your guest. A guest asking about "the rooftop bar" should get your rooftop bar or a refusal — never a neighbouring hotel's, retrieved because the phrasing matched. The correct measurement is blunt: passages returned from a property that isn't yours, and the only acceptable number is zero.
What FluxPMS publishes for Aria
The Aria accuracy report is public, and it's built to the shape described above rather than as a marketing page. The question set is 90 questions written against real hotel content: 72 that have a correct answer to find, and 18 invented on purpose that must be refused. Three things are scored separately — search hit rate (whether the passage holding the answer is among the top passages read before replying), chat answer pass rate, and refusal accuracy — alongside how often Aria wrongly refused something it could have answered, and a count of passages returned from a property that isn't the asking hotel's.
Each of those carries a published deploy gate: below 90% search hit rate, below 95% answer pass rate, or below 90% refusal accuracy, the change doesn't ship. The page also stamps the date of the run and the engine build the figures came from, so a number can be checked for staleness rather than taken on trust. Current figures live on that page — they're the output of a specific run, not a permanent claim, which is exactly why they're dated. Aria itself is designed to the same boundary: it answers from your property's own content, shows the source next to the reply, and escalates to a human when nothing relevant comes back, rather than processing payments or changing reservations on its own.
Run the test yourself, on any vendor
You don't need a lab. Three questions in a live demo do most of the work:
- Ask about something the hotel does not have — an amenity nobody mentioned. Watch whether it refuses or improvises.
- Ask for a rate or policy that was never published. Same test, higher stakes.
- Ask the vendor what its refusal rate is, when it was last measured, and what happens to a release that misses it.
The third question is the one a vendor either answers with a number and a date, or deflects. Both responses are informative.
The short version
Accuracy is a measurement, not an adjective, and the measurement has to include the questions the system can't answer. Demand a fixed question set with deliberate unanswerables, retrieval and answering scored apart, a refusal number read alongside a wrongly-refused number, cross-property isolation proven at zero, a gate that blocks releases, and a date on the whole thing. Then go break the demo yourself — see what an AI can and can't reliably do for a small hotel for where the useful boundary sits, and start a 30-day trial if you'd rather test it on your own content than on a script.
Related
Want the switching checklist?
Email + a one-page comparison. No spam. Or start free — $0 today.
Sent. Check your inbox.