Where we test new AI fairly, and work out how best to solve real problems, before any of it goes near a real business.
How we test
The same rules for every test, whether it’s a new model or a new way to fix a problem. They’re there to stop us fooling ourselves.
The same task, fixed in advance
A real piece of work with a test set and an answer key written before anything is run.
A simple baseline to beat
A plain rule, or how a person does it today. New is only better if it beats that.
Scored blind, run more than once
Scores come from code or from someone who can’t see which system produced the answer. Results are shown as ranges.
Cost and time counted
A result that’s slightly better but ten times the price is written down that way.
A twin with nothing to find
Every test includes data with nothing hidden in it, to catch methods that invent patterns.
Failures published
What didn’t work goes on this page with the same weight as what did.
Track 1
New capabilities
New models, agents and tools arrive every month, each with a bold claim. We test the ones that could matter to a business on a real piece of work, and publish the result, including when the old way wins. None has been run yet. These are next, and each appears here once it’s done.
New models on a real job
When a new model comes out, is it better at drafting a reply from a venue’s price list, without getting a single fact wrong? And is it worth the cost?
Agents that use a screen
Can an agent copy a booking from one system to another with no integration, and how many mistakes does it make in a hundred?
Answers that know their limits
Can an assistant answer from a business’s own documents, show the source, and say “I don’t know” instead of guessing?
The simplest thing that works
For each new capability: would a rule, a spreadsheet or a person have done just as well?
Track 2
Problems, tested on a hotel that doesn’t exist
Why build a fake hotel?
You can’t ask a business for years of customer records to test an idea that might not work. So we made the records ourselves: a country hotel with 80 rooms and two wedding suites, five years of enquiries, bookings, payments and emails, with duplicates, missing reasons, dates typed five ways and a half-finished move from spreadsheets to a CRM. Then we hid patterns in it, sealed the answers away, and built a twin with the same mess and nothing hidden.
wrenholt-hall · zsh
$ python3 -m hotelsim generate --seed 42
Wrenholt Hall (fictional) · 80 rooms · 2 wedding suites · Oct 2021 to Sep 2026
enquiry_spreadsheet/Wrenholt Enquiries.xlsx4,535salesforce/Case.csv5,197bridebook/couples_manager_export.csv1,838function_diary/function_bookings.csv2,642pms/reservations.csv75,116accounts/bank_transactions.csv5,277mailbox/events_mailbox.jsonl55,097calendar/events_calendar.csv3,26512 files · 237,740 rows · answer key sealed · 11.9 s
$ python3 -m hotelsim generate --seed 42 --null
Same hotel, same mess, nothing hidden · 11.1 s
# hidden patterns: found in the planted hotel
# twin with nothing hidden: nothing found
# tests: 41 of 41 pass, on four seeds
Replay of a real run, 28 September 2026. Abridged.
What we built on top of it
A pipeline that takes the raw exports and gives each person a short list for the morning. AI reads the notes and emails and writes the wording. The joining, scoring and checking are ordinary code and statistics, because that’s what they’re good at.
1
Eight exports
Enquiry sheet, CRM, Bridebook, function diary, bookings, bank, mailbox, calendar
2
Clean and join
Duplicates merged, dates fixed, the same couple recognised across systems
3
One history each
Every couple and company, first contact to paid
4
Find and score
What’s still live, and what looks like the enquiries that booked before
5
Check every fact
Reasons built from the record; code confirms every number and date
6
A short list
Up to ten a day per person, with the reason and a next step
Wrenholt Hall · Friday 08:303 of 10
Andrew JacksonWedding · 106 guests
Wedding enquiry still waiting for our first reply
No reply from us is on record since the enquiry arrived on 25 Sep 2026
Planning for 106 guests
Budget given as £11,000
Next: Reply now with availability and prices, and offer a show-round
Kate & ThomasWedding
They have seen the venue and not decided
Came for a show-round on 22 Sep 2026 and has not decided yet
They emailed on 24 Sep 2026: “We’d like to book please!”
Quote on file for £14,123.96
Next: Phone to ask how the show-round felt; offer to hold the date
Craven Care Homes LtdCorporate
Regular corporate account has gone quiet
Last event with us was on 9 Dec 2025
That booking was worth £4,215.43 ex VAT
A regular since their first event on 24 Jul 2023
Next: Call the booker: ask what events they have coming up and offer two dates
Real output from the pipeline. The hotel, the names and the numbers are made up.
Keeping data safe was part of the experiment
Before any AI step sees an enquiry, names, emails and phone numbers are removed. Each client’s data lives in its own encrypted container, created and destroyed by script, with a certificate at the end listing what was deleted. Try a simplified version below.
What an AI model is allowed to see
Hi, it’s [NAME] [NAME] here, we emailed on Friday about Saturday 14 August 2027. You can reach me on [PHONE] or [EMAIL], or my partner [NAME] on [PHONE]. We’re 106 guests, budget around £11,000.
6 removed. The date, guest count and budget stay: the work needs them.
A simplified version running in your browser; nothing you type leaves this page. The real one runs offline and is tested on synthetic guests.
What’s in place and what isn’t yet, line by line: Security.
What it showed
237,740
made-up records across eight systems, generated in 12 seconds
188 of 188
facts on the lists matched the source records
15 of 20
hidden patterns found in the blinded test hotel, with none invented in its twin
0
leaks of personal details on 5,543 planted items
And the result that matters most
We compared our pipeline with the simple rule most teams already use: reply to the newest enquiry first. On the test hotel, it didn’t beat it. A rule we wrote from intuition did worse than both.
0%20%40%60%
Random orderA floor3% · range 0 to 8%3%
A points ruleWritten from intuition22% · range 1 to 46%22%
Newest firstWhat most teams do now33% · range 14 to 56%33%
Our pipelineBuilt with Claude33% · range 15 to 58%33%
A second pipelineBuilt with Codex35% · range 16 to 54%35%
Share of the best realistically achievable result that each method’s daily lists captured, on the blinded test hotel. Dots are the result; lines are 95% ranges. The ranges overlap: on this evidence, nothing beats newest first yet.See the numbers as a table
Method
Result
95% range
Random order (a floor)
3%
0 to 8%
A points rule (written from intuition)
22%
1 to 46%
Newest first (what most teams do now)
33%
14 to 56%
Our pipeline (built with claude)
33%
15 to 58%
A second pipeline (built with codex)
35%
16 to 54%
What it can’t tell us
We wrote Wrenholt’s couples ourselves, so the pipeline can only find what we planted. That makes it a good test of the plumbing, the checks and the security, and no evidence at all that it will win a real venue more bookings. We say that plainly because the first real client deserves to know exactly what’s been proven and what hasn’t.
More experiments on Wrenholt Hall
The best 10 a day, or the best 30?
Would a longer daily list find more of what matters?
Date
25 Sep 2026
Result
The best 10 already held 94% of the value the best 30 could. Precision roughly halved from 10 to 20, and again to 30.
Lesson
Keep lists short. A list of 30 is a list nobody finishes.
A full security rehearsal on a fake client
Can a client’s data go in, be used and be destroyed, with nothing left behind?
Date
25 Sep 2026
Result
Every step passed in about three and a half minutes. It also found a gap the 5,543-item test had missed: fields written like “DateOfBirth” slipped past the filter. Fixed the same day.
Lesson
Rehearse the whole journey on fake data. The test you designed only finds what you thought of.
Did replying to reviews help?
If a hotel starts replying to reviews, does its rating go up, and by how much?
Date
25 Sep 2026
Result
A simple before-and-after comparison said +0.24 stars. A fair comparison, against a group that didn’t change, found +0.11: the true figure we had planted.
Lesson
Before-and-after comparisons flatter. Here, by more than double.
What is a required “lost reason” worth?
If staff had to record why an enquiry was lost, would the analysis get better?
Date
25 Sep 2026
Result
The records got much cleaner, but the ranking barely changed. The ranges overlapped.
Lesson
Better data is worth having for its own sake, but it isn’t magic. Say what it will and won’t change.
Where does a list lose money?
How far below the best possible does each method sit, and why?
Date
25 Sep 2026
Result
Every method, ours included, captured roughly a third of what was realistically achievable. None could be told apart from newest-first on a single day.
Lesson
Measure against the realistic best. Measured against nothing, every method looks good.
What we’re testing next on problems
Real records
Does anything we found in a made-up hotel hold in a real one? Only a first client’s history can say.
Speed before ranking
How much of the value comes simply from replying fast and following up properly, before any ranking at all?
Will AI search recommend you?
Ask the big AI assistants for a venue like yours, every month. Who gets named, and why?