RESEARCH · OUR DATA
We built a system to rank enquiries. It didn’t beat the simplest rule..
Can a pipeline that reads five years of records choose better than “reply to the newest first”?
The test
We generated Wrenholt Hall, a hotel that doesn’t exist, with five years of messy records and a set of hidden patterns. We then tuned our methods on one version of the hotel and scored them, once, on a second version they had never seen. Each method produced daily lists for the coordinators, and we measured how much of the achievable value those lists captured 1.
We compared five approaches: random order, a points rule written from intuition, newest first, our pipeline built with Claude, and a second pipeline built with Codex.
What happened
Newest first captured 33%. Our pipeline captured 33%. The Codex pipeline captured 35%. The points rule captured 22%, and random order 3%. Every range overlaps with its neighbours, so the honest reading is that none of the serious methods beat newest first 1.
Our pipeline was better at other things: it recognised the same couple across systems far more accurately, and every fact it stated matched the records. That matters for trust. It just didn’t translate into better choices on this test.
See the numbers as a table
| Method | Result | 95% range |
|---|---|---|
| Random order (a floor) | 3% | 0 to 8% |
| A points rule (written from intuition) | 22% | 1 to 46% |
| Newest first (what most teams do now) | 33% | 14 to 56% |
| Our pipeline (built with claude) | 33% | 15 to 58% |
| A second pipeline (built with codex) | 35% | 16 to 54% |
Why we’re publishing it
Because the first question any sensible venue should ask is “does it beat what we already do?”, and today the answer is no. A rule that sounds clever is still a guess until it’s tested. We’d rather you heard that from us.
What this means for a business
- If your team replies to the newest enquiries first, that’s a reasonable rule. Speed and consistent follow-up probably matter more than clever ordering.
- Ask any supplier to compare their product with how you work now.
What we don’t know
- Whether the result holds on real records, where behaviour isn’t written by us.
- Whether ranking helps more when there are far more enquiries than a team can handle.
Sources
- Fact Expanza enquiry-engine leaderboard, blinded test hotel (seed 7), 25 Sept 2026 · Our own data; 95% bootstrap ranges