expanza

PROOF OF WORK

Expanza Labs.

Where we test new AI fairly, and work out how best to solve real problems, before any of it goes near a real business.

How we test

The same rules for every test, whether it’s a new model or a new way to fix a problem. They’re there to stop us fooling ourselves.

  1. The same task, fixed in advance

    A real piece of work with a test set and an answer key written before anything is run.

  2. A simple baseline to beat

    A plain rule, or how a person does it today. New is only better if it beats that.

  3. Scored blind, run more than once

    Scores come from code or from someone who can’t see which system produced the answer. Results are shown as ranges.

  4. Cost and time counted

    A result that’s slightly better but ten times the price is written down that way.

  5. A twin with nothing to find

    Every test includes data with nothing hidden in it, to catch methods that invent patterns.

  6. Failures published

    What didn’t work goes on this page with the same weight as what did.

Track 1

New capabilities

New models, agents and tools arrive every month, each with a bold claim. We test the ones that could matter to a business on a real piece of work, and publish the result, including when the old way wins. None has been run yet. These are next, and each appears here once it’s done.

  • New models on a real job

    When a new model comes out, is it better at drafting a reply from a venue’s price list, without getting a single fact wrong? And is it worth the cost?

  • Agents that use a screen

    Can an agent copy a booking from one system to another with no integration, and how many mistakes does it make in a hundred?

  • Answers that know their limits

    Can an assistant answer from a business’s own documents, show the source, and say “I don’t know” instead of guessing?

  • The simplest thing that works

    For each new capability: would a rule, a spreadsheet or a person have done just as well?

Track 2

Problems, tested on a hotel that doesn’t exist

Why build a fake hotel?

You can’t ask a business for years of customer records to test an idea that might not work. So we made the records ourselves: a country hotel with 80 rooms and two wedding suites, five years of enquiries, bookings, payments and emails, with duplicates, missing reasons, dates typed five ways and a half-finished move from spreadsheets to a CRM. Then we hid patterns in it, sealed the answers away, and built a twin with the same mess and nothing hidden.

Replay of a real run, 28 September 2026. Abridged.

What we built on top of it

A pipeline that takes the raw exports and gives each person a short list for the morning. AI reads the notes and emails and writes the wording. The joining, scoring and checking are ordinary code and statistics, because that’s what they’re good at.

  1. 1

    Eight exports

    Enquiry sheet, CRM, Bridebook, function diary, bookings, bank, mailbox, calendar

  2. 2

    Clean and join

    Duplicates merged, dates fixed, the same couple recognised across systems

  3. 3

    One history each

    Every couple and company, first contact to paid

  4. 4

    Find and score

    What’s still live, and what looks like the enquiries that booked before

  5. 5

    Check every fact

    Reasons built from the record; code confirms every number and date

  6. 6

    A short list

    Up to ten a day per person, with the reason and a next step

Wrenholt Hall · Friday 08:303 of 10
  1. Andrew Jackson Wedding · 106 guests

    Wedding enquiry still waiting for our first reply

    • No reply from us is on record since the enquiry arrived on 25 Sep 2026
    • Planning for 106 guests
    • Budget given as £11,000

    Next: Reply now with availability and prices, and offer a show-round

  2. Kate & Thomas Wedding

    They have seen the venue and not decided

    • Came for a show-round on 22 Sep 2026 and has not decided yet
    • They emailed on 24 Sep 2026: “We’d like to book please!”
    • Quote on file for £14,123.96

    Next: Phone to ask how the show-round felt; offer to hold the date

  3. Craven Care Homes Ltd Corporate

    Regular corporate account has gone quiet

    • Last event with us was on 9 Dec 2025
    • That booking was worth £4,215.43 ex VAT
    • A regular since their first event on 24 Jul 2023

    Next: Call the booker: ask what events they have coming up and offer two dates

Real output from the pipeline. The hotel, the names and the numbers are made up.

Keeping data safe was part of the experiment

Before any AI step sees an enquiry, names, emails and phone numbers are removed. Each client’s data lives in its own encrypted container, created and destroyed by script, with a certificate at the end listing what was deleted. Try a simplified version below.

What an AI model is allowed to see

Hi, it’s [NAME] [NAME] here, we emailed on Friday about Saturday 14 August 2027. You can reach me on [PHONE] or [EMAIL], or my partner [NAME] on [PHONE]. We’re 106 guests, budget around £11,000.

6 removed. The date, guest count and budget stay: the work needs them.
A simplified version running in your browser; nothing you type leaves this page. The real one runs offline and is tested on synthetic guests.

What’s in place and what isn’t yet, line by line: Security.

What it showed

237,740
made-up records across eight systems, generated in 12 seconds
188 of 188
facts on the lists matched the source records
15 of 20
hidden patterns found in the blinded test hotel, with none invented in its twin
0
leaks of personal details on 5,543 planted items

And the result that matters most

We compared our pipeline with the simple rule most teams already use: reply to the newest enquiry first. On the test hotel, it didn’t beat it. A rule we wrote from intuition did worse than both.

Share of the best realistically achievable result that each method’s daily lists captured, on the blinded test hotel. Dots are the result; lines are 95% ranges. The ranges overlap: on this evidence, nothing beats newest first yet.
See the numbers as a table
MethodResult95% range
Random order (a floor)3%0 to 8%
A points rule (written from intuition)22%1 to 46%
Newest first (what most teams do now)33%14 to 56%
Our pipeline (built with claude)33%15 to 58%
A second pipeline (built with codex)35%16 to 54%

What it can’t tell us

We wrote Wrenholt’s couples ourselves, so the pipeline can only find what we planted. That makes it a good test of the plumbing, the checks and the security, and no evidence at all that it will win a real venue more bookings. We say that plainly because the first real client deserves to know exactly what’s been proven and what hasn’t.

More experiments on Wrenholt Hall

The best 10 a day, or the best 30?

Would a longer daily list find more of what matters?

Date
25 Sep 2026
Result
The best 10 already held 94% of the value the best 30 could. Precision roughly halved from 10 to 20, and again to 30.
Lesson
Keep lists short. A list of 30 is a list nobody finishes.

A full security rehearsal on a fake client

Can a client’s data go in, be used and be destroyed, with nothing left behind?

Date
25 Sep 2026
Result
Every step passed in about three and a half minutes. It also found a gap the 5,543-item test had missed: fields written like “DateOfBirth” slipped past the filter. Fixed the same day.
Lesson
Rehearse the whole journey on fake data. The test you designed only finds what you thought of.

Did replying to reviews help?

If a hotel starts replying to reviews, does its rating go up, and by how much?

Date
25 Sep 2026
Result
A simple before-and-after comparison said +0.24 stars. A fair comparison, against a group that didn’t change, found +0.11: the true figure we had planted.
Lesson
Before-and-after comparisons flatter. Here, by more than double.

What is a required “lost reason” worth?

If staff had to record why an enquiry was lost, would the analysis get better?

Date
25 Sep 2026
Result
The records got much cleaner, but the ranking barely changed. The ranges overlapped.
Lesson
Better data is worth having for its own sake, but it isn’t magic. Say what it will and won’t change.

Where does a list lose money?

How far below the best possible does each method sit, and why?

Date
25 Sep 2026
Result
Every method, ours included, captured roughly a third of what was realistically achievable. None could be told apart from newest-first on a single day.
Lesson
Measure against the realistic best. Measured against nothing, every method looks good.

What we’re testing next on problems

  • Real records

    Does anything we found in a made-up hotel hold in a real one? Only a first client’s history can say.

  • Speed before ranking

    How much of the value comes simply from replying fast and following up properly, before any ranking at all?

  • Will AI search recommend you?

    Ask the big AI assistants for a venue like yours, every month. Who gets named, and why?

← Back to the homepage