Revenue-based pricing  ·  Our fee is half the savings we project

Ellipse Automation
All posts
ai-implementationoperationsdatabuild-log

Make the AI Show You the Denominator

·Etienne Chanut

The way to verify an AI recommendation is to ask what it divided by. Last week ours recommended we walk away from a paid data source, and every number in the recommendation was correct. The test that produced those numbers had been sabotaged by the model that ran it.

The setup was ordinary. Before buying a data source country-wide for a client's accommodation database, we ran a cheap pilot across the 20 largest cities to see whether it was worth it. Cost of the pilot: $2.32.

The short version

  • A $2.32 pilot returned 816 results and 254 emails, of which 54 were new. The AI recommended skipping the source.
  • The property counts looked impossible for 20 major cities, so we challenged them.
  • The model had capped the pilot itself at 4% sampling to stay inside a $5 budget, then read the small counts as evidence the source was thin.
  • A second defect had silently broken 8 of the 20 cities. Both were invisible in an output where every individual number was correct.

The recommendation looked airtight

The verdict came back clear and quantified: 816 results paid for, 537 new properties, 254 distinct email addresses, of which only 54 were not already in the pool of 29,829 we had built from free sources.

That worked out to $0.043 per net-new email. The sweep's own summary had reported $0.0065 per email, but that counted addresses we already had; measured against the real baseline it was 6.6x more expensive. Extrapolated nationally: $58 for roughly 320 net-new emails, about 1% growth on the pool.

The recommendation was to stop, and to skip the national pass. On those numbers I would have agreed.

Watch out

Every figure in that paragraph is arithmetically correct. That is what makes this failure mode dangerous: there is nothing in the output for a fact check to catch.

The number that did not fit the world

What did not survive was the total. 816 results across the 20 largest and most visited cities in the country implies a market that anyone who has been there knows does not exist.

So I pushed back, in exactly these words: there is no way that out of the 20 most popular cities you only found 800 properties, are you 100% sure?

The model did not lie. It forgot its own test setup, and then analysed the result as if someone else had run it.

It went and checked. It had set a sampling flag to 4% at the start of the pilot, deliberately, to keep the run near a $5 budget ceiling.

One city was capped at 202 results out of more than 5,000 available properties. Others were capped at 100 and 49. Several of the cities that appeared to have worked had also quietly hit their ceiling: 109 returned out of 156, 104 out of 138, 74 out of 95.

So 816 was never what the source contained in those cities. It was what had been asked for. The retraction was immediate and unhedged: the earlier per-email figure was computed from a broken run and is void.

The second defect nobody was looking for

The cap was not the only thing wrong, and the second problem is the one I find more instructive.

Eight of the 20 cities had returned exactly 1 result each. Not zero, which would have been loud. One, which looks like a thin market.

Query sentWhat came back
City name alone1 restaurant
City name + country1 attraction
"Hotels in" + city + country0 results
City + region + country20 hotels

Bare city queries were resolving to a random venue that happened to be named after the city: a restaurant, a historic site. 40% of the pilot had silently failed, and it had failed in the most valuable cities.

Adding the region term to the query fixed it. Re-run on the seven cities that had failed, the source returned 105 hotels out of 105, with 92% email fill.

Where the pilot broke, in order

01

Ask for the denominator

How many were measured, out of how many exist. 816 measured against a market of tens of thousands.

02

Ask who set the limit

If the model set it, the analysis is unreviewed. Ours had set a 4% cap to stay inside a budget.

03

Look for suspiciously round failures

Exactly 1 result in 8 of 20 cities was a broken query, not a thin market.

04

Re-run the corrected test before deciding

105 of 105 hotels at 92% email fill, against a pilot that had reported near-nothing.

Two independent defects, both of which produced small plausible numbers rather than errors. Neither would have been caught by asking the model to double-check its output, because the output was fine.

The honest limit: the conclusion was not entirely wrong

Here is the part that would be easy to leave out. After both defects were fixed, the source still did not pay for itself on email. Corrected, it came out around $0.07 per net-new email, worse than the original plan assumed, because a broad query mostly returns hotels everybody already has.

So the original recommendation to skip the national sweep survived. It survived for reasons that had nothing to do with the reasoning that produced it, which is less reassuring than a correct answer usually is.

What actually changed was the question. Re-tested properly, the source turned out to be excellent at something we had not been evaluating: room counts, with 96% fill, against 786 hotel-class properties stuck on an unreliable room-count floor that risked misfiling real hotels in a client-facing file. Free sources resolved 335 of those at no cost, and the remainder justified a targeted $8.78 pass instead of a $58 national one.

A wrong test risks more than a wrong answer. It also stops you finding out what the thing was actually good for.

How to verify an AI recommendation before you act on it

When an AI hands you a recommendation built on a measurement, ask for the denominator before you ask about the conclusion. How many were checked, out of how many exist, and who set that limit.

If the answer is "I set it", treat the analysis as unreviewed. A model configuring its own test and then interpreting the result is doing the thing we know not to let a person do, which is grading their own homework.

Key takeaways

  • Ask any AI analysis for its denominator: how much of the population was measured, and who chose that limit.
  • A model that configures its own test treats the configuration as background context rather than as a caveat on its conclusion.
  • Accuracy checks cannot catch a scope error, because every individual number in the output can be correct while the conclusion is void.
  • Silent partial failures that return small plausible numbers are more dangerous than errors, which are loud.
  • Separate the party that runs a test from the party that interprets it, for the same reason nothing should grade its own output.

Related reading: what a scraped email column actually contains is the same lesson from the data side, where every row looked complete and the file still did not mean what it appeared to mean, and separating the builder from the grader is the structural version of the same rule. If you want a second pair of eyes on where AI is quietly making decisions inside your operations, that audit is where we start with a company.

Common questions

How do you verify an AI recommendation before acting on it?

Ask for the denominator. Every conclusion drawn from a measurement rests on how much was measured, and a model that configured its own test will report the result without re-checking the limits it set.

What is the failure mode when an AI analyses its own test results?

It treats its own configuration as background rather than as evidence. Ours capped a pilot at 4% sampling to stay inside a budget, then read the resulting small counts as proof the data source was empty.

Do accuracy checks catch this kind of error?

No. Every number in the recommendation was arithmetically correct. The error was in the scope of what had been measured, which no fact check on the output can see.

Ready to see the math

Your bottom line has room. We can show you where.

Book a free 30-minute call. We'll look at your refund rate and growth trajectory, then show you the savings we'd project. No pitch deck, no commitment.

Run your numbers with us

Free 30-minute call. The math is yours to keep either way.