Revenue-based pricing  ·  Our fee is half the savings we project

Ellipse Automation
All posts
ai-implementationdataoperationsbuild-log

A 50-Row Sample Priced a 95,000-Row Run

·Etienne Chanut

How big should a sample be before a full data run? Big enough that a lucky draw cannot survive it, and 50 rows is not that. Ours reported 74% email coverage, the full run of 94,973 records came back at 27%, and everything we had planned downstream had been costed on the first number.

We are building a national accommodation database for a client. Before paying to enrich the whole list, we ran a smoke test on 50 rows to see what the enrichment source would actually return.

The short version

  • A 50-row smoke test returned 74% emails, 80% phones and 10% named decision-makers.
  • The full run across 94,973 records returned 27% emails, 30% phones and 3% named contacts.
  • The pipeline was correct. The sample was too small for a fluke draw to wash out, and roughly 95,000 credits were committed on it.
  • The rule now: a few hundred rows minimum before any full-list spend, drawn per segment, no matter how good the small sample looks.

How big should a sample be before a full data run

The smoke test looked like a clear green light on every field that mattered. Three quarters of the properties came back with an email address, four fifths with a phone number, and one in ten with an actual named person attached.

The full run reported a different source.

Field50-row sample94,973-row run
Email address74%27%
Phone number80%30%
Named decision-maker10%3%

Nothing broke between the two runs. Same source, same fields, same pipeline. The 50 rows we happened to draw were simply not the list.

A smoke test and a coverage estimate are different tests, and we ran one while reading it as the other. Fifty rows is enough to prove a pipeline runs end to end, and nowhere near enough to price what it returns.

Why the error was expensive rather than annoying

A coverage number is never used alone. It is the input to the credit budget, the contactable-record forecast, and the decision about whether the source is worth buying at all, so a wrong coverage number propagates into every one of those before anyone rechecks it.

At 74% email fill, roughly 95,000 credits buys around 70,000 contactable records and the economics are obvious. At 27%, the same spend buys about 25,000, and the named-contact field at 3% is barely a field at all.

The named-contact field shows the shape of it most clearly. At 10% a plan can assume a person to address on one property in ten, which is enough to build a segment around. At 3% that field is not a segment, it is a rounding error, and anything designed on top of it has to be redesigned.

The run itself was not wasted. The records are real and the pipeline is the one we would have built anyway. What was wrong was that we approved the spend against a number that had a much wider swing than anyone treated it as having.

The part I want to keep

The AI running the pipeline was asked directly why the two numbers disagreed, and it did not reach for a reason the data could not support. Its own words: it should have sampled several hundred rows before committing, not 50, and a fluke draw on n equals 50 is exactly the trap it walked into.

That is the behaviour worth designing for. A system that reports its own sampling decision as a cause is one you can correct; one that produces a plausible story about the source degrading is one you cannot.

The pipeline did what it was told. The number it was told to trust was a coin flip nobody had counted.

The rule we run now

01

Sample in the hundreds, not the dozens

Several hundred rows before any full-list commitment. The extra cost is a rounding error against the run it protects.

02

Draw per segment

Coverage varies by property type, region and source. A blended average hides a segment that returns almost nothing.

03

Write the number down as a range

Record the sample size next to the percentage, so the person approving the spend sees what the estimate is standing on.

04

Re-measure on the first slice of the real run

Enrich the first few thousand records, check coverage against the sample, and stop if the gap is wide.

The fourth step is the one that would have caught this specific case for a few hundred credits instead of tens of thousands. A full run is not an atomic action, and treating it as one is a choice.

Why per-segment sampling matters more than raw sample size

Raw sample size protects against a lucky draw. It does not protect against a list whose segments behave differently, and most real lists have segments.

In an accommodation list, a chain hotel and an apartment host are two different data problems. Chain properties have a corporate domain, a website and a switchboard, so contact coverage on that slice is high. Individual hosts publish through a booking platform and often have no independent web presence at all, so coverage on that slice is close to nothing.

Draw 300 rows across a list that is mostly hosts with a minority of hotels, and the hotels can still carry the average high enough to look acceptable. The blended number is then correct and useless at the same time, because nobody buys a blended segment.

The fix is boring: stratify the draw, report coverage per segment, and let the segments with poor coverage be a separate decision rather than an invisible drag on a single percentage.

The honest limit

A bigger sample fixes sampling error and nothing else. It does not catch a source that is silently returning a fraction of what it holds, or a query that resolves to the wrong thing entirely, because those defects are present in the sample too and they scale with it.

We have hit that exact failure on the same project, where a pilot was capped at 4% sampling by the system that ran it and the small result was read as a thin market. Sample size would not have helped there. The question that did was what the numbers had been divided by.

So the two checks are separate. Size the sample against luck, and interrogate the setup against itself.

Key takeaways

  • A 50-row smoke test proves a pipeline runs; it does not price what the pipeline returns.
  • Our 50-row draw reported 74% email coverage against an actual 27% across 94,973 records.
  • Sample in the hundreds and draw per segment, because a blended coverage number hides the segment nobody can use.
  • Check coverage again on the first slice of the real run rather than treating a full run as one irreversible action.
  • Sample size protects against luck only. A capped or misresolved query is wrong in the sample too.

Related reading: make the AI show you the denominator covers the other half of this, where the measurement itself had been limited by the system reporting it, and 60,128 email rows, 3,921 actual inboxes is what happens after enrichment when nobody checks how many of the addresses are distinct. Building the data layer before the automation that sits on it is most of what we do when we work inside a company.

Common questions

How big should a sample be before committing to a full data enrichment run?

A few hundred rows, not a few dozen. At 50 rows a lucky draw survives the test: ours showed 74% email coverage and the full 94,973-row run came back at 27%.

What is wrong with a 50-row smoke test?

A 50-row smoke test proves the pipeline runs. It does not measure coverage, because the swing on a 50-row draw is wide enough to move the number by tens of points in either direction.

How do you sample a list that has segments in it?

Draw the sample per segment rather than across the whole list. Coverage varies by segment, and one segment with good data can carry a blended average that no individual segment supports.

Ready to see the math

Your bottom line has room. We can show you where.

Book a free 30-minute call. We'll look at your refund rate and growth trajectory, then show you the savings we'd project. No pitch deck, no commitment.

Run your numbers with us

Free 30-minute call. The math is yours to keep either way.