← All articles

GTM Engineering

One Vendor Finds 45%. Four in Sequence Find 68%.

Buy a single data provider and you inherit its blind spot as your ceiling. Chain four in sequence, pay only on the miss, and verified coverage climbs to about 68 percent found, 62 percent valid, not the folklore 92. Here is the build order by marginal recovery.

· 14 min read

One provider fills about 45 percent of a cold list and prints 91 percent accuracy on the slide. For a decade the enrichment decision was which vendor to buy: you picked Apollo or ZoomInfo or Cognism, signed the annual contract, and accepted whatever the tool could not find as the natural ceiling of your data. That framing is the mistake. The best single source in an independent test tops out around 45 percent coverage for a mixed list, and stacking four providers in sequence lifts verified coverage to roughly 68 percent found, 62 percent valid (the Aleph and Benchmarkit 2026 multi-provider comparison). The vendor you chose was never your ceiling. Your refusal to chain was.

A bad email column poisons every play downstream, so fix enrichment before you build anything on top of it. The waterfall is how: ask one provider, and when it comes back empty, fall through to the next, keep the first verified answer, and pay mostly for hits. I learned to build this before I learned anything else in Clay, because every play I stacked on a dirty base failed in ways that were hard to trace back to the base. Watch the coverage build before you read the argument. The disc is verified, reachable contacts; the dots are the raw list still waiting to be filled.

Raw list to verified coverageEach layer recovers what the last one missed
0% verifiedRaw listNames from one provider filter. A domain, maybe a title, no verified contact.
45% to 68%
Coverage lift from one source to a four-tool chain (Aleph/Benchmarkit 2026)
~1 credit
Cost per verified cell when the fall-through is wired right
3x
Cost when you run every provider on every row instead

The automation is the cheap part. You can wire three providers and a verify step in an afternoon. The expensive part is the ordering: which source runs first, which one only ever touches the tail, and where the verify sits, because every provider you add on every row is a standing credit cost. Wire the chain wrong and you triple the bill for the same 68 percent. Wire it right and credit spend tracks your actual miss rate, not your row count.

One provider always has a hole

I ran a single-source build for a quarter. Apollo, primary, everything routed through it. Coverage looked fine on the US mid-market rows and collapsed the moment I pushed into EMEA: missing mobiles, stale titles from people who had changed jobs two years back, whole German accounts with nothing but a domain. That is not an Apollo defect. Every vendor has a shape. Apollo is dense in some segments, ZoomInfo reports 65 percent email find rate in the US and about 35 percent globally (ZoomInfo’s own coverage figures), and Cognism holds a 68 percent find rate with 91 percent valid-of-found in EMEA because of how it sources phone-verified panels (Cognism). Pick one and you inherit its blind spot as your ceiling.

ProviderStrong whereWeak whereRough cost per match
ApolloUS mid-market, dense contactsEMEA mobiles, stale senior titles$0.01
CognismEMEA coverage, phone-verified mobilesUS long-tail$0.03
ZoomInfoUS enterprise depthPrice, small-company coverage$0.08

The single-source ceiling is not a coverage problem you buy your way out of by picking a better vendor. It is structural. Every provider builds its data from particular sources (self-reported profiles, scraped signatures, phone-verified panels, partner feeds), and those sources carry geography and segment shapes baked in. A US-enterprise-heavy source stays thin on EMEA mid-market no matter how much you pay, because the data was never collected there. So “which vendor is best” is the wrong question. The right question is “which vendor is best for this row,” and the only honest answer is a fall-through that asks each in turn and keeps the first verified hit. You are not hedging between vendors. You are composing them.

The mechanism A miss falls through; a hit stops the chain
ApolloCognismZoomInfoemptyemptyhit, stophit, stopverify then keep
Each cell only reaches the next provider when the prior one came back empty. Credit spend tracks your actual miss rate, not your row count.

The build order by marginal recovery

Do not turn on four providers at once and hope the coverage lands. Build the chain one rung at a time, and order the rungs by what each layer recovers on top of the layer below it. The first rung is not a provider at all: it is the filter that decides which rows earn a credit. gradient.works’ waterfall analysis puts a single source near 45 percent, three sources near 78 percent, and four near 84 percent on a clean, ICP-filtered list, and the honest cross-provider number on a mixed real-world list lands lower, around 68 percent found. The gap between those two numbers is the ICP filter doing its job before you spend a credit.

Build it bottom to topThe waterfall build order by marginal recovery
  1. L5Mandatory verify, always last62% valid

    An SMTP validation you control runs on the coalesced result and strips catch-alls and stale hits. Found 68 percent becomes valid 62 percent here. Without this rung the coverage number is a story, not a fact.

  2. L4Premium specialist on the tail only+6 pts

    ZoomInfo or Cognism touches the bottom 20 percent nobody else could fill. At roughly $0.08 a match it is your most expensive credit, so it earns its place by running last, on the fewest rows.

  3. L3Mid-tier source on the miss+17 pts

    Apollo or Hunter runs only on what layer 2 left empty. Apollo publishes 91 percent valid-of-found but about 50 percent overall coverage, so it belongs in the middle, not first and not last.

  4. L2Cheap email source first~45%

    Lead with a low-cost specialist (Findymail, LeadMagic, Prospeo) that fills the dense, easy rows. This layer alone recovers about 45 percent, and every hit here is a hit the expensive provider never has to pay for.

  5. L1Filter to ICP before you enrich a single contact-50% rows

    Run a company-level waterfall first and drop the rows that are not your ICP. This kills 40 to 60 percent of a bought list before it costs a contact credit, and it is the cheapest coverage move you will ever make.

The ladder is also the cost story. Sequence on two numbers: match rate for your segment, and cost per hit. Cheapest reliable source first, expensive specialist last, so the specialist only runs on the rows nobody else could fill. Those tail rows are the ones you were missing entirely before, and the ones you would have paid full freight to guess at. Order the chain by database size instead of accuracy and you invert this: the priciest provider runs first on every row, and the Benchmarkit 2026 cost analysis puts that swing at roughly four times the cost per verified record for the same coverage.

5,000 ICP-filtered rows through a three-step email waterfall
Each step only runs on what the prior step missed. The premium provider touches 850 rows, not 5,000, which is the whole cost argument.
View as table
ItemValue
Rows to fill5,000 rows
Cheap source filled2,900 rows
Apollo filled the miss4,150 rows
Specialist filled the tail4,700 rows

Build a separate chain per field

The best email provider is rarely the best for direct dials, and technographics come from a different source than firmographics. Build one chain for email, one for phone, one for company data, and do not fold them into a single blended call.

The per-field split has a cost consequence people miss. Run one blended “enrich this contact” call that returns email, phone, and firmographics together and you pay for all three even when you needed the email, and you inherit whichever provider is weakest on the field you cared about. Mobile lookups run about eight times the cost of email at half the hit rate, so a blended call quietly bills you specialist-mobile rates on rows where you only wanted a verified email. Splitting the chains lets you order each one on its own hit-rate-and-cost curve. Your email chain leads with the cheap source and rarely reaches the specialist. Your mobile chain leads with Cognism, because phone-verified numbers live there. Same three vendors, three different orderings, because the shape of their coverage differs per field.

Ordering is not set-and-forget either. A provider that was your best US-email source last year can degrade quietly after a data-source change on their end, and you will not catch it from the coverage number, because the fall-through silently picks up the slack while your credit spend creeps up instead. Re-check the per-step hit rate every quarter. If step one’s hit rate has fallen and step two’s spend has risen to match, your order is stale and the cheap source is no longer cheap in practice, because you pay for it and then pay again for the fall-through.

The conditional is the whole trick

In Clay each step is a column, and the fall-through is a condition on the enrichment: run this lookup only when the previous column came back empty. That is the mechanism. Wire it wrong, running every provider on every row, and you burn three credits per contact to keep one answer. Wire it right and credit spend tracks your actual miss rate. On a 5,000-row table that difference decides whether the play pencils.

Column 1  Cheap email source
Column 2  Apollo email
          run only if:  `{{Cheap source email}}` is empty
Column 3  ZoomInfo email
          run only if:  `{{Cheap source email}}` is empty AND `{{Apollo email}}` is empty
Column 4  SMTP verify   (runs on the coalesced result)
Column 5  Final email = first non-empty AND verify == "valid"

The condition on columns 2 and 3 is the entire cost story. Without it, every provider runs on every row. With it, the SQL that consolidates the chain is one coalesce over a verified flag:

-- Keep the first VERIFIED hit across the chain, not the first non-empty.
-- A catch-all that "answered" must not win the coalesce.
SELECT
  contact_id,
  domain,
  COALESCE(
    CASE WHEN cheap_status  = 'valid' THEN cheap_email  END,
    CASE WHEN apollo_status  = 'valid' THEN apollo_email END,
    CASE WHEN zoominfo_status = 'valid' THEN zoominfo_email END
  ) AS final_email
FROM contacts_waterfall
WHERE COALESCE(cheap_status, apollo_status, zoominfo_status) IS NOT NULL;
Run every provider (wrong) Empty-cell condition (right)
Provider calls 15,000 (3 per row) ~6,600 (only on misses)
Credits burned ~15,000 ~6,600
Coverage reached Same Same
Cost per verified row Roughly 3x Baseline
Same 5,000 rows, same providers, same coverage. The only difference is a condition on each step.

A worked example that reconciles to the field above

Take a raw 10,000-name list and run it down the full ladder. The numbers below are the ones behind the survivor field at the top of this piece.

StageRows inRecoveredRows outRunning coverage
Raw list10,000filter to ICP5,000filtered
L2 cheap source5,0002,9002,90045% of filtered
L3 Apollo on miss2,1001,2504,15062% of filtered
L4 specialist on tail8505504,70068% of filtered
L5 SMTP verify4,7004,300 valid4,30062% valid

The ICP filter halves the list before a contact credit is spent, so you enrich 5,000 rows, not 10,000. The three-step email chain finds 4,700 of those 5,000, which is the 68 percent found on the third stage of the field. The verify stage then rejects 400 catch-alls and stale hits, dropping the trustworthy set to 4,300, which is the 62 percent valid on the fourth stage. That is why the disc grows on the verify stage while a few dots return: verification consolidates a real core and hands back the addresses that only looked like hits. The premium specialist touched 850 rows, not 10,000, and the whole chain cost around $0.56 per verified email (the Benchmarkit 2026 cost figure), a number you can defend to the person in finance who signs the credit invoice.

Model the cost before you commit an order

Before you lock a sequence, put your own segment’s hit rates and per-match costs into the model and read the blended cost per verified record. Move the cheap source’s hit rate down for EMEA and watch how much work falls to the expensive provider:

Blended cost per verified record

per verified record

Try

Order matters. Put the cheap high-hit provider first and blended cost drops, because the expensive one only ever touches the records nobody else could find.

per verified record: $0.024

Build it in the right order

Your first waterfall, in build order
  1. 1

    Take 200 accounts, not 5,000

    A small table where you can eyeball every row. You are debugging the logic, not running the play yet.

  2. 2

    Filter to ICP before any contact enrichment

    Run the company-level check first and drop the non-ICP rows. This is L1 of the ladder and the cheapest coverage you will buy.

  3. 3

    Chain two providers with an empty-cell condition between them

    Cheapest reliable source in column 1, second source in column 2 gated on "column 1 is empty." Confirm column 2 only fires on the misses.

  4. 4

    Break the first provider on purpose

    Force column 1 to return empty on a few rows and watch column 2 catch them. If it does not catch, your condition is wrong. Fix it at 200 rows, not at 5,000.

  5. 5

    Add the verify step before you trust coverage, then scale

    SMTP validate the coalesced result, require valid rather than present, and re-point the fall-through at the verify flag so catch-alls do not stop the chain. Only then push to the full table.

The habit generalizes

Once the waterfall is in your hands you stop trusting any single point in the system. Sender down, you want a fallback route. API times out, you want a retry and a graceful skip, not a dead run. The enrichment waterfall is the first place you feel the cost of a single dependency, so it is the first place the layered instinct sticks.

One caution before you celebrate the coverage number: a waterfall raises how much you find, not how fresh it is. Every provider in the chain can hand you a contact that was accurate the day they ingested it and wrong today, and the fall-through will keep it because it “answered.” That is why the verify rung is mandatory and why it runs on a schedule, and it is the whole subject of why your enrichment vendor is wrong 30 percent of the time. Coverage is what the deck sells. Valid-of-found, measured by you, is what reaches a human.

Build one this week, at 200 rows, following the order above. The SDR-to-GTM-engineer path puts this in week two for a reason: everything downstream stands on records you trust, and the waterfall is where you first earn that trust cheaply enough to keep earning it every refresh.

clay enrichment data

Keep reading

One email. Every week.

One email a week: a system I built or broke, with the config, the numbers, and what I would change. No roundups, no theory, unsubscribe whenever it stops being useful.

The newsletter opens soon.

Connect a provider in src/config.ts