← All articles

GTM Engineering

GTM Alpha: Firmographics Are the Price of Admission, Not an Edge

Employee count and industry code are what everyone already bought. Alpha is the signal your competitor cannot buy off a list, and defensibility runs in tiers: public, inferred, proprietary. Here is the ladder and the recipes to climb it.

· 15 min read

The old way to target was firmographics: pull a list by industry code, employee band, and revenue range, then send the same “I saw you’re a 500-person SaaS company” opener to all of it. That stopped working, and the reason is arithmetic, not craft. Two SDR teams pull the same list from the same vendor, send the same opener, and get the same 1-3% reply rate (Woodpecker, 26,000 campaigns), because the buyer has seen that exact email forty times this month. Firmographics are not an edge. They are the price of admission, and everyone at the table already paid it.

Alpha is the term I borrowed from trading, and it fits. Alpha is return you earn that the market has not already priced in. In GTM the market has fully priced in employee count, industry, and funding stage, because those fields ship in every enrichment product on earth. The signal that still earns a reply is the one your competitor cannot buy off a list, and defensibility is not binary. It runs in tiers, from public fields anyone can rent, through inferred conclusions you have to read for, up to proprietary signals only your own product or network can see. That ladder is the whole strategy. Climb it before you read the recipes.

Signal defensibility, bottom to topThe signal-sourcing ladder
  1. L5Proprietary first-party signals20-30%

    A usage spike inside your own product, a champion who used you and changed jobs, a support pattern only your network sees. Nobody can buy or infer this. It is the top rung because it is the only signal your competitor cannot reconstruct at all.

  2. L4Stacked inferred signals~18%

    Two independent inferred signals on one account: a pricing change plus a new CFO. Independence, not volume, makes the stack additive. A competitor now has to build and correlate two engines to match one send.

  3. L3Inferred from public sources10-20%

    A conclusion no vendor sells: a repriced pricing page, a JD scoping a build of what you sell, a docs migration in flight. The source is public but the reading is yours. This is where alpha starts, because there is no field to buy.

  4. L2Rented intent~5%

    Third-party intent scores and anonymous-visit feeds. One tier up because it hints at timing, but still a purchase anyone can match, rarely names a human, and rarely implies a job. A full-looking queue that books almost nothing.

  5. L1Public firmographics1-3%

    Employee count, industry code, HQ state, total funding, tech-stack detection from a script tag. All of it on a rate card. The moment a field is for sale, your competitor has it and the buyer has been contacted on it. This is admission, not alpha.

10-20%
Reply rate on a signal-triggered send vs 1-3% for firmographic spray (Woodpecker)
+114%
Win-rate lift when a past champion changes jobs, a top-rung signal (UserGems)
~$0.08
LLM cost to infer one signal per account in my runs

Locate yourself on that ladder. Most teams live on rung one, renting fields everyone else rents and wondering why the reply rate will not move. Every rung up is a step your competitor has to build rather than buy, and the reply rate climbs with the defensibility because a message standing on something the buyer cannot dismiss as list-sourced is a message they answer. The distance from rung one to rung five is the distance from 1-3% to 20-30%, and the middle rungs are where an LLM changed the economics.

Firmographics are commodities, signals are alpha

A field is a commodity the moment two vendors sell it. Employee count, SIC code, total funding, HQ state, tech-stack detection from a script tag: all of it is on a rate card, which is rung one. When a data point is on a rate card, your competitor has it too, and the buyer has been contacted on it already. Contacting a company because it is a 500-person SaaS firm is contacting it on the same axis as everyone else who bought the same list. The message blends into the forty others.

An inferred signal is different because there is no field to buy. Nobody sells “this company just rewrote its pricing page to add a usage-based tier,” or “this company’s engineering blog shifted from Python tutorials to Rust in the last quarter,” or “the new VP of Data they hired came from a company that runs the exact platform you sell.” Those are not fields. They are conclusions, drawn from reading a source a human would have to read manually, at a scale a human cannot. The LLM reads the source, applies your rule, and emits a structured verdict. That inference is the L3 rung, and it is where the reply rate lives.

The twelve signals, ranked by half-life

Once you are climbing past rung one, the next question is cadence, and cadence is set by half-life: the window during which acting on a signal beats acting on nothing. A pricing-page change is hot for weeks. A tech-stack fact is warm for a year. I rank the inferred signals by half-life, because half-life decides how you build the pipeline, not how you write the email.

#Inferred signalSource to readHalf-lifeLadder rung
1New exec hire in a buying roleJob boards, LinkedIn, press2-6 weeksL3
2Pricing page repriced or re-tieredPricing URL diff3-4 weeksL3
3Hiring for a role your product replacesJD text4-8 weeksL3
4Docs shifted to a new frameworkDocs and dev blog1-2 monthsL3
5Support burden visible in public forumsCommunity, Reddit, GitHub issues1-2 monthsL3
6Compliance or cert push (SOC 2, FedRAMP)Trust page, careers2-3 monthsL3
7Product launch adjacent to your categoryChangelog, launch blog1-2 monthsL3
8Leadership language shift in earnings or blogTranscripts, posts1 quarterL3
9Open-source repo momentum changeCommit history, stars1-2 quartersL3
10Two competitors’ logos on their siteCustomer or partner pages2 quartersL3
11Geographic or segment expansionCareers by location, press2 quartersL3
12Headcount composition shift by functionCareers page over time2-3 quartersL3

The top of that list is where I start every build, because a signal with a two-week half-life forces you to build the pipeline right: a scheduled job, a freshness check, and a routing hop that fires while the window is open. A signal with a two-quarter half-life you can batch monthly and still catch. Every row here sits on rung three, inferred from a public source. Rungs four and five come from stacking these or from wiring your own product data, and they are where defensibility peaks.

Show the recipe, not the idea

Naming signals is cheap. The value is the recipe: the exact source, the exact prompt, the exact structured output, and the threshold that turns a verdict into an action. Here is signal number 3, “hiring for a role your product replaces,” as an actual artifact. This is the JSON contract the LLM has to fill, so downstream routing reads structure, not prose.

{
  "signal": "hiring_replaces_us",
  "account_id": "001Rn00000XyZ",
  "source_url": "https://acme.com/careers/senior-data-platform-eng",
  "evidence_quote": "Build and own our internal Python package management and environment reproducibility layer.",
  "verdict": true,
  "confidence": 0.86,
  "reason": "JD scopes an in-house build of the exact capability we sell as a product.",
  "recommended_play": "exec_outbound_build_vs_buy",
  "half_life_days": 42,
  "inferred_at": "2026-07-19T14:03:00Z"
}

The prompt that produces it is short and strict. Give the model the JD text, the definition of what your product does, and a demand for a structured verdict with an evidence quote. The evidence quote is the part most people skip and the part that saves you, because it is the receipt. When a rep asks “why did this account get flagged,” you show the sentence from the JD, not a black-box score.

You are classifying whether a company is about to build in-house
the capability that {{PRODUCT}} sells.

PRODUCT does: {{one_paragraph_definition}}

Read the job description below. Return JSON only:
- verdict (bool): true if the role scopes building this capability internally
- evidence_quote (string): the exact sentence that proves it, verbatim
- confidence (0-1)
- reason (one sentence)
If the JD does not clearly scope an internal build, verdict is false.
Do not infer beyond the text. No evidence quote means verdict false.

JD:
"""
{{job_description_text}}
"""

Two rules make this reliable. First, verbatim evidence or no verdict, which kills the confident hallucination. Second, a confidence floor before it routes to a human. Below the floor it goes to a review queue, not to a rep’s inbox. I use the same harness pattern I described in signal-based outbound: the model proposes, a threshold disposes, a human confirms the edge cases.

The economics: why inferred signals are cheap now

The reason this was not a strategy three years ago is that inference was expensive and slow. It is neither now. Reading one job description and returning a structured verdict costs a fraction of a cent in tokens; even a heavy account with ten sources to read lands around eight cents. Against a firmographic list that costs more per record and returns a 1-3% reply, the math is not close.

The blended cost per verified signal behaves the same way blended enrichment cost does: order your sources cheap-and-broad first, expensive-and-narrow last, and the average drops because the expensive read only ever touches the accounts the cheap read could not resolve. Move the sliders and watch it.

Blended cost per inferred signal, by source order

per verified record

Try

Order matters. Put the cheap high-hit provider first and blended cost drops, because the expensive one only ever touches the records nobody else could find.

per verified record: $0.024

That is the same waterfall logic from building your first enrichment waterfall, applied to inference instead of contact data. The provider names on the sliders stand in for read sources; the shape of the answer is identical. Cheap-first wins.

The reply-rate gap is a defensibility gap

The reason to climb the ladder is one chart, and it maps rung for rung to the ladder at the top. A firmographic spray and a signal-triggered send are not small percentage points apart. They are a different order of magnitude, because one message references something true and current about the buyer and the other references a field the buyer knows is on a list.

Reply rate by ladder rung
Rung one firmographic spray sits at 1-3%. Rung two rented intent nudges to 5%. Rung three inferred signal lands at 10-20%, and rung four stacked signals push higher (reply bands from Woodpecker; rung mapping from the ladder above).
View as table
ItemValue
L1 firmographic spray2%
L2 rented intent5%
L3 single inferred12%
L4 two stacked18%

Stacking two signals beats one, but with a catch: the second signal has to be independent. A pricing change plus a new CFO hire are two roads to the same conclusion and the stack is real. A pricing change plus “detected Stripe on their site” is one signal counted twice. Independence is what makes rung four additive instead of redundant, and it is the difference between a stack that compounds and a stack that double-counts.

The engine that runs the ladder

The pieces have to compose into a pipeline that runs on a schedule, checks freshness, and refuses stale output. Here is the shape of the engine I build.

The signal engine Source to verdict to play, with a freshness gate
SourcesJD, docs, pricingAccount listfrom CRMLLM classifierJSON verdict + quotefreshness gatePlay routerby signal + confidence
The model reads sources and proposes structured verdicts. The freshness gate refuses anything past its half-life. Only fresh, high-confidence signals route to a play.

A worked example that reconciles to the ladder

Take 1,000 target accounts and run them past rung one and rung three, so the reply-rate gap and the cost line both land on real numbers. Rung one rents a firmographic list and sprays all 1,000. Rung three infers one signal, “hiring for a role your product replaces,” and sends only the accounts that clear the confidence floor and the freshness gate.

PathAccounts sentReply rateRepliesBasis costCost per reply
L1 firmographic spray1,0002%20~$0.30/record list = $300$15.00
L3 single inferred signal18012%22~$0.08/account inference = $80$3.64

The rung-one path pays $300 to rent a list, sprays all 1,000, and books 20 replies at $15 each. The rung-three path spends $80 inferring one signal across the same 1,000 accounts, finds 180 that clear the floor, and books 22 replies from a fifth of the sends at $3.64 each. Fewer sends, more replies, a quarter of the cost per reply, and the 12% and 2% reconcile exactly to the L3 and L1 rungs on both the ladder and the chart. The reason the inferred path costs less per reply is not a cheaper email. It is that the message stood on something the buyer could not wave off as list-sourced, so a larger share answered.

Here’s how I’d build it: climb one rung at a time

Four weeks, one signal at a time. Do not try to ship all twelve at once and do not jump to rung five before rung three works. Ship the highest half-life-pressure signal first, prove the reply lift, then climb.

The alpha signal engine
  1. 1

    Start on rung three with one signal

    Pick a short half-life signal that maps to a play you already run. New exec hire or pricing change are good firsts because the play is obvious: build-vs-buy outbound to the new leader.

  2. 2

    Write the JSON contract before the prompt

    Define the exact structured output: verdict, verbatim evidence quote, confidence, reason, half_life_days, inferred_at. Downstream routing reads this shape, so freeze it first.

  3. 3

    Build the classifier with an evidence floor

    Prompt the model to return the contract, refuse any verdict without a verbatim quote, and set a confidence floor below which the signal goes to a review queue instead of a rep.

  4. 4

    Add the freshness gate

    Stamp every verdict with inferred_at and reject routing when age exceeds half_life_days. A stale signal is worse than no signal because it makes you look like you are reading a list.

  5. 5

    Route to a play, then measure reply lift

    Wire fresh high-confidence signals to a specific play, not a generic sequence. Measure reply rate against your firmographic baseline. If it does not clear 3x the baseline, the signal was a commodity in disguise.

  6. 6

    Climb to rung four, then five

    Stack an independent second signal to push reply rate higher, then wire a first-party product signal once you have one. Independence and ownership, not volume, are what make the higher rungs defensible.

Firmographic targeting vs signal targeting

The difference is not the tool. It is what the message stands on, which is another way of saying which rung of the ladder it came from.

Firmographic targeting Signal targeting
Ladder rung L1, a field on a rate card everyone owns L3 and up, a conclusion inferred or a signal you own
Message opener "I saw you are a 500-person SaaS company" "I saw the JD for the platform role scoping X"
Buyer reaction One of forty identical emails this month Someone read something true about us
Reply rate 1-3% (Woodpecker) 10-20% (Woodpecker)
Defensibility Competitor buys the same list tomorrow Competitor has to build the same engine
Cost per record Vendor rate card, per contact ~$0.08 in inference, per account
Same reps, same product. The column on the right is the one that earns a reply, because the buyer cannot tell it came from a list.

The last row is the one that compounds. A list is a purchase your competitor can match in a day. A signal engine is a system your competitor has to build, staff, and tune, and by the time they do, you have twelve signals climbing toward rung five and they have one. Distribution is the moat now, and an inferred-signal engine is how you build distribution nobody can copy off a rate card.

Firmographics tell you who a company is, and everyone has that because it sits on rung one. Signals tell you what a company is about to do, and only the team that read the right source this week has it. Pick one source, write one classifier, wire one play, and start climbing the ladder your competitor cannot buy the bottom of.

signals ai enrichment

Keep reading

One email. Every week.

One email a week: a system I built or broke, with the config, the numbers, and what I would change. No roundups, no theory, unsubscribe whenever it stops being useful.

The newsletter opens soon.

Connect a provider in src/config.ts