Something Inc.Schedule a free consultation
STRATEGY

B2B Data Providers Worth Trying Now That Your Lists Are Dead

Jordan Crawford's team scanned 1.6 million public datasets looking for anything usable for outbound. Twelve survived. Two were outright fakes. Here's what the exercise says about vetting B2B data providers, free or paid, before they touch a live sequence.

TTTyler TruffiManaging Partner · AUG 4, 2026 · 8 MIN READ

Jordan Crawford, who runs Blueprint GTM and has spent the last two years building the pain-based-research end of the outbound world's toolkit, published a piece on July 25, 2026, with a number in the headline that's worth sitting with: 1.6 million. That's how many public datasets his research agent indexed, hunting for anything a B2B team could actually use to build a list. Twelve survived. Two of the fourteen finalists that made it to a final review were outright fabrications.

KEY TAKEAWAYCrawford's team crawled 1,601,200 datasets across HuggingFace, Zenodo, GitHub, and Data Is Plural. A cheap automated filter got that down to 103,081 candidates. An AI grading pass, costing $6.01 total, and a strict provenance check took it to 85. Manually opening files took it to 12. Two of the finalists that looked legitimate on paper were fake, including one that was, quote, '479 bytes of marketing.' If your list-building process doesn't include a step that would have caught that, it isn't a process.

The problem with most B2B data providers

We've written before about how email list decay varies sharply by industry, from 3.2% a year in B2B SaaS to nearly 6% in healthcare and real estate, per LeadMagic's data. That's the decay rate on a list you already own. It says nothing about where you get the next one, and that half of the problem gets far less attention than it should, because most teams default to whichever paid data vendor has the best-designed pricing page rather than asking whether the underlying data is any good.

Crawford's exercise is useful precisely because it treats that question as an empirical one instead of a vendor-trust one. Instead of asking a data provider to vouch for its own data, he pointed a research process at the actual public supply of datasets that exist in the world and asked, systematically, how much of it is real, current, and usable. The answer, at scale, was: less than one percent.

Jordan Crawford's 1.6 million dataset audit

The scope of the crawl is worth stating precisely, because the funnel only means something if you can see how steep it actually was. Crawford's team indexed four sources: HuggingFace (952,837 datasets), Zenodo (620,084), GitHub repositories explicitly tagged as datasets (26,289), and Jeremy Singer-Vine's Data Is Plural project (1,990 curated entries). Total: 1,601,200 candidate datasets.

1,601,200
public datasets indexed across HuggingFace, Zenodo, GitHub, and Data Is Plural
12
datasets that survived the full three-stage filter
$6.01
total cost of the AI grading pass across 103,081 filtered candidates
FILTER STAGECANDIDATES REMAININGWHAT IT REMOVES
Raw index1,601,200Nothing yet — every public dataset across all four sources
Structure filter (cheap, automated)103,081Anything that isn't list-shaped, tabular, contact-relevant data
AI grading + provenance check85Low-usefulness scores and unverifiable data lineage — only 27% of useful-scoring sets could prove it
Manual file inspection12Everything that couldn't survive a human actually opening the file, including two outright fakes

The funnel ran in three stages. First, a cheap database query filtered for list-shaped signals, structured, tabular data that looked like it could plausibly contain names, companies, or contact fields, cutting 1.6 million down to 103,081. Second, an AI grading pass scored each survivor's description for usefulness and data lineage, the entire pass costing $6.01 in model spend. Third, a manual provenance filter rejected anything that couldn't demonstrate where its data actually came from. Per Crawford's own note, only 27% of the datasets that scored as useful could prove their lineage at all. That filter alone took the 103,081 candidates down to 85. From 85, someone had to actually open each file. Twelve made it through.

The two fakes, and why they matter more than the twelve winners

This is the part of Crawford's writeup that should change how a GTM team thinks about vetting a new data source, more than the twelve legitimate finds do. Two of the fourteen finalists that reached the final manual review stage, after already clearing an automated usefulness score, a lineage check, and a validity-date check, turned out to be fabricated. One was a dataset from a vendor claiming to be a Kazakh business-data provider, deposited on Zenodo with what looked, from its metadata, like a real business directory. Opening the actual file revealed a single 479-byte readme pointing back to the vendor's own marketing site. Zero data. The second was a purported US medical-transport provider directory that, on inspection, contained only consumer-facing hotline numbers, no actual provider-level data at all.

One finalist's entire dataset was 'a single file: a 479-byte readme pointing at the vendor's website — zero data.' — from Jordan Crawford's dataset audit, published Jul 25, 2026.

Both fakes made it past an automated usefulness score and a provenance check before a human caught them by actually opening the file. That's the finding worth internalizing. Automated grading and metadata review are necessary, and they cut a huge amount of dead weight cheaply, at $6.01 for over 100,000 candidates in Crawford's case. They are not sufficient. Somewhere in any real data-sourcing pipeline, a human has to open the actual file and look at actual rows before it goes anywhere near an outbound sequence.

A method you can run yourself

You don't need Crawford's exact tooling to apply the same discipline to your own data sourcing. The shape of the funnel is the reusable part, not the specific scale. Most teams skip straight from 'we need more data' to 'let's buy a subscription,' which collapses three genuinely different questions, whether a source exists at all, whether it's actually useful for this specific ICP, and whether it's real, into one purchase decision made under time pressure. Crawford's funnel forces those three questions apart and answers them in order, cheaply, before any money changes hands.

It's also worth noting what this method is not. It isn't a replacement for a paid, reputable enrichment vendor once you've found a real gap in your stack, and it isn't a claim that free public data is generally better than paid data. Plenty of the 1,601,200 datasets Crawford's team indexed were legitimate, just useless for outbound purposes, academic research data, government statistics with no company-level granularity, and so on. The value of the exercise is the discipline of the filter, not a conclusion that free beats paid. Apply the same three-stage funnel to a paid vendor's free sample before you sign, and you'll catch the same category of problem Crawford caught for free.

1Before you pay for any new data sourceCast wide, cheap, and automated first
THE MOVES
Pull candidate sources from public catalogs (HuggingFace, Zenodo, government open-data portals, industry association directories) as well as paid vendors.
Run a fast, cheap filter for structure: does it look like tabular, list-shaped data with the fields you actually need?
Don't spend real time on anything that fails the structure check.
DONE WHENYou have a candidate list an order of magnitude smaller than where you started, at near-zero cost.
2Once you have a manageable candidate listGrade for usefulness and demand a lineage answer
THE MOVES
Score each candidate for how directly it maps to your actual ICP and target fields.
For every candidate, require an answer to one question: where did this data originate, and how would you verify it independently?
Reject anything where the lineage answer is 'unknown' or 'trust us.'
DONE WHENYour candidate list is now small enough that opening files by hand is realistic.
3Before any dataset touches a live sequenceOpen the actual file
THE MOVES
Manually inspect a real sample of rows, not just the description or schema.
Check for the exact failure mode Crawford found twice: a dataset that's mostly metadata, marketing copy, or a redirect back to the vendor, with little or no real underlying data.
Spot-check a handful of records against an independent source (a company website, a LinkedIn profile, a public filing) before trusting the batch.
DONE WHENYou've personally verified real data exists in the file, not just a plausible-looking description of one.

Where B2B data providers fit in a real data stack

This is one layer of a bigger picture we've mapped out in our B2B data stack audit framework: targeting and identity, verification and hygiene, enrichment and signals, consent and compliance, and orchestration. Crawford's exercise sits squarely in the first layer, sourcing, and it's worth noting that ColdIQ's own recent pivot toward publishing B2B data-provider buyer's guides instead of its usual campaign teardowns is a different signal pointing at the same underlying shift: the data layer, not the sequencing or copy layer, is where outbound differentiation is actually happening right now. Crawford's dataset audit is the empirical version of the same observation, run at a scale most teams will never attempt themselves, which is exactly why the method is worth borrowing even if the specific twelve datasets aren't relevant to your ICP.

There's a broader lesson in the 27% lineage-pass-rate figure specifically, separate from the twelve winners and two fakes. If fewer than three in ten datasets that already scored well on usefulness could prove where their data actually came from, that's a statement about the state of the public and semi-public data ecosystem generally, not just about outbound lists. Provenance is the exception, not the default, across a huge share of what gets published and shared as if it were reliable. That should raise your bar for any data source, free or paid, public or vendor-sold, and it's a good reason to build the lineage question into your vendor-evaluation checklist as a standing line item rather than a one-time gut check.

None of this replaces a paid enrichment or verification layer for a serious cold email program running at real volume, and it's worth pairing this sourcing discipline with the same rigor we've applied to cold email agency unit economics elsewhere this year: a cheap, verified data source is only a win if the rest of the funnel, deliverability, sequencing, and follow-up, is disciplined enough to convert it. What sourcing gives you is a cheap, repeatable way to test whether a new data source is worth paying for in the first place, and a concrete reminder that 'it's on Zenodo' or 'it has a professional-looking landing page' are not verification. Opening the file is verification. Everything before that is just filtering.

DO THIS NEXTBefore your next data-provider contract renews, pull a random 50-row sample from the dataset you're paying for and manually verify it against an independent source. If you can't get a clean sample to check, that's information too.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.