The reason is because of spam trap exposure, and scraped files carry the kind of trap that permanently marks a sender as a harvester.
Ask ten deliverability practitioners whether purchased data hurts sending and ten will say yes.
Quick comparison
How the three data sources behave once you actually send to them.
| Parameters | Scraped | Purchased | Enriched |
|---|---|---|---|
| How the address is obtained | Pulled off web pages, directories, and public profiles in bulk | Sold from a pre-built database, often resold repeatedly | Resolved on demand from a person you already selected |
| Typical hard bounce rate | High, often 15% and up | Moderate to high, rises with file age | Low, usually under 3% |
| Spam trap exposure | Highest risk, especially pristine traps | High risk, especially recycled traps | Low, traps are rarely tied to a named ICP target |
| Freshness at send time | Whenever the crawl ran | Whenever the vendor last refreshed | At the moment of the lookup |
| Role and catch-all addresses | Very common | Common | Filterable |
| Relevance to your offer | None, the crawl had no ICP | Weak, filters are broad | Strong, you chose the person first |
Bounce ranges are directional. Your actual numbers depend on vendor, vertical, and how old the file is.
The four things mailbox providers actually score
Before ranking the sources, it helps to know what Gmail and Outlook are measuring. Deliverability is not a quality judgment on your writing. There are four signals.
Notice that three of the four are decided before you write a word.
That is the whole argument.
Scraped data: you are paying in spam traps
Scraping collects addresses from anywhere a string with an @ appears.
Company websites, directories, PDFs, forum signatures, conference pages, old press releases.
The problem is not that scraping is inaccurate. The problem is that scrapers cannot tell a person from a landmine.
Anti-spam organizations seed pristine spam traps in exactly these places.
A pristine trap is an address that has never belonged to a human being and has never opted in to anything.
It exists on a web page for one reason: to catch software that harvests addresses off web pages. If you hit one, the only possible explanation is that you scraped, or bought from someone who scraped. There is no innocent path to that address.
Scraped files also load you up on role accounts.
info@, sales@, hello@, contact@. These sit in shared inboxes, get forwarded to three people, and generate complaints at rates individual mailboxes never touch.
Then there are the guesses. Many scraping stacks do not find the address at all.
They infer a pattern from one known employee and apply it to everyone else at the domain. On a catch-all domain the server accepts every guess, so nothing bounces and you conclude the file is clean.
What scraped data costs you
- Pristine trap hits that mark you as a harvester
- Role accounts that inflate complaint rates
- Pattern guesses that survive validation but never reach a human
- Zero ICP filtering, because the crawler did not know what you sell
Purchased lists: decayed before the file reached you
A purchased list has a different failure mode. The addresses were real. Past tense.
B2B contact data decays at roughly 2% per month, which is about a quarter of any database gone stale every year.
People change jobs, companies rebrand, domains migrate, mailboxes get deprovisioned. A list assembled 14 months ago is not a slightly older version of a good list.
It is a different object.
Here is the part that hurts more than the bounces. When a company deprovisions a mailbox, the address usually starts rejecting mail.
Two more things about purchased data specifically:
You are not the only buyer. Lists get resold. The people on that file have been receiving unsolicited cold emails from every buyer before you, and they are primed to hit the spam button rather than the unsubscribe link.
The filters are broader than they look. "VP of Sales at 50 to 500 employee SaaS companies in North America" sounds targeted. It describes tens of thousands of people who have no idea who you are. Filters narrow a database.
They do not create relevance.
What purchased data costs you
- Recycled traps concentrated in the oldest records
- Bounce rates that climb every month the file sits unused
- Recipients already fatigued by the buyers ahead of you
- Broad filters that feel like targeting and behave like a blast
Enriched data: verified at the moment you use it
Enrichment reverses the order of operations.
With scraped or purchased data, you start with a file and work out who is in it.
With enrichment, you start with a person or a company you have already decided to target, and you resolve their current business contact details on demand.
That sequencing is why enriched data protects deliverability. The address is looked up when you need it, not shipped to you months earlier in a bundle.
3 practical consequences:
Freshness. A lookup run today reflects today. The decay curve that ruins purchased files never gets a chance to start.
Verification before send. A verified result is checked against the domain's real mail configuration rather than returned as a pattern guess. Guessed addresses bounce, and bounces are the fastest way to lose a sending domain.
Trap exposure drops. Spam traps are not attached to named executives at companies you deliberately picked. They live in harvested pools. When you enrich a person you already identified, you are not drawing from that pool.
Where enrichment still gets you into trouble
We would be selling you something if we stopped there. Enrichment is not automatically safe.
Enrichment at purchase time is still purchased data. Some vendors enrich a static database once and sell you the snapshot. Check whether the lookup is live or cached.
"Verified" is not a standard. For some providers it means syntax and MX checks only, which passes almost anything on a catch-all domain. Ask what the verification actually tests.
Catch-all domains still hide failures. No provider can prove delivery to a domain that accepts everything. Treat catch-all results as lower confidence and send to them separately.
Enrichment does not create permission. A verified address belonging to a person with no interest in your product still produces complaints.
The ranking, worst to best
There is a version of this ranking where purchased data outperforms sloppy enrichment: a well-maintained, recently refreshed, compliance-clean B2B database beats an enrichment provider that returns cached pattern guesses.
Data source is a proxy. Freshness and verification methods are what actually matter.
How to protect yourself before you hit send
Knowing which data source is safest does not help if you skip the steps between "list ready" and "campaign live."
These checks apply regardless of whether your data is scraped, purchased, or enriched.
Vet your data provider before you pay
Ask three questions before handing over money or connecting an API:
- Where does the data originate? A provider that cannot explain its collection method is reselling someone else's scrape.
- When was each record last verified, and how? "Verified" with no definition usually means a syntax check, not a live SMTP handshake. Syntax passes catch-all domains and spam traps without blinking.
- What is the refund or credit policy on bounces above a stated threshold? Providers confident in their data put a number on it. Those who will not are telling you something.
Test before you scale
Never load a full list into your sending tool on day one. Send to a small batch of 50 to 100 addresses first and measure three things:
- Hard bounce rate. If it clears 2% on a sample that small, the full list will be worse.
- Spam complaint rate. Anything above 0.1% on a test batch is a stop signal.
- Reply and open patterns. Zero engagement from a supposedly fresh list means the addresses are alive but the people behind them are not expecting your email.
Only scale volume after the test batch clears all three.
Build ongoing list hygiene into your workflow
Data decays whether you bought it, scraped it, or enriched it. People change jobs, companies shut down, and domains expire. Three habits keep your list from rotting between campaigns:
- Re-verify before every campaign, not once at purchase. A quarterly enrichment cycle catches job changes before they become bounces.
- Set a sunset window for unengaged contacts. If someone has not opened or clicked in 90 days across multiple sends, suppress them. Continuing to send to silence trains mailbox providers to filter you.
- Monitor bounce and complaint rates after every send, not just the first one. A spike in week four means something changed in your list, and waiting until week eight to check turns a fixable dip into a domain reputation problem.
Segment by confidence before you send
Not every address in your list deserves the same treatment. Split your list into tiers:
- High confidence: verified with a live SMTP check within the last 30 days, matches your ICP, no prior bounces. Send at full volume.
- Medium confidence: catch-all domains, records older than 30 days, or roles that match your ICP loosely. Send at reduced volume from a secondary domain you can afford to lose.
- Low confidence: unverified, older than 90 days, or outside your ICP. Do not send. Re-enrich or discard.
This costs you ten minutes of filtering and saves you weeks of recovery if the low-confidence segment carries a trap.
What clean data cannot fix
Data quality sets your floor. It does not set your ceiling.
You can enrich every contact perfectly and still land in spam if the sending side is broken.
Google and Yahoo have required SPF, DKIM, and DMARC for senders above 5,000 messages a day since February 2024.
Microsoft matched them on May 5, 2025, and now rejects non-compliant bulk mail to Outlook.com, Hotmail.com, and Live.com with a 550 authentication error rather than filing it in Junk.
Beyond authentication, a new sending domain has no reputation at all. Mailbox providers have nothing to score you on, so they default to caution.
Sending 500 cold emails on day one from a domain registered last week produces spam placement regardless of how clean the list is.
That is what dedicated email warm up tools are for: it builds positive engagement history on the mailbox before your real campaign starts, and keeps building it while you send, so your reputation is not resting on the campaign itself.
The full floor looks like this:
- SPF, DKIM, and DMARC configured and aligned on every sending domain
- A warmed sending domain and mailbox, ramped over weeks rather than days
- Verified recipients, resolved close to send time
- Volume matched to your reputation, not to your quota
- A visible unsubscribe path and immediate opt-out handling
- Postmaster Tools and DMARC reports that somebody actually reads
Clean data gets you to the starting line. It does not run the race.
If you already sent to a bad list
Recoverable, usually. Not instantly.
- Stop sending to that file immediately. Every additional send extends the damage window.
- Separate the wreckage from your good sending. If the bad list went out on your primary domain, move legitimate outreach to a different sending domain while the first one recovers.
- Suppress, do not re-verify. Running a list that already hits traps through a verifier does not undo the trap hits. Suppress every address that bounced or went unanswered.
- Re-warm the affected mailboxes. Reputation recovers through sustained positive engagement at low volume, and that takes weeks rather than days. Tools like TrulyInbox run that engagement automatically across every connected mailbox, which matters when the damage is spread over a dozen sending accounts and you cannot rebuild them by hand.
- Rebuild the list from your ICP. Start from the accounts and roles you actually want, then enrich. Do not try to salvage the file.