All Blogs

How Breakout's Website Deanonymization Works

How Breakout's Website Deanonymization Works

What website deanonymization actually identifies, realistic match rates, why most deanonymization data never gets used, and how Breakout's waterfall across multiple providers plus CRM context fixes that.

Rahul Krishnan

Head of GTM

Published On

Share blog

Content

Website deanonymization is the process of identifying which companies, and in some cases which people, are visiting your website without filling out a form. Most B2B buyers do a good share of their research before they talk to anyone, and almost all of it is anonymous — deanonymization tells you who they are while that research is still happening.

Account-level and contact-level identification

There are two levels of deanonymization, and they work very differently.

Account-level matches the visitor's IP address against databases that map IP ranges to companies — if someone is browsing from inside Acme's network, you see Acme, along with its firmographics (industry, size, revenue, location) and what the company did on your site: which pages, how many visits, how many different people. You don't get the person.

Contact-level matches the visitor to an individual through identity-resolution networks, which link browsers and devices to known people using data shared across many websites. When it works, you get a name, a business email, a title, and a LinkedIn profile — much more useful than account-level data, but for fewer visitors.

Account-level tells you a target account is researching you. Contact-level tells you who to talk to. Most teams need both, since the person who visits is often not the person who buys, and an account that shows up five times in a week from three different people tells you more than any single visit.

What match rates to expect

Match rate is the share of your traffic a provider can identify — it's the first number to ask a vendor for, and the one most often left vague.

On B2B traffic, we typically see 60 to 70% of visitors identified at the account level and 20 to 25% at the contact level. Those numbers move a lot depending on region (contact-level identification works best in the US and Western Europe; coverage in Asia and Latin America is much thinner at both levels), how people connect (a visitor on an office network is easy to place; someone on mobile data or a home connection often isn't), and your traffic mix (if a large share of your visitors come from regions or segments you don't sell into, your overall rate drops — one team we spoke with had 30,000 to 40,000 visitors a month, but most were end users of their product in a country they didn't sell into, so their headline match rate looked very different from the traffic that actually mattered).

The bigger reason match rates vary is that every provider has different coverage, built from different partners and methods, so each has blind spots — one may be strong on large enterprise networks and weaker on smaller companies, another may do well in North America and poorly in Europe. No single provider sees all of your traffic. So when a vendor quotes a match rate, ask two things: was it measured on traffic like yours, and does it come from one provider or several.

Why most de-anonymization data never gets used

Most marketing teams we talk to already pay for some form of de-anonymization. When we ask what they do with it, the answer is usually a version of the same thing: the data shows up somewhere, and if a rep happens to look, they might act on it — there's no process behind it. That's rarely because the team doesn't care; the data is hard to act on, for four reasons:

Gaps from a single source — with one provider, a large share of your traffic is simply missing, and the gaps aren't random. If your best segment is one your provider covers badly, you're missing traffic that matters, and reps notice the misses and stop trusting the feed.

Noise in contact-level data — straight from a provider, it comes with records nobody can use: personal Gmail and Yahoo addresses, duplicates of people already in your CRM, titles two jobs out of date, people who left the company last year. Without a cleaning layer, it's not actionable for reps.

No context — a company name on its own tells a rep very little. Did they read the pricing page or a 2022 blog post? Did they come from a G2 comparison? Are they already a customer, or in an open deal? For the data to be actionable, the visit has to be put in the context of your account list.

No action — most tools stop at identification. The output is a Slack alert, a dashboard, or a field on the account record, with no automated way to act on it — and website visits have a short half-life, so an alert read on Thursday about a Monday visit is mostly useless.

How Breakout does it

De-anonymization in Breakout runs underneath everything else we do on your website, and can also be used on its own, with identified accounts and contacts sent straight to your CRM.

Waterfall matching across multiple providers. Instead of relying on one data provider, we run every visitor through several in sequence — if the first can't identify a visitor, the second tries, then the next. This happens separately at the account level and contact level, with different providers for each, since the best sources for one aren't the best for the other, and each provider's blind spots differ. Say your traffic mixes US enterprise accounts with mid-market companies in Europe: a provider built mostly on enterprise network data does well on the first group and misses much of the second, so a provider stronger in European mid-market sits behind it to start covering that gap. This is how we get to 60 to 70% at the account level and 20 to 25% at the contact level on typical B2B traffic — doing the same thing yourself would mean paying for several data contracts and reconciling the overlap between them; the waterfall handles that before the data reaches you.

Cleaning and enrichment before anything reaches you. Raw contact-level data is noisy, and a lot of what single-source tools return can't be used as-is. We clean every record before it reaches you, attempting to map personal email addresses to work email addresses and filtering out what can't be mapped, so what lands in your CRM is a business contact a rep can actually email. Every record is enriched with firmographics, title, and LinkedIn. If half your contacts have personal email addresses, reps can't do anything with it, and you lose their trust — it also makes the data harder to use across ad platforms.

Context on every visit. Each identified account and contact comes with what they actually did — pages browsed, number of visits, pages per visit, how many people from the account came, and where the traffic came from (paid campaign, G2, email, organic search). We also map every account one-to-one to its Salesforce or HubSpot record, so whatever you track there — account owner, tier, customer status, an open opportunity — comes along with it. That's the difference between "someone from Acme visited" and "a second person from Acme, a tier-one account owned by your enterprise AE with an open opportunity, has read the pricing page twice this week."

Scoring against your ICP. Every account and contact is scored low, medium, or high against your ICP, described in plain language the way you'd explain it to a new SDR — the scoring uses that along with firmographics and everything else attached to the account, since firmographics alone miss things like whether a 200-person company is a likely buyer or a likely partner, which your CRM context can capture. So what the rep sees is the handful of visitors worth their time, not 400 companies sorted by visit count.

What to do with identified traffic

  • Sync it to your CRM. Identified accounts and contacts, with their activity and scores, go into Salesforce or HubSpot so reps work from their own system instead of another dashboard — this is where most teams start.

  • Route key accounts in real time. When a target account is on the site, the AI agent can offer a meeting with the account owner, or bring the owner into the conversation live, instead of sending them through a form.

  • Personalize what they see. A visitor from an Enterprise account gets the relevant case study; a repeat visitor from a target account gets pricing or other bottom-of-funnel content.

  • Reach out after they leave. Identified contacts, or others at an account that keeps coming back, get outreach based on what they looked at — for example, targeting accounts that visit the pricing page.

  • Warm up the ones who aren't ready. Interested visitors who aren't ready to talk get a nurture sequence with useful material, not a "saw you on our site" email.

Privacy and compliance

De-anonymization works from a visitor's IP address and browser fingerprint. Account-level identification matches that data to a company, contact-level matches it to an individual, and the rules vary by region. A few things to check before turning it on:

Your consent banner and privacy policy — because the script relies on IP address and browser fingerprinting, it falls under functional or targeting cookies depending on your legal posture; load it through your consent management platform (CMP) so it respects each visitor's consent choices, and disclose the tracking in your privacy policy.

Where your traffic comes from — European visitors are covered by GDPR, which is stricter about identifying individuals than US law, so check with your legal team on whether contact-level identification is appropriate for that traffic.

What the vendor tracks — Breakout only tracks activity on your own website; it doesn't follow visitors across other sites.

Breakout shares its cookie consent and data-handling documentation as part of any evaluation. None of this is legal advice.

Frequently asked questions

We already have Demandbase or 6sense. Do we need this?

Maybe not, if you only need identification and you're happy with the match rate. Breakout's waterfall includes sources like these and adds others to fill their gaps. The comparison worth making is match rate on your own traffic, and more importantly, how you're acting on that data — if it just sits in a dashboard, that's the problem to solve, whether through Breakout or something else.

We have ZoomInfo. Isn't that the same thing?

No. ZoomInfo, PitchBook, and similar databases tell you about companies and people you already know to look up. Deanonymization tells you which of them are on your site right now. Most teams need both — deanonymization to know who to look at, and a contact database to fill in the rest.

Does it identify visitors on mobile?

Sometimes. A phone on an office Wi-Fi network can usually be matched to the company; a phone on mobile data may or may not, depending on browser fingerprinting.

What happens to visitors you can't identify?

They don't appear in your account list, since there's nothing a rep can do with an unnamed visitor. They still see your site and the AI agent as usual, and if they start a conversation or fill out a form, they're captured like any other lead.

How long does setup take?

You can go live with de-anonymization in less than an hour. Setting up the data flow into your CRM can take a couple of days, depending on filtering and field mapping.

How is it priced?

Pricing is based on your monthly website traffic. Contact-level identification is priced in addition to account-level de-anonymization.

Frequently Asked Questions

Want a smarter, better way to build pipeline?

See how Breakout's AI SDR can run your entire inbound pipeline generation

Want a smarter, better way to build pipeline?

See how Breakout's AI SDR can run your entire inbound pipeline generation

Want a smarter, better way to build pipeline?

See how Breakout's AI SDR can run your entire inbound pipeline generation