Finding accurate email addresses and mobile phone numbers on the open web isn’t as simple as scanning a single page and calling it a day. The best data often lives inside a small network of related pages that surround a person or a company online, and that’s where the idea of “connected web sources” comes into play. Instead of treating each webpage as an isolated record, we treat it as part of a web of related sources that can be linked back to a single contact.
Often, the process starts with what we can call the original web source, this might be a company website, a profile page, or a directory listing where we first see a person’s name, role, and employer. Personal email addresses and mobile numbers can sometimes be found directly on that original page, but that’s usually the exception, not the rule. More often, we need to follow the signals on that page to find additional sources. For example, if we see a Facebook URL on the original source, we’ll follow that link and look for contact details on the Facebook page itself. When we find contact information like a personal email or mobile number there, we associate it back to the original record. We apply the same approach to the employer’s website, scanning it for about pages, team pages, contact pages, author bios, and any other content that might expose contact details. From there, we look for additional connected web sources — social profiles, profile directories, related sites — that clearly tie back to the same individual or organization. Matching these connected web sources to the source URL is our primary method of gathering personal email addresses and mobile phones.
However, we don’t always begin with the source URL. In some cases, we start with the mobile phone number itself and work in reverse. Certain countries have dedicated area codes assigned specifically to mobile carriers. That gives us a powerful way to focus our efforts: we can target websites from these countries and identify phone numbers that match the dedicated mobile ranges. Once we detect a likely mobile number, we examine the pages where it appears to gather context such as name, company, job title, and any social or corporate links on the page. We then follow those connected web sources to trace the number back to a contact record in our database. This approach not only helps us discover more mobile numbers, it also acts as a built-in quality check because we are tying each number to a country, its phone conventions, and the context in which the number appears.
Personal email addresses are generally easier to uncover than mobile phones, largely because they’re not constrained by national dialing conventions. We’re not limited to specific countries or formats, and the universe of personal email domains is well-known — think of the major webmail providers and the long tail of familiar personal domains. That allows us to scan large volumes of web content for these domains at scale. Just last month, we identified 68 million email addresses on websites associated with companies already in our database. Once we detect a personal email on a page, we again rely on connected web sources to match that email to a specific person. A personal email might appear on a blog authored by someone who lists their employer and role, on a portfolio site linked to a LinkedIn profile, or on a directory listing that includes both a personal email and a company name. Each of those cases gives us the context we need to connect that email to the right contact record.
At the heart of all of this is a simple but powerful principle: don’t treat each page as a standalone record. Treat it as part of a network. Whether we’re starting from a source URL and fanning out to social and company pages, or starting from a mobile number or personal email and working backward to a person and company, the real engine is the same. By systematically matching connected web sources and tying them back to a single, unified contact record, we can surface personal email addresses and mobile phones at scale with far greater accuracy and confidence than traditional one-page-at-a-time scanning.
