← Blog Guide · 17 min read

Should You Build Your Own Job Change Detection Pipeline? The Four Systems You Would Have to Own

Short answer: the prototype is a weekend. The system is not. “Detect when our contacts change jobs” decomposes into four subsystems that each need permanent ownership: candidate detection, identity resolution after the move, email discovery with verification, and a defensible legal basis for processing the data. Three of those get harder over time, not easier. If you have an engineer who can own a data pipeline indefinitely, building is viable. If this is going to be someone’s third priority, you are not choosing between build and buy, you are choosing between buy and a pipeline that silently stops working in week nine.

This post is the honest version of that decision. No hand-waving about how hard scraping is, and no pretending the build case never wins.

Why the prototype lies to you

The first version works, and that is the problem.

You pull 500 contacts with LinkedIn URLs, fetch each profile on a weekly cron, diff the current company against what you stored, and post the diffs to Slack. It finds real moves in the first week. Everyone agrees this was easy.

What that prototype does not yet have: a key for the 60% of contacts with no LinkedIn URL, a way to re-identify someone whose name and email both changed, a verified address to actually email, retry and dedupe semantics so one blip does not re-alert 500 people, a reconciliation path for the runs that fail quietly, and an answer for your DPO. Each of those is where the engineering actually lives, and none of them show up in a demo.

Here is each one, with what it genuinely takes.

The naive source is LinkedIn. It is also the one with the clearest contractual exposure, so this deserves precision rather than vibes.

LinkedIn prohibits it in contract and in robots.txt. The User Agreement’s “Don’ts” section bars you from “develop, support or use software, devices, scripts, robots or any other means or processes (such as crawlers, browser plugins and add-ons or any other technology) to scrape or copy the Services, including profiles and other data from the Services” (LinkedIn User Agreement, section 8.2). Its robots.txt opens with the same point: “The use of robots or other automated means to access LinkedIn without the express permission of LinkedIn is strictly prohibited” (LinkedIn robots.txt).

The case everyone cites does not say what people think it says. The shorthand is “the Ninth Circuit legalized scraping public data in hiQ v. LinkedIn.” That is wrong twice. The 2022 opinion was a preliminary injunction appeal, and the panel only went as far as finding that hiQ “raised a serious question” on whether the Computer Fraud and Abuse Act’s “without authorization” concept applies to public pages (hiQ Labs, Inc. v. LinkedIn Corp., No. 17-16783, 9th Cir. Apr. 18, 2022). Then hiQ lost on the merits of the contract claim. The district court held that “the relevant language of the User Agreement unambiguously prohibits hiQ’s scraping and unauthorized use of the scraped data,” and that hiQ “breached LinkedIn’s User Agreement” (hiQ Labs v. LinkedIn, N.D. Cal., Nov. 4, 2022). The matter ended in a consent judgment of $500,000 against hiQ with a permanent injunction requiring it to stop scraping and destroy the derived data (Privacy World, Squire Patton Boggs; Morgan Lewis).

The useful takeaway for an architecture decision: CFAA exposure on public pages is genuinely unsettled, but contract exposure is not unsettled at all. Those are different risks and only one of them is a coin flip.

“Publicly available” is not a GDPR legal basis. If any of your contacts are in the EU, this is the section to send your legal team. Italy’s data protection authority put it directly in its Clearview AI order: “la pubblica disponibilità di dati in Internet non implica, per il solo fatto del loro pubblico stato, la legittimità della loro raccolta da parte di soggetti terzi,” which translates as the public availability of data on the internet does not, by the mere fact of being public, make its collection by third parties lawful. The decision concludes that collecting freely available personal data by web scraping “costituisce un trattamento di dati personali, che deve trovare legittimazione in una delle basi giuridiche previste dall’art. 6 del Regolamento,” meaning it is processing that needs one of the Article 6 legal bases. The order carried a 20 million euro fine (Garante per la protezione dei dati personali, provvedimento n. 50, 10 February 2022, doc. web 9751362; translations ours).

Recent EDPB guidance runs the same way, and closes the loophole people reach for next: “When data subjects make their personal data available online, for example on a web page accessible to everyone, this does not mean that the data subjects gave their consent to the scraping of their personal data for a specific purpose. In particular, the absence or non-applicability of a robots.txt file on a web site does not amount to consent within the meaning of the GDPR” (EDPB Guidelines 03/2026 on web scraping in the context of generative AI, version 1.0, adopted 7 July 2026, paragraph 45). Those guidelines were adopted for public consultation and are framed around AI training, so treat them as direction rather than a final instrument. The same body had already noted that “the mere fact that personal data is publicly accessible does not imply that ‘the data subject has manifestly made such data public’” (EDPB Opinion 28/2024, adopted 17 December 2024).

Worth knowing that the United States runs the opposite rule, which is why a US-only pipeline and an EU-exposed pipeline are not the same compliance project. Under the CCPA, “personal information does not include publicly available information,” including information “a business has a reasonable basis to believe is lawfully made available to the general public by the consumer or from widely distributed media” (California Attorney General, CCPA FAQ). Build once for both jurisdictions and you inherit the stricter one.

None of this makes building impossible. It does mean the detection layer is a legal review with an engineering component, not the reverse, and that the review has to be redone whenever a source’s terms change.

System two: identity resolution, which is the actual hard part

This is the subsystem teams underestimate most, because the prototype hides it behind the LinkedIn URL.

A job change is the one event that invalidates most of your match keys at once. The email changes. The company changes. The title changes. What survives is a name, which is not unique, and a profile URL, which you may not have and which the person can change.

The record linkage literature has named this problem for decades. From the standard reference text: “A major challenge in data matching is the lack of common entity identifiers in the databases to be matched. As a result of this, the matching needs to be conducted using attributes that contain partially identifying information, such as names, addresses, or dates of birth. However, such identifying information is often of low quality. Personal details especially suffer from frequently occurring typographical variations and errors, such information can change over time, or it is only partially available in the databases to be matched” (Peter Christen, Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection, Springer).

“Such information can change over time” is the entire job here, stated by the field that studies it.

Two practical consequences for your design:

Your coverage ceiling equals your key coverage. If you key on LinkedIn URL, your detection rate can never exceed the share of contacts for whom you hold a correct, current URL. For most CRMs that is well under half, and it skews badly: the records with clean URLs are the ones someone prospected last quarter, while the records without them are older closed-won contacts, which is to say your actual champions. You end up watching your newest leads closely and your most valuable relationships not at all. Audit this before anything else. Export closed-won contacts, count populated LinkedIn URLs, and treat that percentage as your hard ceiling.

Naive matching does not scale, and strict matching does not work. Comparing every record in a 100,000-contact database against a candidate set of 1,000,000 profiles is 10^11 comparisons, which is why blocking, indexing and probabilistic scoring exist in the first place. Tighten the thresholds to keep precision and you drop real moves. Loosen them and you alert a rep that their champion moved to a company they have never heard of, because you matched a different person with the same name. One false positive sent to a sales team costs you more trust than ten misses, because the misses are invisible.

Getting this right is a scoring model you tune and re-tune with feedback, not a join. It is the part of the system that is never finished.

System three: email, where a shortcut damages an asset you cannot easily repair

Detection tells you someone moved. It does not tell you how to reach them, and the address in your CRM is now dead. So most builds bolt on pattern guessing: [email protected] and hope.

Do not. Guessed-pattern sending puts your sending domain at risk, and the published thresholds are tight. Amazon SES states it plainly: “For best results, you should maintain a bounce rate below 2%. Higher bounce rates can impact the delivery of your emails,” and escalates from there, with accounts placed under review at 5% and sending potentially paused at 10% (Amazon SES Developer Guide, Sending review process FAQs). Postmark sets a hard ceiling: “Postmark requires that you keep your bounce rate below 10% and spam complaint rate below 0.1%. If the server’s bounce or spam complaint rate exceeds the limit… to avoid having your server’s sending suspended” (Postmark, Servers FAQ). SES is equally direct about where bad addresses come from: “Never rent or buy email lists. These lists may contain large numbers of invalid addresses… If your messages land in a spam trap, your delivery rates and sender reputation could be irrevocably damaged” (Amazon SES, Email program success metrics).

One attribution note, because this gets garbled constantly: Google’s sender guidelines publish spam complaint thresholds, not bounce thresholds. Google’s rule is that “senders should keep their spam rate below 0.1% and should prevent spam rates from ever reaching 0.3%,” with mitigation unavailable above 0.3% (Google Workspace Admin Help, Email sender guidelines). The 2% and 5% bounce numbers are SES’s. Do not put them in Google’s mouth when you write the internal memo.

The obvious engineering workaround, probing mailservers to validate addresses without sending, is itself prohibited on at least one major receiver. Outlook.com’s postmaster policies state: “Senders must not use namespace mining techniques against Outlook.com inbound email servers. This is the practice of verifying email addresses without sending (or attempting to send) emails to those addresses.” The same policies require that after a permanent non-delivery response “the sender must not attempt to retransmit that message to that recipient” (Outlook.com Postmaster, Policies).

So you need a real verification chain with provider fallbacks and a conservative accept threshold, and you need it to stay current as providers change behavior. This matters more on job change outreach than on any other list, because this is your warmest audience and the sender is usually a real rep’s own mailbox. Burning that domain’s reputation to save a verification step is an expensive trade. Our order of operations for finding a new work email covers the manual version of this chain.

System four: continuity, which is the one that actually kills the project

The first three are engineering problems. This one is organizational, and it is why most internal job change trackers are dead within two quarters.

A job change feed produces no visible output when it breaks. If your billing pipeline fails, invoices do not go out and someone escalates within an hour. If your job change cron fails, the Slack channel goes quiet, and a quiet channel looks exactly like a quiet month. Nobody files a ticket. You find out in the quarterly review when someone asks why champion-sourced pipeline went to zero in August.

That failure mode is why the continuity requirement is not negotiable. A real build needs run-level observability, alerting on absence of events rather than presence of errors, dedupe keys so a replay is safe, and a reconciliation job that re-reads a window and backfills whatever the primary path dropped.

And the underlying volume never pauses, which is what makes a stalled pipeline expensive. 2 to 3% of B2B contacts change jobs every month. The labor statistics corroborate the churn independently: US median employee tenure was 4.1 years in January 2026, but only 3.0 years for workers aged 25 to 34, and 20.6% of all wage and salary workers had been with their employer a year or less (U.S. Bureau of Labor Statistics, Employee Tenure, 2026). JOLTS recorded 62.8 million total separations in 2025, of which 38.0 million were voluntary quits, 60.6% of the total (BLS, Job Openings and Labor Turnover Survey). HubSpot’s widely quoted marketing-database figure, derived from MarketingSherpa, puts B2B data decay at about 22.5% per year (HubSpot, Database Decay Simulation) though that one is a vendor restatement of older research, so lean on the BLS numbers when you need something defensible. We work through what the churn costs in CRM data decay.

On a 10,000-contact CRM that arithmetic means roughly 20% surface as hot leads on day one, about 2,000 people, then 200 to 300 new warm leads per month after that. Two to three hundred moves a month is a steady operational flow, not a campaign. Steady flows want a pipeline that runs whether or not anyone is paying attention this quarter.

When building is genuinely the right answer

It happens, and the honest cases are specific:

  • The signal is your product. If you are selling job change data, or it feeds a model that is core intellectual property, own the pipeline. Buying a dependency you resell is a bad structural position.
  • You already have the hard parts. If there is a working identity resolution service and a maintained email verification chain in production for other reasons, you are adding a detection source to existing infrastructure rather than building four systems. The marginal cost is genuinely small.
  • Your population is unusual. A niche where general people-data coverage is thin, or non-LinkedIn geographies, can mean no vendor matches your list well. Run a match-rate test against your real contacts before you accept this as true, because it is often assumed and rarely measured.
  • Data residency forbids the transfer. Some regulated environments cannot send a contact list to a third party at all. That is a hard constraint, not a preference.

Outside those, the thing you would be building is maintenance. Note also the middle option: a general-purpose workbench like Clay can run a scheduled job change check without you writing the pipeline, with its own tradeoffs around keys, cadence and per-check credits. We cover those in Clay job change signals, and how to evaluate a vendor feed on freshness rather than field count in choosing a people data API.

What buying this actually looks like

The reason to buy is not that the endpoint is hard. The endpoint is the easy part. You are buying detection coverage, identity resolution and email verification as a maintained service, with the moves arriving as structured events.

Concretely, with Champions, each detected move is a play.created event carrying the person already matched to the record you track, old_company and old_title, new_company and new_title, the best known email with an email_status so you are not sending to a guess, account_fit and contact_fit scored against your own ICP rules before the event is sent, the owner and crm_record_id so the alert routes itself, and detected_at for your SLAs. Push it to an HTTPS endpoint you register, or pull GET /v1/changes with a since timestamp. Run both: webhooks for latency, polling as the reconciliation job so a delivery missed during a deploy still lands. Deliveries are signed, and event_key is your dedupe key, so a retry is safe to replay. The full field list, delivery modes and signature scheme are on the API page.

Note the shape of what that removes. Identity resolution happens before the event exists, so you receive a match rather than a reconciliation task. Email verification state ships with the record. And the volume is modest by design: the same 2 to 3% monthly rate means a 10,000-contact CRM produces a few hundred events a month, not a firehose, which is why this integrates as a normal consumer and not a streaming project. If you want the event-consumption pattern in more depth, see webhooks for job changes.

Most teams never touch the API at all. Champions writes each move back to the CRM and routes alerts to Slack, with no app for reps to install and no new logins, across Salesforce, HubSpot, Pipedrive, Zoho and others. The how it works page walks the monitoring end to end. Reach for the API when the signal needs to go somewhere that is not the CRM: a warehouse, a lead router, a scoring model, an internal app.

Why the signal is worth either path

Whichever way the decision goes, the reason to have this at all:

  • Champion-sourced deals close at a 2.7x higher probability, in roughly half the sales cycle.
  • Only about 6% of customer contacts tell a vendor they changed jobs. The other 94% never reach out, so without detection the move is invisible to you.
  • You have a 60 to 70% chance of selling to an existing relationship, against 5 to 20% for a cold prospect.
  • 65% of a company’s business comes from existing customers, and acquiring a new customer costs 6x more than retaining one.

For how to report that flow as its own forecast category rather than burying it in general outbound, see champion-sourced pipeline metrics and the RevOps use cases.

Two things to weigh if you are comparing vendors rather than building. Champions monitors intent-to-move signals rather than only confirmed moves, which shifts the outreach window earlier than a post-hoc detection can. And the ROI is contractual: if champion-sourced revenue does not reach 2x the annual service fee, you are covered under our terms. Among the job change tools teams usually shortlist, that guarantee is specific to us. The job change tracking comparison lays the options side by side.

The bottom line

Run the decision on the four systems, not on the prototype:

  1. Detection needs a source whose terms you can live with and a legal basis that survives a DPO review. Contract exposure on LinkedIn scraping is settled law against the scraper, and “publicly available” is not an Article 6 basis.
  2. Identity resolution is a scoring model you tune forever, and your coverage ceiling is your key coverage. Measure that percentage before you decide anything.
  3. Email needs real verification, because the shortcut risks a sending domain at a 2% bounce threshold and your warmest list is the wrong place to test it.
  4. Continuity is the one that kills projects, because a stalled feed produces silence and silence looks like a slow month.

If you have an owner for all four indefinitely, build it. If the honest answer is that it becomes someone’s third priority by next quarter, the maintenance is the thing to buy, and the endpoint was never the hard part.

Want to see the event shape against your own contact list before you commit either way? Book a demo and we will run a match-rate test on your closed-won contacts, which is the number that decides this. If you would rather just ask a question first, email [email protected].

See champion tracking on your own data.

Book a demo and we'll show you how many warm opportunities are already sitting in your CRM.