A $350B+ industry where 4,000+ brokers operate with near-zero regulation. Acxiom has 2.5B consumer records, location data enables warrantless surveillance, and opt-out requires 1,000+ hours of individual effort. 10 pain points per category across the entire surveillance economy.
1. Data Collection Scale & ScopeCritical
1App SDK Supply Chain Leakage▾
Problem
Mobile apps embed third-party SDKs from advertising networks, analytics providers, and data brokers that siphon data without user awareness. A typical free app contains 6-10 SDKs, each independently collecting device identifiers, location, contacts, and behavioral data. Users consent to the app's stated purpose but have no visibility into the SDK supply chain operating behind it.
Current State
The Exodus Privacy project has catalogued SDKs in over 100,000 Android apps, finding that popular apps routinely embed trackers from Facebook (Meta Audience Network), Google (AdMob, Firebase), AppsFlyer, Adjust, Branch, Kochava, and X-Mode. Apple's App Tracking Transparency (ATT) framework reduced iOS tracking rates from ~70% to ~25%, but SDK-level data collection via fingerprinting continues. On Android, Google's Privacy Sandbox for mobile remains incomplete. No platform provides SDK-level consent granularity.
Impact
A weather app sharing GPS coordinates with X-Mode's SDK enabled the US military to purchase movement data of ordinary citizens. The Wall Street Journal's 2019 "Your Apps Know Where You Were Last Night" investigation documented how apps like WeatherBug and GasBuddy sold precise location data to 40+ third-party companies per app. A Muslim prayer app (Muslim Pro) was found sending location data to X-Mode, which sold it to US defense contractors, per Motherboard's 2020 investigation.
References
Exodus Privacy tracker database; Motherboard investigation of X-Mode and Muslim Pro (November 2020); WSJ "Your Apps Know Where You Were Last Night" (December 2018); FTC complaint against Kochava (August 2022); Apple ATT transparency reports.
2Acxiom's 2.5 Billion Consumer Profiles▾
Problem
Acxiom (rebranded as LiveRamp's data marketplace) maintains marketing data on approximately 2.5 billion consumers worldwide and over 700 million consumers in the US alone. Each profile contains up to 3,000 data attributes covering demographics, financial behavior, purchase history, media consumption, political affiliation, health interests, and household composition. This data is collected from public records, surveys, purchase transactions, loyalty programs, and thousands of partnership agreements with retailers and publishers.
Current State
Acxiom rebranded its data marketplace as LiveRamp Data Marketplace after LiveRamp's 2018 acquisition spin-off. The company operates the largest consumer identity graph connecting offline and online identities. Vermont's data broker registry lists Acxiom/LiveRamp, but registration is merely informational with no restrictions on data practices. Acxiom's opt-out page (aboutthedata.com, later deprecated) provided a view of only a fraction of stored attributes and required submitting additional PII (SSN last 4 digits) to verify identity for opt-out.
Impact
Acxiom's data has been used in political microtargeting (Cambridge Analytica sourced seed audiences through Acxiom segments), discriminatory advertising (ProPublica's 2016 investigation showed Facebook allowing Acxiom-sourced "ethnic affinity" targeting for housing ads), and predatory lending (the CFPB documented data brokers selling lists of financially vulnerable consumers). A single Acxiom profile enables targeting an individual across every channel — mail, email, phone, web, social media, connected TV — creating inescapable commercial surveillance.
References
Acxiom corporate filings and investor presentations; FTC "Data Brokers: A Call for Transparency and Accountability" (2014); ProPublica "Facebook Lets Advertisers Exclude Users by Race" (2016); Vermont Secretary of State data broker registry; Senate Commerce Committee hearing testimony (2023).
3Location Data Harvesting at GPS Precision▾
Problem
Location data brokers collect GPS-precision coordinates (accurate to ~3 meters) from mobile devices at intervals of seconds to minutes, creating comprehensive movement histories for hundreds of millions of people. This data reveals home addresses, workplaces, medical visits, religious attendance, political activities, romantic relationships, and daily routines. Companies like Gravy Analytics, SafeGraph, Placer.ai, and Foursquare aggregate location from app SDKs, bidstream data, and direct partnerships.
Current State
The FTC brought its first location data cases in 2024: Kochava (selling geofenced location data including visits to reproductive health clinics, addiction treatment centers, and places of worship), X-Mode/Outlogic (selling precise location data to government contractors without consent), and InMarket (collecting location from 300+ million devices through SDK partnerships). SafeGraph stopped selling data tied to Planned Parenthood visits only after public pressure following the Dobbs decision. Gravy Analytics was breached in January 2025, exposing precise location data for millions.
Impact
The Pillar, a Catholic news outlet, used location data purchased from a broker to identify a Catholic priest using Grindr, leading to his forced resignation (2021). Anti-abortion groups used SafeGraph data to track clinic visitor demographics. The Gravy Analytics breach exposed location histories showing visits to the White House, Pentagon, military bases, and foreign embassies — a national security catastrophe. Location data cannot be effectively anonymized: MIT research demonstrated that four spatiotemporal points are sufficient to uniquely identify 95% of individuals.
References
FTC v. Kochava (2022, amended 2024); FTC v. X-Mode/Outlogic (2024); FTC v. InMarket (2024); The Pillar Grindr investigation (July 2021); de Montjoye et al. "Unique in the Crowd" (Nature, 2013); Gravy Analytics breach reporting (TechCrunch, January 2025).
4Public Records as Bulk Data Source▾
Problem
Data brokers systematically harvest government public records — property deeds, voter registrations, court filings, business licenses, UCC filings, marriage/divorce records, and death certificates — as a foundational data layer. These records, created for specific governmental purposes, become the backbone of commercial profiles. Every home purchase, voter registration, lawsuit, and marriage generates records that brokers ingest within days.
Current State
LexisNexis, Thomson Reuters (CLEAR), and Palantir aggregate public records from 3,000+ county courthouses, 50 state governments, and federal databases. Most jurisdictions have no restrictions on commercial bulk access to public records. The DPPA (Driver's Privacy Protection Act) restricts DMV records, but 14 exemptions render it largely ineffective. Property records are openly available in most US counties, and brokers scrape them continuously via automated systems.
Impact
Domestic violence survivors who obtain protective orders discover that the court filing itself — containing their name and address — is scraped by brokers and appears on people-search sites within weeks. Juror information from court records has been used to identify and intimidate jurors in high-profile cases. A 2023 Duke University study found that data brokers openly sell personal data of US military members, including names, home addresses, financial information, and information about their children, for as little as $0.12 per record.
References
LexisNexis public records aggregation documentation; Duke University "Data Brokers and the Sale of Data on U.S. Military Personnel" (2023); DPPA exemptions analysis (Electronic Privacy Information Center); National Network to End Domestic Violence public records advocacy; r/privacy threads on property records appearing on Spokeo/WhitePages within days of home purchase.
5Purchase Data from Retailers and Financial Institutions▾
Problem
Retailers sell transaction-level purchase data to brokers, and credit card companies sell anonymized (but re-identifiable) spending patterns. Mastercard's data analytics division, Visa's Visa Analytics Platform, and American Express sell aggregated consumer spending insights. Retailers like grocery chains sell loyalty card purchase histories to Acxiom, Nielsen, and IRI. These datasets reveal diet, health conditions, pregnancy status, financial distress, and personal habits.
Current State
Target's predictive pregnancy scoring algorithm (documented by the New York Times in 2012) demonstrated that purchase patterns alone can identify major life events. Nielsen Catalina Solutions (now Circana) links loyalty card purchases to advertising exposure for closed-loop attribution. Amazon Shopper Panel explicitly pays users for purchase data. The FTC has not brought enforcement actions specifically targeting purchase data brokerage, and no federal law restricts the sale of purchase history.
Impact
Insurance companies purchase spending data to identify "risky" behaviors (alcohol purchases, fast food frequency, gun purchases) that correlate with health costs. Employers use purchase data services to screen candidates (a gym membership signals health-consciousness; frequent bar visits signal risk). The Target pregnancy example became emblematic: a father learned his teenage daughter was pregnant because Target's algorithm sent pregnancy-related coupons to their household before she disclosed it to her family.
References
Charles Duhigg, "How Companies Learn Your Secrets" (NYT, 2012); Mastercard Data & Services documentation; FTC workshop on data brokers and consumer scoring; r/privacy discussions on loyalty card data resale; Kashmir Hill, "I Cut the 'Big Five' Tech Giants From My Life" (Gizmodo, 2019).
6IoT and Smart Device Telemetry Harvesting▾
Problem
Smart TVs, connected cars, voice assistants, fitness trackers, smart home devices, and wearables generate continuous telemetry streams that manufacturers and third parties collect, aggregate, and sell. Vizio paid a $2.2 million FTC settlement for collecting second-by-second viewing data from 11 million TVs without consent. Connected car manufacturers collect GPS location, driving behavior, in-car conversations (via voice assistants), and passenger information.
Current State
The Mozilla Foundation's "Privacy Not Included" project found that 25 out of 25 major car brands earned their worst privacy rating. Car manufacturers including Toyota, GM, Honda, and Hyundai collect driving behavior data and share it with insurance companies (LexisNexis Risk Solutions). GM's OnStar collected and sold driving behavior to LexisNexis, which resold it to insurers who raised premiums, as reported by the New York Times in 2024. Samsung, LG, and Vizio smart TVs use ACR (automatic content recognition) to track viewing habits and sell the data to advertisers.
Impact
Consumers who purchased smart TVs discovered their viewing habits were being sold to advertisers without their knowledge — every show watched, every channel switched, every pause and rewind catalogued and monetized. GM drivers discovered their insurance premiums increased after LexisNexis received detailed driving behavior data (hard braking events, late-night driving) collected through their vehicles' OnStar systems, which they believed was merely an emergency roadside service. Ring doorbells created a neighborhood surveillance network accessible to 2,000+ police departments through Amazon's partnership program.
References
FTC v. Vizio ($2.2M settlement, 2017); Mozilla "Privacy Not Included" car reviews (2023); Kashmir Hill, "Your Car May Be Spying On You" (NYT, 2024); Sen. Markey inquiry into connected car data sharing; Ring/Amazon police partnership reporting (EFF).
7Social Media Data Harvesting at Scale▾
Problem
Social media platforms are both data brokers and data sources for brokers. Meta's advertising system processes 2.9 billion user profiles. Social media scraping operations collect public posts, photos, check-ins, relationship status, employment history, and social graphs. Cambridge Analytica demonstrated that app-based collection could harvest data from 87 million Facebook users through 270,000 app installs using the friends permission API.
Current State
After Cambridge Analytica, Facebook restricted API access but continued selling data through its advertising platform's "Custom Audiences" and "Lookalike Audiences" features. LinkedIn allows data enrichment companies to map professional networks. Twitter/X under Musk expanded API data sales while weakening content moderation. TikTok's algorithm collects behavioral data (watch time per video, pause patterns, rewatches) that creates psychometric profiles. Clearview AI scraped 30+ billion images from social media to build its facial recognition database.
Impact
The Cambridge Analytica scandal revealed that a single personality quiz app accessed data on 87 million people — but the underlying data brokerage infrastructure that enabled it remains intact. Clearview AI's scraping demonstrates that public social media posts become permanent biometric surveillance assets. TikTok's granular behavioral data collection creates engagement profiles that Chinese intelligence could theoretically access under China's National Intelligence Law, which was a core argument in the US ban legislation.
References
UK ICO Cambridge Analytica investigation (2018-2020); FTC Meta $5 billion settlement (2019); Clearview AI — ACLU v. Clearview settlement; Senate Intelligence Committee TikTok hearings (2023-2024); The Markup "How We Built a Facebook Ad Library."
8Healthcare Data Broker Pipeline▾
Problem
While HIPAA protects medical records held by covered entities, a massive parallel healthcare data economy operates outside HIPAA's scope. Health apps, pharmacy discount cards (GoodRx), period tracking apps, fitness devices, health-related web searches, and genetic testing services collect sensitive health data and sell it to brokers. These are not "covered entities" under HIPAA and face no health privacy restrictions.
Current State
The FTC fined GoodRx $1.5 million in 2023 for sharing users' health data with Facebook, Google, and other advertising companies — the first enforcement under the Health Breach Notification Rule. Period tracking apps Flo and Premom settled FTC complaints for sharing sensitive reproductive health data with third parties. 23andMe's bankruptcy filing in 2024 raised questions about what happens to the genetic data of 15 million customers when a genomics company fails. HIPAA does not apply to any of these entities.
Impact
After Dobbs, period tracking apps became potential evidence sources for abortion prosecutions, prompting mass deletions documented on r/privacy and r/TwoXChromosomes. GoodRx users discovered that their prescription data — revealing conditions from HIV to mental health to fertility treatments — had been shared with Meta's advertising platform, enabling pharmaceutical companies to target them with ads. 23andMe's financial distress means that 15 million people's genetic data could be acquired by any purchaser in bankruptcy proceedings.
References
FTC v. GoodRx ($1.5M, 2023); FTC v. Flo Health (2021); FTC v. Premom/Easy Healthcare (2023); 23andMe bankruptcy reporting (Wired, 2024); HIPAA coverage gap analysis (The Markup); r/privacy megathreads on period tracker data post-Dobbs.
9Children's Data Collection Through EdTech and Gaming▾
Problem
Children generate extensive data profiles through educational technology, gaming platforms, and connected toys that is collected and brokered despite COPPA protections. School-mandated platforms (Google Classroom, Canvas, Clever) collect behavioral and academic data. Gaming platforms (Roblox, Fortnite, Minecraft) collect behavioral patterns, social interactions, voice chat data, and spending patterns. EdTech companies pivot to selling "insights" derived from student data.
Current State
Epic Games paid a $275 million FTC fine (2022) for COPPA violations related to Fortnite's collection of children's data and use of dark patterns. The FTC fined Microsoft (Minecraft) and Amazon (Alexa/Ring) for children's privacy violations. Despite enforcement, most children's apps violate COPPA according to studies — a 2023 ICSI/AppCensus study found that 72% of children's apps on Google Play shared data with third-party trackers. Schools cannot meaningfully consent on behalf of students to commercial data collection.
Impact
Children's data is uniquely valuable to brokers because it establishes a baseline profile from childhood through adulthood. A child's educational performance data, behavioral patterns, family income indicators (free/reduced lunch), and social interactions create a predictive profile before they reach the age of consent. The generation entering adulthood now has been profiled since birth, with no meaningful ability to access or delete their childhood data trail.
References
FTC v. Epic Games ($275M, 2022); FTC v. Amazon/Alexa ($25M, 2023); ICSI/AppCensus children's app study (2023); EFF "Spying on Students" report; Student Privacy Compass database; r/privacy discussions on children's data permanence.
10Cross-Device and Cross-Platform Identity Linkage▾
Problem
Device identity graphs maintained by companies like LiveRamp, Tapad (acquired by Experian), Drawbridge (acquired by LinkedIn), and The Trade Desk link an individual's phone, tablet, laptop, smart TV, and connected car into a single persistent identity. This cross-device linkage means that a search on a work laptop, a location from a personal phone, and viewing behavior from a smart TV are merged into one profile, even when users deliberately use separate devices to compartmentalize activities.
Current State
LiveRamp's IdentityLink claims to resolve identities across 250+ million US adults. The Trade Desk's Unified ID 2.0 (UID2) aims to replace third-party cookies with email-based deterministic matching plus probabilistic cross-device linkage. Experian's Tapad device graph links 2+ billion devices globally. These identity graphs are the plumbing of the data broker economy — they enable the merger of siloed datasets into comprehensive profiles. No regulation restricts identity graph construction or cross-device linking.
Impact
A user who carefully uses separate devices for work and personal life, separate browsers for different activities, and separate email addresses for different services discovers that identity resolution technologies have linked all of these into a single profile. The compartmentalization strategy recommended by privacy communities (r/privacy, PrivacyGuides) is defeated by probabilistic matching using IP addresses, Wi-Fi networks, Bluetooth proximity, and timing patterns. Cross-device graphs mean there is no effective separation between digital identities.
References
LiveRamp IdentityLink documentation; The Trade Desk UID2 whitepaper; Tapad/Experian device graph specifications; Drawbridge/LinkedIn cross-device research; r/privacy threads on identity graph defeat of compartmentalization strategies; EFF "Behind the One-Way Mirror" (2019).
2. Broker Aggregation & ProfilingCritical
1Identity Resolution Across Fragmented Data▾
Problem
Data brokers use identity resolution — the process of linking records from different sources to the same individual — to merge fragments of data collected across thousands of touchpoints. A voter registration record, a loyalty card transaction, a mobile ad ID, a cookie, an email address, and a physical address are stitched together into a single identity using deterministic matching (exact field matches) and probabilistic matching (statistical inference). This is the foundational technology that makes the broker economy function.
Current State
LiveRamp's RampID is the industry-standard identity resolution platform, linking offline PII (name, address, phone) to online identifiers (cookies, mobile ad IDs, connected TV IDs) for 250+ million US consumers. Experian's identity graph, TransUnion's TrueVision, and Epsilon's CORE ID provide competing resolution services. The technology is so mature that a single email address can unlock an entire profile. NIST and academic research has documented that "anonymized" datasets can be re-identified through identity resolution with 85-99% accuracy.
Impact
A person who provides their email address to a new online retailer discovers that the retailer immediately enriches their profile through LiveRamp or Epsilon, attaching income estimates, home value, political party, marital status, number of children, magazine subscriptions, and predicted interests — all before the first purchase. The email address serves as a universal join key that unlocks years of accumulated data across the broker ecosystem. This enrichment happens in milliseconds, invisibly, at the point of data collection.
References
LiveRamp RampID technical documentation; FTC "Data Brokers: A Call for Transparency" (2014); Sweeney, "Simple Demographics Often Identify People Uniquely" (Carnegie Mellon, 2000); Narayanan & Shmatikov, "Robust De-anonymization of Large Datasets" (2008); Senate Commerce Committee data broker hearing (March 2023).
2Probabilistic Matching Without Consent▾
Problem
When deterministic matching fails (no shared unique identifier), brokers use probabilistic algorithms that infer identity links based on statistical patterns — shared IP addresses, similar device configurations, overlapping location patterns, timing correlations, and behavioral similarities. These algorithms operate on a confidence threshold (typically 70-90%) and inevitably produce both false positives (incorrectly linking different people) and true positives (correctly linking people who deliberately maintained separate identities).
Current State
The Trade Desk, LiveRamp, and Experian all offer probabilistic matching as a core service. Industry accuracy claims range from 85-97%, but independent verification is impossible because the algorithms are proprietary and the ground truth datasets are not shared. The IAB Tech Lab's Addressability working group develops standards for probabilistic ID solutions as the industry prepares for cookie deprecation. No regulatory framework governs the accuracy requirements or error rates of probabilistic matching.
Impact
Probabilistic matching defeats privacy-protective behavior. A user who never provides their real name to a service can be identified through device fingerprinting, IP correlation, and behavioral pattern matching. False positives mean individuals may receive another person's profile attributes — a stranger's medical interests, financial data, or political affiliation attached to their identity. There is no mechanism to discover or correct probabilistic matching errors because individuals do not know they have been matched and brokers do not provide transparency into match logic.
References
IAB Tech Lab Addressability specifications; LiveRamp probabilistic matching patents (US Patent 10,536,468); The Trade Desk cross-device whitepaper; FPF "Understanding Probabilistic Data Linkage" (2022); academic analysis of probabilistic record linkage error rates (Winkler, 2014).
3Data Enrichment From Public Records▾
Problem
Brokers use public records as a foundational layer to enrich commercial data profiles. Property records reveal home value, mortgage amount, and purchase date. Voter records reveal party affiliation, voting frequency, and registration address. Court records reveal lawsuits, divorces, bankruptcies, and criminal history. Vehicle registrations reveal car make, model, and year. These records, collected by governments for specific civic purposes, become the scaffolding on which commercial surveillance profiles are built.
Current State
LexisNexis Risk Solutions aggregates public records from all 3,141 US counties and 50 states into searchable databases marketed to insurance companies, financial institutions, law enforcement, and other data brokers. Thomson Reuters CLEAR provides similar aggregation for investigations and due diligence. Palantir's Gotham platform integrates public records for government intelligence analysis. The cost of bulk public records access varies by jurisdiction — some counties provide free bulk downloads, others charge fees — but no jurisdiction restricts commercial use of bulk records.
Impact
A divorce filing creates a cascade of data broker activity: the record is ingested by LexisNexis within days, triggering updates to Acxiom (marital status change), credit bureaus (address changes), and people-search sites (household composition update). The divorcing parties begin receiving targeted advertising for divorce attorneys, dating apps, apartment rentals, and therapy services — all before they have disclosed the divorce to friends or family. The public record system designed for legal transparency becomes an involuntary broadcast mechanism for life events.
References
LexisNexis public records database documentation; Thomson Reuters CLEAR product specifications; Palantir government contracts (FOIA releases); Duke University data broker military personnel study (2023); National Conference of State Legislatures public records access survey; r/privacy divorce record data broker threads.
4Consumer Scoring Beyond Credit Scores▾
Problem
Data brokers create proprietary consumer scores that go far beyond traditional credit scoring. These include health risk scores (calculated from purchase data, not medical records), fraud risk scores, insurance risk scores, marketing responsiveness scores, "consumer vulnerability" scores, and "consumer stability" scores. Unlike credit scores (regulated by the FCRA), these alternative scores operate in a regulatory vacuum with no accuracy requirements, no dispute rights, and no disclosure obligations.
Current State
LexisNexis Attract (insurance scoring), Sift Science (fraud scoring), and TransUnion's specialized scoring products assign numerical values that determine the prices people see, the offers they receive, and the services available to them. The World Privacy Forum's "The Scoring of America" report identified hundreds of consumer scores. FICO's Ultra FICO and Experian Boost blur the line between credit scoring and alternative data scoring. The CFPB under Director Chopra attempted to extend FCRA-like protections to data brokers, but the regulatory authority remains contested.
Impact
An individual may be denied an insurance quote, shown higher prices for online goods, or excluded from a financial product based on a score they never knew existed, calculated from data they never consented to share, using an algorithm they cannot inspect or challenge. Unlike credit scores — where the FCRA guarantees access, accuracy requirements, and dispute rights — these alternative scores offer no consumer protections. You cannot request your health risk score, challenge its accuracy, or know which decisions it influenced.
References
World Privacy Forum, "The Scoring of America" (2014, updated 2023); CFPB data broker rulemaking proceedings (2023-2024); FTC "Big Data: A Tool for Inclusion or Exclusion?" (2016); Senate Commerce Committee testimony on alternative scoring; Upturn, "Led Astray" (online scoring and decision-making study).
5Household-Level Data Aggregation▾
Problem
Brokers aggregate data at the household level, linking all residents of a physical address into a unified household profile. This merges the data of spouses, parents, children, roommates, and anyone who has ever been associated with the address. Household data includes combined income estimates, total number of residents, presence of children (with age ranges), pet ownership, vehicle count, political affiliations of all voters, and purchase patterns of all household members using shared loyalty cards or payment methods.
Current State
Acxiom's PersonicX clusters 250+ million US adults into 70 lifestyle segments based on household-level attributes. Experian Mosaic classifies every US household into 71 segments and 19 groups. Epsilon's household graph links individuals to addresses and models household-level purchasing power. These household profiles are sold to marketers, real estate companies, and political campaigns. No regulation prevents the inference of one household member's attributes from another's data.
Impact
A college student living at home discovers that their parent's financial data, political donations, and purchase history are attributed to them through household-level aggregation. A roommate's online gambling habits affect the household risk score visible to insurers. More critically, domestic violence survivors who flee an abuser discover that household-level profiles can reveal their new address through the linkage of co-residents — a child's school enrollment, a shared Amazon account, or a forwarded mail record is enough to update the household graph and expose the survivor's location.
References
Acxiom PersonicX methodology; Experian Mosaic segmentation documentation; National Network to End Domestic Violence, "Technology Safety" reports; FTC data broker study household profiling findings; r/privacy threads on household-level data leakage.
6Data Broker-to-Broker Resale Chains▾
Problem
Data brokers sell to each other in layered resale chains that make it impossible to trace the origin or control the flow of personal data. A piece of data collected by an app SDK may pass through 5-10 brokers before reaching its final buyer. Each broker adds, modifies, and recombines data before reselling, creating a supply chain with no transparency, no audit trail, and no mechanism for an individual to determine which brokers hold their data or how many copies exist.
Current State
The FTC's 2014 data broker study documented that the nine studied brokers collectively obtained data from thousands of sources and that many of these sources were other data brokers. Vermont's data broker registry (the only US state that requires registration) lists 500+ registered brokers, but registration does not require disclosure of data sources or resale partners. California's Delete Act (SB 362, signed 2023) creates a single opt-out mechanism but does not address broker-to-broker resale chains. The DPPA, FCRA, and state privacy laws do not restrict broker-to-broker sales.
Impact
When a consumer exercises a deletion right under CCPA or GDPR against one broker, copies of that data persist across dozens of other brokers in the resale chain. The data reappears within weeks as other brokers in the chain resell their copies. The consumer faces an infinite regression: deleting data from Broker A is meaningless if Brokers B through Z still hold copies obtained through resale. This is the structural reason why opt-out is ineffective — the supply chain architecture makes complete deletion technically impossible.
References
FTC "Data Brokers: A Call for Transparency" (2014); Vermont data broker registry (Secretary of State); California Delete Act (SB 362, 2023); The Markup, "The Secret Surveillance Ecosystem" investigation series; Privacy Rights Clearinghouse data broker database.
7Political Microtargeting Infrastructure▾
Problem
Data brokers provide the infrastructure for political microtargeting — creating voter profiles with hundreds of attributes (income, race, religion, media habits, issue positions, donation history, psychological traits) that enable campaigns to deliver personalized messages to individual voters. L2, TargetSmart, and i360 specialize in political data, but mainstream brokers like Acxiom and Experian also sell political segments. The combination of voter files, consumer data, and social media behavior creates persuasion profiles that campaigns use to manipulate individual voters.
Current State
L2 maintains voter files for all 50 states enriched with consumer data, modeled ethnicity, modeled religion, and issue position scores. TargetSmart (Democratic-aligned) and i360 (Koch-affiliated, Republican-aligned) offer competing political data platforms. The FEC does not regulate data broker use by campaigns. Cambridge Analytica's model — psychographic profiling from social media data merged with voter files — was not an aberration but a refinement of standard practices. Political data brokers operate entirely outside election regulation.
Impact
Voters receive hyper-personalized political messaging designed to activate their specific psychological triggers, but they cannot see the messages delivered to their neighbors with different profiles. This creates fragmented information environments where different voters in the same district receive contradictory messages from the same candidate. The privacy harm is compounded by democratic harm: political microtargeting using broker data undermines shared civic discourse by replacing public persuasion with private manipulation.
References
Cambridge Analytica whistleblower testimony (UK Parliament, 2018); L2 political data product documentation; TargetSmart and i360 platform descriptions; Tactical Tech, "Personal Data: Political Persuasion" (2019); ProPublica "Facebook Political Ad Collector" project; FEC advisory opinions on data broker use.
8Tenant and Employment Screening Data Aggregation▾
Problem
Background screening companies — CoreLogic, RealPage, TransUnion SmartMove, Sterling, HireRight — aggregate data from brokers, public records, credit bureaus, and proprietary databases to create screening reports used by landlords and employers. These reports combine criminal records, eviction history, credit data, employment verification, and social media analysis into recommendations that determine whether individuals can rent apartments or get jobs. Errors in broker data cascade into screening reports with life-altering consequences.
Current State
The FCRA theoretically regulates tenant and employment screening, requiring accuracy and dispute rights. In practice, the FTC and CFPB have documented persistent accuracy problems: the National Consumer Law Center found that one in four tenant screening reports contains errors. RealPage's algorithmic pricing was investigated by ProPublica (2022) for potentially facilitating landlord collusion on rent prices. Sterling and HireRight have paid millions in FCRA settlements for reporting inaccurate criminal records. Automated scoring increasingly replaces human review.
Impact
A person with a common name discovers they are being rejected for apartments because a criminal record belonging to a different person with the same name appears in their screening report. By the time they dispute and correct the error with one screening company, three other landlords have already rejected them using reports from different screening companies containing the same error sourced from the same broker data. The FTC documented cases where consumers spent months correcting screening errors that originated from a single data broker's incorrect record propagated through the resale chain.
9Financial Data Aggregation Beyond Credit Bureaus▾
Problem
Beyond the three major credit bureaus (Equifax, Experian, TransUnion), a secondary market of financial data brokers aggregates bank account data, payment histories, and alternative financial data. Companies like Plaid (acquired by Visa, deal later unwound) collect bank transaction data through fintech app connections. Yodlee sells "anonymized" bank transaction data. ChexSystems maintains a banking blacklist. The "alternative data" market uses utility payments, rent payments, and telecom data to create parallel financial profiles outside traditional credit bureau oversight.
Current State
Plaid connects to 12,000+ financial institutions and powers the bank connections for Venmo, Robinhood, Coinbase, and thousands of fintech apps. When a user links their bank account through Plaid, Plaid retains transaction data. Yodlee (Envestnet) was sued by consumers alleging it sold detailed bank transaction data to hedge funds and other buyers. The CFPB's open banking rule (Section 1033) aims to give consumers control over financial data sharing but has faced industry opposition. Fintech data collection operates in a regulatory gap between banking regulation and data protection.
Impact
A user who linked their bank account to a budgeting app via Plaid discovers that their complete transaction history — every purchase, every payment, every transfer — is accessible to Plaid and potentially shared with its partners. Yodlee's sale of "anonymized" transaction data to hedge funds and investment companies means that consumer spending patterns are being used for financial trading, creating a pipeline where ordinary people's financial behavior is monetized by Wall Street without their knowledge or compensation.
References
Plaid consumer data practices lawsuit (Cottle v. Plaid, 2020); Yodlee data sale reporting (Motherboard, 2020); CFPB Section 1033 rulemaking; "Plaid Settles Privacy Lawsuit for $58M" (2022); Senate Banking Committee fintech data hearing (2023); r/personalfinance threads on Plaid data retention.
10Real-Time Data Enrichment at Point of Collection▾
Problem
Modern data enrichment happens in real-time: the moment a user enters an email address, phone number, or physical address on a website, data enrichment APIs from Clearbit (now Breeze by HubSpot), ZoomInfo, FullContact, Pipl, and others instantly return a comprehensive profile containing name, employer, title, social media profiles, estimated income, location, and behavioral attributes. This turns every form fill into a complete dossier before the user even clicks "submit."
Current State
Clearbit's API returns 100+ attributes from an email address in under 200 milliseconds. ZoomInfo maintains 600+ million professional profiles and offers real-time enrichment through its API. FullContact's Identity Resolution API links email, phone, social profiles, and device IDs into unified profiles. These APIs are embedded in thousands of websites through marketing automation platforms (HubSpot, Salesforce, Marketo). Users have no indication that enrichment is occurring at the point of data collection.
Impact
A job applicant who enters only their email address on a company's career page triggers a Clearbit/ZoomInfo enrichment that provides the employer with the applicant's current employer, estimated salary, social media profiles, home location, and professional history — before the applicant has voluntarily shared any of this information. The applicant has no knowledge that enrichment occurred, no ability to see what data was returned, and no mechanism to correct inaccuracies. The hiring decision may be influenced by enriched data the applicant never consented to share.
References
Clearbit (now Breeze) API documentation; ZoomInfo platform documentation; FullContact Identity Resolution API specs; The Markup investigation of real-time data enrichment; HubSpot/Clearbit acquisition (2023); r/privacy threads on real-time enrichment experiences.
3. People-Search Site ProliferationCritical
1Opt-Out Whack-a-Mole Across Hundreds of Sites▾
Problem
There are an estimated 200-400 people-search sites operating in the US, each independently scraping, purchasing, and publishing personal information including home addresses, phone numbers, email addresses, relatives, neighbors, age, and estimated income. Opting out of one site has no effect on the others. New sites appear constantly. Sites that honor opt-outs re-acquire the data within 3-12 months from broker resale chains and re-list it. The process of opting out requires submitting additional PII (government ID, email, physical address) to the very companies you want to stop sharing your data.
Current State
Major people-search sites include Spokeo, BeenVerified, WhitePages, Radaris, TruePeopleSearch, FastPeopleSearch, ThatsThem, USSearch, Intelius, PeopleFinder, and hundreds of smaller operators. Paid opt-out services (DeleteMe, Kanary, Privacy Duck, Optery) charge $100-400/year to automate the whack-a-mole process but cannot guarantee complete removal. California's Delete Act (SB 362) creates a centralized opt-out for data brokers, but implementation details remain contested. No federal law addresses people-search sites specifically.
Impact
A domestic violence survivor spends 40+ hours manually opting out of 100+ people-search sites, submitting government ID and current address to each one, only to discover their information reappears on 60% of those sites within six months. Meanwhile, three new sites launched during that period, already listing their information. The survivor must treat data removal as an ongoing, never-ending maintenance task requiring either significant personal time investment or $300+/year for a removal service — effectively a privacy tax imposed on vulnerable populations.
References
Consumer Reports study on people-search opt-out effectiveness (2023); Privacy Rights Clearinghouse data broker opt-out guide; r/privacy megathread on people-search removal; DeleteMe annual transparency report; California Delete Act (SB 362, 2023); National Network to End Domestic Violence technology safety resources.
2Data Reappearance After Successful Opt-Out▾
Problem
Even when a people-search site honors an opt-out request and removes a listing, the data reappears within weeks to months because the site's upstream data suppliers (brokers, public records aggregators, other people-search sites) continue to feed the same data back into the system. The opt-out removes a single copy but does not address the supply chain. Many sites explicitly state in their privacy policies that they cannot guarantee data will not reappear after an opt-out.
Current State
DeleteMe's internal data shows that 35-40% of successfully removed listings reappear within 6 months. Spokeo's FAQ acknowledges that opt-outs may need to be repeated. BeenVerified's opt-out confirmation states that data may reappear if it is "collected again from public sources." TruePeopleSearch and FastPeopleSearch — which provide free access to records — have particularly high reappearance rates because they aggressively re-scrape public records and broker feeds. The underlying problem is architectural: opt-out is applied at the endpoint, not at the source.
Impact
Users who invest significant time and money in data removal develop a recurring pattern documented across privacy forums: initial relief when listings disappear, followed by frustration when they reappear 3-6 months later, leading eventually to resignation and acceptance that complete removal is impossible within the current system. Privacy communities (r/privacy, PrivacyGuides forums) refer to this as the "opt-out treadmill" — a process designed to exhaust individuals into accepting surveillance as the default.
References
DeleteMe reappearance rate data; Spokeo opt-out FAQ; BeenVerified privacy policy; r/privacy threads documenting reappearance timelines; Consumer Reports, "It's Unreasonably Difficult to Opt Out of Data Broker Sites" (2023); The Markup, "Still Creepy" follow-up investigations.
3Verification Requirements That Demand More PII▾
Problem
People-search sites require individuals to submit additional personal information — government-issued photo ID, current physical address, current email address, date of birth, or phone number — in order to process opt-out requests. This creates a perverse incentive structure where the act of protecting your privacy requires surrendering more data to the very companies profiting from your data. Some sites use this verification data to update and enrich their existing records.
Current State
Radaris requires a selfie photo holding government ID for opt-out verification. Spokeo requires email verification and asks for additional identifying information to locate the correct record. BeenVerified requires an email address and links the opt-out request to that email for tracking. IntelliCheck and other identity verification services used by some people-search sites retain verification data. No regulation prohibits people-search sites from using verification data to update their records, and privacy policies often explicitly permit this.
Impact
An individual attempting to remove their home address from Radaris must photograph themselves holding their driver's license — which contains their home address, full legal name, date of birth, and photo — and upload it to Radaris's servers. They are now providing a verified, current copy of exactly the data they wanted removed, plus biometric data (facial photograph) they never previously shared. Privacy advocates on r/privacy and PrivacyGuides have documented cases where opt-out verification data appears to have been used to refresh stale records.
References
Radaris opt-out requirements documentation; r/privacy threads on opt-out verification paradox; PrivacyGuides forum discussions on ID verification risks; Vice Motherboard, "The Dark Side of Opting Out of Data Broker Sites" (2022); EFF, "How to Remove Yourself from People-Search Sites."
4Free People-Search Sites Monetizing Curiosity▾
Problem
Sites like TruePeopleSearch, FastPeopleSearch, and ThatsThem provide personal information entirely for free, monetized through advertising rather than subscriptions. This eliminates any friction for casual lookups, enabling anyone — ex-partners, stalkers, scammers, doxxers — to access home addresses, phone numbers, and relative lists with zero cost or accountability. Free sites have no financial incentive to honor opt-outs quickly because their revenue comes from advertising impressions, and every page view generates income.
Current State
TruePeopleSearch and FastPeopleSearch consistently rank in the top 10,000 US websites by traffic (per SimilarWeb), generating millions of lookups per month. These sites display Google AdSense and programmatic advertising alongside personal records. Their opt-out processes are deliberately cumbersome — requiring email verification, CAPTCHA solving, and multi-step confirmation — to reduce opt-out completion rates. New free people-search sites appear regularly, often operated by the same entities under different domain names.
Impact
A stalking victim discovers that their ex-partner has been monitoring their address through TruePeopleSearch, which updated their record after they moved to a new location. The ex accessed the information for free, with no account creation, no identity verification, and no audit trail. Law enforcement cannot subpoena access logs because free sites often do not maintain them. The victim's safety was compromised by a site that profits from advertising while externalizing the costs of harm to the individuals whose data it publishes.
References
SimilarWeb traffic data for people-search sites; National Domestic Violence Hotline technology abuse reports; r/stalking and r/legaladvice threads on people-search site misuse; The Markup, "How to Find and Remove Your Personal Information From People-Search Sites"; anti-doxxing resources from EFF and PEN America.
5People-Search Sites Selling to Scammers▾
Problem
People-search data is actively exploited by fraud rings, romance scammers, and social engineering attackers who use the freely or cheaply available personal details to impersonate individuals, craft convincing phishing attacks, and conduct identity theft. The combination of a person's name, age, address, phone number, relatives, and employment history provides everything needed for sophisticated social engineering or synthetic identity fraud.
Current State
The FBI's IC3 reported $10.3 billion in cybercrime losses in 2022, with phishing, personal data breach, and identity theft among the top crime types. Research by Agari (now part of HelpSystems) found that 76% of business email compromise attacks use personal details obtained from public data sources including people-search sites. The AARP documented that elder fraud schemes routinely use people-search data to identify and target vulnerable seniors. No people-search site conducts "know your customer" verification on bulk purchasers, and free sites require no verification at all.
Impact
A senior citizen receives a phone call from someone claiming to be their grandchild, referencing the grandchild's actual name, city of residence, and college — all information available from people-search sites through the relatives and associates section of the senior's listing. The "grandparent scam" costs US seniors an estimated $1 billion annually, and people-search sites provide the raw data that makes these scams convincing. The victims lose life savings while the sites face no liability for enabling the fraud.
References
FBI IC3 Annual Report (2022); AARP Fraud Watch Network elder fraud statistics; Agari/HelpSystems business email compromise research; FTC consumer fraud reports; r/Scams documentation of people-search-enabled fraud; KrebsOnSecurity reporting on people-search data in fraud pipelines.
6Radaris and Foreign-Operated People-Search Sites▾
Problem
Several major people-search sites are operated by entities with opaque corporate structures, offshore registration, or foreign ownership, making regulatory enforcement and legal action extremely difficult. Radaris, one of the largest people-search sites, was investigated by The Markup (2023) and found to have complex ownership connections and a history of making opt-out difficult. Sites operated outside US jurisdiction are not subject to state data broker registration laws, FTC enforcement, or state privacy statutes.
Current State
The Markup's investigation of Radaris revealed connections to a network of people-search and background check sites operated under various corporate entities. Many people-search sites are registered through privacy-protecting domain registrars and hosted on infrastructure that obscures ownership. Vermont's data broker registry and California's Delete Act apply only to entities with a nexus to those states. Offshore operators can clone US public records data and host it on servers in jurisdictions with no data protection enforcement.
Impact
When consumers file complaints about Radaris with the FTC or state attorneys general, enforcement is hampered by corporate opacity. A deletion request sent to an offshore operator may be ignored entirely with no practical legal recourse. Even if one corporate entity is shut down, the same operators can launch new sites under different names within days. The consumer faces a hydra: removing their data from one site triggers no obligation on the ten other sites operated by related entities.
References
The Markup, "This Obscure People-Search Site Has the Most Coverage of Any We've Tested" (Radaris investigation, 2023); Vermont data broker registry foreign operator gaps; GoDaddy/Namecheap privacy registration analysis; r/privacy threads on Radaris opt-out difficulties; FTC jurisdiction limitations for foreign operators.
7Criminal Records Displayed Without Context or Updates▾
Problem
People-search sites display criminal records — arrests, charges, convictions — without context, often without distinguishing between arrests and convictions, without reflecting expungements or dismissals, and without any mechanism for individuals to add context or corrections. A decades-old arrest that was dismissed still appears on these sites, permanently branding individuals with criminal histories that the legal system has determined should not follow them.
Current State
Most people-search sites scrape criminal records from county courts, state repositories, and federal databases (PACER). They display these records alongside current name, address, and photo without indicating whether charges resulted in conviction, were dismissed, or were expunged. Expungement orders, which legally seal records from public access, are frequently not reflected on people-search sites because the sites scraped the data before expungement and have no mechanism to receive or process expungement notifications. The FCRA requires background check companies to maintain accuracy, but people-search sites argue they are not CRAs (Consumer Reporting Agencies).
Impact
A person who was arrested in their twenties for a minor offense that was later dismissed — and eventually expunged — discovers that Spokeo, BeenVerified, and a dozen other sites still display "Criminal Record: 1 offense" on their profile. Prospective employers, landlords, and romantic partners who search their name see this flag. The individual has no practical way to force removal because the sites claim they are not CRAs and therefore not subject to FCRA accuracy requirements. The legal right to expungement is meaningless if commercial databases ignore it.
References
National Employment Law Project, "Ban the Box" research on criminal record employment barriers; SEARCH/National Consortium for Justice Information and Statistics, expungement notification gaps; Legal Action Center, "After Prison: Roadblocks to Reentry"; r/legaladvice threads on expunged records appearing on people-search sites; EFF advocacy on criminal record data broker practices.
8Relative and Associate Networks Exposing Third Parties▾
Problem
People-search sites display "known relatives" and "known associates" sections that expose network connections without any consent from the listed individuals. These sections reveal family relationships (parents, children, siblings, spouses, ex-spouses), roommates, and business associates. This network data enables mapping of an individual's entire social graph and can expose sensitive relationships — estranged family members, undisclosed relationships, or connections individuals have deliberately severed.
Current State
Spokeo, BeenVerified, and WhitePages display lists of 5-30+ relatives and associates derived from shared addresses, shared phone numbers, co-signatures on documents, and public records (marriage, divorce, property). These association lists persist even after relationships end — ex-spouses remain listed for years after divorce, deceased relatives remain listed indefinitely. Opting out of your own listing does not remove you from other people's "relatives" sections. There is no mechanism for an individual to control how they appear in others' profiles.
Impact
An adult child who was estranged from an abusive parent discovers that every people-search site lists them as the parent's "known relative," with their current city and age range visible on the parent's profile. A person in witness protection discovers that their new identity is linked back to family members' unchanged profiles through the "associates" network. Doxxing campaigns use relative lists to expand targeting from a single individual to their entire family — a tactic documented in numerous online harassment cases reported by PEN America and the Anti-Defamation League.
References
PEN America, "Online Harassment Field Manual"; Anti-Defamation League doxxing research; National Network to End Domestic Violence safety planning guides; r/privacy threads on relatives sections exposing estranged family; Spokeo/BeenVerified relatives data persistence documentation.
The people-search industry has consolidated through acquisitions, with a few holding companies controlling dozens of seemingly independent sites. The H.I.G. Capital portfolio includes PeopleConnect (which operates Intelius, USSearch, Classmates.com, and others). System1 operates PeopleSearch, MapQuest, and InfoTracer. This consolidation means that opting out of one brand does not propagate to sister sites owned by the same parent, and the illusion of market competition masks monopolistic control over personal data distribution.
Current State
PeopleConnect (Intelius parent) operates at least 10 people-search brands from the same underlying database. Opt-out requests submitted to Intelius do not automatically propagate to USSearch or other PeopleConnect properties. Similarly, System1's portfolio of people-search sites shares backend infrastructure but maintains separate opt-out processes for each brand. The FTC has not scrutinized people-search industry consolidation as an antitrust concern, and state data broker registries do not require disclosure of corporate relationships between registered brokers.
Impact
A consumer who meticulously opts out of Intelius, USSearch, and Classmates.com — believing they have addressed three separate companies — discovers they were all PeopleConnect brands drawing from the same database, and their data remains on five other PeopleConnect properties they did not know existed. The consolidation creates an information asymmetry where consumers cannot determine which brands share databases, making informed opt-out decisions impossible.
References
PeopleConnect/H.I.G. Capital corporate structure; System1 people-search portfolio; FTC lack of people-search industry scrutiny; Vermont data broker registry corporate relationship analysis; The Markup investigation of people-search ownership networks; r/privacy threads mapping people-search corporate relationships.
10No Liability for Harms Enabled by People-Search Data▾
Problem
People-search sites face no legal liability when their data is used to enable stalking, harassment, doxxing, identity theft, or physical violence. Section 230 of the Communications Decency Act has been interpreted to protect platforms that publish third-party content, and people-search sites argue that public records data constitutes third-party content they merely organize and display. Victims of crimes enabled by people-search data have no civil cause of action against the sites that made targeting possible.
Current State
Multiple stalking cases have involved perpetrators who located victims through people-search sites. The National Network to End Domestic Violence reports that people-search sites are among the top technology-facilitated abuse tools. David Renz, convicted of kidnapping and murder in New York, used people-search sites to identify victims. Despite documented cases of harm, no successful lawsuit has established people-search site liability for downstream criminal use of their data. California's AB 1138 (2024) creates a civil cause of action against individuals who doxx with intent to harass, but does not impose liability on the platforms providing the data. Washington state's anti-doxxing law similarly targets individuals. The Data Broker Accountability and Transparency Act (proposed federal legislation) would create some obligations but has not passed. People-search sites continue to operate in a liability-free zone where the harms of their business model are externalized entirely to the individuals whose data they publish.
Impact
A journalist covering organized crime is doxxed — her home address, phone number, relatives, and daily commute are posted on extremist forums, sourced from Spokeo and WhitePages premium reports. She receives death threats referencing her home address. She cannot sue the people-search sites. She cannot compel them to remove her data beyond their standard opt-out process. She spends thousands of dollars on a deletion service and home security while the sites continue to profit from her data. The people-search industry has externalized 100% of the costs of the harms it enables.
References
National Network to End Domestic Violence, technology-facilitated abuse reports; Renz case documentation; California AB 1138 (anti-doxxing statute, 2024); Section 230 immunity analysis applied to people-search sites; Committee to Protect Journalists, reporter safety resources; PEN America, doxxing case studies; r/privacy and r/legaladvice threads on legal recourse against people-search sites.
4. Ad-Tech Pipeline OpacityCritical
1Real-Time Bidding Broadcasting PII to Hundreds of Companies▾
Problem
Real-time bidding (RTB) is the mechanism through which programmatic advertising works: when a user loads a webpage or app, an auction takes place in milliseconds where the user's data — location, browsing history, device type, demographics, interests, and sometimes sensitive attributes — is broadcast to hundreds of potential advertisers competing to show an ad. The Irish Council for Civil Liberties (ICCL) documented that RTB broadcasts Europeans' data 376 times per day on average and Americans' data 747 times per day, amounting to 178 trillion data broadcasts in the US and 107 trillion in Europe annually.
Current State
Google's authorized buyers program includes 4,700+ companies that receive RTB bid requests. Each bid request contains an OpenRTB protocol data package that can include GPS coordinates, browsing URL, device ID, IP address, demographic segments, and interest categories. The ICCL's 2022 report "The Biggest Data Breach" established that RTB constitutes a systematic data breach because data is broadcast to companies with no contractual relationship with the user and no technical means to verify that losing bidders delete the data. The Belgian DPA found IAB Europe's Transparency and Consent Framework (TCF) itself non-compliant with GDPR.
Impact
Every time a user loads a webpage with programmatic advertising, their personal data is sent to an average of 300-700 companies. These companies include not just advertisers but data aggregators, surveillance companies, and entities in jurisdictions with no data protection laws. The data cannot be recalled after broadcast — there is no technical mechanism to ensure losing bidders delete bid request data. A single day of web browsing results in a person's data being shared with potentially thousands of unique companies, none of which the person has ever heard of or consented to share data with.
References
ICCL, "The Biggest Data Breach" (May 2022); Belgian DPA IAB Europe TCF decision (February 2022); OpenRTB 2.6 protocol specification (IAB Tech Lab); Google authorized buyers list; Dr. Johnny Ryan (ICCL) Senate testimony (2023); r/privacy RTB awareness threads.
2Supply-Side Platform Data Leakage▾
Problem
Supply-side platforms (SSPs) — the technology that publishers use to sell ad inventory — collect and share publisher audience data with demand-side platforms, data management platforms, and ad exchanges. Major SSPs (Google Ad Manager, Xandr/Microsoft, Magnite, PubMatic, OpenX, Index Exchange) process bid requests containing user data for thousands of publishers simultaneously. SSPs have access to the complete browsing behavior across all sites they serve, creating comprehensive user profiles that rival those of the largest data brokers.
Current State
Google Ad Manager (formerly DoubleClick) operates the dominant SSP, serving ads on millions of websites and thus observing users' browsing behavior across the web. The DOJ's antitrust case against Google (2023-2024) documented Google's monopoly position in the ad-tech stack, with internal documents showing Google's awareness that its SSP/ad exchange position gave it data advantages competitors could not match. Magnite (formerly Rubicon Project) processes 6+ trillion ad requests monthly. PubMatic processes 250+ billion ad impressions daily. Each SSP maintains its own user profiles built from bid request data.
Impact
A user who installs an ad blocker on their desktop browser but browses normally on mobile discovers that SSPs have already built a comprehensive profile from their mobile browsing. The SSP data includes every website visited that uses that SSP's ad serving — which, for Google Ad Manager, encompasses the majority of the web. This data is used for retargeting, audience building, and resale. The user has no relationship with the SSP, no knowledge of its existence, and no mechanism to access or delete the profile it maintains.
References
DOJ v. Google antitrust filings (2023); Magnite/Rubicon Project investor disclosures; PubMatic S-1 filing (2020); The Markup, "Google's Secret Offer to Special-Deal Publishers" (2023); Wolfie Christl, "Corporate Surveillance in Everyday Life" (Cracked Labs, 2017); EFF, "Behind the One-Way Mirror" (2019).
3Data Management Platform Profile Depth▾
Problem
Data management platforms (DMPs) — including Oracle BlueKai (shut down 2024), Lotame, Salesforce DMP (Krux), and Adobe Audience Manager — aggregate user data from publishers, advertisers, and third-party data providers into detailed profiles containing thousands of interest segments, behavioral attributes, and inferred demographics. These profiles are the fuel of targeted advertising, and they contain information of extraordinary sensitivity derived from browsing behavior, purchase data, and location history.
Current State
Oracle BlueKai's database leak (reported by TechCrunch in June 2020) exposed billions of records containing names, email addresses, home addresses, browsing history, and purchase intent data for millions of consumers — left unsecured on an internet-facing server. Oracle subsequently shut down its advertising division (Oracle Advertising/BlueKai/Moat/Grapeshot) in June 2024, citing competitive pressures, but the data collected over a decade remains in the profiles of its former customers. Lotame's DMP claims access to 5 billion device IDs. Adobe Audience Manager integrates with Adobe's analytics and marketing cloud, creating profiles that span web analytics, email marketing, and advertising behavior.
Impact
The Oracle BlueKai exposure demonstrated that DMP data is not abstract metadata — it contained records showing specific individuals' browsing behavior on specific websites, including sensitive content categories. A record might show that John.Smith@email.com visited addiction treatment websites, bankruptcy attorney pages, and divorce lawyer sites over a three-week period. This data was unencrypted and internet-accessible. When Oracle exited the advertising business in 2024, the fate of billions of accumulated consumer records remains unclear — the data does not disappear when the business unit closes.
References
TechCrunch, "Oracle's BlueKai tracks you across the web. That data spilled online" (June 2020); Oracle advertising division shutdown (Digiday, June 2024); Lotame platform documentation; Adobe Audience Manager data handling; r/privacy Oracle BlueKai breach threads; Wolfie Christl, Cracked Labs corporate surveillance reports.
4Cookie Syncing Creating Universal Tracking IDs▾
Problem
Cookie syncing (also called cookie matching or pixel syncing) is the process by which ad-tech companies share user identifiers with each other, enabling them to link their independently collected data about the same user. When User A visits Site X, the SSP drops a cookie with ID "abc123." Simultaneously, it fires a pixel to DMP Y, which sees its own cookie "xyz789" for the same user. Both companies now know that abc123 = xyz789, and they can merge their datasets. This process happens billions of times daily and creates a de facto universal tracking ID without user consent.
Current State
A study by Acar et al. (University of Leuven) documented that cookie syncing occurs on 97% of the top 10,000 websites. The average webpage triggers sync events with 5-15 different ad-tech companies simultaneously. Google's syncing infrastructure connects its identifiers with thousands of partner companies. Even as third-party cookies face deprecation (Safari and Firefox already block them; Chrome's cookie plans remain uncertain), cookie syncing has been replaced by alternative identifier sync mechanisms including Universal IDs, email-hashed identifiers, and server-side matching.
Impact
Cookie syncing means that ad-tech companies do not operate in isolation — they form an interconnected surveillance network where every company's data is accessible to every other company through chains of synced identifiers. A user's visit to a health website can be linked through cookie sync chains to their real identity (via email login on another site in the sync network), their physical location (via a location SDK partner), and their purchasing behavior (via a retail data partner). The chain of syncs makes every participant in the ad-tech ecosystem a potential data source for every other participant.
References
Acar et al., "The Web Never Forgets" (ACM CCS, 2014); Papadopoulos et al., "Cookie Synchronization: Everything You Always Wanted to Know But Were Afraid to Ask" (2019); The Markup, "What They Know" investigation series; EFF, cookie syncing analysis in "Behind the One-Way Mirror"; r/privacy and r/degoogle threads on cookie sync tracking.
5Bid Stream Data Harvesting by Non-Advertising Entities▾
Problem
The RTB bid stream — the flow of data in real-time advertising auctions — is accessible to any company that registers as a bidder, including companies whose actual purpose is data collection rather than ad buying. Intelligence agencies, surveillance companies, and data brokers register as demand-side platform participants to passively harvest the bid stream without ever purchasing ads. This turns the advertising ecosystem into a global surveillance infrastructure available to any entity willing to pay the modest cost of participating as a "buyer."
Current State
The Wall Street Journal reported (2023) that Rayzone Group, an Israeli surveillance company, and other intelligence contractors obtained detailed user data through the RTB bid stream. Patternz, a surveillance platform, openly advertised its ability to target mobile devices using bid stream data from ad exchanges. The ICCL's Johnny Ryan documented bid stream exploitation in Senate testimony. RTB participants are not vetted for their actual intent — any company that meets the technical requirements can receive bid requests containing user data, with no obligation to actually bid on ads.
Impact
A government intelligence agency that would normally require a court order to surveil a citizen can instead register as a DSP participant and passively receive that citizen's location, browsing behavior, and device identifiers hundreds of times per day through the bid stream — legally, without a warrant, and at scale. The ad-tech infrastructure has inadvertently created the most comprehensive mass surveillance system ever built, available to any government or private entity for the cost of a DSP license.
References
WSJ, "Intelligence Agencies Tap Ad-Tech" (2023); Patternz surveillance platform advertising materials; ICCL Senate testimony on bid stream surveillance; Cox Media Group "Active Listening" controversy (2024); Sen. Ron Wyden letters to FTC on bid stream surveillance; ISA/Rayzone Group reporting.
6Advertising ID Persistence and Cross-App Tracking▾
Problem
Mobile advertising identifiers — Google's GAID (Google Advertising ID) and Apple's IDFA (Identifier for Advertisers) — are device-level persistent identifiers that enable tracking across all apps on a device. Every app with advertising SDK access can read the same advertising ID, creating a cross-app behavioral profile. While both platforms offer ID reset and opt-out options, the practical effect is limited because apps also collect device fingerprinting signals (IP address, screen resolution, installed apps, battery level) that enable re-identification even after an ID reset.
Current State
Apple's ATT framework requires apps to request permission before accessing the IDFA, reducing opt-in rates to approximately 25%. Google announced GAID deprecation for Android in 2024, replacing it with the Privacy Sandbox Topics API. However, the transition is slow: as of 2025, GAIDs remain active on most Android devices. Both platforms still allow apps to collect fingerprinting signals. The FTC's Kochava complaint specifically addressed the company's use of mobile advertising IDs to build location profiles tied to sensitive locations.
Impact
A user who diligently resets their Google Advertising ID monthly discovers that their ad profile regenerates within days because SDK partners use device fingerprinting to re-link the new ID to the old profile. The mobile advertising ID system was designed to give users the illusion of control while maintaining the underlying tracking infrastructure. Apple's ATT was the most effective intervention, but it protected only iOS users and was motivated partly by Apple's desire to control its own advertising ecosystem rather than pure privacy concern.
References
FTC v. Kochava complaint (GAID/IDFA tracking); Apple ATT framework documentation; Google Privacy Sandbox for Android specifications; Lockdown Privacy study on post-ATT fingerprinting; AppsFlyer opt-in rate data; r/degoogle threads on GAID alternatives.
7Connected TV Advertising Data Collection▾
Problem
Connected TV (CTV) and streaming platforms (Roku, Amazon Fire TV, Samsung TV Plus, LG Channels, Hulu, Peacock) collect second-by-second viewing data through ACR (automatic content recognition) and streaming telemetry, then sell this data through the programmatic advertising pipeline. CTV advertising combines the targeting precision of digital advertising with the persuasive power of television, using household-level data including viewing habits, content preferences, income proxies (inferred from TV model and subscription tier), and increasingly, real-time emotional engagement signals.
Current State
Roku collects viewing data from 80+ million active accounts and sells it through its advertising platform. Samsung Ads leverages ACR data from 50+ million Samsung smart TVs. Vizio's Inscape (now VIZIO Ads) was the subject of the $2.2 million FTC settlement for ACR collection without consent but continues to operate with updated "consent" flows. CTV advertising spend exceeds $30 billion annually, and the data pipeline supporting it is less regulated than traditional web advertising because most CTV privacy disclosures are buried in device setup flows that users click through without reading.
Impact
A family discovers that their Samsung smart TV has been recording every show they watch, every input switch, and every external device connected — data that Samsung sells to advertisers who target the household across their other devices. The viewing data reveals sensitive preferences (political documentaries, religious programming, addiction-related content, children's viewing patterns) that would never be knowingly shared. Turning off ACR in deeply nested TV settings menus is possible but resets with firmware updates and is not discoverable by average consumers.
References
FTC v. Vizio ($2.2M settlement, 2017); Roku advertising platform documentation; Samsung Ads ACR data collection; CTV advertising spend projections (eMarketer/Insider Intelligence); r/privacy smart TV data collection threads; Mozilla "Privacy Not Included" smart TV reviews.
8Retail Media Networks as New Data Silos▾
Problem
Retail media networks — Amazon Ads, Walmart Connect, Target Roundel, Kroger Precision Marketing, Instacart Ads, Albertsons Media Collective — represent a new advertising channel where retailers sell advertising on their properties using first-party purchase data. These networks create closed-loop attribution (connecting ad exposure to purchase) and possess the most commercially valuable data in the advertising ecosystem: what people actually buy. Retail media is a $45+ billion market growing 25%+ annually and operates with even less transparency than traditional programmatic advertising.
Current State
Amazon Ads is the third-largest digital advertising platform (after Google and Meta), generating $46+ billion in advertising revenue annually. Amazon's advertising uses purchase history, browsing behavior, Alexa interactions, Ring footage patterns, and Whole Foods loyalty data. Walmart Connect leverages transaction data from 240+ million weekly customers. These retail media networks operate as walled gardens with no external auditing of data practices. Advertisers who buy retail media ads receive aggregate reporting but the retailers retain and enrich their individual-level data indefinitely.
Impact
A consumer who purchases a pregnancy test at Walmart discovers that their purchase is ingested by Walmart Connect's advertising platform and used to serve baby product ads across Walmart's properties and partner networks. The purchase data is combined with the consumer's Walmart+ membership data, physical store visit patterns (tracked via Walmart app location), and online browsing to create a comprehensive profile. The consumer has no visibility into this data usage and no opt-out mechanism beyond abandoning the retailer entirely.
References
Amazon Ads revenue reports (annual filings); Walmart Connect partner documentation; Kroger Precision Marketing data capabilities; eMarketer retail media forecasts; The Markup, "Amazon Puts Its Own 'Brands' First" investigation; Congressional testimony on Amazon's data practices.
9Header Bidding and Server-Side Tracking Evasion▾
Problem
As client-side tracking faces restrictions from ad blockers and browser privacy features, the ad-tech industry has migrated to server-side architectures that are invisible to users and their privacy tools. Server-side header bidding moves the auction process from the user's browser to the publisher's server, making it invisible to ad blockers. Server-side tag management (server-side Google Tag Manager, Tealium iQ Server-Side) routes tracking through the publisher's first-party domain, defeating third-party cookie blocks. CNAME cloaking disguises trackers as first-party resources.
Current State
Prebid Server (the open-source server-side header bidding solution) is deployed on thousands of publisher sites. Google's server-side tag management has seen rapid adoption as a method to maintain tracking capability despite browser restrictions. A 2023 study found that CNAME cloaking — where a tracker is given a subdomain of the publisher's domain (e.g., track.publisher.com resolving to tracker.thirdparty.com) — is used by 10%+ of top websites to evade Safari's ITP and Firefox's ETP. These server-side techniques are architecturally invisible to the browser and therefore to any client-side privacy tool.
Impact
A privacy-conscious user who installs uBlock Origin, uses Firefox with Enhanced Tracking Protection, and enables Global Privacy Control discovers through network analysis that their browsing data is still being collected via server-side tracking that their privacy tools cannot detect or block. The arms race between browser privacy features and ad-tech evasion techniques has moved decisively to the server side, where users have no visibility or control. Every client-side privacy improvement accelerates the migration to server-side tracking that is technically undetectable.
References
Prebid Server documentation and adoption statistics; Google server-side tag management documentation; Dimova et al., "The CNAME of the Game" (2021 Privacy Enhancing Technologies Symposium); Le Pochat et al., server-side tracking measurement studies; uBlock Origin GitHub issues discussing server-side evasion; r/privacy discussions on the futility of client-side ad blocking.
10Consent Management Platforms as Data Brokers▾
Problem
Consent management platforms (CMPs) — OneTrust, Cookiebot, TrustArc, Didomi, Quantcast Choice — deployed to collect GDPR/CCPA consent are themselves collecting data about users' consent choices, browsing behavior, and device characteristics. The CMP sits on every page load and observes user interactions before any other tracking begins. Some CMPs share consent signals with the ad-tech supply chain through IAB's TCF (Transparency and Consent Framework), creating a system where the tool designed to protect privacy becomes another data collection vector.
Current State
OneTrust is deployed on millions of websites and observes consent interactions for hundreds of millions of users. Quantcast's CMP (Quantcast Choice) is offered free — funded by Quantcast's data business, which uses CMP deployment as a vector for its own tracking pixels. The Belgian DPA's TCF decision found that the consent signal itself constitutes personal data and that IAB Europe's management of TCF is non-compliant with GDPR. CMPs also collect data needed for consent management (IP address, device type, browser, consent history) that constitutes a profile in itself.
Impact
The consent management popup that appears on every EU website — ostensibly a GDPR protection — is itself a data collection mechanism. A user who clicks "Reject All" has still provided their IP address, device fingerprint, geographic location (inferred from IP), and consent preference to the CMP. If the CMP is Quantcast Choice, this interaction data feeds into Quantcast's advertising business. The privacy tool has become a privacy threat — the regulatory requirement designed to protect users has been co-opted as an additional data collection touchpoint.
References
Belgian DPA IAB Europe/TCF decision (2022); Quantcast Choice/Quantcast advertising business relationship; Santos et al., "Consent Management Platforms Under GDPR" (2021); Matte et al., "Do Cookie Banners Respect My Choice?" (2020); noyb CMP compliance analysis; r/privacy threads on CMP data collection.
5. Shadow Profiles & Inferred DataHigh
1Facebook Shadow Profiles for Non-Users▾
Problem
Meta/Facebook builds "shadow profiles" for people who have never created a Facebook account by collecting data about them from existing users' contact uploads, tagged photos, event invitations, and Messenger conversations. When a Facebook user uploads their contact list, every phone number and email address — including those of non-users — is ingested and linked. When photos are uploaded and other users are recognized by facial recognition, non-users accumulate biometric data in Facebook's systems. The non-user has never consented to any of this.
Current State
Facebook acknowledged the existence of shadow profiles during Mark Zuckerberg's Congressional testimony in 2018 but characterized them as necessary for "security" purposes (preventing fake accounts, spam). The company's off-Facebook activity tracker (introduced after the Cambridge Analytica scandal) gives users some visibility into data collected through Facebook Pixel and login-with-Facebook, but shadow profile data for non-users remains entirely inaccessible. GDPR deletion requests from non-users are structurally problematic because Facebook cannot verify the identity of someone without an account. Meta's $5 billion FTC settlement and $1.3 billion EU DPC fine did not specifically address shadow profiles.
Impact
An individual who has deliberately never created a Facebook account — perhaps for deeply held privacy principles — discovers through a GDPR subject access request (filed via paper letter to Meta's Dublin office) that Facebook holds their phone number (uploaded by 17 different contacts), their email address (uploaded by 23 contacts), their physical likeness (tagged in 8 photos by contacts), their workplace (mentioned in 3 contacts' employer fields), and their home address (from a contact's address book entry). Facebook has constructed a detailed profile of someone who never agreed to any relationship with the company.
References
Zuckerberg Congressional testimony on shadow profiles (2018); Ireland DPC Meta investigation; FTC v. Facebook $5B settlement (2019); DPC v. Meta €1.3B fine (2023); Kashmir Hill, "Facebook Is Tracking You Even If You're Not on Facebook" (Gizmodo, 2017); r/privacy shadow profile awareness threads; GDPR subject access request experiences shared on noyb.eu.
2Inferred Sexual Orientation and Gender Identity▾
Problem
Data brokers and ad-tech platforms infer sexual orientation, gender identity, and relationship status from behavioral signals — app usage (Grindr, HER, Taimi), browsing patterns, content consumption, location data (visits to LGBTQ+ venues), purchase data (LGBTQ+ media subscriptions, Pride merchandise), and social network connections. These inferences are attached to profiles and sold or used for targeting without the individual's knowledge. In jurisdictions where LGBTQ+ identity is criminalized, this inferred data poses existential risk.
Current State
The ICCL's RTB investigation documented that Google's advertising taxonomy included categories like "Gay & Lesbian" that were broadcast through bid requests. Grindr was fined $6.5 million by the Norwegian DPA (2021) for sharing users' GPS locations and profile data (including HIV status) with advertising partners. Oracle's BlueKai data leak exposed browsing behavior that implied sexual orientation. IAB's content taxonomy included LGBTQ+ interest categories used for targeting. While some platforms have removed explicit sexual orientation targeting categories, behavioral inference makes the removal cosmetic.
Impact
In the 69 countries where homosexuality is criminalized, inferred sexual orientation data in the advertising ecosystem can be literally life-threatening. A user in Saudi Arabia whose ad profile flags "LGBTQ+ interest" based on browsing behavior, app usage, and location data faces potential criminal prosecution. The data does not need to be accurate to cause harm — a false inference can trigger the same consequences. Even in tolerant jurisdictions, inferred sexual orientation can enable discrimination in employment, housing, and insurance where explicit orientation-based discrimination is illegal but targeting-based exclusion is undetectable.
References
Norwegian DPA v. Grindr ($6.5M fine, 2021); ICCL RTB taxonomy investigation; Oracle BlueKai data exposure (TechCrunch, 2020); OutRight Action International, "The Global State of LGBTIQ Organizing"; IAB content taxonomy sexual orientation categories; r/privacy Grindr data sharing threads; Access Now digital safety for LGBTQ+ communities.
3Predicted Income and Financial Status▾
Problem
Data brokers infer income levels, net worth, investment portfolios, debt levels, and financial stability from proxy signals rather than actual financial records. Property values, car registrations, zip code demographics, purchase patterns, credit card type (inferred from transaction data), subscription services, and even web browsing behavior (luxury brand sites vs. discount sites) are used to generate financial scores and income buckets that are sold to advertisers, insurers, lenders, and landlords.
Current State
Acxiom/LiveRamp offers income estimation in ranges ($15K-$25K, $25K-$35K, up to $250K+) as a standard profile attribute. Experian's income insight products provide estimated income based on credit and public records data. Equifax's Workforce Solutions provides income verification, but its marketing analytics division sells inferred income segments. These income estimates are attached to hundreds of millions of consumer profiles and used to determine which financial products people are offered, what prices they see online, and how they are treated by service providers.
Impact
A person accurately estimated as low-income by data broker algorithms discovers they are systematically shown payday loan advertisements, subprime credit offers, and predatory insurance products — while being excluded from premium credit card offers, investment platform ads, and wealth management services. The income inference creates a feedback loop: being identified as financially vulnerable makes an individual a target for predatory products designed to extract maximum revenue from vulnerable populations, further entrenching financial precarity. The FTC has documented this as "digital redlining."
References
FTC, "Big Data: A Tool for Inclusion or Exclusion?" (2016); CFPB inquiry into data broker credit scoring alternatives; Acxiom/LiveRamp data attribute catalog; Experian income insight products; National Consumer Law Center, "Big Data, Big Discrimination" (2020); r/personalfinance threads on targeted predatory lending ads.
4Health Condition Inference From Non-Medical Data▾
Problem
Data brokers infer health conditions from non-medical data that is not protected by HIPAA — purchase patterns (buying glucose test strips, joint supplements, anti-nausea medication), browsing behavior (visiting WebMD pages for specific conditions, reading cancer treatment articles), location data (visiting oncology clinics, methadone clinics, fertility centers), and app usage (calorie tracking, mental health apps, sobriety trackers). These inferences create health profiles that are sold to insurers, employers, and pharmaceutical marketers.
Current State
The data broker industry maintains health-related audience segments including "Diabetes Interest," "Arthritis Sufferers," "Expectant Parents," "Weight Loss Interest," and "Mental Health." Oracle's BlueKai leak exposed browsing data that revealed health conditions. The FTC's Health Breach Notification Rule was used against GoodRx but covers only entities that collect actual health data — it does not address inference of health conditions from behavioral signals. No federal law prevents a data broker from inferring that someone has cancer based on their browsing history and selling that inference to an insurance company.
Impact
A user who googles symptoms of depression, reads articles about antidepressant medications, and visits a therapist's website (tracked via the therapist's Google Analytics implementation) discovers that pharmaceutical companies are now targeting them with antidepressant ads across all platforms. More consequentially, the inferred mental health profile — visible to insurers through broker data — may influence life insurance underwriting, disability insurance pricing, and long-term care insurance decisions. The user never disclosed a mental health condition to anyone, but their browsing behavior created a profile that functionally substitutes for a medical disclosure.
References
The Markup, "How We Analyzed Patient Data" (health data broker investigation); FTC Health Breach Notification Rule enforcement; Oracle BlueKai health data exposure; World Privacy Forum health scoring analysis; Senate Finance Committee health data broker inquiry (2023); r/privacy health data inference discussions.
5Predictive Life Event Scoring▾
Problem
Data brokers predict major life events — pregnancy, marriage, divorce, retirement, home purchase, job change, death of a family member — before the individual has publicly disclosed them or sometimes before the individual is fully aware. These predictions are based on pattern matching across purchase data, browsing behavior, location changes, social media activity, and financial transaction patterns. Predicted life events are among the most commercially valuable broker data products because they identify consumers at moments of maximum purchasing activity and vulnerability.
Current State
Acxiom, Experian, and Oracle (before its ad division shutdown) all offered "life event triggers" as advertising targeting segments. These include "New Mover," "Expectant Parent," "Recently Divorced," "New Empty Nester," "Recently Bereaved," and "Pre-Retiree." The segments are updated in near real-time as behavioral signals accumulate. Target's pregnancy prediction algorithm (using 25 products whose purchase patterns predict pregnancy with high accuracy) was documented by the New York Times in 2012 and remains the canonical example, but every major broker now offers equivalent capabilities across dozens of life events.
Impact
A woman who has told no one she is pregnant begins receiving baby product catalogs, prenatal vitamin advertisements, and maternity clothing targeted ads. The prediction was triggered by a combination of signals: she stopped buying alcohol, purchased folic acid supplements, searched for OB-GYN offices, and her period tracking app (sharing data with a broker through an SDK) detected a missed period. The data broker infrastructure identified her pregnancy before her first prenatal appointment. She has been denied the fundamental human experience of choosing when and how to share this information.
References
Duhigg, "How Companies Learn Your Secrets" (NYT, 2012); Acxiom life event trigger products; Experian life stage segmentation; The Markup, "How Your Pharmacy Records Get Exploited"; r/privacy pregnancy prediction anecdotes; FTC workshop on predictive analytics and consumer privacy.
6Political Ideology and Belief Inference▾
Problem
Data brokers infer political ideology, religiosity, and social values from behavioral signals far beyond voter registration records. Media consumption patterns (Fox News vs. MSNBC, podcast subscriptions), donation history (via FEC records), bumper sticker and yard sign detections (via satellite and street view imagery), social media behavior, consumer brand preferences, and even grocery purchases (organic vs. conventional, gun shop proximity) feed algorithms that assign political and ideological scores to consumer profiles.
Current State
L2, TargetSmart, and i360 assign partisan scores and issue-position predictions to every registered voter. Acxiom and Experian sell "political interest" and "social values" segments to non-political advertisers. Cambridge Analytica demonstrated that psychological profiles (OCEAN/Big Five personality traits) could be predicted from Facebook likes with significant accuracy. Post-Cambridge Analytica, explicit psychographic targeting was restricted on some platforms, but the underlying inference capabilities remain available through the broker ecosystem.
Impact
An individual discovers that their entire information environment has been shaped by an inferred political profile they cannot see or correct. They receive news articles, product recommendations, and social media content calibrated to their predicted political orientation. If the prediction is wrong — a moderate classified as extreme, or a politically evolving individual locked into a stale profile — the information bubble reinforces a political identity they may not actually hold. Political inference also enables discrimination: a documented 2016 Bloomberg investigation showed that political affiliation data from brokers was used in employment screening.
References
Cambridge Analytica psychographic profiling documentation; L2/TargetSmart political scoring methodologies; Acxiom political interest segments; Bloomberg, "They Know What You Did" (employment screening using political data, 2016); Tactical Tech, "Data and Elections" research; r/privacy political inference threads.
7Behavioral Biometric Profiling▾
Problem
A new category of inferred data captures behavioral biometrics — typing patterns, mouse movements, touchscreen gestures, gait analysis, voice patterns, and interaction rhythms — to create persistent identifiers that cannot be changed because they are intrinsic to the individual's physiology. Companies like BioCatch (banking fraud detection), TypingDNA, and BehavioSec (now part of LexisNexis) build behavioral biometric profiles that identify users even when they use different devices, clear cookies, or use VPNs.
Current State
BioCatch profiles are deployed by banks to detect fraud through behavioral biometrics — measuring how a user types, swipes, and moves their mouse to distinguish legitimate users from imposters. This same technology creates persistent behavioral identifiers. TypingDNA can identify individuals from their typing cadence with 99%+ accuracy. LexisNexis acquired BehavioSec in 2022 to add behavioral biometrics to its identity verification stack. These systems create biometric data — as immutable and sensitive as fingerprints — from ordinary interactions with devices, often without explicit notification.
Impact
A user who has taken extreme privacy measures — using Tor, changing devices regularly, never providing real identifying information — can still be identified by their typing pattern, mouse movement style, or touchscreen interaction habits. Behavioral biometrics represent the ultimate defeat of pseudonymity: you cannot change how you type or move a mouse. Unlike a username or IP address, a behavioral biometric cannot be reset. If this data is breached or misused, the individual cannot adopt a new behavioral pattern to recover their anonymity.
References
BioCatch technology documentation; TypingDNA academic publications; LexisNexis/BehavioSec acquisition (2022); EDPB guidelines on biometric data processing; Mondal et al., "Continuous Authentication Using Behavioral Biometrics" (IEEE, 2017); r/privacy behavioral biometric tracking threads.
8Social Graph Inference for Non-Participating Individuals▾
Problem
Data brokers and platforms construct social graphs for individuals based on other people's data — contact lists uploaded by their acquaintances, co-location signals (two devices frequently appearing at the same GPS coordinates), co-transaction patterns (frequently purchasing from the same merchant at the same time), network analysis of communication metadata, and social media connections of their contacts. An individual who shares no data themselves can have their social network fully mapped through the data shared by everyone around them.
Current State
Facebook's "People You May Know" feature demonstrated the power — and danger — of social graph inference, famously surfacing connections that users wanted to keep private (a psychiatrist's patients were suggested to each other, a sperm donor's biological children were connected). LinkedIn's social graph maps professional relationships. Data brokers like FullContact and Pipl construct relationship networks from public and purchased data. The people-search "relatives and associates" feature described in Category 3 is a visible manifestation of social graph inference, but the underlying graph is far more detailed than what is displayed publicly.
Impact
A therapist discovers that her clients are being suggested as connections to each other by Facebook's algorithm, potentially revealing their mental health treatment to other patients. An undercover law enforcement officer's cover is compromised when social graph algorithms connect them to other officers through co-location patterns. An anonymous domestic violence hotline counselor's identity is inferred from the pattern of calls between their personal phone and the hotline's phone number, mapped through contact list uploads by mutual acquaintances.
References
Kashmir Hill, "People You May Know: The Secrets Facebook's Algorithm Hides" (Gizmodo, 2017); Facebook "People You May Know" privacy concerns reporting; FullContact social graph API documentation; Pipl identity resolution social network features; r/privacy PYMK exposure anecdotes; EFF social graph surveillance analysis.
9Emotional State and Mental Health Inference▾
Problem
Platforms and data companies infer emotional states and mental health conditions from behavioral signals — posting frequency, language sentiment, sleep patterns (inferred from device usage times), social withdrawal (reduced messaging), content consumption shifts (from entertainment to crisis-related content), and physiological signals from wearables (heart rate variability, skin conductance). Facebook's internal research (leaked by Frances Haugen) demonstrated that the company could identify teens experiencing emotional vulnerability and potentially target advertising to them during these states.
Current State
Facebook's leaked internal documents (the "Facebook Papers," 2021) included research showing the company could identify when teenagers felt "insecure," "worthless," or "need a confidence boost" and that this information was presented to advertisers. Instagram's internal research acknowledged that the platform worsened body image for 1 in 3 teen girls. Fitbit/Google Health collects physiological data that can indicate depression (changes in sleep, activity, heart rate variability). Affective computing companies like Affectiva and Realeyes analyze facial expressions through webcams for "emotional AI" advertising optimization.
Impact
A teenager going through a depressive episode generates behavioral signals across multiple platforms — reduced social media posting, late-night scrolling patterns, searches for "am I depressed," and changes in music streaming to sadder content. These signals converge in ad-tech profiles that classify the teen as "emotionally vulnerable" — a high-value advertising target for products promising self-improvement, beauty enhancement, or mood improvement. The advertising ecosystem has monetized mental illness, targeting people at their lowest moments not to help them but to sell them products.
References
Frances Haugen/Facebook Papers whistleblower disclosures (2021); Facebook internal research on teen emotional states; Instagram body image internal study; Affectiva emotional AI documentation; Realeyes advertising emotion measurement; WSJ "The Facebook Files" investigation series; r/privacy emotional targeting discussions.
10Synthetic Identity Assembly From Inferred Data▾
Problem
The ultimate expression of shadow profiling is the synthetic assembly of comprehensive identity profiles for individuals who have never directly provided data to any broker. By combining inferred data (from contacts' uploads), public records (property, voter, court), observed behavioral signals (IP addresses, device fingerprints, location from apps used by household members), and purchased data from the resale chain, brokers construct profiles that are almost entirely inferred rather than voluntarily disclosed. These profiles are indistinguishable from profiles built on directly collected data in the broker marketplace.
Current State
LiveRamp, Acxiom, and Experian maintain profiles on 250+ million US adults — effectively the entire adult population, including individuals who have never directly interacted with any data broker. The FTC's 2014 study documented that brokers create profiles for "virtually every US consumer." For privacy-conscious individuals who minimize their digital footprint, brokers fill gaps through inference: income estimated from zip code and property records, political affiliation modeled from neighborhood demographics, interests inferred from household members' data, and social graph constructed from contacts' uploaded address books.
Impact
An individual who has spent years practicing digital minimalism — using cash, avoiding social media, using a VPN, never providing real information to commercial services — discovers through a CCPA data access request that Acxiom holds a profile containing their name, address, estimated income range, political party, number of household members, vehicle type, home value, and 200+ inferred interest categories. None of this was directly provided. The profile was assembled entirely from public records, neighbors' data, household members' activities, and statistical inference. Digital minimalism reduces the accuracy of the profile but cannot prevent its creation.
References
FTC "Data Brokers: A Call for Transparency" (2014); Acxiom/LiveRamp data access portal experiences; CCPA data access request results shared on r/privacy; Privacy Rights Clearinghouse, "Data Brokers and Your Personal Information" (updated 2023); The Markup, "What Data Brokers Know About You" investigation; PrivacyGuides forum threads on data minimalism limitations.
6. Government Procurement of Broker DataCritical
1Warrantless Location Surveillance via Commercial Purchase▾
Problem
Federal agencies including ICE, CBP, the FBI, the Secret Service, the DEA, and the IRS purchase commercial location data from brokers like Venntel (now Babel Street), Locate X, and previously X-Mode Social (now Outlogic) to track individuals' movements without obtaining a warrant. This practice directly circumvents the Supreme Court's 2018 Carpenter v. United States ruling, which held that accessing historical cell-site location information requires a warrant. By purchasing the same data commercially, agencies argue they are buying "commercially available information" rather than conducting a search.
Current State
A 2023 ODNI (Office of the Director of National Intelligence) declassified report acknowledged that the government purchases commercially available data that could reveal sensitive information about Americans, including location tracking, and that this data "can be misused to pry into private lives." DHS signed contracts worth millions with Venntel between 2018-2022. The Fourth Amendment Is Not For Sale Act, introduced repeatedly in Congress by Senators Wyden and Paul, has not passed as of early 2026. Executive Order 14086 (2022) addresses signals intelligence but does not restrict commercial data purchases.
Impact
CBP used Venntel data to track individuals near the US-Mexico border without warrants, including US citizens. The Wall Street Journal reported in 2020 that DHS used Venntel to identify and track undocumented immigrants via their phone location data obtained from weather and gaming apps. Individuals have no notice, no opportunity to contest, and no recourse when their commercially purchased location data is used for government surveillance.
References
ODNI declassified report "Senior Advisory Group Report on Commercially Available Information" (Jan 2022, declassified June 2023); Carpenter v. United States, 585 U.S. 296 (2018); WSJ investigation "Federal Agencies Use Cellphone Location Data for Immigration Enforcement" (Feb 2020); EFF analysis of Venntel contracts via FOIA.
2ICE and CBP Procurement of Surveillance Tools▾
Problem
Immigration and Customs Enforcement (ICE) and Customs and Border Protection (CBP) have built a comprehensive surveillance apparatus through commercial data broker contracts. ICE has purchased access to LexisNexis Accurint (identity and address data), Thomson Reuters CLEAR (comprehensive person search), Babel Street (location analytics), Clearview AI (facial recognition), and Palantir (data integration platform). These purchases enable mass surveillance of immigrant communities without judicial oversight, probable cause, or individualized suspicion.
Current State
Georgetown Law's Center on Privacy and Technology documented over $2.8 billion in ICE surveillance technology spending between 2008-2021. The ACLU obtained records showing ICE used Thomson Reuters CLEAR to identify targets for enforcement actions. Contract records show CBP spent over $1.1 million on Babel Street's Locate X tool for phone location tracking between 2020-2022. Internal DHS Inspector General reports have found inadequate privacy impact assessments for these procurements.
Impact
ICE's surveillance tools enable dragnet monitoring of entire communities. Advocates report a chilling effect on immigrant communities, with individuals avoiding medical care, reporting crimes, or participating in civic life for fear that any interaction generating data could be funneled through commercial channels to immigration enforcement. The ACLU documented cases where utility connection records purchased from commercial databases were used to identify deportation targets.
References
Georgetown Law Center on Privacy & Technology "American Dragnet: Data-Driven Deportation in the 21st Century" (2022); ACLU FOIA on ICE-Thomson Reuters contracts; DHS OIG reports on privacy assessments; Mijente #NoTechForICE campaign documentation.
3FBI Purchases of Geolocation and Ad Data▾
Problem
The FBI purchased access to commercial geolocation data from Venntel to track Americans' movements without warrants, as confirmed by FBI Director Christopher Wray in Senate testimony in 2023. The FBI also uses commercially acquired advertising data, social media monitoring tools (including from Babel Street and Dataminr), and open-source intelligence platforms that aggregate broker-sourced data. The agency has acknowledged that it previously purchased netflow data (internet metadata) from Team Cymru without legal process.
Current State
In March 2023, FBI Director Wray confirmed under questioning by Senator Wyden that the FBI had purchased Americans' location data from commercial brokers. Wray stated the program was subsequently ended due to "budget" concerns, not legal ones — implying the FBI considered the practice lawful. The FBI continues to purchase social media monitoring tools and other commercially available datasets. An internal FBI policy memo reportedly restricts but does not prohibit commercial data purchases for investigative purposes.
Impact
The FBI's admission confirmed what civil liberties organizations had long alleged: domestic law enforcement uses the commercial data marketplace as a backdoor around the warrant requirement. Even after the specific Venntel contract ended, the FBI retains access to location-relevant data through other commercial tools and fusion center arrangements. The precedent signals to other federal and state agencies that commercial data purchases for surveillance face no legal barrier.
References
Senate Judiciary Committee hearing testimony, FBI Director Wray (March 2023); Sen. Wyden letter to DOJ regarding FBI location data purchases; Vice Motherboard "The FBI Just Admitted It Bought US Location Data" (March 2023); Team Cymru netflow data controversy reporting.
4Military and Intelligence Community Data Purchases▾
Problem
The Department of Defense, NSA, DIA, and other intelligence community agencies purchase commercially available data including location data, web browsing data, and app usage data from commercial brokers. A declassified ODNI report revealed that intelligence agencies consider commercially available information a valuable supplement to traditional signals intelligence, and that the volume and sensitivity of this data has grown beyond what existing oversight frameworks anticipated.
Current State
The ODNI's Senior Advisory Group report (declassified June 2023) warned that commercially available information "can reveal sensitive and intimate information about individuals" and that "in the wrong hands, [it] could facilitate blackmail, stalking, harassment, and public shaming." Despite this internal acknowledgment of risk, no binding restrictions have been imposed. DIA confirmed purchasing smartphone location data from commercial brokers. The NSA has purchased internet browsing records from data brokers, as reported by the New York Times in January 2024 following Senator Wyden's disclosure.
Impact
Military and intelligence agencies operate with even less transparency than domestic law enforcement. The scale of data purchases is classified, the purposes are classified, and the oversight is conducted by classified courts and committees. Senator Wyden's office revealed that the NSA's purchase of internet browsing data included records of Americans' web visits, effectively creating a warrantless browsing history surveillance program through commercial channels.
References
ODNI declassified Senior Advisory Group report (June 2023); NYT "N.S.A. Buys Americans' Internet Data Without Warrants" (Jan 2024); Sen. Wyden disclosure on NSA data purchases; DIA smartphone location data confirmation; ACLU analysis of intelligence community commercial data procurement.
5IRS Criminal Investigation Data Broker Access▾
Problem
The IRS Criminal Investigation division purchased access to commercial location data from Venntel to track suspects' movements and identify potential tax evasion without warrants or court orders. The IRS also contracts with LexisNexis, Palantir, and other data aggregators for person-search and financial profiling capabilities. These purchases blur the line between lawful tax enforcement and warrantless surveillance of financial behavior.
Current State
Contract records obtained by the ACLU and reported by Vice Motherboard revealed IRS-CI purchases of Venntel location data in 2019-2020. The IRS Inspector General reviewed the purchases but did not find they violated existing IRS policy — because no policy specifically addressed commercial location data procurement. The IRS uses Palantir's Investigative Case Management platform, which integrates commercially purchased data with IRS records. Senator Wyden has specifically called out IRS data broker purchases as requiring legislative restriction.
Impact
The IRS's tax enforcement mission gives it access to some of the most sensitive financial data in existence. Adding commercially purchased location and behavioral data creates a comprehensive profile of individuals' financial lives, physical movements, and daily patterns — all without the judicial oversight that would be required if the IRS sought this information directly from telecom providers.
References
Vice Motherboard "The IRS Bought Location Data from a Data Broker" (2021); ACLU FOIA on IRS-Venntel contracts; IRS-Palantir contract documentation; Sen. Wyden correspondence with IRS Commissioner.
6State and Local Law Enforcement Broker Access▾
Problem
State and local police departments increasingly purchase commercial surveillance tools including Fog Data Science (phone location tracking), Clearview AI (facial recognition), social media monitoring platforms (Geofeedia, Media Sonar, Babel Street), and automated license plate reader data (Vigilant/Motorola Solutions, Flock Safety). These purchases are typically made without city council oversight, public debate, or privacy impact assessments, and are often funded through federal grants or asset forfeiture funds that bypass normal procurement scrutiny.
Current State
Fog Data Science, exposed by the AP and EFF in 2022, sold phone location tracking to at least 40 state and local agencies, many of which had no formal policy governing location surveillance. Clearview AI sold facial recognition access to over 3,100 law enforcement agencies by 2022, many of which signed up using individual officers' email addresses without departmental authorization. The ACLU has documented social media monitoring tool purchases by police departments in dozens of cities. Community surveillance ordinances (enacted in Oakland, San Francisco, Seattle, and others) require public disclosure and approval of surveillance technology purchases, but most US jurisdictions have no such requirement.
Impact
Small-town police departments with budgets under $1 million can purchase the same surveillance capabilities that were previously available only to federal intelligence agencies. A 2022 AP investigation found Fog Data Science was used by local police to track individuals visiting abortion clinics, attending protests, and visiting specific homes — all without warrants. The lack of oversight means misuse is discovered only through investigative journalism or FOIA requests, long after the surveillance has occurred.
References
AP/EFF investigation "Fog Revealed" (2022); BuzzFeed News Clearview AI customer list investigation; ACLU reports on police surveillance technology purchases; surveillance technology oversight ordinances database.
7Social Media Monitoring and Predictive Policing Contracts▾
Problem
Government agencies at all levels purchase social media monitoring and analysis tools from companies like Babel Street, Dataminr, Media Sonar, ShadowDragon, and ZeroFox. These tools scrape, aggregate, and analyze social media posts, sometimes integrating with data broker datasets to connect online identities to real-world individuals. DHS has used social media monitoring for "situational awareness" at protests, the FBI has used it for counter-terrorism and domestic threat assessments, and local police departments have used it for gang monitoring that disproportionately targets Black and Brown communities.
Current State
DHS's Social Media and Situational Awareness program monitors social media during "events of national significance." The Brennan Center for Justice documented DHS social media monitoring of Black Lives Matter protests in 2020. Dataminr, which has a special partnership with Twitter/X for real-time data access, has sold its tools to police departments despite Twitter's stated policy prohibiting the use of its data for surveillance. The FBI's use of social media monitoring tools was detailed in an Inspector General report that found insufficient policies governing their use.
Impact
Social media monitoring creates chilling effects on First Amendment-protected speech and assembly. Individuals who know or suspect their social media activity is monitored by law enforcement self-censor, avoid organizing, and withdraw from public discourse. The Brennan Center documented cases where individuals were placed on watchlists based on social media activity that constituted protected political speech.
References
Brennan Center for Justice "Monitoring Social Media" (2019); Brennan Center analysis of DHS protest monitoring (2020); Twitter/Dataminr surveillance controversy; FBI OIG social media monitoring report; ShadowDragon product documentation.
8Data Fusion Centers and Broker Integration▾
Problem
The 80+ DHS-supported state and local fusion centers combine government databases with commercially purchased data broker datasets to create comprehensive surveillance profiles. Fusion centers aggregate criminal justice records, motor vehicle data, financial records, utility records, and commercially purchased data including location tracking, people-search results, and social media monitoring. This creates a government surveillance capability that exceeds what any single agency could legally obtain through direct collection, by laundering the information through commercial intermediaries.
Current State
A 2012 Senate Permanent Subcommittee on Investigations report found fusion centers produced "predominantly useless information," violated civil liberties, and lacked adequate privacy protections. Despite these findings, fusion center funding and data broker integration have expanded. The Government Accountability Office has reported inadequate oversight of fusion center data practices. Individual fusion centers sign their own data broker contracts with minimal transparency, making comprehensive accounting of government data purchases nearly impossible.
Impact
Fusion center intelligence products — combining government records with commercial data — are shared across law enforcement agencies through systems like the FBI's eGuardian and DHS's Homeland Security Information Network. An individual flagged by a fusion center based partly on commercially purchased data can face law enforcement scrutiny without ever knowing the basis for that scrutiny or having the opportunity to challenge inaccurate commercial data.
References
Senate PSI "Federal Support for and Involvement in State and Local Fusion Centers" (2012); GAO fusion center oversight reports; ACLU "What's Wrong with Fusion Centers" report; EFF fusion center FOIA documents.
9Customs and Immigration Biometric Data Commercialization▾
Problem
CBP collects biometric data (facial images, fingerprints) from international travelers and has shared this data with commercial entities through partnerships and contracts. The CBP Traveler Verification Service processes hundreds of millions of facial comparisons annually. Airlines and airports collect biometric data under CBP programs and may retain or share it for commercial purposes. The reverse also occurs: commercial facial recognition companies (Clearview AI) scrape billions of public photos and sell identification services back to government agencies.
Current State
CBP's facial recognition program has been deployed at over 250 airports, processing virtually all international departures. Opt-out mechanisms for US citizens exist in theory but are inconsistently implemented and often not communicated to travelers. A 2020 DHS Privacy Impact Assessment acknowledged that biometric data collected at airports could be retained for up to 75 years. Clearview AI scraped over 30 billion images from public sources and sold facial recognition services to over 3,100 law enforcement agencies and multiple federal agencies.
Impact
The biometric data pipeline between government collection and commercial availability creates a permanent identification infrastructure. Once your face is in CBP's system and Clearview AI's database, you can be identified in any public space where cameras feed into either system. The 2019 CBP data breach exposed traveler photos and license plate images from a subcontractor, demonstrating the security risks of this data sharing.
References
DHS Privacy Impact Assessment for Traveler Verification Service; CBP 2019 biometric data breach disclosure; Clearview AI investigation by NYT (2020); ACLU v. Clearview AI litigation; GAO reports on CBP facial recognition program.
10Executive Order Gaps and Congressional Inaction▾
Problem
Despite years of investigative journalism, civil liberties litigation, congressional hearings, and even internal government reports acknowledging the problem, no binding legal restriction prevents government agencies from purchasing commercially available personal data to circumvent warrant requirements. Executive Order 14086 (Oct 2022) addressed signals intelligence collected from non-US persons but did not restrict commercial data purchases. The Fourth Amendment Is Not For Sale Act has been introduced in multiple congressional sessions but has not passed. Agency-level policies are voluntary, inconsistent, and unenforceable.
Current State
As of early 2026, there is no federal law prohibiting government agencies from purchasing commercially available location data, browsing history, or other personal information without a warrant. The ODNI report recommending restrictions led to no binding policy changes. Individual agencies have adopted varying internal policies — the FBI reportedly ended its Venntel contract, while other agencies continue similar purchases through different vendors. The GAO has not been tasked with comprehensive auditing of government commercial data purchases. Congressional attempts to legislate have stalled due to national security concerns raised by intelligence community lobbyists.
Impact
The absence of legal restriction creates a permanent loophole in Fourth Amendment protections. As commercial data collection expands (through IoT devices, connected cars, health apps, and smart home devices), the volume and intimacy of data available for warrantless government purchase grows continuously. Each new consumer technology category creates a new surveillance vector available to any government agency with a procurement budget. The constitutional right to be free from unreasonable searches is effectively nullified for any information that passes through a commercial intermediary.
References
Executive Order 14086 text and analysis; Fourth Amendment Is Not For Sale Act bill text (multiple sessions); ODNI Senior Advisory Group recommendations; Brennan Center legislative tracker on surveillance reform; EFF "Government Use of Commercial Data" policy analysis.
7. Data Marketplace Regulation GapsHigh
1No Comprehensive US Federal Privacy Law▾
Problem
The United States has no comprehensive federal data privacy law comparable to the EU's GDPR, despite decades of advocacy and multiple legislative attempts. The American Data Privacy and Protection Act (ADPPA) passed the House Energy and Commerce Committee in 2022 with bipartisan support but died before reaching the House floor due to disputes over federal preemption of state laws and private right of action provisions. Subsequent attempts have similarly stalled. This leaves data brokers operating in a regulatory environment where collection, aggregation, and sale of personal data is legal by default.
Current State
Federal privacy regulation remains sectoral: HIPAA covers health data, FERPA covers education records, COPPA covers children under 13, GLBA covers financial data, and FCRA covers credit reporting. None of these laws comprehensively regulate data brokers. The FTC uses its Section 5 "unfair or deceptive practices" authority for enforcement but can only act when companies violate their own stated privacy policies or engage in practices that meet the legal standard for unfairness. The FTC cannot write rules establishing baseline data protection requirements without new legislation or lengthy rulemaking proceedings.
Impact
Data brokers operate in the gaps between sectoral laws. A broker that collects location data (not covered by HIPAA), aggregates it with purchase history (not covered by GLBA), appends social media activity (not covered by any federal law), and sells the combined profile for advertising, employment screening, or government surveillance faces no federal legal restriction on any of these activities — as long as it does not make deceptive promises about privacy.
References
ADPPA bill text and committee markup (2022); FTC Section 5 authority analysis; Brookings Institution "Why America needs a federal data privacy law" series; IAPP federal privacy legislation tracker; comparison analyses of failed federal privacy bills (2012-2025).
2State Privacy Law Patchwork Creates Compliance Arbitrage▾
Problem
In the absence of federal legislation, states have enacted their own privacy laws: California (CCPA/CPRA), Virginia (VCDPA), Colorado (CPA), Connecticut (CTDPA), Utah (UCPA), Texas (TDPSA), Oregon (OCPA), Montana (MCDPA), and others — with each law using different definitions, different thresholds, different rights, and different enforcement mechanisms. This patchwork creates compliance arbitrage opportunities where brokers structure their operations to minimize regulatory exposure. A broker incorporated in a state without a privacy law, processing data on residents of multiple states, faces a complex jurisdictional calculation that often resolves in the broker's favor.
Current State
As of early 2026, approximately 20 US states have enacted comprehensive privacy laws, but they differ on fundamental questions: What constitutes a "sale" of data? What thresholds trigger applicability (revenue, data volume, percentage of revenue from data sales)? Do consumers have a private right of action? What is "sensitive data"? Only California's law provides a dedicated data broker registration requirement. Only a handful of states grant a private right of action. Most state laws exempt "publicly available information" without defining the term precisely enough to prevent broker exploitation.
Impact
Data brokers maintain compliance with the most permissive applicable state law while doing business nationally. A broker processing data on California residents must comply with CPRA, but the same broker processing data on residents of states without privacy laws faces no restrictions. Brokers have relocated corporate registration, data processing facilities, and legal entities to minimize state law exposure. The patchwork also burdens legitimate businesses that must comply with 20+ different frameworks.
References
IAPP US state privacy legislation tracker; CPRA implementing regulations (California Privacy Protection Agency); state-by-state privacy law comparison matrices; National Conference of State Legislatures privacy law database; industry compliance cost analyses.
3Vermont Data Broker Registry Limitations▾
Problem
Vermont enacted the first US data broker registration law in 2018 (Act 171), requiring companies that collect and sell data about consumers with whom they have no direct relationship to register annually with the Secretary of State, pay a $100 fee, and disclose basic practices. While groundbreaking in concept, the registry has proven toothless: registration is self-reported with no verification, non-compliance penalties are minimal, the registry does not restrict any actual data practices, and the law has no extraterritorial enforcement mechanism for out-of-state brokers who ignore the requirement.
Current State
The Vermont registry lists approximately 500-600 registered data brokers, but researchers estimate the actual number of companies meeting the statutory definition exceeds 4,000 nationally. Many brokers simply do not register, and Vermont lacks the enforcement resources to identify and compel compliance from out-of-state companies. The registry provides transparency about which companies acknowledge being data brokers but imposes no substantive restrictions on their data collection, aggregation, or sale practices. California enacted its own broker registration requirement (effective 2024 via the Delete Act/SB 362), which adds the requirement of participating in a universal deletion mechanism.
Impact
The Vermont registry is cited by industry as evidence that regulation exists, while providing virtually no consumer protection. An individual who discovers they are in 200 brokers' databases gains no actionable right from the Vermont registry — it tells you who the brokers are but provides no mechanism to make them stop. Privacy researchers use the registry as a research tool, but its consumer protection value is negligible.
References
Vermont Act 171 (2018) text; Vermont Secretary of State data broker registry; Duke Sanford School of Public Policy analysis of Vermont registry effectiveness; California Delete Act (SB 362) text and implementation timeline; Privacy Rights Clearinghouse broker registry analysis.
4FTC Enforcement Actions Are Infrequent and Insufficient▾
Problem
The FTC is the primary federal agency with authority over data broker practices, but its enforcement actions are sporadic, narrowly scoped, and impose penalties that amount to a rounding error on broker revenue. The FTC brought actions against Kochava (location data), X-Mode Social/Outlogic (location data sold to military contractors), InMarket (location data without consent), and data broker Epsilon (deceptive data practices), but these cases take years to resolve, cover only the most egregious practices, and result in consent orders rather than structural industry reform.
Current State
The FTC's January 2024 order against X-Mode Social/Outlogic prohibited the sale of sensitive location data (near medical facilities, religious sites, domestic violence shelters) but allowed the company to continue selling other location data. The FTC's action against Kochava (filed 2022) alleged the company sold precise geolocation data that could track visits to reproductive health clinics, places of worship, and homeless shelters. The FTC's proposed settlement with InMarket (March 2024) required consent for location data collection. These actions address individual bad actors but do not establish industry-wide rules.
Impact
FTC enforcement addresses the worst abuses while leaving the business model intact. A broker that sells location data tracking people to grocery stores, workplaces, and homes faces no FTC action — only those selling data specifically tied to sensitive locations face scrutiny. The industry adapts by avoiding the specific practices named in consent orders while continuing all others. The pace of enforcement (5-10 cases per year across all industries) versus the scale of the industry (4,000+ brokers) means the probability of any individual broker facing action is negligible.
References
FTC v. Kochava complaint (2022); FTC v. X-Mode Social/Outlogic order (Jan 2024); FTC v. InMarket proposed settlement (March 2024); FTC data broker enforcement action compilation; FTC budget and staffing constraints analysis.
5CCPA/CPRA "Sale" Definition Loopholes▾
Problem
The California Consumer Privacy Act (CCPA) and its successor California Privacy Rights Act (CPRA) define "sale" of personal information as "selling, renting, releasing, disclosing, disseminating, making available, transferring, or otherwise communicating" personal information for "monetary or other valuable consideration." Data brokers exploit ambiguities in this definition by characterizing data transfers as "sharing" (a separate CPRA category with different rules), "service provider" arrangements, or "business purpose" transfers — each of which has different consent and opt-out requirements.
Current State
The California Privacy Protection Agency (CPPA) has issued implementing regulations clarifying some definitional issues, but enforcement is still maturing. Data brokers restructure contracts to characterize data transfers as "sharing for cross-context behavioral advertising" rather than "sales," which triggers different consumer rights under CPRA. Some brokers argue that providing data access through an API (rather than a file transfer) does not constitute a "sale." Others claim that aggregated or de-identified data falls outside the definition entirely, even when re-identification is trivially possible.
Impact
Consumers exercising their CCPA/CPRA "Do Not Sell" right discover that their data continues to flow through channels characterized as "sharing," "service provider" relationships, or "business purpose" transfers. The legal distinction between "sale" and "sharing" is meaningless from the consumer's perspective — their data is still being transferred to third parties for purposes they did not consent to — but it determines which legal protections apply.
References
CCPA/CPRA statutory text; CPPA implementing regulations (2023); California AG enforcement actions under CCPA; IAPP analysis of "sale" vs. "sharing" under CPRA; industry compliance guides on CCPA data transfer characterization.
6Broker "Publicly Available Information" Exemptions▾
Problem
Most state privacy laws exempt "publicly available information" from their coverage, and data brokers exploit this exemption aggressively. Brokers argue that data scraped from social media profiles, court records, property records, voter rolls, professional licenses, and other public sources is "publicly available" and therefore exempt from privacy law requirements including opt-out rights, deletion requests, and consent requirements. The aggregation of multiple "publicly available" data points creates profiles far more revealing than any individual source.
Current State
The definition of "publicly available information" varies by state law. CPRA defines it as information "lawfully made available from federal, state, or local government records" but broadens it to include information the consumer has made available to the general public. Brokers stretch this to include any data posted on social media, mentioned in a news article, or appearing in a public record — even if the individual had no meaningful choice about the data's publication. The aggregation problem is unaddressed: combining a public court record with a public property record with a public voter registration creates a comprehensive profile that is arguably not "publicly available" as a combined dataset.
Impact
People-search sites like Spokeo, BeenVerified, Whitepages, and Intelius build comprehensive profiles entirely from "publicly available" sources and claim exemption from privacy law obligations. An individual who has never consented to data collection finds their home address, phone number, family members, estimated income, political affiliation, and court records aggregated and sold — all from "publicly available" sources. Stalking victims, domestic violence survivors, and individuals in witness protection find their current addresses published because the underlying data is technically "public."
References
CPRA "publicly available information" definition and exemptions; Spokeo v. Robins litigation; Vermont AG consumer guidance on people-search sites; National Network to End Domestic Violence reports on data broker risks; Privacy Rights Clearinghouse people-search site analysis.
7No Fiduciary Duty or Loyalty Obligation for Data Holders▾
Problem
Unlike attorneys, doctors, or financial advisors, companies that hold personal data owe no fiduciary duty or duty of loyalty to the individuals whose data they possess. Data brokers can legally act against their data subjects' interests — selling data to entities that will use it to deny employment, insurance, housing, or credit. The concept of an "information fiduciary" has been proposed by legal scholars (notably Jack Balkin at Yale) but has not been enacted into law. Without a loyalty obligation, data holders face no legal consequence for using data in ways that harm the people it describes.
Current State
The information fiduciary concept would impose duties of care, loyalty, and confidentiality on entities holding personal data, analogous to the duties professionals owe their clients. Several federal privacy bills have included weakened versions of this concept, but none has passed. The FCRA imposes something like a fiduciary duty on credit reporting agencies (requiring accuracy, dispute resolution, and permissible purpose limitations), but this model has not been extended to data brokers generally. The FTC's "unfairness" doctrine can address some harms but does not impose an affirmative duty to act in data subjects' interests.
Impact
A data broker can simultaneously sell an individual's data to a marketing firm (generating revenue) and to a debt collector targeting that same individual (generating additional revenue) — profiting from both sides of a transaction that harms the data subject. Without a loyalty obligation, the data broker's legal duty runs to its shareholders and customers (data buyers), not to the people whose data it trades.
References
Balkin, "Information Fiduciaries and the First Amendment" (2016); proposed Data Care Act; FCRA permissible purpose framework; FTC unfairness doctrine analysis; academic proposals for information fiduciary legislation.
8Data Broker Opacity and Corporate Structure Obfuscation▾
Problem
Data brokers deliberately obscure their corporate identities, ownership structures, and data practices through holding companies, subsidiaries, frequent name changes, and corporate restructuring. Acxiom rebranded to LiveRamp. X-Mode Social became Outlogic. Exact Data became Stirista. Near Intelligence went through bankruptcy. Oracle shut down its advertising data division (Oracle Data Cloud/BlueKai/AddThis/Moat) in 2024 but the data assets were redistributed. Consumers attempting to exercise privacy rights cannot determine which corporate entity holds their data, which entity to send opt-out requests to, or which entity is responsible for data practices.
Current State
No law requires data brokers to maintain consistent corporate identities, disclose subsidiary relationships, or inform consumers when corporate restructuring affects their data. Merger and acquisition activity in the data broker space is frequent, with data assets transferring between entities without consumer notification. Bankruptcy proceedings (like Near Intelligence's 2023 Chapter 11 filing) raise questions about whether personal data is a corporate asset that can be sold to satisfy creditors. The FTC has limited authority to track data through corporate transformations.
Impact
A consumer who successfully opts out of Acxiom discovers their data persists in LiveRamp (which is what Acxiom rebranded its data connectivity business to). A consumer who opted out of X-Mode discovers Outlogic (same company, new name) has the same data. Corporate opacity makes individual privacy rights unexercisable because the target of those rights keeps changing identity. The Near Intelligence bankruptcy revealed that the company had amassed location data on over a billion devices, and this data became an asset in bankruptcy proceedings.
References
FTC comments on data broker transparency; Near Intelligence Chapter 11 filing and data asset disposition; Acxiom/LiveRamp corporate restructuring; Oracle Data Cloud shutdown (June 2024); corporate genealogy of major data broker entities.
9Children's Data Broker Economy Persists Despite COPPA▾
Problem
COPPA prohibits the collection of personal information from children under 13 without verifiable parental consent, but data brokers routinely hold and sell data on children through indirect collection channels. Children's data enters broker databases through family profiles (parent-child household inference), school records sold by EdTech companies, app SDK data collected from devices used by children, and public records (birth announcements, sports league registrations). The FTC has increased COPPA enforcement but cannot address data broker acquisition of children's data through indirect channels.
Current State
The FTC fined Epic Games $275 million in 2022 for COPPA violations related to Fortnite's data collection from children. The proposed COPPA 2.0 (Kids Online Safety Act, or KOSA, and Children and Teens' Online Privacy Protection Act) would extend protections to teens aged 13-16 and restrict targeted advertising to minors. However, these bills address direct collection by online services, not the secondary data broker market where children's data is packaged and sold as part of family-level profiles. Data brokers like Acxiom/LiveRamp, Experian, and Epsilon maintain household-level databases where individual opt-outs create incomplete household records but do not erase the individual from relational connections.
Impact
Children's data in broker databases follows them into adulthood, creating pre-existing digital profiles before individuals are old enough to understand or consent to data collection. Credit bureaus have reported cases of children having credit files created through identity theft facilitated by broker data. Data broker profiles of children have been used to target advertising for age-inappropriate products to minors.
References
FTC v. Epic Games COPPA enforcement (2022); COPPA 2.0 and KOSA legislation; Acxiom household segmentation product documentation; FTC reports on children's online privacy; Common Sense Media data broker analysis.
10First Amendment Weaponization Against Privacy Regulation▾
Problem
The data broker industry argues that the collection, aggregation, and sale of personal data constitutes protected speech under the First Amendment. In multiple legal challenges, industry groups have argued that data is speech, data processing is expression, and privacy regulations that restrict data flows are content-based restrictions subject to strict scrutiny. The Supreme Court's decision in Sorrell v. IMS Health (2011) struck down a Vermont law restricting the sale of pharmacy prescriber data, finding that data sales restrictions were subject to heightened First Amendment scrutiny.
Current State
The Sorrell precedent casts a shadow over all data broker regulation. After Sorrell, any law that singles out data sales for restriction must survive "heightened scrutiny" — a standard that favors data brokers' commercial interests over individual privacy. Industry trade groups (NetChoice, Computer & Communications Industry Association, US Chamber of Commerce) routinely cite the First Amendment in opposing privacy legislation and challenging state privacy laws. The ADPPA's failure to pass was partly due to concerns that it could face First Amendment challenges. Courts have not definitively resolved whether comprehensive privacy regulation can survive Sorrell-level scrutiny.
Impact
The First Amendment argument creates a constitutional shield for an industry that profits from the collection and sale of information about unconsenting individuals. The framing of commercial data trafficking as protected speech elevates corporate interests above individual privacy in ways the Constitution's framers could not have anticipated. Privacy advocates argue that Sorrell was wrongly decided and that commercial data transactions are conduct, not speech, but this argument has not been adopted by the Supreme Court.
References
Sorrell v. IMS Health Inc., 564 U.S. 552 (2011); NetChoice and CCIA legal challenges to state privacy laws; Balkin "Information Fiduciaries" First Amendment analysis; academic analysis of data-as-speech doctrine; industry amicus briefs in privacy law challenges.
Browser fingerprinting creates a unique identifier for each user by combining dozens of browser and device attributes: screen resolution, installed fonts, WebGL rendering characteristics, audio processing fingerprint, Canvas API output, timezone, language settings, hardware concurrency, and more. Unlike cookies, fingerprints cannot be deleted, blocked through browser settings, or controlled through consent mechanisms. The EFF's Panopticlick (now Cover Your Tracks) project demonstrated that 83.6% of browsers have a unique fingerprint, rising to 94.2% when Flash or Java is enabled. Fingerprinting makes cookie consent banners irrelevant because tracking persists regardless of consent choices.
Current State
FingerprintJS (now Fingerprint.com), a commercial fingerprinting company, serves billions of API calls monthly and markets 99.5% visitor identification accuracy. The company positions fingerprinting as a "fraud detection" tool, but the same technology enables persistent tracking. Major advertising networks use fingerprinting as a fallback when cookies are blocked or consent is denied. The W3C's Privacy Community Group has proposed mitigations, but browser vendors have implemented them inconsistently. Firefox's Enhanced Tracking Protection blocks some known fingerprinting scripts, but the technique evolves faster than blocklists. GDPR and ePrivacy Directive technically cover fingerprinting (as it creates a "unique identifier"), but enforcement is nearly nonexistent.
Impact
Users who carefully manage cookies, enable tracking protection, and deny consent on cookie banners are still tracked through fingerprinting. The Tor Browser is one of the few browsers that effectively resists fingerprinting (by making all users look identical), but its usability tradeoffs make it impractical for daily use. The fingerprinting industry has grown to serve as a complete cookie replacement, rendering the entire consent infrastructure of GDPR's cookie regime performative.
References
EFF Cover Your Tracks project; AmIUnique.org research dataset; Fingerprint.com documentation; Laperdrix et al. "Browser Fingerprinting: A Survey" (2020); ENISA fingerprinting analysis; W3C Privacy Community Group fingerprinting mitigations.
2Probabilistic Cross-Device Identity Matching▾
Problem
Data brokers and AdTech companies use probabilistic algorithms to link devices belonging to the same person without any explicit identifier. By analyzing patterns — devices on the same WiFi network, at the same GPS location, used at the same times, visiting the same websites — companies like Tapad (acquired by Experian), Drawbridge (acquired by LinkedIn/Microsoft), Oracle Data Cloud, and LiveRamp build "device graphs" that link smartphones, tablets, laptops, smart TVs, and IoT devices to individual identity profiles. These probabilistic links operate without user knowledge, consent, or any opt-out mechanism.
Current State
Cross-device identity resolution is a $4+ billion market segment. Tapad's device graph claims to connect over 3 billion devices globally. LiveRamp's IdentityLink connects offline identity to online devices through deterministic (email-based) and probabilistic (behavioral) matching. The NAI (Network Advertising Initiative) and DAA (Digital Advertising Alliance) self-regulatory programs nominally cover cross-device tracking, but their opt-out mechanisms are device-specific — opting out on your phone does not affect your laptop's cross-device profile. No privacy law specifically addresses probabilistic device linking.
Impact
An individual who maintains separate devices for work and personal use, uses different browsers, and avoids logging into the same accounts discovers that their devices have been linked anyway through behavioral patterns. A person researching sensitive medical conditions on their personal laptop finds related advertising appearing on their work computer and shared family tablet, revealing private information to coworkers and family members. The probabilistic nature of the matching means errors link unrelated individuals' devices, creating phantom profiles that combine strangers' data.
3Email-Based Identity Graphs and Unified ID Systems▾
Problem
The advertising industry has built identity systems that use hashed email addresses as persistent cross-platform identifiers, replacing third-party cookies as the backbone of online tracking. The Trade Desk's Unified ID 2.0 (UID2), LiveRamp's RampID (formerly IdentityLink), and ID5 all create encrypted but deterministic identifiers from email addresses. Because users provide email addresses to log into most online services, these systems create a universal tracking identifier that persists across websites, apps, and devices — with the user's "consent" obtained through login screens that bury tracking permissions in terms of service.
Current State
UID2 has been adopted by hundreds of publishers, advertisers, and AdTech platforms as a cookie replacement. The system claims to be "privacy-conscious" because email addresses are hashed (using SHA-256), but hashing is not anonymization — the same email always produces the same hash, creating a deterministic link. Apple's iCloud Private Relay and Hide My Email features partially disrupt email-based tracking, but only for Apple users who activate these features. Google's Privacy Sandbox proposals do not address email-based identity systems. No privacy law specifically regulates the use of hashed emails as cross-platform identifiers.
Impact
Email-based identity systems are more persistent and harder to evade than cookies. A user can delete cookies, but they cannot change their email address without significant disruption to their digital life. Every site login becomes a tracking event. The system creates a comprehensive cross-platform activity log tied to a single identifier that follows the user across the web, apps, connected TV, and offline purchases (through loyalty programs linked to the same email). Users who provide their email to access an article or create an account unknowingly enable cross-platform surveillance.
References
The Trade Desk UID2 documentation and adoption metrics; LiveRamp RampID technical specifications; IAB Tech Lab identity framework; Apple Hide My Email and iCloud Private Relay documentation; privacy analyses of hashed-email identity systems.
4Connected TV and Streaming Platform Surveillance▾
Problem
Smart TVs and streaming devices (Roku, Amazon Fire TV, Apple TV, Chromecast) collect detailed viewing data including what content is watched, when, for how long, and which ads are viewed. This data is sold to advertisers and data brokers through Automatic Content Recognition (ACR) technology, which identifies content on screen by matching audio or visual fingerprints against a reference database. ACR operates even when users watch over-the-air broadcast TV, cable, or content from external devices — the TV itself is surveilling what appears on its screen regardless of the source.
Current State
Vizio paid $17 million in 2017 to settle FTC and New Jersey AG charges that it collected viewing data from 11 million smart TVs without adequate disclosure or consent. Despite this precedent, ACR remains standard on smart TVs from Samsung, LG, Vizio, and others, with consent buried in initial setup flows that most users click through. Samba TV, iSpot.tv, and Inscape (Vizio's data subsidiary) monetize viewing data from tens of millions of TVs. Roku's platform business (advertising) generates more revenue than hardware sales, making every Roku TV a surveillance device subsidized by advertising revenue. Amazon Fire TV integrates viewing data with Amazon's broader shopping and device ecosystem.
Impact
The living room television — traditionally a passive entertainment device — has become an always-on surveillance sensor. ACR captures viewing habits with second-by-second granularity, revealing political preferences (news channels), health concerns (medical show content), financial status (financial programming), and personal interests. This viewing data, linked to household identity through IP address and device registration, is integrated into data broker profiles and used for targeted advertising across all platforms.
References
FTC v. Vizio settlement (2017); Samba TV privacy analysis; Roku privacy policy and advertising business model; Samsung Smart TV privacy controversy; Consumer Reports smart TV tracking investigation; iSpot.tv and Inscape data products documentation.
5Ultrasonic Cross-Device Beacons▾
Problem
Ultrasonic beacons embed inaudible sound signals in television commercials, radio ads, web pages, and retail environments that are picked up by microphones in smartphones and other devices. These beacons create a covert cross-device and cross-environment link: a TV ad containing an ultrasonic beacon is picked up by a nearby phone, linking the TV viewing to the phone's identity. Retail stores use ultrasonic beacons to track in-store movement and link it to mobile device identifiers. The technology operates entirely without user awareness — the signals are inaudible, and the SDK processing them runs in the background.
Current State
Research by Mavroudis et al. (2017) at University College London identified ultrasonic tracking in 234 Android apps from the Google Play Store, with beacons found in retail locations in European cities. The SilverPush SDK was one of the most prominent ultrasonic tracking platforms before public exposure led to FTC warnings in 2016. While SilverPush claimed to discontinue the practice, the underlying technology persists in less visible forms. Shopkick, Lisnr, and Signal360 have used ultrasonic or near-ultrasonic signals for proximity detection. Android and iOS have tightened microphone permissions, but apps with legitimate microphone access (voice assistants, communication apps) can still process ultrasonic signals.
Impact
Ultrasonic beacons create a tracking channel that users cannot detect, block, or opt out of without revoking all microphone permissions from all apps. A user watching a TV commercial in their living room has their phone covertly identify the specific ad, the time of viewing, and the viewing location — linking their TV consumption to their mobile identity without any visible interaction. The technology can also de-anonymize Tor users by linking their anonymous browsing to their physical location through ambient ultrasonic signals.
References
Mavroudis et al. "On the Privacy and Security of the Ultrasound Ecosystem" (PETS 2017); FTC warning letter to SilverPush (2016); Arp et al. "Privacy Threats through Ultrasonic Side Channels on Mobile Devices" (IEEE EuroS&P 2017); Shopkick ultrasonic beacon patents; Android/iOS microphone permission evolution.
6Retail and In-Store WiFi and Bluetooth Tracking▾
Problem
Retailers and shopping centers track shoppers' physical movements through WiFi probe requests and Bluetooth beacons emitted by smartphones. When a phone searches for WiFi networks, it broadcasts its MAC address, which can be captured by sensors throughout a retail environment to track movement patterns, dwell times, and store visits. Bluetooth Low Energy (BLE) beacons placed throughout stores interact with retail apps to track precise indoor positioning. Companies like RetailNext, Euclid Analytics (acquired by Aruba/HPE), Shopperception, and InMarket aggregate this data across retail locations.
Current State
Apple and Google have implemented MAC address randomization in iOS 14+ and Android 10+ to mitigate WiFi tracking, but research shows that randomization is imperfect — devices often reveal their real MAC address when connecting to known networks, and behavioral patterns (movement sequences, timing) can re-link randomized addresses to individuals. Bluetooth tracking continues to be effective through retail apps that request Bluetooth permissions. InMarket, which operates a location data platform through SDK integrations in popular apps, was subject to an FTC enforcement action in March 2024 for collecting location data without adequate consent.
Impact
Shoppers are tracked throughout malls and retail environments without their knowledge. Movement data reveals store preferences, shopping duration, product interest (based on department-level positioning), and visit frequency. This physical-world tracking data is linked to online identity through app SDKs and sold to advertisers, landlords, and investment firms. Hedge funds have used retail foot traffic data from tracking platforms like Placer.ai and SafeGraph to make trading decisions based on store visit trends — profiting from physical-world surveillance of unwitting shoppers.
References
FTC v. InMarket proposed order (March 2024); RetailNext and Euclid Analytics product documentation; MAC address randomization effectiveness research; Vanhoef "Why MAC Address Randomization is not Enough" (2016); Placer.ai and SafeGraph retail analytics products; hedge fund use of location data reporting.
7Mobile Advertising ID Tracking Ecosystem▾
Problem
Every smartphone has a mobile advertising identifier — Apple's IDFA (Identifier for Advertisers) and Google's AAID/GAID (Google Advertising ID) — that serves as a persistent tracking beacon for the app ecosystem. Apps embed SDKs from data brokers (formerly X-Mode, Kochava, SafeGraph, Placer.ai, Foursquare/Factual) that collect the MAID along with GPS location, app usage, and device data. This creates a continuous stream of timestamped location data linked to a persistent identifier, which is aggregated and sold. The MAID ecosystem has been described as "the largest mass surveillance system ever built" by privacy researchers.
Current State
Apple's App Tracking Transparency (ATT) framework, introduced in iOS 14.5 (2021), requires apps to ask permission before accessing the IDFA. Approximately 75-80% of users opt out when asked, dramatically reducing IDFA availability on iOS. However, the location data ecosystem has adapted: apps collect location through alternative permissions (imprecise location, WiFi-based positioning), and data brokers use probabilistic methods to link data without the IDFA. Google announced similar AAID restrictions but has implemented them more gradually. Android still provides the GAID by default, and the Android user base (75% global market share) remains largely trackable.
Impact
Despite Apple's ATT intervention, the mobile advertising ID ecosystem continues to function. A 2022 investigation by The Markup found that data broker SafeGraph was selling location data derived from apps used by people visiting Planned Parenthood clinics, including visit duration and routes taken. Kochava's data was used to track visits to addiction recovery centers, mental health facilities, and houses of worship. The data is available for purchase by anyone — advertisers, insurance companies, bail bond agencies, or stalkers — with no verification of buyer intent.
References
The Markup "How We Built a Tool to Track the Location Data Industry" series; FTC v. Kochava complaint; Apple ATT documentation and opt-out rate data; Narseo Vallina-Rodriguez et al. "Are These Ads For You?" (CCS 2019); SafeGraph/Placer.ai data products; MAID ecosystem mapping by Wolfie Christl/Cracked Labs.
8Smart Speaker and Voice Assistant Surveillance▾
Problem
Smart speakers (Amazon Echo/Alexa, Google Home/Nest, Apple HomePod) are always-listening devices that process voice commands through cloud services. While manufacturers claim devices only record after hearing a wake word, investigations have revealed that recordings are frequently triggered by false wake-word detections. Amazon, Google, and Apple employ human reviewers to listen to recordings for quality improvement. Voice data reveals household composition, daily routines, health conditions (spoken symptoms), relationship dynamics, and private conversations. This data is integrated into each company's broader advertising and data ecosystem.
Current State
Bloomberg revealed in 2019 that Amazon employs thousands of workers worldwide who listen to Alexa recordings, including recordings made without intentional activation. Google and Apple made similar admissions. All three companies have since added opt-out options for human review, but continue cloud processing of all voice commands. Amazon's Alexa division lost $10 billion in 2022, suggesting the business model depends on data value rather than hardware margins. Amazon's "Alexa Hunches" feature proactively monitors household patterns. Ring and other Amazon smart home devices create additional data streams that complement voice data.
Impact
A household with a smart speaker has effectively installed a corporate-operated listening device. False activations capture private conversations, arguments, medical discussions, and financial deliberations. A 2020 study by Northeastern University and Imperial College London found that smart speakers activated without the wake word between 1.5 and 19 times per day, recording up to 43 seconds of audio per false activation. Voice biometrics derived from smart speaker data can identify individual household members, creating per-person profiles within a shared device.
References
Bloomberg "Amazon Workers Are Listening to What You Tell Alexa" (2019); Edu et al. "Hey Alexa, Is This Skill Safe?" (NDSS 2020); Choffnes et al. smart speaker false activation study (2020); Amazon Alexa privacy settings documentation; Google Assistant data handling disclosures; Apple Siri quality review controversy.
9Connected Vehicle Data Collection and Sale▾
Problem
Modern vehicles collect massive amounts of data: GPS location (continuous), driving patterns, speed, braking, destinations, cabin conversations (through hands-free systems), paired phone contacts, text messages read aloud, music preferences, and biometric data (seat position, weight). Automakers including GM (OnStar), Toyota, Honda, Ford, and Hyundai have been found selling or sharing this data with data brokers, insurance companies, and advertisers. A Mozilla Foundation study found that cars are "the worst category of products for privacy" with 25 of 25 car brands collecting more data than needed.
Current State
A March 2024 New York Times investigation revealed that GM's OnStar Smart Driver program collected detailed driving data and shared it with LexisNexis Risk Solutions, which in turn sold "driving behavior" scores to insurance companies. This affected millions of drivers who did not knowingly consent to insurance-relevant data sharing. Senator Wyden's office documented that automakers sell location data to data brokers who aggregate it with other data sources. Verisk, LexisNexis Risk Solutions, and other insurance-adjacent data companies purchase driving data from automakers and offer risk-scoring services to insurers.
Impact
Drivers have seen insurance premiums increase by 20-30% based on driving behavior data they did not know was being collected or sold. The NYT investigation found GM customers whose data was shared with LexisNexis faced insurance rate increases or policy non-renewals. Connected vehicles effectively make every road trip a surveillance event, with the route, speed, duration, and destination recorded and available for commercial exploitation. Reproductive rights advocates have raised concerns about vehicle location data tracking visits to healthcare facilities.
References
Mozilla Foundation "Privacy Not Included: Cars" study (2023); NYT "Automakers Are Sharing Consumers' Driving Behavior With Insurance Companies" (March 2024); Sen. Wyden connected vehicle data investigation; GM OnStar Smart Driver data sharing controversy; Verisk and LexisNexis driving data products; The Markup vehicle tracking investigations.
10IoT and Smart Home Device Data Aggregation▾
Problem
The proliferation of Internet of Things (IoT) devices — smart thermostats (Nest/Google, Ecobee), smart light bulbs (Philips Hue, LIFX), robot vacuums (iRobot Roomba, Roborock), fitness trackers (Fitbit, Garmin), smart scales (Withings), sleep trackers (Oura), and hundreds of other categories — creates an intimate data layer about daily life. Each device category reveals specific behavioral patterns: thermostat data shows when you are home, light patterns reveal sleep schedules, robot vacuum maps reveal home layouts, fitness data reveals health status. This data is aggregated through smart home platforms (Google Home, Amazon Alexa, Samsung SmartThings) that serve as central collection points.
Current State
iRobot's proposed acquisition by Amazon (announced 2022, abandoned 2024 after regulatory concerns) highlighted the value of home mapping data — Roomba vacuums create detailed floor plans of users' homes. Amazon already possesses data from Ring cameras (exterior surveillance), Echo speakers (audio), and Alexa-connected devices, and the iRobot acquisition would have added interior home layouts. Google's acquisition of Fitbit (completed 2021) combined health and fitness data with Google's existing behavioral profile. The FTC imposed conditions on the Fitbit acquisition but enforcement of those conditions relies on self-reporting. No comprehensive IoT privacy regulation exists in the US.
Impact
The aggregation of IoT data creates a comprehensive behavioral model of daily life: when you wake up (sleep tracker, lights, thermostat), your health status (fitness tracker, smart scale), who is home (motion sensors, camera), what you eat (smart fridge, grocery delivery), and your stress levels (heart rate variability data). A single data breach or sale exposes the entire pattern of daily existence. Insurance companies have expressed interest in smart home data for underwriting decisions, and the lack of regulation means this data can flow freely to any buyer.
References
iRobot/Amazon proposed acquisition and FTC scrutiny; Google/Fitbit acquisition FTC conditions; Mozilla IoT privacy analysis; Apthorpe et al. "A Smart Home is No Castle" (2017); ENISA IoT security and privacy guidelines; r/privacy and r/degoogle community discussions on smart home surveillance.
9. Opt-Out Mechanism FailuresHigh
1Impossible Scale of Individual Broker Opt-Outs▾
Problem
Privacy rights organizations estimate there are 4,000+ data brokers operating in the US alone. Exercising opt-out rights requires an individual to identify each broker that holds their data, navigate each broker's unique opt-out process, verify their identity (often requiring submission of additional personal data), and monitor for re-inclusion. At an average of 15-30 minutes per broker (including research, form completion, identity verification, and follow-up), opting out of all known brokers would require 1,000-2,000 hours of labor per person — a task that must be repeated regularly as data reappears.
Current State
The Privacy Rights Clearinghouse maintains a database of approximately 500 data brokers with opt-out links, but this represents a fraction of the industry. Each broker has different opt-out procedures: some require email, some require postal mail, some require notarized identity documents, some require creating an account (providing additional data to opt out of data collection), and some have no opt-out mechanism at all. There is no central registry of all data brokers, no standardized opt-out protocol, and no legal requirement that opt-outs be easy or effective. Vermont's registry lists ~500 brokers; California's Delete Act aims to create a universal deletion mechanism but implementation is still underway.
Impact
The opt-out burden falls entirely on individuals who have the least information (which brokers have their data), the least power (no enforcement mechanism), and the least time (the process is extraordinarily labor-intensive). Privacy-conscious individuals who invest dozens of hours opting out discover they have addressed perhaps 10-15% of brokers holding their data. The system is designed to fail at scale: it works for the rare individual willing to make privacy a full-time project but provides no meaningful protection for the general population.
References
Privacy Rights Clearinghouse data broker database; California Delete Act (SB 362) implementation timeline; Vermont data broker registry; r/privacy opt-out experience threads; EFF guide to data broker opt-outs; Yael Grauer's Big Ass Data Broker Opt-Out List.
2Identity Verification Paradox in Opt-Out Processes▾
Problem
To opt out of a data broker's database, the individual must prove their identity — which typically requires providing the very personal information they are trying to remove. Brokers require some combination of full legal name, date of birth, current and former addresses, email addresses, phone numbers, and sometimes government ID or notarized documents. This creates a paradox: the opt-out process itself feeds more personal data to the broker and confirms the accuracy of data they already hold. Some brokers use the identity verification data to update and enrich their records.
Current State
There is no standardized privacy-preserving identity verification protocol for opt-out requests. Brokers set their own verification requirements, and some deliberately make them burdensome to discourage opt-outs. Whitepages requires an account creation (with email and phone verification) to process a removal request. Spokeo requires an email address and the specific URL of the listing to be removed. Some brokers require a photo of government-issued ID. No regulator has mandated that opt-out verification be proportionate to the data being removed or that verification data cannot be retained or used for other purposes.
Impact
Individuals who attempt to opt out of one broker may find their data appearing at new brokers — because the opt-out verification data has been processed, sold, or used to confirm records at affiliated entities. A person who provides their current address to opt out of Spokeo may find that address appearing at BeenVerified, Whitepages, and Intelius within weeks. The opt-out process itself becomes a data collection event, defeating its purpose.
References
Privacy Guides community discussions on opt-out data harvesting; Hacker News threads on broker opt-out paradoxes; Consumer Reports "What Happens When You Try to Delete Your Data" investigation; CCPA opt-out implementation analysis; noyb complaints about excessive identity verification in GDPR deletion requests.
3Automated Removal Services Have Limited Effectiveness▾
Problem
Commercial data removal services — DeleteMe (by Abine), Privacy Duck, Kanary, Optery, EasyOptOuts, and others — automate the process of opting out from data broker databases. These services charge $100-300/year and process opt-outs from 50-200+ brokers. However, they cover only a fraction of the 4,000+ brokers, they can only process opt-outs where a public-facing mechanism exists, they have no legal authority to compel compliance, and their effectiveness varies dramatically by broker. Testing by Consumer Reports and privacy researchers shows removal rates of 30-70% across targeted brokers, with data frequently reappearing within 3-6 months.
Current State
DeleteMe (the largest service, with claimed 100,000+ subscribers) processes removals from approximately 750+ data broker sites. Optery covers a similar range with a different methodology. Consumer Reports' Permission Slip app attempted to automate CCPA data deletion requests. Independent testing by journalists and privacy researchers consistently finds that no service achieves complete removal: some brokers ignore automated requests, others re-ingest data from public records, and many simply do not have automatable opt-out processes. The services also cannot address data held by brokers that sell only to businesses (B2B data brokers) with no consumer-facing presence.
Impact
Users of removal services experience a false sense of protection. They pay $129-299/year believing their data is being removed, but 30-50% of their data broker presence persists. The services' quarterly re-scan cycles mean data can exist in broker databases for months between checks. Users who cancel their subscription find their data reappears at previously cleaned brokers within weeks. The removal service market itself creates a perverse incentive: these companies profit from the data broker ecosystem's continued existence.
References
Consumer Reports removal service testing; Privacy Duck vs. DeleteMe comparison analyses; Optery effectiveness documentation; Yael Grauer evaluation of data removal services; r/privacy threads on DeleteMe experiences; CNET "Do Data Removal Services Actually Work?" analysis.
4Data Re-Ingestion After Successful Opt-Out▾
Problem
Even when a data broker successfully processes an opt-out request and removes an individual's data, the data typically reappears within 1-6 months because brokers continuously ingest new data from public records, commercial data exchanges, partner data sharing agreements, and scraping operations. An opt-out removes a single record at a single point in time but does not prevent the broker from re-collecting the same data from its sources. No opt-out creates a permanent prohibition on future collection of that individual's data.
Current State
CCPA/CPRA's "Do Not Sell" right creates an ongoing obligation, but it applies only to data sales, not collection or aggregation. The California Delete Act (SB 362) is intended to create a single deletion mechanism with ongoing effect, but implementation is still in progress and the mechanism's ability to prevent re-collection is legally untested. Most data brokers outside California have no legal obligation to maintain opt-out status permanently. Some brokers explicitly state in their privacy policies that opt-outs apply only to data currently held and do not prevent future collection.
Impact
Individuals who invest hours opting out discover their data reappearing months later, requiring the entire process to be repeated indefinitely. This Sisyphean dynamic is a feature, not a bug: data brokers' business models depend on comprehensive coverage, and permanently honoring opt-outs would create growing gaps in their databases. The re-ingestion cycle transforms opt-out from a one-time action into a perpetual maintenance obligation that most individuals cannot sustain.
References
CCPA "Do Not Sell" right implementation analysis; California Delete Act re-collection provisions; consumer complaints to California AG about data reappearance; Privacy Guides forum threads on opt-out persistence; DeleteMe re-scan findings; Spokeo/BeenVerified data re-ingestion patterns.
5Dark Patterns in Opt-Out User Interfaces▾
Problem
Data brokers deliberately design opt-out processes to discourage completion through dark patterns: multi-step processes that reset if the browser is closed, CAPTCHAs that fail repeatedly, confirmation emails that arrive hours later (or not at all), opt-out pages that are not linked from the main website, forms that require information the user cannot easily provide, processing times of 30-45 days, and confirmatory "are you sure?" interruptions. These design choices are not accidental — they exploit behavioral economics to minimize the number of users who successfully complete the opt-out process.
Current State
The FTC has identified dark patterns in opt-out processes as a priority enforcement area, and the CPRA specifically requires that "the process for submitting a request to opt-out shall not require the consumer to provide more information than necessary." However, enforcement is complaint-driven and slow. noyb (the European privacy organization led by Max Schrems) has filed complaints against cookie consent dark patterns, establishing precedents that could apply to opt-out processes. The California Privacy Protection Agency has begun rulemaking on opt-out process requirements, but rules are not yet finalized.
Impact
Behavioral research shows that each additional step in an opt-out process reduces completion rates by 20-40%. A broker with a 6-step opt-out process that includes email verification, CAPTCHA, and a 10-day waiting period will see 90-95% of opt-out attempts abandoned before completion. The brokers with the most data (and therefore the most to lose from opt-outs) invest the most in designing difficult opt-out processes. Users who abandon opt-out attempts believe they have "tried" to exercise their rights but were defeated by the process.
References
FTC dark patterns report (2022); CPRA opt-out process requirements; noyb cookie consent complaints; Mathur et al. "Dark Patterns at Scale" (CHI 2019); California Privacy Protection Agency rulemaking proceedings; Harry Brignull darkpatterns.org documentation.
6No Universal Opt-Out Mechanism Exists▾
Problem
Despite years of advocacy, no functioning universal opt-out mechanism covers the data broker industry. The Global Privacy Control (GPC) signal, recognized by CCPA/CPRA, communicates opt-out preferences via browser headers, but it only applies to websites the user visits and does not reach data brokers that have no direct consumer interaction. California's Delete Act (SB 362) mandates a universal deletion mechanism, but it is still being built and applies only to brokers registered in California. The NAI and DAA opt-out tools cover advertising networks but not data brokers. No mechanism allows a single action to opt out of all data broker data collection, sale, and sharing.
Current State
GPC is supported by Firefox, Brave, and DuckDuckGo browsers and is legally binding under CCPA/CPRA, but compliance among websites is spotty and the signal does not reach the data broker layer. The California Delete Act requires the CPPA to establish a universal deletion mechanism by January 2026, but implementation has been delayed and the mechanism's technical architecture is still being finalized. The mechanism will apply only to registered California data brokers, leaving thousands of non-registered and out-of-state brokers unaffected. The Do Not Track (DNT) header, proposed in 2009, was abandoned as a standard after industry refused to honor it.
Impact
An individual who enables GPC in their browser, uses the DAA opt-out tool, submits removal requests to the top 50 brokers, and subscribes to a deletion service has still only addressed a fraction of their data broker exposure. There is no equivalent of the Do Not Call Registry for data brokers — no single action that communicates "stop collecting, selling, and sharing my data" to the entire industry. Each opt-out mechanism covers a different slice of the ecosystem, and the gaps between them are where most data broker activity occurs.
References
Global Privacy Control specification; CCPA/CPRA GPC recognition; California Delete Act (SB 362) implementation status; Do Not Track header history and abandonment; DAA WebChoices and AppChoices tools; NAI opt-out page; Privacy Guides GPC discussion.
7Opt-Out Does Not Equal Deletion▾
Problem
Most data broker opt-out processes suppress data from public-facing search results but do not delete the underlying data from the broker's databases. The broker retains the data for internal use, re-sale to business customers, analytics, and model training. "Opting out" of Spokeo removes your listing from spokeo.com but does not delete your data from Spokeo's underlying database or prevent it from being sold through Spokeo's enterprise API. The distinction between suppression and deletion is not disclosed to consumers and is not addressed by most privacy laws.
Current State
CCPA/CPRA provides a "right to delete" that is stronger than mere suppression, but brokers argue that data obtained from public records is exempt from deletion requirements under the "publicly available information" exception. GDPR's "right to erasure" is more comprehensive but faces enforcement challenges with US-based brokers. People-search sites typically offer "suppression" (removing the listing from search results) rather than "deletion" (removing the data entirely). The technical difference is invisible to consumers but critical to privacy outcomes.
Impact
A domestic violence survivor who opts out of Whitepages sees her listing removed from the website but her data remains in Whitepages' enterprise database, accessible to institutional customers including skip-tracing services used by debt collectors and, potentially, her abuser working through a private investigator. The suppression-not-deletion model means that "opting out" creates an illusion of privacy while the data continues to circulate through commercial channels invisible to the consumer.
References
Whitepages/Spokeo enterprise API documentation; CCPA right to delete vs. right to opt-out distinction; GDPR right to erasure implementation; National Network to End Domestic Violence data broker safety planning; investigative reporting on people-search site data retention after opt-out.
8Household and Relational Data Persistence▾
Problem
Even if an individual successfully removes their own data from data brokers, their information persists through household associations, family relationships, and social connections in other people's records. A person who opts out of all brokers can be re-identified through their spouse's record (which lists household members), their adult children's records (which list parents), their property records (which list co-owners), and their social media connections' data. Data brokers build relationship graphs that make individual opt-outs ineffective because identity can be reconstructed from surrounding connections.
Current State
No opt-out mechanism addresses household or relational data. An individual can request deletion of their own record but cannot compel deletion of references to themselves in other people's records. Acxiom/LiveRamp, Experian, and TransUnion all maintain household-level databases where individual opt-outs create incomplete household records but do not erase the individual from relational connections. People-search sites list "known associates" and "possible relatives" — information derived from address co-residency, shared last names, and social network analysis — that persists even after the individual's own record is removed.
Impact
A person in witness protection who meticulously opts out of all data brokers can be located through their relative's BeenVerified listing, which shows "possible relatives" at "previous addresses." A domestic violence survivor who removes their data from people-search sites can be found through their ex-partner's record, which still lists the survivor as a "known associate." Individual opt-out rights are structurally incapable of addressing the relational nature of data broker profiles.
References
NNEDV safety planning guides for data broker exposure; Privacy Rights Clearinghouse household data analysis; Acxiom/LiveRamp household segmentation products; people-search site "known associates" feature analysis; academic research on re-identification through social connections.
9Mobile and App-Level Opt-Outs Do Not Propagate▾
Problem
Opting out of tracking on a mobile device (resetting advertising ID, revoking app permissions, enabling Apple's ATT opt-out) does not propagate to data brokers that have already collected the data. Historical location data, behavioral profiles, and device graphs built from previously collected MAID data persist in broker databases indefinitely. The opt-out prevents future collection from that specific app/device combination but does not address the years of already-collected data. Additionally, many app SDKs circumvent mobile-level opt-outs through server-side tracking, hashed identifiers, and probabilistic matching.
Current State
Apple's ATT requires apps to ask permission before tracking, but data already collected before ATT was enabled (pre-iOS 14.5) remains in broker databases. Google's GAID restrictions allow users to delete their advertising ID, but brokers retain historical data linked to the old ID. App SDK providers (Kochava, Adjust, AppsFlyer, Branch) have developed server-side attribution methods that circumvent client-side opt-outs. The FTC's action against X-Mode/Outlogic required deletion of previously collected data, but this was an extraordinary enforcement action, not a general requirement.
Impact
A user who resets their advertising ID today has no effect on the 3-5 years of location data already held by data brokers. That historical data — showing every place they have been, every store they have visited, every doctor they have seen — remains commercially available. The forward-looking nature of mobile opt-outs means they protect future privacy while leaving the past fully exposed. Data brokers maintain historical databases as their most valuable asset precisely because this data cannot be "opted out of" retroactively.
References
Apple ATT documentation and adoption timeline; Google GAID deletion feature; Kochava server-side attribution documentation; FTC v. X-Mode/Outlogic data deletion requirement; AppsFlyer and Adjust SDK documentation; Privacy Guides mobile tracking discussion.
10Deceased, Minor, and Vulnerable Population Opt-Out Gaps▾
Problem
Data brokers maintain records on deceased individuals, minors, incapacitated adults, and other populations that cannot exercise opt-out rights on their own behalf. Deceased individuals' records persist in databases indefinitely, enabling identity theft using dead people's information. Minor children have no legal capacity to submit opt-out requests, and parents may not know which brokers hold their children's data. Elderly individuals with diminished capacity cannot navigate complex opt-out processes. These populations represent systematic gaps in an opt-out model that assumes a competent adult can and will advocate for their own privacy.
Current State
No data broker proactively removes records of deceased individuals; estates must submit individual opt-out requests to each broker with death certificate documentation. COPPA restricts collection from children under 13 but provides no mechanism for parents to opt out of data already in broker databases. No law specifically addresses data broker obligations regarding incapacitated adults. The California Delete Act's universal mechanism is intended to allow authorized agents to submit requests on behalf of others, but the agent verification process is still being designed and may be impractical for estate executors, parents, and guardians.
Impact
The Social Security Death Master File is itself sold as a data product, and the gap between a person's death and the processing of their records across thousands of brokers creates a window for identity theft. The FTC has documented cases of tax refund fraud and credit account opening using deceased individuals' data obtained from broker databases. Children who reach adulthood discover pre-existing data broker profiles built from household data, school records, and app usage data, with no way to determine the provenance of the data or comprehensively delete it.
References
FTC identity theft reports involving deceased individuals; Social Security Death Master File access and sale; COPPA parental rights limitations; California Delete Act authorized agent provisions; AARP analysis of data broker exploitation of elderly populations; Privacy Rights Clearinghouse vulnerable populations guidance.
10. International Data Broker ArbitrageHigh
1EU Personal Data Export via Non-Adequate Countries▾
Problem
GDPR restricts transfers of EU personal data to countries without "adequate" data protection (adequacy decisions), requiring safeguards like Standard Contractual Clauses (SCCs) or Binding Corporate Rules (BCRs). Data brokers circumvent these restrictions by routing data through intermediary countries or corporate entities. An EU data broker subsidiary exports data to a holding company in Singapore, which transfers it to a processing facility in India, which makes it available to US purchasers. Each hop adds legal distance from the original GDPR obligation, and enforcement across multiple jurisdictions is practically impossible.
Current State
The Schrems II decision (CJEU, 2020) invalidated the EU-US Privacy Shield and imposed stricter requirements on SCCs, but the practical effect has been to increase creative compliance rather than stop data flows. The EU-US Data Privacy Framework (DPF), adopted in 2023, restored a legal basis for EU-US transfers but only for companies that self-certify with the US Department of Commerce. Data brokers that do not self-certify, or that route data through non-DPF channels, continue to transfer EU personal data to the US and other jurisdictions without adequate protection. noyb has filed multiple complaints challenging data transfers that rely on inadequate safeguards.
Impact
EU residents' data reaches US data brokers through chains of transfers that each appear individually compliant but collectively defeat GDPR's purpose. A German citizen's data collected through an app with a European subsidiary, processed by a contractor in India, and aggregated by a broker in the US is subject to German data protection law at collection but effectively unprotected by the time it reaches the US broker. The individual has no practical ability to trace their data through the transfer chain or exercise GDPR rights against entities in foreign jurisdictions.
References
CJEU Schrems II (Case C-311/18); EU-US Data Privacy Framework adequacy decision (2023); noyb data transfer complaints; EDPB guidance on supplementary measures for international transfers; Cracked Labs / Wolfie Christl analysis of AdTech data transfers.
2Regulatory Arbitrage Between US States▾
Problem
Data brokers strategically locate corporate entities, data processing infrastructure, and legal domicile to minimize exposure to state privacy laws. A broker incorporated in Wyoming or Delaware, with servers in Texas, processing data on California residents, faces a complex jurisdictional calculation. The broker may argue that CCPA applies only to "businesses that do business in California" and that its limited California nexus falls below applicability thresholds (annual revenue under $25 million, data on fewer than 100,000 California consumers, less than 50% of revenue from data sales). This interstate arbitrage exploits the fragmented regulatory landscape.
Current State
California's CPRA has the broadest applicability thresholds but can only enforce against companies with sufficient California nexus. Small and mid-size data brokers deliberately structure operations to fall below CCPA/CPRA thresholds while still processing California residents' data. States without privacy laws (including major economies like Pennsylvania, Ohio, and Michigan as of early 2026) serve as regulatory havens. The lack of federal preemption means brokers can exploit gaps between state laws indefinitely. Cross-state enforcement cooperation is limited and ad hoc.
Impact
Consumers in states without privacy laws have no data broker rights at all — no opt-out, no deletion, no access. Even consumers in states with privacy laws face brokers that have structured their operations to avoid applicability. The regulatory arbitrage dynamic means that the most protective state laws are undermined by the least protective states, creating a race to the bottom where brokers seek the most permissive jurisdiction. A national data broker can effectively choose which state's law applies by structuring its corporate presence.
References
CCPA applicability threshold analysis; state incorporation and privacy law nexus requirements; IAPP state privacy law comparison matrix; National Conference of State Legislatures privacy law tracker; industry compliance strategies for multi-state privacy law landscape.
3Offshore Data Processing and Server Location Exploitation▾
Problem
Data brokers process personal data in jurisdictions with minimal privacy regulation — including countries with no data protection law at all — to reduce compliance obligations and enforcement risk. Processing facilities in countries like Malaysia, Philippines, Vietnam, and various offshore jurisdictions handle personal data from US and EU residents. The physical location of servers determines which country's law enforcement has jurisdiction, and hosting data in a country with weak privacy laws or limited international cooperation makes enforcement of foreign privacy rights effectively impossible.
Current State
Major cloud infrastructure providers (AWS, Azure, GCP) offer server regions globally, making it trivial to process data in any jurisdiction. Data brokers use cloud regions in countries with favorable regulatory environments. Some brokers maintain servers in jurisdictions that do not respond to foreign regulatory inquiries or mutual legal assistance requests. The Budapest Convention on Cybercrime facilitates some cross-border data access, but privacy enforcement cooperation is far less developed than criminal cooperation. GDPR's territorial scope claims jurisdiction over processing of EU residents' data regardless of processor location, but enforcement against entities with no EU presence is impractical.
Impact
A US data broker processing data on EU servers is subject to GDPR enforcement, but the same broker processing the same data on Malaysian servers faces only Malaysian regulatory authority — which has limited resources and no obligation to enforce EU law. Individuals whose data is processed offshore have no practical recourse: they cannot identify where their data is processed, which country's law applies, or which regulator to complain to. The offshore processing model makes privacy rights paper promises that cannot be enforced across jurisdictions.
References
GDPR territorial scope (Article 3); Budapest Convention on Cybercrime; UNCTAD data protection legislation worldwide map; cloud infrastructure provider region availability; EDPS commentary on offshore data processing enforcement challenges.
4UK Post-Brexit Data Protection Divergence▾
Problem
Following Brexit, the UK enacted the UK GDPR (Data Protection Act 2018 as amended) which initially mirrored EU GDPR. However, the UK government has pursued regulatory divergence through the Data Protection and Digital Information Act (2024), which relaxes certain GDPR requirements including automated decision-making restrictions, legitimate interest assessments, and international transfer mechanisms. This creates an arbitrage opportunity: data brokers can establish UK operations to process EU data under the UK's EU adequacy decision, then benefit from the UK's more relaxed domestic rules for onward transfers and processing. If the EU revokes the UK's adequacy determination due to divergence, the resulting data transfer chaos will benefit brokers who have already moved data.
Current State
The EU granted the UK an adequacy decision in June 2021, enabling free data flow from the EU to the UK. However, this decision must be renewed and can be revoked if UK data protection standards diverge too far from GDPR. The UK's Data Protection and Digital Information Act introduced departures from EU GDPR that some EU commentators argue could jeopardize adequacy. Data brokers with UK operations can currently receive EU data freely and process it under the UK's increasingly distinct rules. The ICO (UK Information Commissioner's Office) has signaled a more "business-friendly" approach to data protection enforcement.
Impact
The UK risks becoming a data laundering jurisdiction — a GDPR-adequate country that processes EU data under progressively weaker standards. Data brokers establishing UK subsidiaries can receive EU data legally, process it under UK rules that may not provide equivalent protection, and potentially transfer it onward to non-adequate countries under the UK's more permissive international transfer rules. EU residents lose the protections they assumed GDPR provided when their data transits through the UK.
References
UK-EU adequacy decision (June 2021); UK Data Protection and Digital Information Act (2024); ICO regulatory approach statement; European Parliament assessment of UK adequacy; noyb analysis of UK GDPR divergence; EDPB commentary on UK adequacy sustainability.
5Israel as a Data Broker Jurisdiction▾
Problem
Israel has an EU adequacy decision (since 2011) that enables free data flow from the EU to Israel, combined with a domestic privacy law (Protection of Privacy Law, 1981) that is significantly less comprehensive than GDPR. Israel's thriving surveillance technology sector (NSO Group, Cellebrite, Cognyte, Cobwebs, Candiru) leverages this regulatory position: companies can receive EU personal data legally through the adequacy decision, process it under Israel's less restrictive domestic law, and develop surveillance products that would face greater legal challenge if developed within the EU. Israel is also home to data brokers that aggregate global datasets.
Current State
Israel's EU adequacy decision is under periodic review, and the European Commission has raised concerns about Israel's data protection modernization timeline. Israel's Privacy Protection Authority has limited enforcement resources compared to EU DPAs. The Israeli surveillance technology industry has faced international criticism (NSO Group Pegasus scandal, EU Parliamentary inquiry), but this has not triggered adequacy revocation. Data brokers and surveillance companies incorporated in Israel benefit from the adequacy decision's permission to process EU data while operating under a legal framework that does not impose GDPR-equivalent restrictions on their use of that data.
Impact
Israeli data and surveillance companies occupy a privileged regulatory position: they can receive EU data legally, have access to a sophisticated technology ecosystem for processing it, and face less regulatory scrutiny than their EU-based competitors. The Pegasus spyware revelations demonstrated that Israeli-developed surveillance tools were used against EU citizens, journalists, and politicians, raising questions about whether the adequacy decision enables a data flow that undermines EU privacy protections.
References
EU adequacy decision for Israel; Israel Protection of Privacy Law (1981); European Parliament Pegasus inquiry; NSO Group litigation; Cellebrite data extraction products; Israeli Privacy Protection Authority enforcement statistics; adequacy review documentation.
6India's Emerging Data Broker Hub Status▾
Problem
India processes vast quantities of personal data through its Business Process Outsourcing (BPO) industry, IT services sector, and growing domestic data broker market. The Digital Personal Data Protection Act (DPDPA), enacted in 2023, provides a framework for data protection but allows broad government exemptions, has weaker enforcement provisions than GDPR, and permits data transfers to countries notified by the central government (a whitelist approach that has not yet been implemented). India does not have an EU adequacy decision, meaning EU data transfers to India require SCCs or other safeguards, but the scale of data processing in India makes enforcement of transfer restrictions impractical.
Current State
Indian IT services companies (TCS, Infosys, Wipro, HCL) process personal data from US, EU, and global clients as part of outsourcing arrangements. India's own data broker ecosystem is growing, with companies aggregating data from India's 1.4 billion population including Aadhaar (national ID) linked data, UPI (Unified Payments Interface) transaction data, and mobile data from the world's second-largest smartphone market. The DPDPA's implementing rules are still being finalized, and the Data Protection Board has not yet begun enforcement. The gap between the law's enactment and operational enforcement creates a regulatory vacuum.
Impact
Data flows to India for outsourced processing, but Indian data protection enforcement is nascent. EU personal data processed by Indian BPOs is nominally protected by SCCs but practically subject to Indian domestic law once on Indian infrastructure. India's government exemptions under the DPDPA mean that data accessible to Indian government agencies faces fewer restrictions than data in the EU or even the US. The combination of massive processing capacity, growing domestic data broker activity, and immature enforcement makes India an increasingly significant jurisdiction for data broker arbitrage.
References
India Digital Personal Data Protection Act (2023); DPDPA implementing rules status; Indian BPO industry data handling practices; Aadhaar data privacy controversies; EDPB guidance on transfers to India; Indian Data Protection Board establishment timeline.
7China's Data Outflow and Broker Landscape▾
Problem
Chinese data protection law (PIPL, enacted 2021) restricts outbound data transfer from China but says little about Chinese companies' collection and processing of non-Chinese individuals' data. Chinese data brokers and AdTech companies (including TikTok's parent ByteDance, Tencent, and numerous smaller entities) collect data on users worldwide and process it under Chinese law, which grants broad government access rights. The reciprocal problem also exists: US and EU individuals' data processed by Chinese-affiliated companies may be accessible to Chinese government entities under China's national security laws.
Current State
The TikTok controversy has made the Chinese data processing question politically salient, but TikTok is only the most visible example. Chinese-developed apps across categories (Temu, Shein, various utility and gaming apps) collect data from US and EU users and process it on infrastructure accessible to Chinese corporate entities. China's Data Security Law and PIPL create a framework where data deemed important to national security must be processed domestically and is accessible to government agencies. US government bans on TikTok on federal devices and the proposed TikTok divestiture/ban legislation reflect concerns about Chinese data access, but no comprehensive policy addresses the broader Chinese data processing ecosystem.
Impact
The China-US data flow creates a bilateral surveillance concern: Chinese government entities may access US users' data through Chinese-affiliated apps and services, while US intelligence agencies purchase commercially available data on Chinese nationals and others through US data brokers. The resulting dynamic is one of mutual surveillance enabled by data broker ecosystems in both countries, with individuals in both countries bearing the privacy costs.
References
China PIPL (Personal Information Protection Law, 2021); TikTok-related legislation and CFIUS review; Temu/Shein data collection analysis; US-China Economic and Security Review Commission reports on Chinese data practices; Project Texas (TikTok data localization effort); ByteDance internal data access reporting.
8Data Broker Activity in Privacy Haven Countries▾
Problem
Countries that market themselves as privacy-respecting jurisdictions — Switzerland, Iceland, and to some extent Germany and the Netherlands — attract both privacy-conscious individuals and data brokers seeking to exploit the trust associated with these jurisdictions. A data broker incorporated in Switzerland can market its services as "Swiss privacy protected" while Switzerland's Federal Data Protection Act (revised 2023) does not restrict international data sales in the same way consumers might assume. The association between a country's privacy reputation and the actual privacy protections available to non-residents whose data is processed there creates misleading expectations.
Current State
Switzerland's revised FADP (effective September 2023) modernized Swiss data protection but maintains differences from GDPR, particularly regarding enforcement mechanisms and penalties. Swiss data processing is often marketed as a premium privacy feature (by VPN providers, email services, and data storage companies), but Swiss law does not prevent a Swiss-incorporated company from selling non-Swiss residents' data internationally. Iceland and other Nordic countries are similarly marketed as privacy-friendly, but their data protection laws primarily protect their own residents, not data subjects globally. The privacy haven marketing creates a mismatch between brand perception and legal reality.
Impact
Consumers who choose Swiss-based services believing Swiss privacy law protects their data may discover that their data can be transferred, sold, or processed internationally under Swiss rules that differ from what they expected. Data brokers that incorporate in "privacy haven" countries benefit from the jurisdictional trust while exploiting regulatory differences that favor their business model. The privacy haven effect also attracts cryptocurrency and financial services companies whose data practices may not align with the jurisdiction's privacy reputation.
References
Swiss Federal Act on Data Protection (revised 2023); EDPB adequacy assessment of Switzerland; ProtonMail/Proton AG Swiss jurisdiction analysis; Icelandic Data Protection Authority guidance; comparative analysis of Swiss and EU data protection; jurisdiction shopping in privacy services.
9Latin American Data Broker Emergence and Regulatory Gaps▾
Problem
Latin American countries are experiencing rapid growth in both data collection (driven by smartphone penetration, fintech adoption, and digital government services) and data broker activity, but regulatory frameworks vary dramatically. Brazil's LGPD (Lei Geral de Protecao de Dados) is the most comprehensive, but enforcement by the ANPD (National Data Protection Authority) is still maturing. Argentina has an EU adequacy decision and a data protection law, but enforcement is inconsistent. Mexico's data protection law has weak enforcement mechanisms. Other countries — Colombia, Chile, Peru — have enacted laws of varying strength. This creates a patchwork that data brokers exploit by processing Latin American data in the least regulated jurisdiction.
Current State
Brazil's ANPD has begun enforcement actions but lacks the resources and institutional maturity of European DPAs. Data brokers targeting Latin American populations operate across borders, collecting data in multiple countries and processing it in whichever jurisdiction offers the least resistance. US-based data brokers (including people-search sites) increasingly cover Latin American individuals, particularly those with US connections (immigrant communities, cross-border business relationships). The lack of coordinated enforcement between Latin American data protection authorities means brokers face a fragmented regulatory landscape with minimal cross-border cooperation.
Impact
Latin American individuals, particularly those in countries without strong data protection enforcement, find their data collected and sold with few restrictions. Immigrant communities in the US face double exposure: their US data is collected by US brokers while their home-country data is collected by local and international brokers, with the two datasets merged through identity resolution to create cross-border profiles. The regulatory fragmentation means no single authority has jurisdiction over the complete data lifecycle.
References
Brazil LGPD and ANPD enforcement actions; Argentina data protection adequacy decision; Mexico LFPDPPP enforcement analysis; OAS Inter-American Juridical Committee data protection standards; IAPP Latin American privacy law tracker; data broker activity in Latin American markets.
10African Data Sovereignty and Broker Exploitation▾
Problem
Africa's 54 countries present the most extreme regulatory fragmentation globally, with data protection laws ranging from comprehensive (South Africa's POPIA, Kenya's DPA 2019) to nonexistent. Data brokers — primarily US and European — collect data on African populations through mobile network operators, fintech apps, social media platforms, and development/aid organization data sharing. Africa's rapidly growing mobile internet population (approaching 600 million smartphone users) represents a massive data collection opportunity with minimal regulatory constraint. The African Union's Convention on Cyber Security and Personal Data (Malabo Convention), adopted in 2014, has been ratified by only a minority of AU member states.
Current State
South Africa's POPIA (effective 2021) is the most mature African data protection law, but the Information Regulator has limited enforcement capacity. Kenya's Data Commissioner has begun enforcement activities. Nigeria's NDPR (now replaced by the Nigeria Data Protection Act 2023 and the Nigeria Data Protection Commission) is developing institutional capacity. Most other African countries either lack data protection laws or have enacted them without creating functional enforcement bodies. The Malabo Convention requires 15 ratifications to enter into force and has not yet achieved this threshold. International data brokers collect African data with near-complete impunity in countries without functioning data protection enforcement.
Impact
African populations' data is extracted by international brokers with virtually no regulatory constraint, no opt-out mechanism, and no enforcement authority to appeal to. Mobile money transaction data (M-Pesa and competitors), mobile network location data, and fintech app data from hundreds of millions of Africans enters global data broker databases. This data is used for credit scoring, insurance underwriting, and risk assessment in ways that affect access to financial services, with no transparency about how the data was collected or how algorithms use it. The digital colonialism critique — that African data is extracted by foreign companies for foreign benefit — has gained traction in African policy circles.
References
South Africa POPIA and Information Regulator; Kenya Data Protection Act 2019; Malabo Convention ratification status; Nigeria Data Protection Act 2023; Access Now Africa digital rights reports; CIPESA data protection in Africa analysis; Research ICT Africa data governance publications; digital colonialism discourse in African policy forums.
This page is part of the anonym.community PII pain point research project, which documents 1,478 distinct pain points generated by 98 irreducible structural drivers across 14 research tracks and 240 jurisdictions. The research synthesizes privacy legislation analysis, enforcement decisions, technical literature, and real-world case studies to explain why PII privacy problems persist despite technological and regulatory advances. The complete research corpus is freely available at anonym.community.