AI Can Now Doxx Your Anonymous Accounts? Here's What’s Going On
Researchers from ETH Zurich, MATS, and Anthropic published a paper demonstrating an AI pipeline that can link pseudonymous online accounts to real identities using only public web search, embeddings, and reasoning models. In controlled tests the system matched Hacker News users to LinkedIn profiles with 90% precision at an estimated cost of $1–$4 per target, though the authors withheld their code and the real identities they discovered and say the study passed an ethics review.

Why It Matters
The study shows that widely available LLM capabilities can meaningfully increase the risk of deanonymization without any data breach, shifting privacy threat models for people who rely on pseudonymity online. That has implications for journalists, privacy-conscious users, and communities where exposure can lead to harassment or physical danger, as documented by prior incidents cited in the paper.
Key Facts
- Paper title: Large-scale online deanonymization with LLMs
- Institutions: ETH Zurich, MATS, and Anthropic (researcher Nicholas Carlini)
- Core claim: AI pipeline can match anonymized accounts to real profiles using only web search, embeddings, and reasoning models
- Reported precision/cost: 90% precision in some tests; cost estimated at $1–$4 per target
- Hacker News test: 338 users with LinkedIn in bio; AI named 226 correctly (67% recall); wrong on ~1 in 10 of its guesses that it made
A team from ETH Zurich, the AI safety group MATS, and researchers at Anthropic published a study showing that large language models can be used to deanonymize pseudonymous internet users by combining content extraction, vector search, and reasoning. The researchers describe a four-step pipeline—Extract, Search, Reason, and Calibrate—where an AI summarizes a user’s posts, converts that summary into embeddings to retrieve candidate profiles, uses a stronger model to cross-check matches, and then reports confidence scores to avoid low-certainty guesses. In a controlled experiment on Hacker News, the authors assembled 338 accounts that had linked to LinkedIn in their bios, removed identifying information, and asked their AI agent to find matches via public web search. The system produced correct names for 226 accounts (about 67% recall) and operated at high precision in its confident guesses, with roughly one incorrect prediction for every ten guesses it actually made. In another small test using Anthropic interview transcripts, the AI identified at least nine scientists from descriptions of their work. A central contribution of the paper is the low operational cost: the authors estimate running the pipeline costs between $1 and $4 of model subscription fees per target, and it does not rely on hacked or leaked data. The team emphasizes that the method stitches together commonplace model abilities—summarization, embedding-based retrieval, and reasoning—so there is no single feature to switch off that would fully prevent this chain of inference. The authors qualify their headline results with important caveats. To evaluate performance they used ground-truth subjects (accounts that had already linked to LinkedIn, or split a single user’s history), which represents a best-case scenario rather than a random-sample test of all pseudonymous accounts. As the candidate pool grows, recall falls: against 89,000 candidates the study’s strongest method achieved about half the correct matches at 90% precision. The researchers withheld code, prompts, and the real identities they found, and report that their work underwent ETH Zurich’s ethics review prior to publication.
Keep Reading

Australian PM warns of AI’s ‘furious pace’ after agent breached government site

Meta's New AI Toy Is a Keychain That Watches, Listens, and Never Blinks
3 things Micron investors need to watch as the stakes get higher
