At least for Linkedin and Hacker News. Accuracy on Reddit etc. is lower. Still worrying.
We collect 338 Hacker News (HN) users who linked a LinkedIn profile in their publicly-visible HN bio, providing verified real-world identities as ground truth. We first summarize each user’s HN activity (comments and stories) into a structured profile. Then we create a search prompt and anonymize it (see Appendix˜A for details), and pass it to the agent. The agent correctly identifies 226 of 338 targets (67%) at 90% precision (95% CI: 86–93%; 25 incorrect identifications, 86 abstentions). We note that these edited profiles are much easier to identify than most pseudonymous accounts; we discuss this bias in Section˜3.3.
Stylometry has been a recognised field for a long time. I'd argue LLMs offer an advantage against stylometric analysis, you can use a local LLM to rewrite your posts and remove (or add?) writing quirks to prevent deanonymization.
qualia wrote: Mon Aug 10, 2026 4:23 pm
Stylometry has been a recognised field for a long time. I'd argue LLMs offer an advantage against stylometric analysis, you can use a local LLM to rewrite your posts and remove (or add?) writing quirks to prevent deanonymization.
It's not that simple though. LLMs can mass analyze profiles, see if there are any commonalities - not just stylometry. Places that one mentioned being visited, dog's names, interests, political affiliations, etc.
Well, if you post your dog's name or mulitple specific places you actually visit on an account you don't want to get deanonymized then you're just a moron. I suppose this may become an issue for normies who just post whatever on reddit without caring about opsec but it's not really a concern for people here.