The Deduplication Problem: Why You See the Same AI News 8 Times
Last Tuesday, OpenAI announced a new model. Within four hours, you'd seen coverage of the same announcement on Hacker News, TechCrunch, The Verge, Ars Technica, three newsletters, and your Twitter feed. Eight sources. One story. Forty-five minutes of your life reading slight variations of the same information.
This isn't a new problem, but it's getting dramatically worse. As AI coverage explodes across more outlets, the duplication rate has gone from annoying to actively harmful. Here's why deduplication is harder than it looks and what it takes to actually solve it.
The Scale of the Problem
The duplication isn't always obvious. Consider how the same story appears differently:
- Hacker News: "Show HN: Open-source alternative to X" (links to GitHub repo)
- TechCrunch: "Startup Y launches open-source competitor to X with $5M seed" (startup angle)
- The Verge: "A new open-source tool wants to replace X" (consumer angle)
- Newsletter A: "This week's top repos: Y (2,400 stars)" (curated list mention)
- Newsletter B: "Five tools that could replace X" (thematic roundup)
Same underlying story. Five completely different headlines, angles, and framings. Simple title matching catches none of these. Even URL deduplication fails because they all point to different articles.
Why Naive Approaches Fail
Title matching is useless here. "Show HN: Open-source alternative to X" and "Startup Y launches open-source competitor" share zero words besides common ones.
URL deduplication only works when sources link to the same original. Most don't — they link to their own articles about the same thing.
Keyword overlap produces too many false positives. Two articles mentioning "Python," "AI," and "open-source" might be about completely different projects.
Publication time proximity helps narrow candidates but can't confirm duplicates. Plenty of genuinely different stories publish within the same hour.
What Actually Works: Semantic Clustering
The approach that works is semantic similarity clustering — computing how similar two stories are in meaning, not just in words. Here's the pipeline:
- Extract core entities — What company, project, person, or technology is this about? Entity extraction catches that "Y" in the HN post and "Startup Y" in TechCrunch refer to the same thing.
- Compute embeddings — Convert each story's title + first paragraph into a vector embedding. Stories about the same topic cluster tightly in embedding space.
- Cluster by similarity threshold — Group stories with cosine similarity above 0.82 (tuned through testing). This catches semantic duplicates while avoiding false merges.
- Pick the canonical version — From each cluster, select the story with the most information (usually the longest, with the most technical detail).
- Preserve unique angles — If a duplicate adds genuinely new information (like fundraising details), note it as supplementary context rather than discarding it entirely.
The Edge Cases That Break Everything
Deduplication sounds clean in theory. In practice, these edge cases eat your lunch:
Follow-up stories. "Company X launches product" on Monday. "Company X's product crashes on launch day" on Tuesday. These are related but distinct stories — merging them would hide critical information. The system needs temporal awareness.
Same technology, different announcements. "Google releases Gemini 2.5 Pro" and "Google releases Gemini 2.5 Flash" are two separate launches that share 90% of their keywords. Naive similarity would merge them.
Roundup articles. A "top 10 tools this week" post mentions five different stories you've already seen individually. Do you deduplicate the entire roundup? Each mention within it? Neither answer is obviously correct.
Opinion vs. news. "OpenAI releases new model" (news) and "Why OpenAI's new model changes everything" (opinion) cover the same event but serve different purposes. Aggressive deduplication loses the analysis.
Our Approach at ScanBrief
We handle deduplication as a three-stage process:
- Exact dedup — Remove stories linking to the same canonical URL. Fast, zero false positives.
- Entity + embedding clustering — Group by shared entities and semantic similarity. This catches 80% of duplicates.
- AI-assisted merge — For borderline cases, an LLM reviews the cluster and decides: same story (merge), related story (link), or different story (keep separate). This handles the edge cases that rules-based systems miss.
The result: a typical morning scan of 200+ raw items compresses to 15-25 unique stories. You read each story once, with the best available context from all sources that covered it.
Why This Matters More Than You Think
Deduplication isn't a nice-to-have feature — it's the difference between a useful brief and another firehose. Without it, any aggregator eventually becomes as noisy as reading each source individually. The value of aggregation without deduplication is approximately zero.
The developers who stay current without burning hours do so by consuming deduplicated, scored information. Not more sources. Not faster scrolling. Better filtering.
Read Each Story Once
ScanBrief deduplicates across 20+ sources so you never read the same story twice. Five minutes every morning.
Try ScanBrief Free