Tutorial
Building an AI News Pipeline in Python: From 20 Sources to 1 Brief
Published March 2026 · 8 min read
You want a daily tech brief tailored to your interests. Newsletters are too generic. RSS readers are too noisy. What if you built your own pipeline? Here's how to build an AI-powered news aggregation system in Python that pulls from 20 sources, deduplicates, scores for relevance, and outputs a concise daily brief.
We'll cover the full architecture, with real code for each stage. By the end, you'll have a working pipeline you can run via cron every morning.
1 Define the Data Model
Every item in the pipeline needs a consistent shape. Start with a dataclass:
from dataclasses import dataclass, field
from datetime import datetime
@dataclass
class NewsItem:
title: str
url: str
source: str
timestamp: datetime
content: str = ""
score: float = 0.0
tags: list[str] = field(default_factory=list)
cluster_id: str | None = None
This structure carries each item through every pipeline stage. The cluster_id field gets populated during deduplication — items with the same cluster_id are duplicates.
2 Build Source Collectors
Each source gets its own collector function. Use asyncio + aiohttp to pull all sources in parallel:
import aiohttp
import asyncio
import feedparser
from typing import List
async def collect_hackernews(session: aiohttp.ClientSession) -> List[NewsItem]:
"""Pull top stories from Hacker News API."""
async with session.get("https://hacker-news.firebaseio.com/v0/topstories.json") as resp:
story_ids = await resp.json()
items = []
for sid in story_ids[:100]: # Top 100
async with session.get(f"https://hacker-news.firebaseio.com/v0/item/{sid}.json") as resp:
story = await resp.json()
if story and story.get("score", 0) > 30:
items.append(NewsItem(
title=story.get("title", ""),
url=story.get("url", f"https://news.ycombinator.com/item?id={sid}"),
source="hackernews",
timestamp=datetime.fromtimestamp(story.get("time", 0)),
score=story.get("score", 0),
))
return items
def collect_rss(feed_url: str, source_name: str) -> List[NewsItem]:
"""Pull items from an RSS/Atom feed."""
feed = feedparser.parse(feed_url)
items = []
for entry in feed.entries[:20]:
items.append(NewsItem(
title=entry.get("title", ""),
url=entry.get("link", ""),
source=source_name,
timestamp=datetime(*entry.published_parsed[:6]) if hasattr(entry, 'published_parsed') else datetime.now(),
content=entry.get("summary", ""),
))
return items
The key pattern: each collector returns List[NewsItem]. Every source, regardless of its API shape, normalizes into the same dataclass. This is what makes the rest of the pipeline source-agnostic.
3 Run All Collectors in Parallel
RSS_SOURCES = {
"techcrunch": "https://techcrunch.com/feed/",
"ars_technica": "https://feeds.arstechnica.com/arstechnica/index",
"the_verge": "https://www.theverge.com/rss/index.xml",
"openai_blog": "https://openai.com/blog/rss/",
"anthropic_blog": "https://www.anthropic.com/feed",
}
async def collect_all() -> List[NewsItem]:
async with aiohttp.ClientSession(timeout=aiohttp.ClientTimeout(total=30)) as session:
# API sources (async)
hn_items = await collect_hackernews(session)
# RSS sources (sync, but fast)
rss_items = []
for name, url in RSS_SOURCES.items():
rss_items.extend(collect_rss(url, name))
return hn_items + rss_items
With 20 sources and a 30-second timeout, the entire collection phase runs in under a minute. Sources that time out return empty lists — the pipeline continues without them.
4 Deduplicate with Embeddings
This is where most aggregators fail. Simple URL or title matching misses semantic duplicates. Use sentence embeddings to find stories about the same topic:
from sentence_transformers import SentenceTransformer
from sklearn.cluster import AgglomerativeClustering
import numpy as np
model = SentenceTransformer('all-MiniLM-L6-v2')
def deduplicate(items: List[NewsItem], threshold: float = 0.75) -> List[NewsItem]:
"""Cluster items by semantic similarity, keep the best from each cluster."""
if len(items) < 2:
return items
texts = [f"{item.title} {item.content[:200]}" for item in items]
embeddings = model.encode(texts)
clustering = AgglomerativeClustering(
n_clusters=None,
distance_threshold=1 - threshold,
metric='cosine',
linkage='average'
)
labels = clustering.fit_predict(embeddings)
# Keep the highest-scored item from each cluster
clusters = {}
for item, label in zip(items, labels):
if label not in clusters or item.score > clusters[label].score:
clusters[label] = item
return list(clusters.values())
The all-MiniLM-L6-v2 model is fast (runs on CPU) and surprisingly good at catching semantic duplicates. The 0.75 threshold is tuned to catch "same story, different angle" while keeping genuinely distinct stories separate. Adjust based on your tolerance for duplicates vs. missed stories.
5 Score for Relevance
After deduplication, score each remaining item by how much it matters to your interests:
MY_INTERESTS = ["python", "ai agents", "mcp", "devops", "rust", "llm"]
def score_relevance(items: List[NewsItem]) -> List[NewsItem]:
"""Score items by relevance to configured interests."""
interest_embeddings = model.encode(MY_INTERESTS)
for item in items:
item_embedding = model.encode(f"{item.title} {item.content[:200]}")
similarities = np.dot(interest_embeddings, item_embedding)
relevance = float(np.max(similarities))
# Composite score: source score + relevance bonus
item.score = (item.score * 0.3) + (relevance * 100 * 0.7)
return sorted(items, key=lambda x: x.score, reverse=True)
This gives you personalized ranking. A Python library release scores higher than a Java one — not because Java is bad, but because your pipeline knows what you care about.
6 Generate the Brief with an LLM
import anthropic
client = anthropic.Anthropic()
def generate_brief(items: List[NewsItem], max_items: int = 20) -> str:
"""Use Claude to generate a concise morning brief."""
top_items = items[:max_items]
items_text = "\n".join(
f"- [{item.source}] {item.title} ({item.url})"
for item in top_items
)
response = client.messages.create(
model="claude-sonnet-4-20250514",
max_tokens=2000,
messages=[{
"role": "user",
"content": f"""Summarize these tech news items into a morning brief.
Group by theme. Each item gets 1-2 sentences. Include URLs.
Be concise — the reader has 5 minutes.
Items:
{items_text}"""
}]
)
return response.content[0].text
7 Wire It Together
async def run_pipeline():
print("Collecting from sources...")
items = await collect_all()
print(f"Collected {len(items)} items")
print("Deduplicating...")
unique = deduplicate(items)
print(f"Deduplicated to {len(unique)} items")
print("Scoring relevance...")
scored = score_relevance(unique)
print("Generating brief...")
brief = generate_brief(scored)
# Save or send
with open(f"briefs/{datetime.now():%Y-%m-%d}.md", "w") as f:
f.write(brief)
print("Done!")
if __name__ == "__main__":
asyncio.run(run_pipeline())
Add this to a cron job (0 6 * * * cd /path/to/pipeline && python main.py) and you have a daily automated brief.
What to Optimize Next
- Persistent storage — Store items in SQLite to track trends over time and avoid re-fetching
- Email delivery — Use
smtplib or a service like Resend to email the brief
- Discord integration — Post to a channel via webhook for team consumption
- Feedback loop — Track which items you click to improve relevance scoring over time
This pipeline handles 80% of what you need. The remaining 20% — better dedup edge cases, source reliability, formatting polish — is where ScanBrief puts its engineering effort so you don't have to.
Skip the Build, Get the Brief
ScanBrief runs this pipeline (and a lot more) every morning so you don't have to maintain it yourself.
Start Free