Tutorial

Building an AI News Pipeline in Python: From 20 Sources to 1 Brief

Published March 2026 · 8 min read

You want a daily tech brief tailored to your interests. Newsletters are too generic. RSS readers are too noisy. What if you built your own pipeline? Here's how to build an AI-powered news aggregation system in Python that pulls from 20 sources, deduplicates, scores for relevance, and outputs a concise daily brief.

We'll cover the full architecture, with real code for each stage. By the end, you'll have a working pipeline you can run via cron every morning.

1 Define the Data Model

Every item in the pipeline needs a consistent shape. Start with a dataclass:

from dataclasses import dataclass, field from datetime import datetime @dataclass class NewsItem: title: str url: str source: str timestamp: datetime content: str = "" score: float = 0.0 tags: list[str] = field(default_factory=list) cluster_id: str | None = None

This structure carries each item through every pipeline stage. The cluster_id field gets populated during deduplication — items with the same cluster_id are duplicates.

2 Build Source Collectors

Each source gets its own collector function. Use asyncio + aiohttp to pull all sources in parallel:

import aiohttp import asyncio import feedparser from typing import List async def collect_hackernews(session: aiohttp.ClientSession) -> List[NewsItem]: """Pull top stories from Hacker News API.""" async with session.get("https://hacker-news.firebaseio.com/v0/topstories.json") as resp: story_ids = await resp.json() items = [] for sid in story_ids[:100]: # Top 100 async with session.get(f"https://hacker-news.firebaseio.com/v0/item/{sid}.json") as resp: story = await resp.json() if story and story.get("score", 0) > 30: items.append(NewsItem( title=story.get("title", ""), url=story.get("url", f"https://news.ycombinator.com/item?id={sid}"), source="hackernews", timestamp=datetime.fromtimestamp(story.get("time", 0)), score=story.get("score", 0), )) return items def collect_rss(feed_url: str, source_name: str) -> List[NewsItem]: """Pull items from an RSS/Atom feed.""" feed = feedparser.parse(feed_url) items = [] for entry in feed.entries[:20]: items.append(NewsItem( title=entry.get("title", ""), url=entry.get("link", ""), source=source_name, timestamp=datetime(*entry.published_parsed[:6]) if hasattr(entry, 'published_parsed') else datetime.now(), content=entry.get("summary", ""), )) return items

The key pattern: each collector returns List[NewsItem]. Every source, regardless of its API shape, normalizes into the same dataclass. This is what makes the rest of the pipeline source-agnostic.

3 Run All Collectors in Parallel

RSS_SOURCES = { "techcrunch": "https://techcrunch.com/feed/", "ars_technica": "https://feeds.arstechnica.com/arstechnica/index", "the_verge": "https://www.theverge.com/rss/index.xml", "openai_blog": "https://openai.com/blog/rss/", "anthropic_blog": "https://www.anthropic.com/feed", } async def collect_all() -> List[NewsItem]: async with aiohttp.ClientSession(timeout=aiohttp.ClientTimeout(total=30)) as session: # API sources (async) hn_items = await collect_hackernews(session) # RSS sources (sync, but fast) rss_items = [] for name, url in RSS_SOURCES.items(): rss_items.extend(collect_rss(url, name)) return hn_items + rss_items

With 20 sources and a 30-second timeout, the entire collection phase runs in under a minute. Sources that time out return empty lists — the pipeline continues without them.

4 Deduplicate with Embeddings

This is where most aggregators fail. Simple URL or title matching misses semantic duplicates. Use sentence embeddings to find stories about the same topic:

from sentence_transformers import SentenceTransformer from sklearn.cluster import AgglomerativeClustering import numpy as np model = SentenceTransformer('all-MiniLM-L6-v2') def deduplicate(items: List[NewsItem], threshold: float = 0.75) -> List[NewsItem]: """Cluster items by semantic similarity, keep the best from each cluster.""" if len(items) < 2: return items texts = [f"{item.title} {item.content[:200]}" for item in items] embeddings = model.encode(texts) clustering = AgglomerativeClustering( n_clusters=None, distance_threshold=1 - threshold, metric='cosine', linkage='average' ) labels = clustering.fit_predict(embeddings) # Keep the highest-scored item from each cluster clusters = {} for item, label in zip(items, labels): if label not in clusters or item.score > clusters[label].score: clusters[label] = item return list(clusters.values())

The all-MiniLM-L6-v2 model is fast (runs on CPU) and surprisingly good at catching semantic duplicates. The 0.75 threshold is tuned to catch "same story, different angle" while keeping genuinely distinct stories separate. Adjust based on your tolerance for duplicates vs. missed stories.

5 Score for Relevance

After deduplication, score each remaining item by how much it matters to your interests:

MY_INTERESTS = ["python", "ai agents", "mcp", "devops", "rust", "llm"] def score_relevance(items: List[NewsItem]) -> List[NewsItem]: """Score items by relevance to configured interests.""" interest_embeddings = model.encode(MY_INTERESTS) for item in items: item_embedding = model.encode(f"{item.title} {item.content[:200]}") similarities = np.dot(interest_embeddings, item_embedding) relevance = float(np.max(similarities)) # Composite score: source score + relevance bonus item.score = (item.score * 0.3) + (relevance * 100 * 0.7) return sorted(items, key=lambda x: x.score, reverse=True)

This gives you personalized ranking. A Python library release scores higher than a Java one — not because Java is bad, but because your pipeline knows what you care about.

6 Generate the Brief with an LLM

import anthropic client = anthropic.Anthropic() def generate_brief(items: List[NewsItem], max_items: int = 20) -> str: """Use Claude to generate a concise morning brief.""" top_items = items[:max_items] items_text = "\n".join( f"- [{item.source}] {item.title} ({item.url})" for item in top_items ) response = client.messages.create( model="claude-sonnet-4-20250514", max_tokens=2000, messages=[{ "role": "user", "content": f"""Summarize these tech news items into a morning brief. Group by theme. Each item gets 1-2 sentences. Include URLs. Be concise — the reader has 5 minutes. Items: {items_text}""" }] ) return response.content[0].text

7 Wire It Together

async def run_pipeline(): print("Collecting from sources...") items = await collect_all() print(f"Collected {len(items)} items") print("Deduplicating...") unique = deduplicate(items) print(f"Deduplicated to {len(unique)} items") print("Scoring relevance...") scored = score_relevance(unique) print("Generating brief...") brief = generate_brief(scored) # Save or send with open(f"briefs/{datetime.now():%Y-%m-%d}.md", "w") as f: f.write(brief) print("Done!") if __name__ == "__main__": asyncio.run(run_pipeline())

Add this to a cron job (0 6 * * * cd /path/to/pipeline && python main.py) and you have a daily automated brief.

What to Optimize Next

This pipeline handles 80% of what you need. The remaining 20% — better dedup edge cases, source reliability, formatting polish — is where ScanBrief puts its engineering effort so you don't have to.

Skip the Build, Get the Brief

ScanBrief runs this pipeline (and a lot more) every morning so you don't have to maintain it yourself.

Start Free