Behind the Scenes

How We Scan ArXiv, GitHub, and Hacker News Every Morning

Published March 2026 · 7 min read

Every morning at 6:30 AM, ScanBrief's pipeline wakes up and starts pulling from over 20 sources. By 7:00 AM, you have a deduplicated, scored brief in your inbox. Here's exactly how that pipeline works — the sources, the challenges, and the engineering decisions behind each one.

The Source Stack

We pull from three categories of sources, each with different APIs, rate limits, and data quality characteristics.

Tier 1: Structured APIs

These sources provide clean, structured data through official APIs. They're the most reliable part of the pipeline.

Hacker News (Firebase API)

HN's API is beautifully simple. We pull /v0/topstories.json for the top 200 story IDs, then fetch each story's metadata. We filter by score (>50 points) and age (<24 hours). The API has no rate limits and returns clean JSON. It's the gold standard for source integration.

GitHub Trending

GitHub doesn't have an official trending API, so we scrape the trending page and cross-reference with the GraphQL API for star counts, descriptions, and language data. We track daily, weekly, and monthly trending separately — a repo trending daily is more likely to be genuinely new versus a perennial favorite.

ArXiv API

ArXiv provides an Atom feed API with category-based queries. We pull from cs.AI, cs.CL (NLP), cs.LG (machine learning), and cs.SE (software engineering). The challenge is volume: 100+ papers per day in AI alone. We use abstract analysis and citation velocity to surface the 3-5 that matter practically.

Tier 2: RSS/Atom Feeds

Most tech publications still maintain RSS feeds, and they're the most efficient way to pull content at scale:

RSS feeds are simple to parse with Python's feedparser library, but they have quirks. Some feeds include full content, others only excerpts. Publication timestamps vary in format and timezone handling. We normalize everything to UTC and store raw + parsed versions.

Tier 3: Web Scraping

A few high-value sources don't offer APIs or RSS. For these, we use targeted scrapers with careful rate limiting and caching. We won't detail which sources fall here — scraper endpoints are fragile by nature and sharing specifics invites breakage.

The Pipeline Architecture

The pipeline runs as a linear sequence of stages, each feeding the next:

6:30 AM [COLLECT] Pull from all sources in parallel (30s timeout per source) 6:31 AM [NORMALIZE] Standardize format: title, url, source, timestamp, content 6:31 AM [DEDUP] Exact URL dedup + semantic clustering (embeddings) 6:32 AM [SCORE] Relevance scoring: HN points, GH stars, entity importance 6:33 AM [RANK] Sort by composite score, select top 15-25 items 6:34 AM [BRIEF] Generate summary text with AI, format as brief 6:35 AM [DELIVER] Push to email, Discord, web dashboard

Source-Specific Challenges

Hacker News: The recency problem. HN's top stories list is a mix of stories from the last 1-36 hours. A story at position 30 might be 2 hours old (still climbing) or 20 hours old (dying). We weight by velocity — points per hour — not just absolute score.

GitHub: Stars vs. substance. A repo with 2,000 stars in 24 hours might be a genuine breakthrough or a toy demo that went viral on Twitter. We check: Does it have tests? Documentation? Is the author credible (previous repos, commit history)? Star count alone is a terrible signal.

ArXiv: Papers vs. practical impact. Most ArXiv papers are incremental improvements on benchmarks. We look for papers that: (a) introduce a new technique with code, (b) achieve a significant improvement on a task developers actually care about, or (c) come from a lab with a track record of shipping to production. A paper improving BLEU scores by 0.3 points doesn't make the cut.

RSS feeds: The timezone nightmare. Some feeds use UTC, some use the author's local timezone, some omit timezone information entirely. We've built a timezone normalization layer that handles 14 different datetime formats we've encountered across sources.

Failure Modes and Resilience

Any source can fail on any given day. Our approach:

In practice, the full pipeline runs in under 5 minutes, and we've had zero complete outages in the last 30 days. Individual sources fail 2-3 times per month (usually GitHub trending, which changes page structure occasionally).

Why We Built This Instead of Using Existing Tools

There are plenty of RSS readers and news aggregators. The gap we saw was in the intelligence layer: deduplication, relevance scoring, and project-aware filtering. An RSS reader shows you everything. A newsletter shows you what one curator thinks matters. ScanBrief shows you what matters to you, deduplicated across every source, ranked by actual impact.

The technical challenge isn't pulling from sources — it's making 200 raw items into 20 actionable insights. That's the pipeline we've spent the most time refining.

See the Pipeline in Action

ScanBrief delivers a deduplicated, scored brief from 20+ sources every morning at 6:30 AM.

Start Free