How We Scan ArXiv, GitHub, and Hacker News Every Morning
Every morning at 6:30 AM, ScanBrief's pipeline wakes up and starts pulling from over 20 sources. By 7:00 AM, you have a deduplicated, scored brief in your inbox. Here's exactly how that pipeline works — the sources, the challenges, and the engineering decisions behind each one.
The Source Stack
We pull from three categories of sources, each with different APIs, rate limits, and data quality characteristics.
Tier 1: Structured APIs
These sources provide clean, structured data through official APIs. They're the most reliable part of the pipeline.
Hacker News (Firebase API)
HN's API is beautifully simple. We pull /v0/topstories.json for the top 200 story IDs, then fetch each story's metadata. We filter by score (>50 points) and age (<24 hours). The API has no rate limits and returns clean JSON. It's the gold standard for source integration.
GitHub Trending
GitHub doesn't have an official trending API, so we scrape the trending page and cross-reference with the GraphQL API for star counts, descriptions, and language data. We track daily, weekly, and monthly trending separately — a repo trending daily is more likely to be genuinely new versus a perennial favorite.
ArXiv API
ArXiv provides an Atom feed API with category-based queries. We pull from cs.AI, cs.CL (NLP), cs.LG (machine learning), and cs.SE (software engineering). The challenge is volume: 100+ papers per day in AI alone. We use abstract analysis and citation velocity to surface the 3-5 that matter practically.
Tier 2: RSS/Atom Feeds
Most tech publications still maintain RSS feeds, and they're the most efficient way to pull content at scale:
- TechCrunch — Full article content in RSS, no scraping needed
- Ars Technica — Excellent technical depth, reliable feed
- The Verge — Consumer tech angle, good for broader context
- OpenAI Blog — First-party announcements, always high signal
- Google AI Blog — Research updates, product launches
- Anthropic Blog — Model releases, safety research
RSS feeds are simple to parse with Python's feedparser library, but they have quirks. Some feeds include full content, others only excerpts. Publication timestamps vary in format and timezone handling. We normalize everything to UTC and store raw + parsed versions.
Tier 3: Web Scraping
A few high-value sources don't offer APIs or RSS. For these, we use targeted scrapers with careful rate limiting and caching. We won't detail which sources fall here — scraper endpoints are fragile by nature and sharing specifics invites breakage.
The Pipeline Architecture
The pipeline runs as a linear sequence of stages, each feeding the next:
Source-Specific Challenges
Hacker News: The recency problem. HN's top stories list is a mix of stories from the last 1-36 hours. A story at position 30 might be 2 hours old (still climbing) or 20 hours old (dying). We weight by velocity — points per hour — not just absolute score.
GitHub: Stars vs. substance. A repo with 2,000 stars in 24 hours might be a genuine breakthrough or a toy demo that went viral on Twitter. We check: Does it have tests? Documentation? Is the author credible (previous repos, commit history)? Star count alone is a terrible signal.
ArXiv: Papers vs. practical impact. Most ArXiv papers are incremental improvements on benchmarks. We look for papers that: (a) introduce a new technique with code, (b) achieve a significant improvement on a task developers actually care about, or (c) come from a lab with a track record of shipping to production. A paper improving BLEU scores by 0.3 points doesn't make the cut.
RSS feeds: The timezone nightmare. Some feeds use UTC, some use the author's local timezone, some omit timezone information entirely. We've built a timezone normalization layer that handles 14 different datetime formats we've encountered across sources.
Failure Modes and Resilience
Any source can fail on any given day. Our approach:
- 30-second timeout per source — If a source is slow, we skip it rather than delay the entire brief
- Cached fallback — If a source fails completely, we use yesterday's data marked as "from cache"
- Source health tracking — If a source fails 3 days in a row, we get alerted and investigate
- Graceful degradation — A brief with 15 sources is still valuable even if 5 are down
In practice, the full pipeline runs in under 5 minutes, and we've had zero complete outages in the last 30 days. Individual sources fail 2-3 times per month (usually GitHub trending, which changes page structure occasionally).
Why We Built This Instead of Using Existing Tools
There are plenty of RSS readers and news aggregators. The gap we saw was in the intelligence layer: deduplication, relevance scoring, and project-aware filtering. An RSS reader shows you everything. A newsletter shows you what one curator thinks matters. ScanBrief shows you what matters to you, deduplicated across every source, ranked by actual impact.
The technical challenge isn't pulling from sources — it's making 200 raw items into 20 actionable insights. That's the pipeline we've spent the most time refining.
See the Pipeline in Action
ScanBrief delivers a deduplicated, scored brief from 20+ sources every morning at 6:30 AM.
Start Free