NewsLake
Bronze/Silver/Gold
A news lakehouse that runs itself. Every day, Airflow pulls fresh articles into object storage, Spark cleans and aggregates them through three medallion layers, dbt models the result, and everything below is read live from the warehouse — no static exports.
At a glance
What is in the lake
Counted straight from the warehouse at page build time — not a figure typed into the markup.
Articles
0
rows in fct_articles
Sources
0
distinct publishers
Topics
0
tracked categories
Days of history
0
grows once per run
What the world is talking about
Topics by volume
Every article carries one or more topic tags. Spark explodes those into a bridge table so a single article can count toward several topics at once — these are the busiest, ranked across all history.
Who is publishing
Most active sources
1,319 distinct publishers have appeared in the lake so far. Each gets a stable source_id derived from its name plus a hash — so publishers writing in non-Latin scripts stay distinct instead of collapsing together.
Yahoo Sports
yahoo_sports_be24346f
BBC
bbc_22586817
The Times of India
the_times_of_india_c308379d
The Guardian
the_guardian_5c46904f
Goal.com
goal_com_6cda118c
CBS Sports
cbs_sports_0bc91bd5
Yahoo
yahoo_db95f8b8
USA Today
usa_today_bfcadc0c
Sports Illustrated
sports_illustrated_7f9784c6
The Independent
the_independent_91e3555d
Straight from the warehouse
Latest articles
The most recent rows in fct_articles, queried live on page load. Titles link out to the original publisher.
How it works
Four layers, once a day
Airflow runs the whole chain on a schedule — fetch, validate, transform, check, aggregate, load, model, test. If any step fails its quality gate, the run stops there rather than publishing bad data.
Raw ingestion
The article exactly as the API returned it, written to MinIO as JSON and partitioned by ingest time. Nothing is cleaned here — Bronze exists so any downstream bug can be replayed against the original response.
- Python ingestion hits the listing endpoint, then enriches each article
- Rate-limited to stay inside the free API's daily budget
- Partitioned source / year / month / day / hour
Cleaned & validated
PySpark flattens the nested JSON into a typed, article-grain table, drops rows that fail validation into a quarantine path, and deduplicates so a re-run can't double-count.
- Rows missing an id, title, url or timestamp are quarantined, not silently dropped
- Stable source_id = ASCII slug + hash, so non-Latin publishers stay distinct
- Written as Snappy-compressed Parquet
Aggregated metrics
Business-level rollups computed once, so the dashboard never pays for a full scan: daily article and source metrics, topic trends, and publishing activity.
- Topics pre-exploded before counting distinct sources
- Four aggregate tables, each rewritten per run
- Independent quality checks gate the run before it proceeds
Postgres + dbt
Spark lands article-grain data in a raw schema on Neon, then dbt builds staging, intermediate and mart models on top — with 22 tests that fail the run if the data is wrong.
- Loader truncates rather than drops, so dependent dbt views survive
- unique / not_null / relationships / accepted_values tests
- The marts behind this page are the dbt output, read live
The other half
There is a cockpit behind this
This site is the public face of the lake. The operational side — watching a run execute, poking at raw files, checking a Parquet schema — lives in a companion Streamlit dashboard, also public.
It mirrors the same warehouse data shown above, plus a full walkthrough of every pipeline stage. Two panels below stay local-only by design: they reach MinIO and the Airflow API over Docker-internal hostnames that have no route from the public internet.
Live pipeline view
Local onlyThe Airflow DAG as an animated flow diagram, polling task states every few seconds — plus a button to trigger a run and the next scheduled run time.
Data Explorer
Local onlyA read-only browser for every storage layer: raw Bronze JSON straight out of object storage, Silver and Gold Parquet with full schemas, and both Postgres schemas.








