Researcher turned analytics engineer. I spent 10+ years studying local news ecosystems, and somewhere along the way, I became more interested in building the systems that make the research possible than in the research itself.
Cleaned, normalized, and structured 16,000+ local news stories from 100 U.S. communities into a queryable database for research into local journalism health and content patterns.
An R pipeline for analyzing survey data without touching the local file system. Raw data lives in Google Sheets, is processed in R, and results are written back to Google Sheets for use in downstream tools like Datawrapper.
Podcast Metadata Pipeline (in progress)
An analytics engineering pipeline transforming Podcast Index API data into content strategy and market opportunity data marts using R, dbt, and PostgreSQL.
A generalizable, locally-run LLM pipeline for categorizing news content at scale.
- Languages: SQL, R, Python
- Databases: PostgreSQL, MySQL, SQLite
- Transformation: dbt
- Analysis & Visualization: RMarkdown, Datawrapper, Tableau
- Other: SPSS, Stata