data engineering & analytics · milwaukee, wi

Dean Kuhn

I build data pipelines that end in a live, interactive product you can actually use.

Last pipeline run: 2026-09-10  ·  3 projects live, 1 in development
Music Growth Pipeline
data engineering · live
● live
37,435 artists tracked · 19 weeks
median growth by listener size
under 69k
3.25%
69k–137k
2.94%
137k–236k
2.51%
236k–471k
2.10%
471k+
2.04%
KitchenSync
data engineering · live
● live
96.9% ML service level
95.7% baseline
+-3.0pp waste trade-off
34 days run
ML service level · last 14 days
PharmaWatch
data engineering
◌ in development
phase 1 in progress

Drug safety signal platform over FDA adverse event data (FAERS/openFDA). Cleaning and deduplicating two decades of shifting schemas, landing on Parquet + DuckDB, with a RAG layer over drug labels planned on top.

Package Delivery Routing
algorithms · complete
✓ complete
38993.9 final fitness score · 10000 generations
165 packages
10 trucks
4 refrigerated
Market Cynic Pipeline
data engineering
◌ retired
archived · no longer running

Bronze→Silver→Gold Delta Lake pipeline that correlated Reddit sentiment against Yahoo Finance price and volume data to flag S&P 500 divergence signals, surfaced in a Streamlit dashboard. Broken since Reddit shut down the public .json endpoints its ingestion relied on — kept here as a relic, not an active project.

CS graduate based in the Milwaukee area. I enjoy solving problems that involve cleaning and transforming difficult data, and displaying those insights in the form of live, queryable sites.

Each project runs the full stack: ingestion, transformation, warehousing, and, at times, machine learning.

stack
languages
PythonSQL
data & storage
PostgreSQLSnowflakedbt CorepandasFastAPI
ml
LightGBMscikit-learn
viz
StreamlitPower BI
infra
GitHub ActionsAWS EC2DockerLinux/WSL
learning
DatabricksAzure Data Lake Storage Gen2Delta LakePySparkApache Airflow