I'm a data engineer, building real-time streaming systems that power machine learning and AI. Over the past decade I've worked extensively with Kafka, Flink, and Spark on AWS and Google Cloud, designing high-throughput pipelines and lakehouse architectures that stay resilient, observable, and maintainable at scale. Along the way I've built data transformations with dbt, orchestrated them with Airflow, and deployed platforms on Kubernetes with Terraform.
What excites me most is where streaming meets intelligence: feeding live data into models for online learning, real-time detection, and adaptive decision-making. I'm increasingly focused on applying ML and AI to streaming data, turning fast-moving events into systems that learn and respond in the moment.
As a passionate engineer and writer, I share practical insights on real-time analytics, data architectures, data lakehouse patterns, and data lineage. I recently presented "Building End-to-End Data Lineage with Kafka, Flink, and Spark" at Current and Programmable in 2026, and I write regularly on my blog.
Focus areas:
- Stream processing: Kafka, Flink, Kafka Streams and Flink SQL
- Streaming machine learning: online learning, contextual bandits and real-time features
- MLOps: feature stores, model registries and pipelines with Feast, MLflow and Airflow
- AI engineering: agents with tools, semantic layers and conversational analytics over a lakehouse
- Data engineering: dbt, Spark and Airflow, on data lakes, lakehouses and warehouses
- Lakehouse and lineage: Iceberg and end-to-end data lineage
- Cloud and platforms: AWS (MSK, EMR, Glue, Athena, Lambda) and Google Cloud (BigQuery), on Kubernetes with Terraform
- Simulation and digital twins: discrete-event models that drive live systems
Blog: https://jaehyeon.me
Open source:
- dynamic-des (dynamic discrete event simulation & digital twins)
- odctl (Open Data Stack)
- benchtop (hands-on data streaming and ML projects that run locally)
- nicegui-fastapi-template (full-stack Python web app template)





