Skip to content
View jaehyeon-kim's full-sized avatar

Organizations

@beam-pyio

Block or report jaehyeon-kim

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
jaehyeon-kim/README.md

About Me

Linkedin Badge Website

I'm a data engineer, building real-time streaming systems that power machine learning and AI. Over the past decade I've worked extensively with Kafka, Flink, and Spark on AWS and Google Cloud, designing high-throughput pipelines and lakehouse architectures that stay resilient, observable, and maintainable at scale. Along the way I've built data transformations with dbt, orchestrated them with Airflow, and deployed platforms on Kubernetes with Terraform.

What excites me most is where streaming meets intelligence: feeding live data into models for online learning, real-time detection, and adaptive decision-making. I'm increasingly focused on applying ML and AI to streaming data, turning fast-moving events into systems that learn and respond in the moment.

As a passionate engineer and writer, I share practical insights on real-time analytics, data architectures, data lakehouse patterns, and data lineage. I recently presented "Building End-to-End Data Lineage with Kafka, Flink, and Spark" at Current and Programmable in 2026, and I write regularly on my blog.

Focus areas:

  • Stream processing: Kafka, Flink, Kafka Streams and Flink SQL
  • Streaming machine learning: online learning, contextual bandits and real-time features
  • MLOps: feature stores, model registries and pipelines with Feast, MLflow and Airflow
  • AI engineering: agents with tools, semantic layers and conversational analytics over a lakehouse
  • Data engineering: dbt, Spark and Airflow, on data lakes, lakehouses and warehouses
  • Lakehouse and lineage: Iceberg and end-to-end data lineage
  • Cloud and platforms: AWS (MSK, EMR, Glue, Athena, Lambda) and Google Cloud (BigQuery), on Kubernetes with Terraform
  • Simulation and digital twins: discrete-event models that drive live systems

Blog: https://jaehyeon.me

Open source:

Pinned Loading

  1. dynamic-des dynamic-des Public

    Real-time SimPy control plane to dynamically update parameters and stream outputs via external systems like Kafka, Redis, or Postgres. Built for event-driven digital twins.

    Python 4

  2. odctl odctl Public

    A curated collection of open source technologies and the odctl CLI for running a local Docker-based data platform to experiment with modern data architecture and MLOps.

    Python 4

  3. nicegui-fastapi-template nicegui-fastapi-template Public

    A template for rapidly building full-stack web applications in Python, featuring a FastAPI backend, a NiceGUI frontend, PostgreSQL, and Docker.

    Python 120 17

  4. benchtop benchtop Public

    Small, hands-on demos for data engineering, stream processing, machine learning, AI engineering, and MLOps, each running locally on the odctl stack from a cold clone.

    Python 1

  5. kafka-pocs kafka-pocs Public

    Runnable Kafka examples: Docker Compose and Kubernetes clusters, Python producers and consumers, Kafka Connect pipelines, Glue Schema Registry integration, TLS and SASL security, and MSK ingestion …

    Shell 32 14

  6. flink-demos flink-demos Public

    Apache Flink examples in PyFlink and Kotlin: local Docker clusters, AWS Managed Flink deployments with Kafka, a Flink SQL cookbook, and worked exercises from Confluent courses and the Stream Proces…

    Python 47 15