Real-time Data & Streaming
Kafka, Pub/Sub, and CDC pipelines that keep the warehouse in sync in seconds.
Event and CDC pipelines from product/marketing/IoT sources to BigQuery or Snowflake — Kafka, Pub/Sub, Debezium, materialized views.
Problems I solve
- Analytics stuck on daily batch
- Operational DB and warehouse out of sync
- No real-time signals for ops or product
- Kafka set up but no clear ownership
What you get
- Kafka / Pub/Sub topic design
- CDC from Postgres / MySQL with Debezium
- Streaming into BigQuery / Snowflake
- Materialized views + streaming SQL
Use cases
Real-time product analytics
Events land in BigQuery within seconds; dashboards refresh continuously.
CDC from Postgres
Every operational write mirrored to the warehouse with schema evolution.
Fraud & anomaly signals
Streaming windowed aggregations feeding an alerting agent.
Live inventory / pricing
Sub-second inventory updates across stores + storefront.
Examples I've shipped
DTC live analytics
Outcome — Marketing dashboard latency: 24h → 30s.
Fintech CDC fabric
Outcome — Sub-minute mirror of core banking tables.
How I work
- 1
Event modeling
Design topic taxonomy, schemas, and ownership.
- 2
Pipeline build
Producers, consumers, sinks with schema registry.
- 3
Warehouse land
Streaming inserts + dbt models for consumption.
- 4
Monitor
Lag, error rate, and cost dashboards.
Deliverables
- Streaming pipelines
- Schema registry + docs
- Warehouse models
- Monitoring
Benefits
- Decisions on live data
- Warehouse mirrors reality
- Foundation for real-time AI features
Frequently asked questions
Kafka or Pub/Sub?+
Pub/Sub if you're on GCP and want zero infra; Kafka (or MSK / Confluent) when you need replay semantics, wide ecosystem, or on-prem.