Chapter 7A: Data Engineering & Stream Processing
Chapter 7A: Data Engineering & Stream Processing
Data engineering turns raw events and files into trustworthy, queryable data. Stream processing reacts to records as they arrive, batch processing works on bounded collections, and lakehouse engines give analytical workloads a managed table surface over object storage.
- Stateful Stream & Batch Processing Frameworks: Apache Spark, Apache Flink, Apache Beam, Watermarking, Event-Time vs Processing-Time, and Windowing Paradigms (Tumbling, Sliding, Session)
- Data Architecture & Lakehouse Engines: ETL vs ELT, Data Lake vs Data Warehouse vs Data Lakehouse (Apache Iceberg, Delta Lake, Apache Hudi)
- Data Serialization & In-Memory Formats: Protobuf, Apache Avro, Apache Thrift, Apache Arrow Zero-Copy Memory Mapping, and Feather
- Chapter 7A References