Chapter 7A References
Chapter 7A References
These sources support the shared concepts in Chapter 7A: stream and batch execution, lakehouse table formats, and serialization boundaries.
Books
- Tyler Akidau, Slava Chernyak, and Reuven Lax, Streaming Systems: The What, Where, When, and How of Large-Scale Data Processing, O’Reilly, 2018.
- Martin Kleppmann, Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Resilient, and Maintainable Systems, O’Reilly, 2017.
Websites
- Apache Spark Structured Streaming guide — event-time processing, watermarks, windows, and stateful queries.
- Apache Flink stateful stream processing and event-time watermarks — keyed state, timers, checkpoints, and watermark behavior.
- Apache Beam programming guide — the runner-independent stream and table model.
- Apache Iceberg documentation — table snapshots, manifests, schema evolution, and time travel.
- Delta Lake documentation — the transaction log, ACID table operations, schema evolution, and time travel.
- Apache Hudi documentation — incremental ingestion, file versions, and copy-on-write or merge-on-read layouts.
- Apache Arrow columnar format and Feather documentation — column buffers, validity data, IPC exchange, and Feather V2.
- Protocol Buffers encoding guide — field numbers, wire types, and compatibility behavior.
- Apache Avro specification and Apache Thrift specification — schema models, encodings, and interface contracts.