This will serve as the index for all my notes, brain dumps, and the links between them, allowing me to interrelate different concepts.
This is the origin of my in-depth Technical Newsletter: The Pragmatic Data Engineer’s Playbook
DataEngineering Spark internals Optimization Programming Airflow BestPractices Python Debug
Index
- Dec 20, 2025: Range Joins Optimization
- Nov 30, 2025: HashPartitioning in Apache Spark
- Nov 29, 2025: HashPartitioner in Apache Spark
- Nov 05, 2025: What are MapOutputTracker, BlockManager, and ESS in Spark Architecture?
- Nov 03, 2025: yarn top - details
- Nov 02, 2025: Internals of Shuffle Hash Join
- Nov 01, 2025: Using Context Objects in Python
- Oct 25, 2025: Why Do Parquet and ORC Store Data That Way?
- Oct 24, 2025: Sum Types in SQL
- Oct 23, 2025: PyArrow - Read-Write into Hive External Tables
- Oct 22, 2025 (Updated): Right Sizing Spark Executors
- Sep 13, 2025: Dataclasses - When to and not to use
- Sep 13, 2025: What the heck are Percentiles?
- Sep 07, 2025: Making Spark UI Jobs more Readable
- Aug 31, 2025: Understanding FetchFailedException Message
- Aug 31, 2025: Iceberg - Data and Delete File Relationship
- Aug 31, 2025: Why does Spark convert Columnar Data into Row Orientation?
- Aug 31, 2025: Data-aware scheduling from callbacks in Airflow
- Aug 31, 2025: Querying Airflow Meta DB
- Aug 29, 2025: 9 Clean Code Principles