Curated Spark resources: PySpark fundamentals, performance tuning, and Spark for large-scale ML feature engineering.
The textbook. Covers DataFrames, SQL, Structured Streaming, and MLlib. Read chapters 1-7 before anything else.
The API reference you will tab-switch to constantly. Bookmark the DataFrame and functions modules.
Adaptive Query Execution, partition pruning, join strategies. This is where most production Spark pain lives.
ACID transactions on Spark. Almost every serious Spark data lake uses Delta or Iceberg. Start here.
Some course links above are affiliate links. If you enroll, we may earn a small commission at no extra cost to you.
New resources and perspective on building AI-ready data systems, a few times a month. No spam.