Dynamic allocation
With spark.dynamicAllocation.enabled=true, use the executor count as maxExecutors and keep cores and memory as calculated. Spark then scales between your minimum and that ceiling.
Spark configuration calculator · YARN · EMR · Kubernetes
Enter your cluster and get executors, cores, executor memory, overhead and shuffle partitions, with the reasoning behind every number and ready-to-paste spark-submit flags.
vCPU and RAM as listed by AWS. On EMR, YARN gets less than the full RAM: enter the YARN memory of the instance type and set reserved memory to 0.
Kept free on each node for the OS and the Hadoop/YARN daemons.
5 cores is the usual sweet spot for I/O throughput. Raise the overhead to 20–40% for PySpark jobs with heavy Python UDFs or pandas.
Size of the largest shuffle stage (see "Shuffle Write" in the Spark UI). Leave 0 if you don't know it yet.
With spark.dynamicAllocation.enabled=true, use the executor count as maxExecutors and keep cores and memory as calculated. Spark then scales between your minimum and that ceiling.
Databricks runs one executor per worker and sizes it for you, so the executor part does not apply. The partition sizing does: use it to set spark.sql.shuffle.partitions, or let Adaptive Query Execution coalesce them.
On Spark 3.2+ AQE is on by default and merges small shuffle partitions at runtime. A generous partition count is safe: AQE shrinks it, it cannot grow it.
I'm Francesco, a freelance data engineer based in Italy. I work on Data Vault warehouses on SQL Server, PySpark pipelines and data migrations, and I built this tool for a part of that work I kept redoing by hand.
It is free and stays free. If your team is starting a Data Vault, migrating to Spark, or fighting a slow job, I can help.
Also by me: HubLinkSat, a Data Vault 2.0 code generator for SQL Server and Spark SQL.
Data Vault design, Spark tuning, SQL Server to lakehouse migrations.
Contact me on LinkedIn