Executors need your modules and JVM libraries too. --py-files ships .py, .zip or .egg files onto every worker's Python path; --packages resolves Maven 129 coordinates with Ivy, dependencies included; --jars ships local JARs. Here jobs/export_avro.py imports top_titles from a small booknest package and writes Avro 129 , which needs the separate spark-avro module.
(cd jobs && zip -qr ../booknest.zip booknest) # a package for the executors
spark-submit --py-files booknest.zip \
--packages org.apache.spark:spark-avro_2.13:4.2.0 \
jobs/export_avro.py out/orders_daily out/top_titles_avroOutput
:: resolving dependencies :: org.apache.spark#spark-submit-parent-...;1.0 found org.apache.spark#spark-avro_2.13;4.2.0 in central ... +------------------------+----------+ |title |revenue | +------------------------+----------+ |Patterns of the Deep Web|8878336.00| |Salt and Saffron |6791664.00| |The Quiet Harbor |6550749.92| +------------------------+----------+
The Scala suffix (_2.13) and the version must match your Spark 129 build exactly. Libraries with compiled code need more than --py-files: pack a virtual environment with venv-pack and pass it with --archives.