The save mode decides what happens when the target already holds data: errorifexists (the default) raises an error, ignore skips the write like CREATE TABLE IF NOT EXISTS, append adds part files beside the old ones, and overwrite deletes the old data and writes.
from pyspark.errors import AnalysisException
daily = (spark.read.parquet("data/order_lines.parquet")
.groupBy(F.to_date("order_ts").alias("day")).agg(F.sum("qty").alias("copies")))
path = "out/daily_copies"
daily.write.mode("overwrite").parquet(path) # start with one copy
for mode in ("errorifexists", "ignore", "append", "overwrite"):
try:
daily.write.mode(mode).option("compression", "zstd").parquet(path)
print(f"{mode:<13} rows now {spark.read.parquet(path).count()}")
except AnalysisException as e:
print(f"{mode:<13} {e.getCondition()}")Output
errorifexists PATH_ALREADY_EXISTS ignore rows now 546 append rows now 1092 overwrite rows now 546
append doubled the 546 daily rows, so a retried append job double-counts. Make reruns safe by overwriting a whole output or only the partitions you rebuilt (Partitioned Table Writes), or use a table format with MERGE (Lakehouses, Data Quality and Governance). On plain files overwrite deletes first and is not atomic. Options are per format: compression (Parquet 129 defaults to snappy), CSV's header, sep and dateFormat, and maxRecordsPerFile.