Writing Files

Writing DataFrames Back to CSV, JSON and Parquet

df.write mirrors the reader: a format, options, a save mode and a path. The path names a directory, not a file, because every task writes its own part file in parallel.

Writing the customer table as CSV, JSON and ParquetJavaScript
import os, re
UUID = r"[0-9a-f]{8}(-[0-9a-f]{4}){3}-[0-9a-f]{12}"           # shortened for printing
customers = spark.read.parquet("data/customers.parquet")
for fmt in ("csv", "json", "parquet"):
    path = f"out/customers_{fmt}"
    customers.write.format(fmt).mode("overwrite").option("header", True).save(path)
    files = sorted(n for n in os.listdir(path) if not n.startswith("."))   # skip .crc
    size = sum(os.path.getsize(os.path.join(path, n)) for n in files)
    print(f"{fmt:<8} {size:>9,} bytes", files[0], re.sub(UUID, "<uuid>", files[-1]),
          f"({len(files) - 1} parts)")
Output
csv      2,834,534 bytes _SUCCESS part-00001-<uuid>-c000.csv (2 parts)
json     5,984,448 bytes _SUCCESS part-00001-<uuid>-c000.json (2 parts)
parquet    831,528 bytes _SUCCESS part-00001-<uuid>-c000.snappy.parquet (2 parts)

Each of the DataFrame's two partitions became a part file, named by partition number and a per-job UUID, and the commit protocol wrote the empty _SUCCESS marker last, so downstream jobs can wait for it. Parquet 129 came out at under a third of the CSV and a seventh of the JSON; CSV also lost the types, and the other two formats silently ignored header. For a single file, coalesce(1) first (Repartition and Coalesce), and only for small results.