spark.read returns a DataFrameReader: set a format and options, then load a file, a glob or a directory (.json(path) is short for .format("json").load(path)). The CSV here is a copy of the customer file written with Python's csv module.
orders_json = spark.read.json("data/raw/orders.jsonl") # JSON Lines (Section 3.2)
customers = (spark.read.option("header", True).option("inferSchema", True)
.csv("data/raw/customers.csv")) # CSV (Section 3.7)
orders = spark.read.parquet("data/orders.parquet") # Parquet (Section 3.13)
catalog = spark.read.option("multiLine", True).json("data/raw/books.json") # one document
for name, df in [("orders_json", orders_json), ("customers", customers),
("orders", orders), ("catalog", catalog)]:
print(f"{name:<12} {df.count():>9,} rows {len(df.columns):>3} columns "
f"{df.rdd.getNumPartitions()} partitions")Output
orders_json 1,000,000 rows 10 columns 4 partitions customers 50,000 rows 5 columns 1 partitions orders 1,000,000 rows 10 columns 4 partitions catalog 1 rows 1 columns 1 partitions
The JSON reader expects one record per line; the pretty-printed catalog document needs multiLine and arrives as one row holding a books array to explode. Without header and a schema, CSV columns are named _c0, _c1 and typed as strings. The 238 MB JSON file became four partitions because Spark 129 sizes splits from the total bytes and the core count, capped at spark.sql.files.maxPartitionBytes (128 MB).