Reading Files

Reading CSV, JSON and Parquet into DataFrames

spark.read returns a DataFrameReader: set a format and options, then load a file, a glob or a directory (.json(path) is short for .format("json").load(path)). The CSV here is a copy of the customer file written with Python's csv module.

Reading JSON Lines, CSV and Parquet with one reader API
orders_json = spark.read.json("data/raw/orders.jsonl")              # JSON Lines (Section 3.2)
customers = (spark.read.option("header", True).option("inferSchema", True)
             .csv("data/raw/customers.csv"))                       # CSV (Section 3.7)
orders = spark.read.parquet("data/orders.parquet")                  # Parquet (Section 3.13)
catalog = spark.read.option("multiLine", True).json("data/raw/books.json")  # one document
for name, df in [("orders_json", orders_json), ("customers", customers),
                 ("orders", orders), ("catalog", catalog)]:
    print(f"{name:<12} {df.count():>9,} rows {len(df.columns):>3} columns "
          f"{df.rdd.getNumPartitions()} partitions")
Output
orders_json  1,000,000 rows  10 columns 4 partitions
customers       50,000 rows   5 columns 1 partitions
orders       1,000,000 rows  10 columns 4 partitions
catalog              1 rows   1 columns 1 partitions

The JSON reader expects one record per line; the pretty-printed catalog document needs multiLine and arrives as one row holding a books array to explode. Without header and a schema, CSV columns are named _c0, _c1 and typed as strings. The 238 MB JSON file became four partitions because Spark 129 sizes splits from the total bytes and the core count, capped at spark.sql.files.maxPartitionBytes (128 MB).