Spark 129 needs the delta-spark_4.1_2.13 and delta-storage jars plus Delta's session extension and catalog, set up in demos/ch08/delta/dspark.py. The script writes JSON, Columnar and Binary Formats's orders with a list-price gross column, partitioned by month, then deletes the four pending orders:
"""Chapter 3's orders as a Delta table: write, delete with a deletion vector, read history."""
from pyspark.sql import functions as F
from dspark import spark
P = "/home/dev/v7-l1/ch08/delta/orders" # a local path; delta-rs uses MinIO (8.7.4)
src = spark.read.parquet("/mnt/d/Books/Data Engineering/demos/ch03/out/orders.parquet")
line = lambda acc, i: (acc + i.qty * i.unit_price).cast("decimal(12,2)") # list-price total
orders = (src.withColumn("gross", F.aggregate("items", F.lit(0).cast("decimal(12,2)"), line))
.drop("items").withColumn("order_month", F.date_format("order_ts", "yyyy-MM")))
(orders.write.format("delta").mode("overwrite").partitionBy("order_month")
.option("delta.enableDeletionVectors", "true").save(P)) # version 0
spark.sql(f"DELETE FROM delta.`{P}` WHERE status = 'pending'") # version 1
spark.sql(f"""SELECT v, count(*) AS orders, sum(gross) FILTER (WHERE status <> 'cancelled')
AS gross FROM (SELECT 0 AS v, * FROM delta.`{P}` VERSION AS OF 0
UNION ALL SELECT 1, * FROM delta.`{P}`) GROUP BY v ORDER BY v""").show()Output
+---+------+----------+ | v|orders| gross| +---+------+----------+ | 0|100000|3303427.30| | 1| 99996|3303254.12| +---+------+----------+
Version 0 reproduces the 3,303,427.30 baseline in a second table format. Delta partitions by real columns such as order_month; generated columns let it derive them and prune on order_ts filters, its nearest equivalent to hidden partitioning. SQL adds MERGE INTO, OPTIMIZE ... ZORDER BY and liquid clustering (CLUSTER BY), which replaces fixed partitions.