Comparing DataFrames in Tests

pyspark.testing (since Spark 3.5 129 ) provides assertDataFrameEqual and assertSchemaEqual. The first compares rows ignoring order unless checkRowOrder=True, with a relative tolerance (rtol, default 1e-5) for floating-point columns, and prints a readable diff.

assertDataFrameEqual: order, tolerance and a failing diffPython
from pyspark.testing import assertDataFrameEqual
schema = "genre STRING, revenue DOUBLE"
actual = spark.createDataFrame([("Fiction", 29.98), ("Travel", 18.75)], schema)
close = spark.createDataFrame([("Travel", 18.75), ("Fiction", 29.9800001)], schema)
assertDataFrameEqual(actual, close)               # row order ignored, rtol=1e-5 by default
print("equal, ignoring row order and tiny float differences")
wrong = spark.createDataFrame([("Fiction", 29.98), ("Travel", 18.70)], schema)
try:
    assertDataFrameEqual(actual, wrong)
except AssertionError as e:
    print(e)
Output
equal, ignoring row order and tiny float differences
[DIFFERENT_ROWS] Results do not match: ( 50.00000 % )
*** actual ***
  Row(genre='Fiction', revenue=29.98)
! Row(genre='Travel', revenue=18.75)
*** expected ***
  Row(genre='Fiction', revenue=29.98)
! Row(genre='Travel', revenue=18.7)

The ! marks (red in a terminal) flag differing rows. The tolerance applied only when the expected value was a DataFrame: an expected list of tuples with 29.9800001 failed on both rows here, so compare floats against a DataFrame and money as Decimal.