Creating a SparkSession

SparkSession is the single entry point for DataFrames, SQL and the catalog. Its builder collects settings, and getOrCreate() either starts the driver's SparkContext or returns the session already running, so settings such as the master only count on the call that creates it.

Building and inspecting a SparkSessionPython
from pyspark.sql import SparkSession
spark = (SparkSession.builder
         .appName("booknest-dataframes")
         .master("local[4]")                          # leave out when using spark-submit
         .config("spark.sql.session.timeZone", "UTC")  # BookNest timestamps are UTC
         .config("spark.ui.port", "33040")
         .getOrCreate())
again = SparkSession.builder.appName("ignored").getOrCreate()   # returns the running one
print(spark.version, again is spark, spark.sparkContext.appName, spark.sparkContext.uiWebUrl)
print(spark.conf.get("spark.sql.session.timeZone"), spark.conf.get("spark.sql.ansi.enabled"),
      spark.conf.get("spark.sql.shuffle.partitions"))
spark.stop()
Output
4.2.0 True booknest-dataframes http://10.255.255.254:33040
UTC true 200

The second app name was ignored; spark.sql.* settings can still change through spark.conf.set. The time zone matters: this workstation runs at UTC+8, and without the setting Spark 129 shows BookNest's UTC timestamps, and dates derived from them, in local time. Listings up to Partitioning and Caching run in a session like this one, with pyspark.sql.functions imported as F.