Structured Streaming treats a stream as an unbounded input table to which new rows are appended. You write an ordinary query against it; the engine runs it incrementally, by default as a series of micro-batches (about 100 ms end to end at best), keeping running aggregates as state and recording progress in a checkpoint directory so that a restarted query resumes where it stopped, with exactly-once results for replayable sources and idempotent sinks. Three choices define a query:
Output mode: append emits only new result rows, complete the whole result table each time, update only the rows that changed.
Trigger: run micro-batches back to back (the default), on a processingTime interval, or availableNow=True to process everything present and stop, which turns a stream into an incremental batch job.
Event time: withWatermark("ts", "10 minutes") tells the engine how late data may arrive, so it can finalize windows and drop old state.
Sources include files (CSV, JSON, Parquet 129 , ORC, text), Kafka 129 , and the test-only rate and socket sources; sinks include files, Kafka, foreachBatch (any batch writer, such as JDBC or a lakehouse table), and the debugging console and memory sinks.