Kappa Architecture

The Kappa Architecture's Single Streaming Path

Jay Kreps, one of Kafka 129 's creators, answered in "Questioning the Lambda Architecture" (O'Reilly Radar, 2 July 2014). His argument: instead of maintaining a batch system to correct the streaming one, make the streaming system good enough to do both jobs. He named the result the Kappa architecture, noting that it might be "too simple of an idea to merit a Greek letter".

Kappa keeps all events in a retained log, such as a Kafka topic configured to keep data for as long as you need it. One stream processing job, with one code base, produces the serving tables. Reprocessing, the main reason Lambda had a batch layer, becomes a replay: when the logic changes, start a second copy of the new job reading from the beginning of the log, let it write to a new output table, and switch readers over once it has caught up. Then stop the old job and drop its table.

Kappa has its own costs. Replaying months of events through a streaming engine can be slower than a batch scan of the same data in columnar files, and the log must actually retain that history. Some computations, such as a model retrained over all customers, are naturally batch-shaped and awkward to express as a stream.