The Vectorized Execution Model

PostgreSQL 1,289 runs plans with the iterator (Volcano) model: each operator's next() returns one tuple, and every row pays for function calls, type dispatch and branches. A vectorized engine passes batches: DuckDB 61,228 's operators exchange vectors of 2,048 values per column, paying the overhead once per vector and running tight loops over plain arrays. The idea comes from the MonetDB/X100 research engine (2005), later Vectorwise.

Tuple-at-a-time iteration versus vector-at-a-time execution
Tuple-at-a-time iteration versus vector-at-a-time execution

Python shows the same contrast. A loop interprets bytecode and boxes an object per element; a NumPy or pandas 16,086 column expression runs one compiled loop over a contiguous array of machine numbers:

vectorize.py: loop versus vectorized in NumPy and pandasPython
import timeit
import numpy as np, pandas as pd
best = lambda fn: min(timeit.repeat(fn, number=1, repeat=5)) * 1000   # fastest run, in ms
data = list(range(1_000_000))
def loop():
    result = []
    for x in data:
        result.append(x * 2)
    return result
arr = np.array(data)                          # one contiguous block of int64
print(f"NumPy  loop  {best(loop):6.1f} ms, vectorized {best(lambda: arr * 2):5.2f} ms")
df = pd.read_csv("csv-out/orders.csv", usecols=["total"]).rename(columns={"total": "amount"})
print(f"pandas apply {best(lambda: df['amount'].apply(lambda x: x * 0.06)):6.1f} ms, "
      f"vectorized {best(lambda: df['amount'] * 0.06):5.2f} ms")
df["tax"] = df["amount"] * 0.06
df["total"] = df["amount"] + df["tax"]
Output
NumPy  loop    67.6 ms, vectorized  1.26 ms
pandas apply   21.0 ms, vectorized  0.11 ms

Over nine runs on the shared host, doubling a million numbers was 50 to 175 times faster vectorized, and a 6 percent tax on 100,000 order amounts 145 to 620 times faster than with apply, which looks columnar but still calls Python once per row. When a DataFrame step is slow, look for a per-row Python function first.