Why an Embedded Database

Why an Embedded Analytical Database

An embedded (in-process) database is a library linked into your program: no server, no port, no network hop between query and data. SQLite 4,756 made that model the default for transactional storage on phones; DuckDB 61,228 applies it to analytics, a columnar, vectorized engine inside a Python process, notebook, CLI or dbt 37,942 run that reads Parquet 129 , CSV and JSON directly (Querying Files with DuckDB) and hands results to pandas 16,086 or Arrow 129 without a socket in between.

DuckDB began in 2018 in CWI's Database Architectures group in Amsterdam (Mark Raasveldt and Hannes Mühleisen) and reached 1.0.0 on 3 June 2024, promising that later versions read 1.0 files. It is MIT licensed; the non-profit DuckDB Foundation holds the intellectual property and DuckDB Labs leads development (github.com/duckdb/duckdb (https://github.com/duckdb/duckdb 41,877 )).

Embedded and server databases compared
SQLite DuckDB PostgreSQL 1,289 ClickHouse 29,491 server
Runs as Library Library Server Server (cluster optional)
Storage Row pages Columnar row groups Row heap Columnar parts
Built for Small transactions Scans and aggregates Mixed, many users Large, real-time analytics
Concurrency One writer One writing process Many writers Many writers

Use DuckDB where one machine holds the data, one process does the work, and the job is to scan and aggregate: notebooks, pipeline transformation steps, tests of SQL models, and querying files in a data lake. Do not make it the shared backend of a web application with many writers; that stays PostgreSQL's job (PostgreSQL or DuckDB?).