dlt

dlt: A Python-Native Load Library

dlt 876,044 (data load tool, github.com/dlt-hub/dlt (https://github.com/dlt-hub/dlt 5,925 ), Apache 2.0, by dltHub; 1.30.0 on 11 August 2026; pip 21,050 install "dlt[postgres]") is a library, not a service: you write a generator that yields Python dictionaries, and dlt infers a schema, normalizes nested data into child tables, loads to a destination and keeps state between runs. It runs wherever Python runs, including inside an Airflow 129 , Dagster 177,056 or Prefect 74,615 task, and supports Python 3.10 to 3.14.

dlt's building blocks
Concept What it does
@dlt.resource A generator of records, with a table name, primary key and write disposition
dlt.pipeline Binds a destination and a dataset (schema); run() extracts, normalizes, loads
Write disposition append, replace, or merge (upsert on the primary key)
dlt.sources.incremental A cursor field whose last value is saved as pipeline state

Under the hood each run has three stages: extract writes the yielded items to local files, normalize unnests lists into child tables linked by _dlt_parent_id and _dlt_list_idx and infers column types, and load copies the files into the destination in a transaction, recording each load in _dlt_loads. State lives in the destination too (_dlt_pipeline_state), so a pipeline can resume on a new machine. dlt also ships verified sources (SQL databases, REST APIs, cloud storage), groups resources into a @dlt.source, and lets a schema contract freeze or evolve columns when the data changes.