PyAirbyte (pip 21,050 install airbyte, 0.71.1) runs Airbyte 128,715 connectors from Python with no platform: it installs a connector into its own virtual environment (or runs its Docker 514 image), reads its streams and writes them to a cache, DuckDB 61,228 by default or here PostgreSQL 1,289 . The File source supports only Python 3.10 and 3.11, so ingest/airbyte/ runs in a python:3.11-slim image (l2/pyairbyte:0.71.1):
source = ab.get_source("source-file", config={
"dataset_name": "orders",
"format": "jsonl",
"url": "/data/source/orders.jsonl", # the order service's export
"provider": {"storage": "local"},
})
source.check()
source.select_all_streams()
cache = PostgresCache(host="l2-pg", port=5432, username="postgres", password="booknest",
database="booknest", schema_name="airbyte_raw")
result = source.read(cache=cache)
for name, dataset in result.streams.items():
print(f"stream {name}: {len(dataset)} records")Connection check succeeded for `source-file`. • Read 100,000 records over 39 seconds (2,552.8 records/s, 0.67 MB/s). ... Sync completed at 03:44:53. Total time elapsed: 59 seconds stream orders: 100000 records
The cache table airbyte_raw.orders shows Airbyte's conventions: the source's columns plus _airbyte_raw_id, _airbyte_extracted_at and _airbyte_meta (per-record errors), with nested items kept as one json column and order_ts as character varying, because the File source inferred a string. Typing and unnesting are left to dbt 37,942 , the ELT split of ETL and ELT. Stream state goes to _airbyte_state, so incremental sources resume; schedule sync.py from any orchestrator task.