Streaming and Appending

Streaming and Appending Records

Appending is one write() of one line in append mode (>>, or "a" in Python), whose O_APPEND flag puts every write at the current end even with several writers. A crash can leave a half-written last line, so a robust reader treats a malformed final line without its newline as an unfinished write, but fails on bad lines elsewhere. Reading streams too. This script totals revenue over the sample orders, once from one JSON array (made with jq 133,477 -s . data/orders.jsonl > orders-array.json) and once line by line:

peak_memory.py: parsing a whole array versus one line at a timePython
import json, resource, sys
mode, path = sys.argv[1:3]
with open(path, encoding="utf-8") as f:
    if mode == "array":                  # one document: parse everything at once
        total = sum(o["total"] for o in json.load(f))
    else:                                # JSON Lines: one record at a time
        total = sum(json.loads(line)["total"] for line in f)
mib = resource.getrusage(resource.RUSAGE_SELF).ru_maxrss / 1024
print(f"{mode:6} revenue {total:,.2f}  peak RSS {mib:,.0f} MiB")
Output
array  revenue 3,474,495.41  peak RSS 146 MiB
lines  revenue 3,474,495.41  peak RSS 11 MiB

The array needs 146 MiB for 37 MB of text; the stream stays at 11 MiB at any file size. Likewise split -l 25000 yields four valid parts for four workers. S3 general purpose buckets cannot append, so pipelines write a new part-NNNN.jsonl.gz object per batch.