Ray Data

Ray Data: Distributed Data for ML Pipelines

Ray (github.com/ray-project/ray (https://github.com/ray-project/ray 43,965 ), Apache-2.0, 2.58.0 on 23 Aug 2026, pip 21,050 install "ray[data]") is a distributed runtime for Python, and Ray Data is its library for loading and transforming data on the way into machine learning: it streams Arrow 129 blocks through a pipeline of operators, schedules them across CPUs and GPUs, and hands batches to training or inference code (map_batches with a model on each GPU). It is built for last-mile ML preprocessing and batch inference rather than SQL-style joins and aggregations, which it supports only in basic form. Anyscale sells the managed version. Ray Data is not run here: its installation failed six times on this workstation's network (SSL: RECORD_LAYER_FAILURE while downloading the wheels), so it is described from its documentation.