SIMD and Batch Processing

SIMD Instructions and Batch-at-a-Time Processing

Batches pay off again inside the CPU. SIMD (single instruction, multiple data) instructions apply one operation to several values in a wide register: four 32-bit floats with SSE, eight with AVX2, sixteen with AVX-512. A loop over a plain typed array, such as a 2,048-value vector, can use them; a loop over Python objects or heap tuples cannot. NumPy chooses SIMD kernels at run time and can be told not to, on this host's Core i5-6500:

simd.sh: the same NumPy kernels with and without AVX2Shell
python simd.py                                     # NumPy dispatches to AVX2 kernels
NPY_DISABLE_CPU_FEATURES="X86_V3" python simd.py   # fall back to the SSE-level baseline
Output
X86_V2 X86_V3* X86_V4? AVX512_ICL? AVX512_SPR? | exp 1.47 ms, max 0.15 ms, multiply 0.65 ms
X86_V2 X86_V3? X86_V4? AVX512_ICL? AVX512_SPR? | exp 3.49 ms, max 0.21 ms, multiply 0.66 ms

simd.py times exp, max and a multiplication over a million floats. With AVX2 (X86_V3*, the star marks the level in use) exp ran 2.4 to 2.7 times faster and max about 1.4 times; the memory-bound multiplication did not change. Vectorization itself is one of several ways to go faster:

Ways to speed up a data pipeline and when each helps
Technique What it does Helps when
Vectorization Many values per optimized operation Numeric work in Python loops
Multithreading Threads share one process Waiting on I/O; native code releasing the GIL
Multiprocessing Processes use several cores CPU-bound pure Python
Async I/O One thread juggles many waits Many network calls at once
Distributed processing Work spread over machines Data beyond one machine (Batch Processing with Apache Spark)