Batches pay off again inside the CPU. SIMD (single instruction, multiple data) instructions apply one operation to several values in a wide register: four 32-bit floats with SSE, eight with AVX2, sixteen with AVX-512. A loop over a plain typed array, such as a 2,048-value vector, can use them; a loop over Python objects or heap tuples cannot. NumPy chooses SIMD kernels at run time and can be told not to, on this host's Core i5-6500:
python simd.py # NumPy dispatches to AVX2 kernels
NPY_DISABLE_CPU_FEATURES="X86_V3" python simd.py # fall back to the SSE-level baselineX86_V2 X86_V3* X86_V4? AVX512_ICL? AVX512_SPR? | exp 1.47 ms, max 0.15 ms, multiply 0.65 ms X86_V2 X86_V3? X86_V4? AVX512_ICL? AVX512_SPR? | exp 3.49 ms, max 0.21 ms, multiply 0.66 ms
simd.py times exp, max and a multiplication over a million floats. With AVX2 (X86_V3*, the star marks the level in use) exp ran 2.4 to 2.7 times faster and max about 1.4 times; the memory-bound multiplication did not change. Vectorization itself is one of several ways to go faster:
| Technique | What it does | Helps when |
|---|---|---|
| Vectorization | Many values per optimized operation | Numeric work in Python loops |
| Multithreading | Threads share one process | Waiting on I/O; native code releasing the GIL |
| Multiprocessing | Processes use several cores | CPU-bound pure Python |
| Async I/O | One thread juggles many waits | Many network calls at once |
| Distributed processing | Work spread over machines | Data beyond one machine (Batch Processing with Apache Spark) |