ONNX Runtime 231,552 (microsoft/onnxruntime (https://github.com/microsoft/onnxruntime 21,938 ), MIT) runs models in the ONNX format, which PyTorch 11,702 , TensorFlow 7,068 and scikit-learn can export. Its browser build, onnxruntime-web (1.30.0), picks a backend by execution provider: 'wasm' on the CPU, 'webnn', or 'webgpu', which implements each operator (MatMul, Conv, Softmax and so on) as a WGSL compute shader. Load ort.webgpu.min.js and ask for it:
<script src="https://cdn.jsdelivr.net/npm/onnxruntime-web@1.30.0/dist/ort.webgpu.min.js"></script>
<script type="module">
const session = await ort.InferenceSession.create('booknest-score.onnx',
{ executionProviders: ['webgpu'] });
const books = [[14.99, 4.6], [39.5, 4.3], [24, 4.8], [18.75, 4.1], [16.2, 4.5], [21.3, 4.4]];
const features = new ort.Tensor('float32', books.flat(), [6, 2]); // (price, rating) rows
const { score } = await session.run({ features });
console.log('scores:', [...score.data].map((s) => s.toFixed(3)).join(' '));
console.log('WebGPU device in use:', ort.env.webgpu.device instanceof GPUDevice);
window.__done = true;
</script>scores: 0.657 0.236 0.646 0.369 0.596 0.484 WebGPU device in use: true
The model, 191 bytes, was written with the onnx Python package (demos/ch04/make_score_model.py): a "recommend" score, sigmoid(-0.05 * price + 2 * rating - 7.8), as MatMul, Add and Sigmoid nodes. The scores match ONNX Runtime's Python CPU build to three decimals, and with verbose logging the session reported "All nodes placed on [WebGpuExecutionProvider]". Operators the provider lacks fall back to the CPU, each fallback costing a copy between GPU and CPU memory, so check that log for real models. The whole page took about 5.6 seconds, most of it downloading the 28 MB WebAssembly runtime, which browsers cache.