Transformers.js 1,113 (huggingface/transformers.js (https://github.com/huggingface/transformers.js 16,329 ), Apache-2.0, @huggingface/transformers 4.3.0) wraps ONNX Runtime Web 231,552 in the pipeline() API of Python's Transformers: name a task and a model from the Hugging Face 1,113 Hub, and it downloads the ONNX weights, tokenizes the input and runs the model. device: 'webgpu' selects the WebGPU provider (the default is WebAssembly on the CPU), and dtype picks the weight precision ('fp32', 'fp16', 'q8', 'q4' and others). This page turns BookNest's titles into sentence embeddings with all-MiniLM-L6-v2, a 22-million-parameter model, and ranks them against a search:
<script type="module">
import { pipeline } from 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@4.3.0/+esm';
let start = performance.now();
const embed = await pipeline('feature-extraction', 'Xenova/all-MiniLM-L6-v2',
{ device: 'webgpu', dtype: 'fp32' }); // ONNX Runtime Web, WebGPU provider
const load = (performance.now() - start) / 1000;
const titles = ['The Quiet Harbor', 'Patterns of the Deep Web', 'Salt and Saffron',
'Small Steps to Big Summits', "The Clockmaker's Paradox", 'Gardens in Glass'];
start = performance.now();
const vectors = await embed(['a cookbook of spicy recipes', ...titles],
{ pooling: 'mean', normalize: true }); // one unit vector per text
const [query, ...books] = vectors.tolist();
const run = performance.now() - start;
const cosine = (a, b) => a.reduce((sum, x, i) => sum + x * b[i], 0);
const ranked = titles.map((t, i) => [cosine(query, books[i]), t]).sort((a, b) => b[0] - a[0]);
console.log(`load ${load.toFixed(1)} s, 7 texts ${run.toFixed(0)} ms, ${query.length} dims`);
for (const [score, title] of ranked.slice(0, 3)) console.log(score.toFixed(3), title);
window.__done = true;
</script>load 28.7 s, 7 texts 434 ms, 384 dims 0.267 Salt and Saffron 0.154 Patterns of the Deep Web 0.145 Gardens in Glass
"Salt and Saffron" shares no word with the query, yet its vector lies closest: the model has learned that saffron is a spice. The first run includes building the model's pipelines: a second run of the same seven texts took 23 ms. Loading took about 30 seconds (28.7 to 33.2 in three runs) because the fp32 weights, about 90 MB, came from the Hub on a cold cache; the browser's Cache API keeps them for the next visit, and dtype: 'q8' or 'fp16' cuts the download to a quarter or a half. Keep models this small for interactive pages. Image Generation APIs goes on to image generation, which needs models of gigabytes and, on this machine, runs on the CPU in Python rather than in the browser.