BullMQ Queues

Background Job Queues with BullMQ

A cron task runs on a clock; a queue runs on demand, survives a crash, retries, and spreads work across as many machines as you start. When a request wants to resize an image or send mail, enqueue the work, return 202 Accepted with a job id, and let a worker do it. BullMQ 9,449 (MIT, 6.3.6, github.com/taskforcesh/bullmq (https://github.com/taskforcesh/bullmq 9,449 ), npm 2,036 i bullmq) is the standard Node implementation, built on Redis 2,763 and on Lua scripts that make "take a job and mark it active" atomic even with fifty workers competing.

A producer and a worker sharing one Redis queueJavaScript
import { Queue, Worker, QueueEvents } from 'bullmq';
import IORedis from 'ioredis';
const connection = new IORedis({ maxRetriesPerRequest: null });   // required by BullMQ
const thumbnails = new Queue('thumbnails', { connection });
await thumbnails.add('resize', { id: 4821, width: 320 }, {
  attempts: 5, backoff: { type: 'exponential', delay: 1000 },
  removeOnComplete: 1000, removeOnFail: 5000,
});
const worker = new Worker('thumbnails', async (job) => {
  return { bytes: await resize(job.data.id, job.data.width) };
}, { connection, concurrency: 8 });
new QueueEvents('thumbnails', { connection })
  .on('completed', ({ jobId }) => console.log('done', jobId));

Running this needs a Redis server, which the machine used for the measurements in this section does not have, so no output is printed here; every other listing in this section was run as shown.

attempts with exponential backoff turns a flaky API into a job that eventually succeeds, doubling the delay from one second per retry. removeOnComplete and removeOnFail bound history; without them a busy queue fills Redis with finished jobs. concurrency is how many jobs one worker runs at once — raise it for I/O-bound work, keep it low for CPU-bound work and add processes instead, because a job that blocks the event loop stalls every other job in the same worker. Point the worker at a separate file and BullMQ runs it in a sandboxed child process.

Two design rules save the most pain. Make jobs idempotent: at-least-once delivery means a job that timed out may run twice, so give each one a deterministic id (jobId: 'thumb:4821:320'). Keep payloads small — an id, not the image; a megabyte per job turns your cache into a slow message bus. For visibility, bull-board 3,493 (MIT, @bull-board/api 9.10.1, github.com/felixmosh/bull-board (https://github.com/felixmosh/bull-board 3,493 )) mounts a web UI over the queue.