Sample Data

Generating Realistic Sample Data

Five books never expose a broken paginator. Faker 15,502 (github.com/faker-js/faker (https://github.com/faker-js/faker 15,502 ), MIT, npm 2,036 i -D @faker-js/faker, version 10.6.0) supplies the volume, and a factory per collection keeps generation in one place: one function per document, every field overridable, no database access.

seed/factory.js — one factory per collectionJavaScript
export const makeBook = (authorId, over = {}) => ({
  _id: new Types.ObjectId(), authorId, title: faker.book.title(),
  year: faker.number.int({ min: 1950, max: 2026 }),
  genre: faker.helpers.arrayElement(['sci-fi', 'satire', 'history']),
  ...over              // the caller states what it cares about, faker fills the rest
});

seed/bulk.js then calls faker.seed(4213), builds 200 authors, 5,000 books pointing at them and 0-4 reviews per book with flatMap, and loads each collection with one insertMany(docs, { ordered: false }), which keeps batching past a rejected document: 200 authors, 5000 books, 10008 reviews, inserted in 613 ms.

faker.seed(4213) makes the strings and numbers deterministic — the same call sequence yields the same titles on every machine, which is why the review count is exactly 10,008 every run — but it does not fix the _id values, because new ObjectId() mixes in a timestamp and a random counter, so take stable ids from the fixture when a test needs them. Ten thousand documents in under a second is the point: cheap to rebuild, and big enough that a missing index shows up in explain() (Reading an explain Plan).