Relational design starts from the entities and normalizes until every fact lives in exactly one place. MongoDB 1,815 design starts from the other end — from the queries the application will run and how often it runs each one. There is no single correct document shape for a blog, an order book or a course catalog, only shapes that make your frequent queries cheap and your rare ones expensive. Schema design is a judgment call, which is why it should never be argued in the abstract.
Everything below is measured. Two collections hold the same 2,000 posts and 49,640 comments — posts_emb with the comments inside each post, posts_ref with them in a separate collection — so the same question goes to both shapes and the cost comes off explain('executionStats'), on MongoDB 8.3.11 and the machine of Indexes and Query Performance.
db = db.getSiblingDB('blog');
db.posts_ref.drop(); db.comments.drop(); db.posts_emb.drop();
let z = 20260922; // mulberry32: same data every run
const r = () => { z = (z + 0x6D2B79F5) | 0;
let t = Math.imul(z ^ (z >>> 15), 1 | z); t = (t + Math.imul(t ^ (t >>> 7), 61 | t)) ^ t;
return ((t ^ (t >>> 14)) >>> 0) / 4294967296; };
const t0 = Date.parse('2025-01-01'), span = 600 * 86400 * 1000, posts = [], com = [];
for (let i = 1; i <= 2000; i++) {
const createdAt = new Date(t0 + Math.floor(r() * span)), n = Math.floor(r() * 51);
posts.push({ _id: i, slug: 'post-' + i, title: 'Post number ' + i, commentCount: n,
body: 'b'.repeat(1080), authorId: 1 + Math.floor(r() * 50), createdAt,
tags: ['mongo', 'node', 'react'].slice(0, 1 + Math.floor(r() * 3)) });
for (let k = 0; k < n; k++) com.push({ postId: i, body: 'c'.repeat(150),
authorId: 1 + Math.floor(r() * 4000),
createdAt: new Date(createdAt.getTime() + Math.floor(r() * 3e9)) });
}
db.posts_ref.insertMany(posts);
for (let i = 0; i < com.length; i += 10000) db.comments.insertMany(com.slice(i, i + 10000));
db.posts_ref.aggregate([{ $lookup: { from: 'comments', localField: '_id',
foreignField: 'postId', as: 'comments' } }, { $out: 'posts_emb' }]);
db.posts_emb.createIndex({ slug: 1 }); db.posts_emb.createIndex({ 'comments.authorId': 1 });
db.posts_ref.createIndex({ slug: 1 }); db.comments.createIndex({ authorId: 1 });
db.comments.createIndex({ createdAt: -1 });
db.comments.createIndex({ postId: 1, createdAt: -1 });Deriving the embedded shape from the referenced one with $lookup plus $out guarantees both collections hold identical facts. Each post ends up with 0 to 50 comments, 24.8 on average; the busiest is post-28 with 50. stats() reports posts_emb documents averaging 7,003 bytes for 13.4 MB, against 1,239 bytes and 2.4 MB for posts_ref plus 10.8 MB of comments and 1.9 MB of index for the join.