Embedding puts the child inside the parent: one document, one seek, one atomic write. Referencing puts the child in its own collection with a field pointing back: independent indexes and writes, and a join on every read that needs both.

cost, bytes and ex below are reused throughout this section.
const st = e => e.executionStats || e.stages[0].$cursor.executionStats;
const cost = (label, e) => { const s = st(e); print(label.padEnd(24) + 'keys='
+ s.totalKeysExamined + ' docs=' + s.totalDocsExamined + ' ms=' + s.executionTimeMillis); };
const bytes = c => { let n = 0; c.forEach(d => { n += bsonsize(d); }); return n; };
const ex = q => q.explain('executionStats');
cost('Q1 emb one findOne', ex(db.posts_emb.find({ slug: 'post-28' })));
cost('Q1 ref its comments', ex(db.comments.find({ postId: 28 }).sort({ createdAt: -1 })));
cost('Q2 ref indexed sort', ex(db.comments.find().sort({ createdAt: -1 }).limit(10)));
cost('Q2 emb unwind + sort', db.posts_emb.explain('executionStats').aggregate([
{ $unwind: '$comments' }, { $sort: { 'comments.createdAt': -1 } }, { $limit: 10 }]));
cost('Q3 emb multikey index', ex(db.posts_emb.find({ 'comments.authorId': 77 })));
print('Q3 bytes: ref ' + bytes(db.comments.find({ authorId: 77 })) + ', emb '
+ bytes(db.posts_emb.find({ 'comments.authorId': 77 })));Q1 emb one findOne keys=1 docs=1 ms=0 Q1 ref its comments keys=50 docs=50 ms=0 Q2 ref indexed sort keys=10 docs=10 ms=0 Q2 emb unwind + sort keys=0 docs=2000 ms=57 Q3 emb multikey index keys=11 docs=11 ms=0 Q3 bytes: ref 2508, emb 106674
Q1, the post page, favors embedding — but not on document counts. Both shapes seek once; the embedded shape then reads 1 document where the referenced shape reads 51 across two queries. On a 1 ms hop that second round trip costs more than the 50 extra reads, and $lookup does not save you: it is still two index seeks server-side.
Q2, the site-wide feed, is where embedding collapses. A createdAt inside an array cannot serve a top-level sort, so the pipeline scans all 2,000 posts, unwinds 49,640 comments and sorts them in memory; keys=0 is the tell. The referenced shape reads 10 off the createdAt index — 200x fewer examined, a gap that widens with the collection.
Q3 is a draw on counters and a rout on bytes. The multikey index locates the 11 matching posts in 11 examined documents, exactly as the authorId index locates the 11 comments — but each embedded match is a whole 7 KB post, so 106,674 bytes cross the wire instead of 2,508. And { 'comments.$': 1 } returns only the first match per document, so someone who commented twice on one post loses a comment.