Before you draw a single document, write down the operations: what each reads or writes, how often it runs, and which fields it filters and sorts on. The list is short — most applications have fewer than twenty distinct queries — and it is the only input that matters, since a document shape is a bet on which operations get to be one seek. For this section's blog it is five lines long:
Render a post page — very high rate, filters on slug; needs the post, the author's name and the first twenty comments.
Post a comment — medium rate, filters on postId; appends one comment and bumps a counter.
Site-wide recent comments — low rate, no filter, sorts on createdAt; needs the ten newest.
A commenter's history — rare, filters on authorId.
Moderate one comment — rare, filters on the comment's _id.
The first line is the whole design: it runs orders of magnitude more often than the rest and wants a post, an author name and twenty comments together, so the shape that returns all three in one seek wins even if the last three — rare enough to pay for with an index or a round trip — turn awkward.
If the application already exists, measure the rates instead of guessing. Turn on the profiler from The Profiler, let a load test run, then group system.profile by the set of filter field names — the keys of command.filter, via $objectToArray and $map — summing docsExamined per group. Filter values differ on every call, so grouping on them is useless; grouping on names collapses 300 lookups of 300 slugs into one readable line.
Then ask three questions of every relationship. How many are on the many side, and is it bounded? Is the child ever needed without its parent? Do parent and child change at different rates, given that MongoDB 1,815 rewrites the whole document either way?