A student takes many courses and a course holds many students. MongoDB 1,815 offers three ways to say that: an array of ids on one side, arrays on both sides, or a join collection with one document per pair. The data is 200 courses, 20,000 students and ten enrollments each — 200,000 pairs, rosters averaging 1,000 students.
const line = (label, q) => { const s = ex(q()).executionStats;
print(label.padEnd(20) + 'docs=' + String(s.totalDocsExamined).padStart(4) + ' returned='
+ String(s.nReturned).padStart(4) + ' bytes=' + String(bytes(q())).padStart(6)); };
print('course 7: ' + db.courses.findOne({ _id: 7 }).studentIds.length + ' ids in '
+ bsonsize(db.courses.findOne({ _id: 7 })) + ' bytes; one enrollment '
+ bsonsize(db.enrollments.findOne()) + ' bytes');
line('array on student', () => db.students.find({ _id: 12345 }, { courseIds: 1 }));
line('array on course', () => db.courses.find({ studentIds: 12345 }));
line('join collection', () => db.enrollments.find({ studentId: 12345 }));
db.enrollments.find({ courseId: 7 }, { _id: 0 }).sort({ grade: -1, studentId: 1 }).limit(2)
.forEach(d => print('student ' + d.studentId + ' grade ' + d.grade + ' since '
+ d.enrolledAt.toISOString().slice(0, 10)));course 7: 1059 ids in 9531 bytes; one enrollment 82 bytes array on student docs= 1 returned= 1 bytes= 100 array on course docs= 10 returned= 10 bytes= 89523 join collection docs= 10 returned= 10 bytes= 820 student 1924 grade 100 since 2026-02-18 student 5606 grade 100 since 2026-01-09
Keeping courseIds on the student answers "which courses?" in a single 100-byte read, which is why arrays on the short side are so attractive. Asking the same question of the roster side costs 89,523 bytes for ten answers: the multikey index on studentIds finds the right courses, but each drags 9.5 KB of roster along. A projection fixes the bytes and not the writes — enrolling one student rewrites a 9.5 KB document and its multikey keys, and a roster has no ceiling.
The last two output lines settle the argument. grade and enrolledAt belong to the pair, and an array of ids has nowhere to put them. Once a relationship carries its own fields the join collection is the only correct model: index it both ways ({ studentId: 1, courseId: 1 } and { courseId: 1, grade: -1 }) and each side becomes one indexed seek at 82 bytes per pair.