Many-to-Many

Modeling Many-to-Many Relationships

A student takes many courses and a course holds many students. MongoDB 1,815 offers three ways to say that: an array of ids on one side, arrays on both sides, or a join collection with one document per pair. The data is 200 courses, 20,000 students and ten enrollments each — 200,000 pairs, rosters averaging 1,000 students.

Three shapes for the same many-to-many relationshipJavaScript
const line = (label, q) => { const s = ex(q()).executionStats;
  print(label.padEnd(20) + 'docs=' + String(s.totalDocsExamined).padStart(4) + '  returned='
    + String(s.nReturned).padStart(4) + '  bytes=' + String(bytes(q())).padStart(6)); };
print('course 7: ' + db.courses.findOne({ _id: 7 }).studentIds.length + ' ids in '
  + bsonsize(db.courses.findOne({ _id: 7 })) + ' bytes; one enrollment '
  + bsonsize(db.enrollments.findOne()) + ' bytes');
line('array on student', () => db.students.find({ _id: 12345 }, { courseIds: 1 }));
line('array on course', () => db.courses.find({ studentIds: 12345 }));
line('join collection', () => db.enrollments.find({ studentId: 12345 }));
db.enrollments.find({ courseId: 7 }, { _id: 0 }).sort({ grade: -1, studentId: 1 }).limit(2)
  .forEach(d => print('student ' + d.studentId + ' grade ' + d.grade + ' since '
    + d.enrolledAt.toISOString().slice(0, 10)));
Output
course 7: 1059 ids in 9531 bytes; one enrollment 82 bytes
array on student    docs=   1  returned=   1  bytes=   100
array on course     docs=  10  returned=  10  bytes= 89523
join collection     docs=  10  returned=  10  bytes=   820
student 1924 grade 100 since 2026-02-18
student 5606 grade 100 since 2026-01-09

Keeping courseIds on the student answers "which courses?" in a single 100-byte read, which is why arrays on the short side are so attractive. Asking the same question of the roster side costs 89,523 bytes for ten answers: the multikey index on studentIds finds the right courses, but each drags 9.5 KB of roster along. A projection fixes the bytes and not the writes — enrolling one student rewrites a 9.5 KB document and its multikey keys, and a roster has no ceiling.

The last two output lines settle the argument. grade and enrolledAt belong to the pair, and an array of ids has nowhere to put them. Once a relationship carries its own fields the join collection is the only correct model: index it both ways ({ studentId: 1, courseId: 1 } and { courseId: 1, grade: -1 }) and each side becomes one indexed seek at 82 bytes per pair.