Counting and Distinct

Counting, Distinct, and Cursor Options

Counting has two methods with very different contracts. countDocuments() takes a filter and wraps an aggregation — $match then $group — so the number is exact, even after an unclean shutdown and with orphaned documents on a sharded cluster. estimatedDocumentCount() reads collection metadata, takes no filter, and returns in constant time.

Exact counts, fast counts, distinct values, and a cursor chain
db.books.countDocuments({ tags: 'scifi' })
db.books.estimatedDocumentCount()
db.books.distinct('tags')
db.books.distinct('editions.format', { 'author.country': 'US' })
db.books.find({ tags: 'scifi' }, { title: 1, _id: 0 }).sort({ year: -1 }).skip(1).limit(2)
Output
5
8
[ 'classic', 'cyberpunk', 'fantasy', 'scifi' ]
[ 'ebook', 'hardcover', 'paperback' ]
[ { title: 'Snow Crash' }, { title: 'Neuromancer' } ]

countDocuments() also accepts limit, skip, hint and maxTimeMS as a second argument; hint is worth remembering, because an empty filter otherwise counts by scanning. It rejects $where, $near and $nearSphere, and the older cursor.count() and collection.count() are deprecated.

distinct() unwinds arrays, which is why tags returns four values rather than eight arrays, and it takes an optional filter and collation. Its returned array must fit in a 16 MB BSON document, so a high-cardinality field needs aggregate([{ $group: { _id: '$field' } }]), which streams. On a sharded collection distinct() may include orphaned documents and is barred from transactions — $group again.

Everything else you ask of a read is a cursor method, chained before iteration begins. Chaining order is for the reader, not the server: skip and limit always apply after sort. Put maxTimeMS on every query an HTTP handler issues — without it a slow scan holds a connection until the client gives up, while the server keeps working on a result nobody is waiting for.

Cursor methods that tune a read
Method Effect
.sort() .skip() .limit() order and window the results (Cursors and Sorting)
.hint({ field: 1 }) force an index instead of letting the planner choose
.batchSize(n) documents per network round trip
.maxTimeMS(n) abort the operation server-side after n milliseconds
.collation({ locale, strength }) language-aware, case-insensitive comparison
.explain('executionStats') report the plan instead of the data (Reading an explain Plan)