Monitoring

Monitoring, Metrics, and Capacity Planning

db.adminCommand({ serverStatus: 1 }) is the single source for a server's counters and returns several hundred of them. A handful tell you almost everything, and a clusterMonitor user may read them all.

A health check built from serverStatusJavaScript
const s = db.adminCommand({ serverStatus: 1 });
const wt = s.wiredTiger.cache, q = s.queues.execution, o = s.opcounters;
print(`opcounters q=${o.query} i=${o.insert} u=${o.update} userAsserts=${s.asserts.user}`);
print(`cache=${(wt['maximum bytes configured'] / 1073741824).toFixed(1)}GB ` +
      `pagesRead=${wt['pages read into cache']}`);
print(`tickets r=${q.read.available}/${q.read.totalTickets}` +
      ` w=${q.write.available}/${q.write.totalTickets}` +
      ` queued=${s.globalLock.currentQueue.total}`);
Output
opcounters q=72 i=19 u=68 userAsserts=352
cache=63.3GB pagesRead=104
tickets r=36/36 w=36/36 queued=0

Tickets are the concurrency limiter and the best early warning here: every read and write takes one to enter the storage engine, and when available sits near zero with a non-zero queue, operations are waiting for the engine rather than the disk — almost always missing indexes or bulk scans. WiredTiger's cache takes half of RAM minus 1 GB by default, 63.3 GB here; watch pagesRead, because a working set that fits in cache barely reads pages while one that does not turns every query into disk I/O. userAsserts counts rejected operations — 352 here, each a failure deliberately caused earlier in this section.

For a live view, mongostat prints one line per second per server and mongotop ranks namespaces by time spent reading and writing, while db.currentOp() lists operations in flight and db.killOp(opid) ends a stuck one. For capacity, the ratio that matters is index size against available RAM, reported by db.stats() as indexSize: every lookup into an index larger than cache costs a disk read, and the cliff teams report is almost always that line being crossed. Alert on symptoms users feel — replication lag, available tickets, queue depth — because a server under index pressure shows modest CPU while every query waits on I/O. Meanwhile mongod writes FTDC into diagnostic.data inside your dbPath, so the data explaining last night's incident exists even if nothing was scraping it.