Bucket and Outlier

Bucket, Computed, and Outlier Patterns

The bucket pattern groups a stream of small, time-ordered records into one document per time window: 144,000 one-minute sensor readings become 2,400 hourly buckets of 60 samples, each carrying its own rolled-up count, sum, min and max. The computed pattern stores the result of an aggregation beside the data and updates it on write instead of recomputing it on every read.

Appending to a bucket, and the cost of not precomputingJavaScript
const at = new Date('2026-09-02T08:17:00Z'), temp = 21.4;
const hour = d => new Date(Math.floor(d.getTime() / 3600000) * 3600000);
db.readings_bucket.updateOne(                    // append; create the bucket if needed
  { sensorId: 's7', hour: hour(at), count: { $lt: 60 } },
  { $push: { samples: { m: at.getUTCMinutes(), t: temp } },
    $inc: { count: 1, sum: temp }, $min: { min: temp }, $max: { max: temp } },
  { upsert: true });
const lo = new Date('2026-09-01T00:00:00Z'), hi = new Date('2026-09-01T12:00:00Z');
cost('flat, 12 hours', ex(db.readings.find({ sensorId: 's7', at: { $gte: lo, $lt: hi } })));
cost('bucketed, 12 hours', ex(db.readings_bucket.find({ sensorId: 's7',
  hour: { $gte: lo, $lt: hi } })));
cost('grades, computed', db.enrollments.explain('executionStats').aggregate([
  { $group: { _id: '$courseId', avg: { $avg: '$grade' } } }]));
cost('grades, stored', ex(db.courses.find({}, { avgGrade: 1 })));
Output
flat, 12 hours          keys=720 docs=720 ms=1
bucketed, 12 hours      keys=12 docs=12 ms=0
grades, computed        keys=0 docs=200000 ms=127
grades, stored          keys=0 docs=200 ms=0

Bucketing paid three times over. collStats puts the flat collection at 2,879,488 bytes of data and 3,022,848 of index against 442,368 and 86,016 for the buckets: 6.5 times less data and 35 times less index, since the index holds one key per hour instead of one per minute. A twelve-hour window costs 12 examined documents instead of 720. The count: { $lt: 60 } in the filter is what makes it safe — the update lands in the current bucket while it has room and the upsert opens a new one when it does not, so no array can pass 60 elements. Time series collections (Time Series Collections) apply the whole pattern for you.

The computed pattern is the same bargain reversed: 200,000 documents and 127 ms to average grades on demand, against 200 documents and no measurable time to read averages a write already maintained. The outlier pattern covers the one document that breaks the design — the post with 900,000 comments. Keep the ordinary shape for everyone, flag that one (hasOverflow: true), and put its overflow in a side collection.