The Aggregation Pipeline

find() answers questions about documents you already store. Aggregation answers questions about the data between documents: revenue per category, the median basket size, a report joining orders to the customers who placed them. You express all of it as a pipeline — an array of stages, each a small transformation feeding the next, like grep | sort | uniq -c. Stages share no private protocol, so you can insert, delete or reorder one without rewriting the rest.

Every example below runs against a shop database on a MongoDB 8.3 1,815 server: 14 orders and 6 customers (_id, name, city, country, tier, since). Short string _id values keep the output narrow.

A document from the orders collection
db.orders.insertOne({
  _id: 'o-1003', customerId: 'c1', placedAt: ISODate('2026-02-02T11:40:00Z'),
  status: 'delivered', channel: 'web', shipping: 0,
  items: [
    { sku: 'AU-31', name: 'Studio Headphones', category: 'audio', qty: 1, price: 180 },
    { sku: 'AU-32', name: 'USB Microphone', category: 'audio', qty: 1, price: 75 },
    { sku: 'AC-51', name: 'Laptop Stand', category: 'accessories', qty: 1, price: 45 }
  ]
})

MongoDB 8.3 ships more than forty stages; you will use eight constantly. The distinction that matters most is streaming against blocking. Streaming stages — $match, $limit, $skip, $project, $addFields, $set, $unset, $replaceWith, $unwind — pass documents through one at a time and cost almost nothing in memory. Blocking stages — $group, $sort without an index, $bucket, $bucketAuto, $sortByCount, $setWindowFields — must collect their input before they can emit anything, and each is capped at 100 MB of RAM (Pipeline Performance). $lookup, $unionWith, $graphLookup and $facet read a second stream; $out and $merge write a collection and must come last.

Subsections