Zero-Downtime Data Migrations

Taking the API down while a script rewrites ten million documents is not an option, and a rename is never one step: while it runs, old and new code talk to the same collection. Make the change additive first and destructive much later — expand, migrate, contract.

Renaming reviews.body to reviews.text without downtime
Renaming reviews.body to reviews.text without downtime

Each phase is its own deploy, and only the last destroys anything. The backfill walks the collection in _id order and remembers where it stopped, so an interrupted run resumes:

scripts/backfill.js — a resumable backfill over 10,008 reviewsJavaScript
let lastId = null, done = 0;
for (;;) {
  const filter = { text: { $exists: false }, ...(lastId && { _id: { $gt: lastId } }) };
  const batch = await reviews.find(filter).sort({ _id: 1 }).limit(2000)
    .project({ body: 1 }).toArray();
  if (batch.length === 0) break;
  await reviews.bulkWrite(batch.map((r) => ({
    updateOne: { filter: { _id: r._id }, update: { $set: { text: r.body } } }
  })), { ordered: false });
  console.log(`${done += batch.length} copied, resume point ${lastId = batch.at(-1)._id}`);
}

It printed 2000 copied, resume point 6ab2087e1ca38de2cc93e333 through 10008 copied and finished in 1088 ms; a second run took 20 ms and printed nothing, the $exists filter matching no document. Batching keeps the working set in cache and leaves the replication stream room; on a busy cluster sleep between batches and watch replication lag (Monitoring).

A transaction can wrap a migration on a replica set, and its rollback is real: a withTransaction that updated 5,005 documents and then failed left zero of them changed. Index builds are the exception — the same transaction answered OperationNotSupportedInTransaction: Cannot create new indexes on existing collection ... in a multi-document transaction. Transactions hold locks and default to a 60-second limit (Transaction Retries), so they suit a small atomic fix-up, not a large backfill.