Retryable Operations

Error Handling and Retryable Operations

Every driver error descends from MongoError, and the branch tells you what to do:

Since driver 4, retryWrites and retryReads default to true: a single-document write that fails with a retryable error is resent once to a freshly selected server — or repeatedly while a timeoutMS budget lasts — and the session's txnNumber, recorded by the server, keeps the repeat from applying twice. The run below forces one failure with the failCommand fail point.

A duplicate key, a survived stepdown, and a timeoutJavaScript
let trace = false;
const client = new MongoClient(uri, { monitorCommands: true });
client.on("commandStarted", (e) => {
  if (trace && e.commandName === "insert") console.log(" insert #", `${e.command.txnNumber}`);
});
const books = client.db("bookshelf").collection("books");   // unique index on isbn
try {
  await books.insertOne({ isbn: "978-0-441-01359-3", title: "Dune (again)" });
} catch (err) {
  console.log(err.constructor.name, err.code, JSON.stringify(err.keyValue),
    "| retryable?", err.hasErrorLabel("RetryableWriteError"));
}
await client.db("admin").command({ configureFailPoint: "failCommand", mode: { times: 1 },
  data: { failCommands: ["insert"], errorCode: 189 } });   // 189 = PrimarySteppedDown
trace = true;
const res = await books.insertOne({ isbn: "978-1-83-546543-4", title: "Node Cookbook" });
console.log("insert survived the stepdown:", res.acknowledged);
try { await books.findOne({ $where: "sleep(900) || true" }, { timeoutMS: 300 }); }
catch (err) { console.log(err.constructor.name, "|", err.message); }
Output
MongoServerError 11000 {"isbn":"978-0-441-01359-3"} | retryable? false
 insert # 2
 insert # 2
insert survived the stepdown: true
MongoOperationTimeoutError | Timed out during socket read (299ms)

Read the retry carefully: two insert commands, one txnNumber, and code that never saw an error. An election or a rolling restart then costs one round trip instead of a failed request. The rules have edges: retries need a replica set or a sharded cluster, since a standalone server keeps no session history to deduplicate against; updateMany and deleteMany are not retryable; and an error that survives the second attempt carries the RetryableWriteError label, the signal to fail the request rather than loop.

timeoutMS, set per operation, per collection or on the client, bounds the whole operation — server selection, pool checkout, retries and execution — replacing the older tangle of socketTimeoutMS, maxTimeMS and wtimeoutMS. One shorter than your HTTP timeout keeps a slow scan from piling up connections.