GridFS

Storing Large Files with GridFS

GridFS is the convention drivers implement for files that will not fit in 16 MiB. It is not a server feature: the driver splits a byte stream into chunks of 255 KiB, writes each as an ordinary document in fs.chunks, and writes one metadata document in fs.files. Because chunks are numbered, you can seek into the middle of a 2 GB video without loading the rest.

The two collections form a bucket, named fs by default; pass bucketName to keep videos and invoices apart. On first upload the driver creates unique { files_id: 1, n: 1 } on the chunks and { filename: 1, uploadDate: 1 } on the files.

Uploading and streaming back a file with GridFSJavaScript
import { MongoClient, GridFSBucket } from 'mongodb';
import { createReadStream, createWriteStream } from 'node:fs';
import { pipeline } from 'node:stream/promises';
const client = await MongoClient.connect('mongodb://127.0.0.1:27017');
const bucket = new GridFSBucket(client.db('bookshelf'), { bucketName: 'media' });
const upload = bucket.openUploadStream('cover.png', { metadata: { bookId: 42 } });
await pipeline(createReadStream('cover.png'), upload);
const [file] = await bucket.find({ filename: 'cover.png' }).toArray();
console.log(file.length, 'bytes in', Math.ceil(file.length / file.chunkSize), 'chunks');
await pipeline(bucket.openDownloadStream(file._id), createWriteStream('copy.png'));
await bucket.delete(file._id);       // removes the file document and every chunk
await client.close();

That listing needs a running mongod, which was not available here, so it is printed without output; run it against your own server from Getting MongoDB Running and inspect db.media.files.findOne(). Put your own fields under metadata: the legacy top-level md5, contentType and aliases fields are deprecated and newer drivers no longer write them. GridFS does not participate in multi-document transactions, and its chunks count against your storage, working set and backups like application data.

Ask first whether you need it. Under 16 MiB a single Binary field is simpler and one round trip faster, and for files served to browsers, object storage behind a CDN — S3, R2, or MinIO 30,943 (github.com/minio/minio (https://github.com/minio/minio 61,350 ), AGPL-3.0) — is cheaper per gigabyte and better at range requests. GridFS earns its place when you want one backup, one replication stream and one access-control story covering data and files together.

GridFS splits one stream into 255 KiB chunk documents plus a metadata document
GridFS splits one stream into 255 KiB chunk documents plus a metadata document