Lesson 17 of 18

MongoDB Best Practices

Schema Decisions Are the Expensive Ones

Almost everything else in this list can be fixed in an afternoon. A missing index takes one command. A slow query can be rewritten. But changing the shape of documents in a live collection needs a migration script, a deployment that copes with both old and new shapes at once, and a plan for what happens if it fails halfway. Spend your careful thinking here.

The single question that guides every choice is the one from the data modeling lesson: what does the application read most often, and can that read be satisfied from one document? Everything below is a consequence of taking that question seriously.

  • Design around your queries, not around a diagram of entities
  • Embed data that is always read with its parent and cannot grow without limit
  • Reference data that has its own lifecycle, is shared, or grows unboundedly — and index the reference field
  • Never let an array grow for ever inside a document; cap it with $slice or move it to its own collection
  • Store the correct BSON types from day one — real dates, real numbers, decimals for money
  • Give every collection a createdAt, and ideally an updatedAt; you will want them for debugging long before you want them for features
  • Copy a field into another document only when you have decided whether that copy is a cache to refresh or a snapshot to freeze
Notes
  • Write your schema decisions down — even three lines in a README saying "comments are referenced because they need their own moderation screen". Six months later, neither you nor anyone else will remember why, and the reasoning is what tells the next person whether the decision still holds.

Reading Efficiently

Most performance problems in a young application are not exotic. They are a query with no index behind it, a query that fetches far more than it needs, or a loop that runs one query per row. All three are easy to find if you look, and all three get dramatically worse as data grows — which is why they are invisible during development and obvious in production.

Make explain("executionStats") a normal part of writing a query rather than something you do when firefighting. If the number of documents examined is far larger than the number returned, you have found the problem before your users did.

  • Project only the fields you need; never send a whole document to build a list of names
  • Always limit what a page renders, and paginate by a range rather than a deep skip
  • Index the fields you filter, sort and join on, ordered by the ESR rule — equality, sort, range
  • Prefer $in over $ne and $nin, which cannot use an index selectively
  • Anchor regular expressions at the start (/^abc/) or use a text index; an unanchored, case-insensitive pattern scans
  • Do grouping and totals with the aggregation pipeline in the database, not with JavaScript over a full collection
  • Fetch related documents in one query with $in or $lookup instead of looping — the N+1 problem is the most common cause of a slow page
Notes
  • Test with a realistic amount of data. Fifty documents make every query look fast and every index look unnecessary. Generate a lakh of rows with a loop in mongosh before you decide your queries are fine.

Writing Safely

Writes are where mistakes become permanent. The habits below cost seconds each and prevent the categories of error that end with restoring a backup — or discovering there is no backup to restore.

The most valuable of them is also the simplest: read the result object. matchedCount, modifiedCount and deletedCount are the database telling you what it actually did, and code that ignores them cheerfully reports success for an update that changed nothing.

  • Target updates and deletes by _id, or by a field with a unique index
  • Run the filter through countDocuments() and find() before any updateMany or deleteMany
  • Treat deleteMany({}) as a command that empties a collection — because it is
  • Prefer a soft delete for anything a user might want back, or that you might need for an audit
  • Use $inc and conditional filters instead of read-modify-write on counters and stock
  • Make scripts idempotent with upserts, so running one twice does not duplicate your data
  • Use a transaction only when two or more documents must change together, and always handle its failure path
Notes
  • Never run an exploratory command against production while connected as an administrator. Keep a separate, read-only database user for looking around, and switch to a privileged one deliberately when you intend to change something.

Security You Cannot Skip

Two rules cover the basics. Authentication must be switched on, and credentials must live in environment variables rather than source code. An unauthenticated MongoDB server reachable from the internet is found by automated scanners quickly, and the usual outcome is that the data is deleted and replaced with a ransom note. Atlas enables authentication by default, which is one more reason to start there.

Beyond that, give each application user the least privilege it needs — readWrite on one database, not administrator on the cluster — so a leaked credential does limited damage. Rotate any secret that has ever been committed to git, even a private repository.

The subtle one is operator injection. A MongoDB filter is a document, so if you place a value from a request body directly into a filter, an attacker can send an object instead of a string. A login handler that builds { email, password } from req.body can be handed { "$ne": null } as the password, which matches any account. The defence is to check types before they reach a query: confirm the value is a string, or run the request body through a validation library first. Mongoose's casting helps here, because a schema-typed field rejects an object where a string was declared — but do not rely on that alone for fields you query directly.

Example
// Vulnerable: whatever JSON arrives becomes part of the filter
// POST { "email": "admin@site.com", "password": { "$ne": null } }
const user = await db.collection('users').findOne({
  email: req.body.email,
  password: req.body.password        // matches ANY password
});

// Safer: insist on strings, and never compare passwords in a query
const { email, password } = req.body;
if (typeof email !== 'string' || typeof password !== 'string') {
  return res.status(400).json({ error: 'Invalid credentials format' });
}
const account = await db.collection('users').findOne({ email });
const ok = account && await bcrypt.compare(password, account.passwordHash);
if (!ok) return res.status(401).json({ error: 'Invalid email or password' });
Notes
  • Return the same message for "no such email" and "wrong password". Distinguishing them tells an attacker which email addresses are registered, which is the first step of a targeted attack.

Operations: Backups, Monitoring, Migrations

A backup you have never restored is not a backup — it is a file you hope is correct. Practise the restore into a scratch database at least once, and note how long it takes, because that number is your recovery time when something goes wrong for real. On the Atlas free tier, where automated backups are not included, run mongodump yourself on a schedule and keep the output somewhere other than the machine that produced it.

Schema changes deserve the same care as code changes. The safe sequence for renaming or reshaping a field is to write both shapes for a while, backfill the old documents in batches, switch reads to the new shape, and only then stop writing the old one. Doing it in one step means any request in flight during the deployment sees a document it does not understand.

  • Create indexes as an explicit step in your deployment, not as a side effect of the application starting
  • Keep an eye on the slow-query log or the Atlas Metrics tab after every release
  • Reuse one connection pool for the process; opening a connection per request exhausts the server's limit
  • Migrate large collections in batches with a filter and a limit, not in one enormous updateMany
  • Version your schema changes alongside your code so any developer can rebuild the database from scratch
  • Test the restore path, and keep at least one backup copy off the machine that holds the data
Notes
  • The habits in this lesson are worth more than any single trick. A project with sensible documents, the right handful of indexes, credentials in environment variables and a tested backup will outlast a cleverer one without them.
Ask AI