Lesson 13 of 18

Mongoose ODM

What an ODM Is, and Why You Might Add One

MongoDB itself does not care what shape your documents are. That is a feature, but it means every rule about your data — that an email is required, that a role must be one of three values, that a password is hashed before it is stored — has to live somewhere in your code, applied consistently everywhere. Mongoose is a library that gives those rules a home.

ODM stands for Object Data Modeling. You declare a schema describing the fields of a document, compile it into a model, and from then on you work with the model instead of the raw collection. Mongoose then enforces the schema, converts values to the declared types, runs your validators, and lets you hook code into events such as "before save".

It is not compulsory. The official mongodb driver is perfectly pleasant to use, and plenty of production systems use it directly, especially in TypeScript where types already describe the documents. Mongoose earns its place when several people work on the same codebase and you want the rules in one file rather than scattered across twenty route handlers.

  • Schemas — a single definition of what a document contains
  • Type casting — the string "25" arriving from a form becomes the number 25
  • Validation — required fields, ranges, allowed values, and your own rules
  • Middleware — run code before or after save, update or delete
  • populate() — follow references to other collections without writing the join
  • Query helpers and virtuals — computed fields and reusable query fragments
Notes
  • Mongoose 7 and later dropped callback support entirely. Every method returns a promise, so all the examples here use await. If a tutorial passes a function as the last argument, it is written for Mongoose 6 or earlier and will simply not work.

Connecting Once, at Start-Up

mongoose.connect() opens a connection pool and stores it globally on the mongoose object. That is why you call it once when your application boots and never again — every model you define anywhere in the project uses that same connection. Calling it inside a route handler creates a new pool on every request, and the server runs out of connections in minutes.

The pool matters more than beginners expect. Opening a TCP connection to a database is slow, so the driver keeps a set of them open and hands them out as queries need them. Your code never manages this; it just means the first query after start-up is a little slower than the rest.

Keep the connection string in an environment variable, never in the source file. A connection string contains a username and password, and committing one to a public repository is one of the fastest ways to have a database emptied by a stranger.

Example
// npm install mongoose
import mongoose from 'mongoose';

async function start() {
  try {
    await mongoose.connect(process.env.MONGODB_URI);   // once, at boot
    console.log('MongoDB connected');
    app.listen(3000);
  } catch (err) {
    console.error('Could not connect:', err.message);
    process.exit(1);        // fail loudly rather than serve broken pages
  }
}

// Watch the connection for later problems
mongoose.connection.on('error', err => console.error('Mongo error:', err));
mongoose.connection.on('disconnected', () => console.warn('Mongo disconnected'));

start();
Notes
  • Exit on a failed initial connection instead of continuing. A server that starts up without a database answers every request with a confusing error; one that refuses to start tells you exactly what is wrong in the first line of the log.

Schemas and Models

A schema lists your fields and their types, plus any rules attached to them. A model is the class Mongoose builds from the schema — you call methods on it to query, and calling it with new creates a document you can save.

Field definitions can be a bare type (name: String) or an object with options (name: { type: String, required: true }). Common options are required, default, unique, enum, trim and lowercase. The last two are worth adopting as a habit on email fields, because they remove the entire class of bugs where " Ananya@Example.com " and "ananya@example.com" become two different accounts.

One detail surprises everyone: the collection name. Mongoose takes your model name, lower-cases it and makes it plural, so model('User', ...) reads and writes a collection called users, and model('Category', ...) uses categories. If you go looking in mongosh for a collection named User and find nothing, this is why. You can override it by passing a third argument to model().

Example
import { Schema, model } from 'mongoose';

const userSchema = new Schema({
  name:  { type: String, required: true, trim: true },
  email: { type: String, required: true, unique: true, lowercase: true, trim: true },
  age:   { type: Number, min: 0, max: 120 },
  role:  { type: String, enum: ['student', 'teacher', 'admin'], default: 'student' },
  skills: [String],
  address: {
    city: String,
    pincode: String
  }
}, { timestamps: true });      // adds createdAt and updatedAt for you

const User = model('User', userSchema);       // -> collection "users"

const user = await User.create({
  name: '  Ananya Sharma ',
  email: 'Ananya@Example.com',
  age: 22
});
console.log(user.name, user.email);
// "Ananya Sharma" "ananya@example.com"   <- trim and lowercase applied
Notes
  • unique: true is not a validator, despite sitting among them. It tells Mongoose to build a unique index on that field, and the rejection comes from MongoDB as an E11000 duplicate key error, not as a Mongoose validation error. Your error handler must recognise both, or duplicate signups will return a 500 instead of a helpful message.

CRUD, and Two Options You Must Remember

Mongoose's query methods mirror the driver's, with friendlier names. create inserts, find and findOne and findById read, findByIdAndUpdate and updateMany modify, findByIdAndDelete and deleteMany remove. Queries are chainable in the same way as cursors: .select() projects, .sort(), .limit() and .skip() behave as you would expect.

Two options catch people out on every project. The first is that findByIdAndUpdate returns the document as it was before the update. Pass { new: true } to get the updated version — without it, you send the user the old values and it looks like nothing saved.

The second is more serious. Mongoose validators do not run on update methods by default; they run on save(). So findByIdAndUpdate(id, { age: -5 }) happily stores a negative age even though the schema says the minimum is zero. Pass { runValidators: true } to enforce the schema on updates too. Get into the habit of writing both options together.

.lean() is worth knowing early. It returns plain JavaScript objects rather than full Mongoose documents, skipping the work of building document instances. Those objects have no .save() and no virtuals, so use it only for read-only paths — but for a list endpoint that returns fifty records as JSON, it is measurably faster and simpler.

Example
// Create
const user = await User.create({ name: 'Ananya', email: 'ananya@example.com' });

// Read
const admins = await User.find({ role: 'admin' }).select('name email').lean();
const one    = await User.findById(id);
const byMail = await User.findOne({ email: 'ananya@example.com' });

// Update — remember BOTH options
const updated = await User.findByIdAndUpdate(
  id,
  { $set: { age: 23 } },
  { new: true, runValidators: true }
);

// Without them:
// await User.findByIdAndUpdate(id, { age: -5 });   // stored! min:0 never checked

// Delete
await User.findByIdAndDelete(id);
await User.deleteMany({ role: 'student', active: false });
Notes
  • Mongoose is in strict mode by default: any field you set that is not in the schema is silently dropped rather than saved. This is usually protection against junk from a request body, but if a field you carefully set never appears in the database, check the spelling against your schema first.

Middleware and populate()

Middleware, also called hooks, lets you run code around an operation. The classic use is hashing a password before it is stored, written once in the schema so no route handler can forget it. Inside a pre('save') hook, this is the document about to be saved, and isModified stops you re-hashing an already-hashed password every time the user edits their name.

There is a sharp edge here that causes real security bugs. pre('save') runs for save() and for create(), but not for findByIdAndUpdate, updateOne or updateMany — those talk to the database directly and never build a document. If a "change password" route uses findByIdAndUpdate, the new password is stored in plain text. Either load the document, assign the field and call save(), or register a matching pre('findOneAndUpdate') hook.

populate() resolves references. Declare a field as an ObjectId with a ref naming the other model, and Mongoose can replace the id with the real document on demand. Behind the scenes it runs one extra query using $in, so it avoids the N+1 problem — but it is still an extra round trip per populated path, and you can restrict which fields come back with a second argument.

Example
import bcrypt from 'bcrypt';

userSchema.pre('save', async function (next) {
  if (!this.isModified('password')) return next();
  this.password = await bcrypt.hash(this.password, 10);
  next();
});

// Runs the hook
const u = await User.findById(id);
u.password = plainText;
await u.save();               // hashed

// Does NOT run the hook — stores plain text
// await User.findByIdAndUpdate(id, { password: plainText });

// References and populate
const postSchema = new Schema({
  title: String,
  author: { type: Schema.Types.ObjectId, ref: 'User', required: true }
});
const Post = model('Post', postSchema);

const posts = await Post.find()
  .populate('author', 'name email')   // only these fields of the author
  .limit(20);

posts[0].author.name;   // "Ananya Sharma"
Notes
  • By default Mongoose tries to create every index in your schemas when the application starts. That is convenient in development and unwelcome on a large production collection, where an unexpected index build can slow the database down. Production applications usually set autoIndex: false and create indexes deliberately as part of a deployment step.
Ask AI