Lesson 11 of 18

Data Modeling

Model for Your Queries, Not for Your Entities

If you have studied relational databases, you were taught to start with normalisation: identify entities, give each one a table, and remove every duplicated value. The structure comes first and the queries adapt to it. MongoDB inverts that. You start by listing the questions your application will ask most often, and then design documents that answer those questions with as little work as possible.

That sounds vague until you make it concrete. Write down the screens of your application. A product page needs the product, its images, its specifications and its average rating. An order confirmation needs the order, its line items and the delivery address. A dashboard needs totals per month. Each of those is a read, and each read should ideally be satisfied by fetching one document, or a small number of them, without stitching together five collections.

This is why two MongoDB schemas for the same business can look completely different and both be correct. The right shape depends on what you read most, how often things change, and how large the data can grow. There is no single normalised answer to copy.

Notes
  • Schema design is the decision in a MongoDB project that is hardest to change later. A missing index can be added in a minute; reshaping documents across a live collection needs a migration script, a deployment plan, and code that copes with both shapes at once.

Embedding: Putting Related Data Inside

Embedding means the related data lives inside the parent document, as a nested object or an array of objects. Fetching the parent brings everything with it in a single read, with no join and no second round trip. Because MongoDB guarantees that an update to a single document is atomic, changing several embedded values at once cannot leave the record half-updated.

An order is the standard example. The order document holds its line items, the delivery address and the payment status. Nothing about an order is useful without them, they are always read together, and their number is naturally bounded — nobody buys ten thousand distinct products in one order.

Embedded line items also solve a problem that catches beginners who reference everything. An order must record what the customer paid, not what the product costs today. If the order merely references the product and reads its current price, then every historical invoice silently changes the next time you run a sale. Copying the name, price and tax rate into the line item at the moment of purchase is not duplication to be ashamed of — it is a deliberate snapshot, and it is the correct design.

Example
// One read gives you the whole order
{
  _id: ObjectId("..."),
  orderNo: "ORD-1001",
  customerId: ObjectId("..."),
  placedAt: ISODate("2026-02-01T10:15:00Z"),
  status: "completed",

  // snapshot of what was bought, at the price paid
  items: [
    { sku: "AB-1", name: "Wireless Mouse", qty: 2, price: NumberDecimal("799.00") },
    { sku: "CD-9", name: "USB-C Cable",   qty: 1, price: NumberDecimal("299.00") }
  ],

  shippingAddress: {
    line1: "12 MG Road", city: "Pune", state: "Maharashtra", pincode: "411001"
  },

  total: NumberDecimal("1897.00")
}
Notes
  • Embed when the child data is always read with its parent, belongs to exactly one parent, and cannot grow without limit. If all three are true, embedding is almost always the right call.

Referencing: Storing an id Instead

Referencing means keeping the related data in its own collection and storing only its _id. This is the right choice when the related data has a life of its own — a customer exists independently of any order, and is read and updated on its own screens.

It is also the answer whenever the number of related records can grow without limit. A customer might place five orders or five thousand; embedding orders inside the customer would make that document grow for ever, until it becomes slow to read and eventually hits the 16 MB ceiling. Instead the order carries a customerId, and you index that field so "all orders for this customer" stays fast no matter how many there are.

The cost of referencing is that reading the whole picture takes more than one operation: either a $lookup in an aggregation, or a second query from your application. That extra work is usually a fair price when the alternative is an unbounded document.

Example
// users collection — exists independently
{ _id: ObjectId("64a..."), name: "Ananya Sharma", email: "ananya@example.com" }

// orders collection — the "many" side carries the reference
{ _id: ObjectId("71b..."), customerId: ObjectId("64a..."), total: 1897 }
{ _id: ObjectId("71c..."), customerId: ObjectId("64a..."), total:  499 }

// Make the reference usable
db.orders.createIndex({ customerId: 1 })

// All orders for one customer
db.orders.find({ customerId: ObjectId("64a...") }).sort({ placedAt: -1 })
Notes
  • Store the reference as an ObjectId, not as a string version of one. They are different BSON types, so a query for the string will not match a document holding the ObjectId, and the mistake produces empty results rather than an error.

Rules of Thumb by Relationship Size

Most decisions collapse into a single question: roughly how many related items are there, and can that number grow? A useful way to think about it is in three sizes — a few, many, and an unbounded number.

The middle case is where judgement is needed. A blog post with fifty comments could go either way, and the deciding factor is usually whether the comments are ever read or written on their own. If they only appear underneath the post, embed them. If they need their own moderation screen, their own pagination, or their own notifications, they deserve their own collection.

  • One-to-one — embed. A user and their address are one read; splitting them just creates work.
  • One-to-few (a handful, with a natural ceiling) — embed. Order items, a user's phone numbers, a product's specifications.
  • One-to-many (dozens to a few hundred, bounded in practice) — either. Embed if always read together, reference if the children have their own screens.
  • One-to-unbounded (orders per customer, log lines per server, messages per chat) — always reference, with an index on the reference field.
  • Many-to-many — reference from the side with the smaller, bounded list, or use a separate collection that joins the two.
  • Data shared by many parents, such as a category or a tax rate — reference it, so one edit updates everything.

Deliberate Duplication

Relational training says duplicated data is a defect. In MongoDB it is a tool, used on purpose to avoid a join on a read you perform constantly. The common form is the extended reference: alongside the id, you copy the one or two fields you always display.

An order list that shows the customer's name is the classic case. Storing customerId alone means a $lookup every time the list is drawn. Storing customerId plus customerName means the list renders from one collection. You have duplicated the name, and that is fine — provided you decide what should happen when the name changes.

There are two honest answers, and the mistake is not choosing one. Either the copy is a cache, in which case renaming a customer must also update the copies, and you accept that work; or the copy is a snapshot of a moment, like the price on an invoice, in which case it must deliberately never be updated. Write down which one it is in a comment next to the field, because the next person to read your schema cannot tell by looking.

Example
// Extended reference: id plus the fields the list screen needs
{
  _id: ObjectId("..."),
  orderNo: "ORD-1001",
  customer: {
    _id: ObjectId("64a..."),
    name: "Ananya Sharma",   // cached copy — refresh when the user renames
    city: "Pune"
  },
  items: [ /* price fields here are snapshots — never updated */ ],
  total: NumberDecimal("1897.00")
}

// If it is a cache, keep it honest when the source changes
db.orders.updateMany(
  { "customer._id": ObjectId("64a...") },
  { $set: { "customer.name": "Ananya Iyer" } }
)
Notes
  • Duplicate only fields that rarely change and that you display very often. Copying a field that changes every day into a million documents turns every edit into a mass update, and you have traded a cheap read for an expensive write.

Limits and Common Antipatterns

MongoDB enforces a hard limit of 16 MB per document. That sounds enormous, and it is — but the point of the limit is not the ceiling itself. Long before a document approaches it, every read of that document pulls megabytes across the network, every small update rewrites a large record, and memory fills with data nobody asked for. Treat a document that has grown past a few hundred kilobytes as a design problem, not as a success.

The patterns below are the ones that most often cause trouble in beginner projects. Each of them looks reasonable at the start and only reveals itself once there is real data.

  • Unbounded arrays — pushing every comment, event or log line into one document. This is the most common serious mistake in MongoDB schemas.
  • Massive documents — storing an image or a PDF as base64 text inside a record. Put files in object storage and keep the URL in the document.
  • Deep nesting — objects inside objects inside objects. BSON permits many levels, but queries and updates become hard to write and read past two or three.
  • A collection per user or per tenant — thousands of collections each hold their own indexes and metadata. Use one collection with a tenantId field instead.
  • Referencing everything — a faithful copy of a relational schema, where every page needs four $lookup stages. If you wanted that, a relational database would have served you better.
  • Embedding everything — the opposite error, where one document tries to hold an entire subsystem and grows without limit.
Notes
  • A good self-check: for each screen in your application, count how many collections a single page load has to touch. If the answer is one or two for your busiest screens, the model is working. If your home page needs six, revisit what should have been embedded.
Ask AI