One-to-One: Usually Just Embed
A one-to-one relationship is a user and their address, a product and its dimensions, an employee and their salary details. In a relational database these often end up in separate tables for tidiness. In MongoDB, splitting them costs you a join and buys you almost nothing, so the default is to embed.
There are two cases where you should still split. The first is size: if one half is large and rarely needed — a long article body, a stored resume, a blob of raw import data — keeping it out of the main document means every ordinary read stays small. The second is access control: if one half is sensitive and read by different code paths, a separate collection makes it much harder to leak it by forgetting a projection.
When you do split, the child document normally uses the parent's _id as its own _id. That gives you a guaranteed one-to-one relationship for free — the unique index on _id makes a second child impossible — and no extra index is needed.
// Default: embed
{
_id: ObjectId("64a..."),
name: "Ananya Sharma",
address: { city: "Pune", state: "Maharashtra", pincode: "411001" }
}
// Split when one half is large or sensitive.
// Reusing the parent's _id enforces one-to-one automatically.
// users
{ _id: ObjectId("64a..."), name: "Ananya Sharma", email: "ananya@example.com" }
// userSecrets
{ _id: ObjectId("64a..."), passwordHash: "...", twoFactorSecret: "..." }
db.userSecrets.findOne({ _id: user._id }) - Never store a plain password in either document. Store only a hash produced by a password-hashing function such as bcrypt or argon2, and keep it out of every projection your API returns.
One-to-Many: Which Side Holds the Link?
When the relationship is referenced rather than embedded, you must decide which document stores the link. The answer is nearly always the many side stores the id of the one side. An order stores customerId; a comment stores postId; a chapter stores courseId.
The alternative — an array of child ids on the parent — looks tidier and is a trap. That array grows every time a child is added, so it is unbounded by definition, and every new child requires an update to the parent document as well as an insert. The parent becomes a contention point that every write has to touch. Keeping the link on the child means adding a child is a single insert that nothing else has to know about.
Whichever field holds the reference must be indexed. Without an index on customerId, "show this customer's orders" scans the entire orders collection, and that is the query your customer dashboard runs on every visit.
// The many side points back
// posts
{ _id: ObjectId("p1"), title: "Learn MongoDB", author: "Ananya" }
// comments
{ _id: ObjectId("c1"), postId: ObjectId("p1"), user: "Rahul", text: "Helpful!" }
{ _id: ObjectId("c2"), postId: ObjectId("p1"), user: "Meera", text: "Thanks!" }
db.comments.createIndex({ postId: 1, createdAt: -1 })
// Paginated comments for one post — index-friendly
db.comments.find({ postId: ObjectId("p1") }).sort({ createdAt: -1 }).limit(20)
// A useful hybrid: embed the newest few, reference the rest
{
_id: ObjectId("p1"),
title: "Learn MongoDB",
commentCount: 248,
recentComments: [ { user: "Rahul", text: "Helpful!" } ] // capped at 3
} - The hybrid in the last block is worth remembering. The post page renders instantly from one document because the newest three comments are embedded, and "view all comments" falls back to the comments collection. Keep the embedded copy capped with
$pushand$sliceso it can never grow.
Many-to-Many: Two Workable Shapes
Students and courses is the standard example: a student takes many courses, and a course has many students. There are two reasonable ways to model it, and the choice depends on whether the relationship itself carries information.
The simpler shape stores an array of ids on one side — normally the side whose list is smaller and naturally bounded. A student takes perhaps forty courses in a degree; a popular course may have thousands of students. So courseIds on the student is safe, while studentIds on the course is an unbounded array waiting to cause trouble. With an index on courseIds, both directions are still fast, because a multikey index makes "which students have this course id in their array?" an indexed lookup.
The second shape is a separate join collection, one document per pair. Use it when the relationship has its own fields — enrolment date, grade, attendance, payment status — because those belong to the pairing rather than to the student or the course. It is also the shape that scales when both sides are large, since neither parent document grows at all.
// Shape 1: array on the bounded side
// students
{ _id: "s1", name: "Ananya", courseIds: ["c1", "c2", "c3"] }
db.students.createIndex({ courseIds: 1 }) // multikey index
db.students.find({ courseIds: "c1" }) // students in course c1
db.courses.find({ _id: { $in: student.courseIds } }) // that student's courses
// Shape 2: a join collection, when the pairing carries data
// enrolments
{
_id: ObjectId("..."),
studentId: "s1",
courseId: "c1",
enrolledAt: ISODate("2026-01-05"),
grade: "A",
feePaid: true
}
db.enrolments.createIndex({ studentId: 1, courseId: 1 }, { unique: true })
db.enrolments.createIndex({ courseId: 1 }) - The unique compound index on
{ studentId, courseId }is what stops the same student being enrolled twice. Enforcing it in the database is far more reliable than checking first and then inserting, which two simultaneous requests can both pass.
Reading References Without the N+1 Problem
Once data is referenced, you have to fetch it. The naive approach is to load the list, then loop over it fetching each related document. Twenty orders become one query for the list plus twenty queries for the customers — twenty-one round trips where two would do. This is known as the N+1 query problem, and it is the most common performance bug in applications built on referenced data.
There are two good fixes. The first is to collect the ids and fetch them in a single query with $in, then match them up in memory. Two round trips regardless of how many rows you are showing. The second is to let the database do the join with $lookup inside an aggregation, which is one round trip.
Which to choose depends on the shape of the page. $lookup is neat when you want the joined data attached to each row anyway. The $in approach is often better when several rows share the same related document — twenty orders from five customers means fetching five customers, not twenty. Either is fine; looping is not.
// The bug: one query per row
// for (const order of orders) {
// order.customer = await users.findOne({ _id: order.customerId }) // N+1
// }
// Fix 1: one extra query for all of them
const orders = db.orders.find({ status: "pending" }).toArray()
const ids = [...new Set(orders.map(o => o.customerId))]
const users = db.users.find({ _id: { $in: ids } }).toArray()
const byId = new Map(users.map(u => [String(u._id), u]))
orders.forEach(o => { o.customer = byId.get(String(o.customerId)) })
// Fix 2: let the database join
db.orders.aggregate([
{ $match: { status: "pending" } },
{ $lookup: { from: "users", localField: "customerId", foreignField: "_id", as: "customer" } },
{ $unwind: { path: "$customer", preserveNullAndEmptyArrays: true } }
]) - Mongoose's
populate(), covered in the next lesson, implements the$inapproach for you: it collects the referenced ids from the documents it just loaded and fetches them in one extra query.
References Can Break, and Nothing Will Tell You
MongoDB has no foreign keys. A customerId is an ordinary field that happens to contain a value you intend to look up elsewhere; the database does not check that the target exists, and it will not stop you deleting it. Delete a customer and their orders keep pointing at an id that no longer resolves. These are called orphans, and their symptom is quiet: a $lookup returns an empty array, and your page shows a blank name instead of throwing an error.
Because nothing is enforced for you, cleanup is your application's responsibility, and you have three sensible options. Delete the related documents yourself, ideally inside a transaction so that either everything goes or nothing does. Refuse the delete while children still exist, the way a relational RESTRICT constraint would. Or — most commonly in real systems — do not delete at all: mark the parent as inactive and keep the history intact.
Whichever you choose, write your read code defensively. Assume a referenced document may be missing, and render something sensible when it is.
// Option A: cascade the delete yourself, atomically
const session = db.getMongo().startSession()
session.startTransaction()
try {
const sdb = session.getDatabase("shopDB")
sdb.orders.deleteMany({ customerId: userId })
sdb.users.deleteOne({ _id: userId })
session.commitTransaction()
} catch (e) {
session.abortTransaction()
throw e
} finally {
session.endSession()
}
// Option B: refuse while children exist
if (db.orders.countDocuments({ customerId: userId }) > 0) {
throw new Error("Customer has orders and cannot be deleted")
}
// Option C: never delete — deactivate
db.users.updateOne({ _id: userId }, { $set: { active: false, closedAt: new Date() } }) - Auditing for orphans is a useful exercise on any project that has been running a while: a
$lookupfrom the child collection followed by{ $match: { parent: { $size: 0 } } }lists every child whose reference no longer resolves.
