What You Are Building
The project is the back end of a small online bookstore. It is deliberately ordinary, because an ordinary application is where every idea in this course shows up: books and authors are related, reviews arrive without limit, orders must record what was paid rather than what things cost today, and the shop owner wants a sales report.
You do not need a front end to finish this. Everything below can be built and tested with mongosh, or as an Express API with Mongoose if you want something to show. What matters is that each decision is made for a reason you can explain — that is what turns a tutorial into a project you can put your name on.
- Four collections:
authors,books,reviewsandorders - Full create, read, update and delete for books, with validation on every write
- Search books by title and by author, with pagination that still works on page fifty
- Reviews with a rating, plus an average rating shown on the book page without a slow calculation
- Placing an order: reduce stock and create the order together, or do neither
- A monthly sales report built with the aggregation pipeline
- Work through it in the order given. Each step depends on the modelling decision made in the step before, and skipping to the code is how projects end up with a schema nobody can explain.
Step 1: Decide the Shape, and Say Why
Before writing a schema, list the reads. The book page needs the book, its author's name and its average rating. The catalogue page needs a page of books with title, price, cover and rating. The account page needs a customer's past orders. The owner's report needs revenue grouped by month.
Now the decisions follow almost by themselves. The author is referenced, because an author has many books and their own page — but the author's name is also copied onto the book, because every catalogue row displays it and re-joining for a name on every row is wasteful. That copy is a cache: if an author is renamed, the books must be updated too.
Reviews get their own collection. They are unbounded — a popular book can collect thousands — and they need their own moderation and pagination. But the book keeps a small rollup, ratingSum and ratingCount, updated whenever a review is added, so the average can be shown without aggregating anything at read time.
Order items are embedded, and they store the title and price at the moment of purchase. This is a snapshot, not a cache: when a book's price changes next month, old invoices must keep showing what the customer actually paid.
authors— referenced from books; each author is read on their own pagebooks— carries a cachedauthorNameand aratingSum/ratingCountrollupreviews— its own collection, referencingbookId, because the count is unboundedorders— line items embedded as a frozen snapshot of title, price and quantity- Every collection gets timestamps, so anything can be traced later
- Notice the two different kinds of copied data in one schema.
authorNameon a book is a cache that must be refreshed when the author changes. The price inside an order item is a snapshot that must never be refreshed. Confusing the two is how invoices start changing after the fact.
Step 2: The Schemas
These are Mongoose schemas, but the shape is the same if you use the native driver with a $jsonSchema validator instead. Note the small things: prices are stored as whole paise to avoid floating point trouble, isbn is unique, enum restricts the genre, and every schema turns on timestamps.
The rollup fields on the book start at zero and are only ever changed with $inc, never by reading and rewriting them. That is what keeps two reviews arriving at the same moment from losing one another.
import { Schema, model } from 'mongoose';
const authorSchema = new Schema({
name: { type: String, required: true, trim: true },
bio: { type: String, maxlength: 2000 },
country: String
}, { timestamps: true });
const bookSchema = new Schema({
title: { type: String, required: true, trim: true },
isbn: { type: String, required: true, unique: true },
authorId: { type: Schema.Types.ObjectId, ref: 'Author', required: true },
authorName: { type: String, required: true }, // cached for list screens
genre: { type: String, enum: ['Fiction', 'Non-Fiction', 'Sci-Fi', 'Tech'], required: true },
pricePaise: { type: Number, required: true, min: 0 }, // whole paise, never a float
stock: { type: Number, required: true, min: 0, default: 0 },
publishedOn: Date,
ratingSum: { type: Number, default: 0 }, // rollup, updated with $inc
ratingCount: { type: Number, default: 0 }
}, { timestamps: true });
bookSchema.virtual('avgRating').get(function () {
return this.ratingCount ? +(this.ratingSum / this.ratingCount).toFixed(1) : 0;
});
const reviewSchema = new Schema({
bookId: { type: Schema.Types.ObjectId, ref: 'Book', required: true },
user: { type: String, required: true },
rating: { type: Number, required: true, min: 1, max: 5 },
text: { type: String, maxlength: 2000 }
}, { timestamps: true });
const orderSchema = new Schema({
customerId: { type: Schema.Types.ObjectId, ref: 'Customer', required: true },
items: [{ // snapshot at purchase time
bookId: { type: Schema.Types.ObjectId, ref: 'Book', required: true },
title: { type: String, required: true },
pricePaise: { type: Number, required: true },
qty: { type: Number, required: true, min: 1 }
}],
totalPaise: { type: Number, required: true },
status: { type: String, enum: ['placed', 'shipped', 'cancelled'], default: 'placed' },
placedAt: { type: Date, default: Date.now }
}, { timestamps: true });
export const Author = model('Author', authorSchema);
export const Book = model('Book', bookSchema);
export const Review = model('Review', reviewSchema);
export const Order = model('Order', orderSchema); Step 3: Indexes, With a Reason for Each
Create indexes deliberately, and be able to name the query each one serves. If you cannot name it, do not create it — you would be paying on every write for nothing.
Notice the compound index on the catalogue query. The genre is an equality match, the sort is by price, so genre comes first and price second, following the ESR rule from the indexing lesson. Verify each one with explain("executionStats") after loading a realistic amount of test data, and check that totalDocsExamined is close to nReturned.
// Uniqueness — enforced by the database, not by application checks
db.books.createIndex({ isbn: 1 }, { unique: true })
// Catalogue: filter by genre, sort by price (E then S)
db.books.createIndex({ genre: 1, pricePaise: 1 })
// "All books by this author"
db.books.createIndex({ authorId: 1 })
// Title search
db.books.createIndex({ title: "text" })
// Reviews for one book, newest first
db.reviews.createIndex({ bookId: 1, createdAt: -1 })
// One review per user per book
db.reviews.createIndex({ bookId: 1, user: 1 }, { unique: true })
// A customer's order history, and the monthly report
db.orders.createIndex({ customerId: 1, placedAt: -1 })
db.orders.createIndex({ placedAt: -1 })
// Prove it
db.books.find({ genre: "Tech" }).sort({ pricePaise: 1 }).explain("executionStats") - The unique index on
{ bookId, user }is doing real work: it makes "one review per person per book" impossible to violate, even if two requests arrive at the same instant. A check-then-insert in your route handler cannot make that promise.
Step 4: The Reads
These are the queries behind the screens listed in step one. Each one projects only what the screen needs and limits what it returns — the two habits that keep an application fast as data grows.
The report is the interesting one. It groups completed orders by month and sums revenue entirely inside the database, so no matter how many orders exist, only a handful of result documents travel back to your application. Building the same report by fetching every order into Node.js and looping would work today and fail next year.
// Catalogue page, index-friendly and paginated
await Book.find({ genre: 'Tech' })
.select('title authorName pricePaise ratingSum ratingCount')
.sort({ pricePaise: 1, _id: 1 }) // tiebreaker keeps paging stable
.limit(20)
.lean();
// Title search
await Book.find({ $text: { $search: 'mongodb' } }).limit(20).lean();
// One book page: the book, plus its most recent reviews
const book = await Book.findById(bookId).lean();
const reviews = await Review.find({ bookId })
.sort({ createdAt: -1 }).limit(10).lean();
// Monthly sales report
await Order.aggregate([
{ $match: { status: { $ne: 'cancelled' } } },
{ $group: {
_id: { $dateToString: { format: '%Y-%m', date: '$placedAt' } },
revenuePaise: { $sum: '$totalPaise' },
orders: { $sum: 1 }
}},
{ $sort: { _id: 1 } }
]);
// Best sellers, from the embedded line items
await Order.aggregate([
{ $match: { status: { $ne: 'cancelled' } } },
{ $unwind: '$items' },
{ $group: { _id: '$items.bookId', title: { $first: '$items.title' },
unitsSold: { $sum: '$items.qty' } } },
{ $sort: { unitsSold: -1 } },
{ $limit: 10 }
]); Step 5: The Writes That Must Not Half-Happen
Placing an order touches two collections: stock goes down on the book, and an order document is created. Either both happen or neither should, so this is a genuine use for a transaction — unlike adding a review, which is a single-document rollup plus one insert and is handled more simply.
The stock update uses a conditional filter, stock: { $gte: qty }, so it only succeeds if there is enough stock. If it matches nothing, we throw, and withTransaction rolls back everything already done in the block. There is no window between checking and reserving for another customer to slip into.
Note also what the order stores: the title and price copied from the book at this moment. That is what makes the invoice permanent.
async function placeOrder(customerId, lines) { // lines: [{ bookId, qty }]
const session = await mongoose.startSession();
try {
let order;
await session.withTransaction(async () => {
const items = [];
let totalPaise = 0;
for (const { bookId, qty } of lines) {
const book = await Book.findOneAndUpdate(
{ _id: bookId, stock: { $gte: qty } }, // reserve only if available
{ $inc: { stock: -qty } },
{ new: true, session }
);
if (!book) throw new Error(`Not enough stock for book ${bookId}`);
items.push({ bookId, title: book.title, pricePaise: book.pricePaise, qty });
totalPaise += book.pricePaise * qty;
}
const created = await Order.create(
[{ customerId, items, totalPaise, placedAt: new Date() }],
{ session }
);
order = created[0];
});
return order;
} finally {
await session.endSession();
}
}
// Adding a review: insert, then move the rollup with $inc (no transaction needed)
async function addReview(bookId, user, rating, text) {
await Review.create({ bookId, user, rating, text });
await Book.updateOne(
{ _id: bookId },
{ $inc: { ratingSum: rating, ratingCount: 1 } }
);
} - Every operation inside the transaction block is passed
{ session }. Miss it on one line and that write commits immediately and will not roll back — the mistake described in the transactions lesson, and the first thing to check when a rollback does not behave.
Step 6: Finish It Properly
A project is finished when someone else can clone it, run one command, and see it work. That means seed data, indexes created by a script rather than by hand, and a short README explaining the modelling decisions from step one.
Then try to break it. Order more copies than exist in stock and confirm nothing was changed. Submit a review with a rating of 9 and confirm validation rejects it. Submit the same review twice and confirm the unique index stops it. Delete an author and decide what should happen to their books — then implement that decision instead of leaving orphans behind.
- Write a seed script that is safe to run twice, using upserts rather than plain inserts
- Create every index from a script that runs as part of setup
- Handle
ValidationError,CastErrorand duplicate key error11000with proper 400-class responses - Keep the connection string in an environment variable and add
.envto.gitignore - Load a lakh of test books and run
explain("executionStats")on your catalogue query again - Write down, in the README, why reviews are referenced and why order items are embedded
- Extensions worth attempting once the core works: a wishlist as an array on the customer capped with
$slice; a TTL index on abandoned carts so they clean themselves up; soft deletes for books that are out of print but still appear in old orders; and a$lookupreport joining orders to authors to find which author earns the most.
