Lesson 19 of 20

Node.js Best Practices

Project Structure That Survives Growth

There is no single correct layout for a Node project, and most of the arguments about it are noise. What actually matters is that somebody who has never seen your code can work out where a thing belongs within a minute, and that changing one part of the application does not require touching five others.

The arrangement most projects converge on separates by layer, and you have met all of it already: routes for HTTP, services for the rules of the application, models for storage, middleware for cross-cutting concerns, config for configuration, and utils for helpers that are genuinely generic. server.js starts the process; src/app.js builds the Express application and exports it.

Once a project has more than three or four resources, grouping by feature beats grouping by layer. All the book files together, all the auth files together, each folder holding its own routes, service and model. The reason is entirely practical: a change is almost always about one feature and almost never about "all routes at once", so one folder to open beats four.

Two rules do the real work whichever layout you pick. A module should have one reason to change, which is what prevents a utils.js from growing into four hundred lines that every other file depends on. And dependencies should point in one direction — routes may import services, services may import models, and nothing points back. Circular imports do not crash; they hand you a half-initialised object at runtime, which is considerably worse than crashing.

Do not build structure you do not need. A project with one resource does not need a repository layer, an interface, and a folder of five-line files; that is ceremony which makes a small thing harder to read. Start with the smallest arrangement that keeps the HTTP layer separate from the logic, and add structure when the pain arrives rather than in anticipation of pain that may not.

Example
// Recommended project structure
my-api/
├── src/
│   ├── config/        # Configuration files
│   │   └── index.js
│   ├── controllers/   # Request handlers
│   │   └── userController.js
│   ├── middleware/     # Custom middleware
│   │   ├── auth.js
│   │   └── errorHandler.js
│   ├── models/        # Database models
│   │   └── User.js
│   ├── routes/        # Route definitions
│   │   ├── index.js
│   │   └── userRoutes.js
│   ├── utils/         # Helper functions
│   │   └── AppError.js
│   └── app.js         # Express app setup
├── tests/             # Test files
├── .env               # Environment variables (never committed)
├── .env.example       # Template for .env (always committed)
├── .gitignore
├── package.json
└── server.js          # Entry point: config, DB connect, app.listen


// Once there are several resources, group by FEATURE rather than by layer:
src/
├── books/
│   ├── book.routes.js
│   ├── book.service.js
│   └── book.model.js
├── auth/
│   ├── auth.routes.js
│   └── auth.service.js
├── middleware/
├── config/
└── app.js

// A change is nearly always about one feature and almost never about
// "all routes", so one folder to open beats four.
  • Routes handle HTTP · services hold the rules · models handle storage
  • server.js starts the process; src/app.js builds and exports the app
  • Group by layer while small; group by feature once there are several resources
  • One reason to change per module — this is what kills the ever-growing utils.js
  • Dependencies point one way; circular imports fail silently at runtime
  • Config in one module, error handling in one place, both imported everywhere
Notes
  • A good check on your structure: pick a feature you have not touched for a month and ask where you would add a field to it. If the answer is "three files, and I know which three", the layout is doing its job. If it is "somewhere in routes, probably utils as well, and I would have to search", the structure is costing you more than it saves.

The Security Checklist

Security in a Node application is not one feature; it is a short list of defaults, and almost every item on it is a line or two of code. What follows is the difference between a project that would survive being made public and one that would not, and none of it takes an afternoon.

Start with input, because everything else is downstream of it. Body, query, params, headers and uploaded filenames all come from outside and are all untrusted. Validate types and lengths, name the fields you accept rather than spreading a body, parameterise every database query, and cap the request body size. Those four habits cover injection, mass assignment and the simplest denial-of-service attempt in one pass.

Then authentication and transport. Hash passwords with a function designed for passwords, keep secrets in the environment with no fallback value, expire tokens, and rate-limit login. Serve everything over HTTPS in production — a token sent over plain HTTP can be read by anything sitting between the user and your server, which on shared public Wi-Fi is not a theoretical concern.

Headers and dependencies come next and are cheap. helmet for response headers and cors configured to a specific origin cost one line each. Then run npm audit, keep the lock file committed, and remember that every package you install runs with exactly your privileges — a quiet project with a hundred dependencies has a larger attack surface than a busy one with ten.

Finally, give everything the least access that works. The database user your application connects as does not need permission to drop tables. The API key you test with does not need production scope. And nothing you log should contain a password, a token, a card number or a full address, because logs get copied to laptops, pasted into chats and shipped to third-party services far more casually than databases ever are.

Example
// npm install helmet cors express-rate-limit

const express = require('express');
const helmet = require('helmet');
const cors = require('cors');
const rateLimit = require('express-rate-limit');

const app = express();

// Security headers
app.use(helmet());

// CORS
app.use(cors({ origin: process.env.CORS_ORIGIN }));

// Rate limiting
const limiter = rateLimit({
  windowMs: 15 * 60 * 1000, // 15 minutes
  max: 100, // Max 100 requests per window
  message: { error: 'Too many requests, try again later' }
});
app.use('/api', limiter);

// Body size limit
app.use(express.json({ limit: '10kb' }));

// Much tighter limits on the routes that get attacked first
const authLimiter = rateLimit({ windowMs: 15 * 60 * 1000, max: 10 });
app.use('/api/auth/login', authLimiter);
app.use('/api/auth/forgot-password', authLimiter);

// Behind a proxy, this is what lets rate limiting see real client addresses
// app.set('trust proxy', 1);

// Dependency hygiene
//   npm audit            read it; judge each finding by where that package runs
//   npm ci               reproducible installs, straight from the lock file
//   npm ci --omit=dev    production installs skip devDependencies

// Least privilege: the application's database user should not be able to drop
// tables, and your test API keys should not carry production scope.
  • Validate every input — its type, its length, and which field names you accept
  • Parameterised queries everywhere; never build SQL or a Mongo filter from a body
  • express.json({ limit: '100kb' }) — cap the body before it reaches memory
  • helmet for response headers, cors configured to a specific origin
  • Rate-limit globally, and far more tightly on login and password reset
  • HTTPS in production; secrets from the environment with no fallback value
  • npm audit, a committed lock file, and as few dependencies as you can manage
  • Least privilege for database users and API keys; never log secrets or personal data
Notes
  • Most of this list is enforced by things you install, but validation is not — and validation is the item that actually holds. helmet, cors and a rate limiter are perimeter measures; a route that trusts req.body is a hole inside the perimeter. If you only have time for one thing on this page, make it checking what arrived before you use it.

Logging You Can Actually Use

console.log is perfectly good while you are building, and it stops being enough the moment your application runs somewhere you cannot watch. A real log has levels, structure and enough context to answer a question, and you feel the difference on the first day something goes wrong that you did not personally witness.

Levels let you turn the volume down without losing the ability to turn it up. error for something that failed, warn for something suspicious that was handled, info for the small number of events that genuinely matter — started, connected, shutting down — and debug for detail you want locally and never in production. Read the threshold from an environment variable so you can raise it temporarily without a deployment.

In production, log JSON rather than sentences. A line such as {"level":"error","requestId":"a1b2","userId":7,"msg":"payment failed"} can be searched and filtered by tooling; "Payment failed for user 7" can only be read by a human who already knows what to look for. Libraries like pino or winston do this for you, and are faster than console.log under load into the bargain.

The single most valuable field is a request id, generated per request and attached to every line that request produces. Without one, a busy log is many stories interleaved and none of them readable. With one, a single identifier from a user's complaint reconstructs exactly what happened, in order, across every layer of your application.

And what never to log: passwords, tokens, OTPs, card numbers, or full personal records. Logs are copied to laptops, pasted into chat and shipped to third-party services, so a password in a log file is a password in many more places than a password in a database. Be particularly wary of logging an entire request body or an entire user document — that is how these values arrive in a log without anybody deciding to put them there.

Example
// A structured logger, configured once
const pino = require('pino');
const logger = pino({ level: process.env.LOG_LEVEL || 'info' });

// A request id, attached to everything this request produces
const crypto = require('node:crypto');
app.use((req, res, next) => {
  req.id = req.headers['x-request-id'] || crypto.randomUUID();
  req.log = logger.child({ requestId: req.id });
  res.setHeader('X-Request-Id', req.id);
  next();
});

app.post('/api/orders', async (req, res, next) => {
  req.log.info({ userId: req.user.id }, 'creating order');
  try {
    const order = await createOrder(req.user.id, req.body);
    req.log.info({ orderId: order.id }, 'order created');
    res.status(201).json(order);
  } catch (err) {
    req.log.error({ err }, 'order failed');   // logged once, where it is handled
    next(err);
  }
});

// NEVER:
// logger.info({ body: req.body });   // passwords, tokens, card numbers
// logger.info({ user });             // the whole document, hash included
// console.log('token', token);

// Log an error where you handle it, not at every level it passes through,
// or one failure produces six near-identical entries and none of them helps.
  • Levels — error, warn, info, debug — with the threshold read from the environment
  • Structured JSON in production; readable output in development
  • A request id on every line, and returned to the client in a header
  • Log an error once, where it is handled, not at every level it passes
  • Never log passwords, tokens, OTPs, card numbers or whole request bodies
  • pino or winston — both more useful and faster than console.log
Notes
  • The test of a log line is whether it would help somebody at midnight who does not have your context. "Error" helps nobody. "payment failed" with a request id, a user id and the provider's error code turns a support message into a two-minute investigation — and it costs about the same number of keystrokes to write.

Testing: the Minimum That Pays for Itself

Testing is where most student projects stop, and a small amount of it is worth far more than the effort suggests. You do not need full coverage. You need enough tests that a change breaking something obvious fails before you deploy it rather than afterwards, in front of somebody.

Node ships with a test runner — node --test — which is entirely sufficient for a project and requires no dependency at all. Jest and Vitest add more machinery if you want it. The piece you genuinely want for an API is supertest, which makes requests against your Express app in memory, with no port opened and nothing to clean up afterwards.

That is precisely why app.js and server.js were separated back in lesson 8. Because app is exported without calling listen, a test can import it directly and fire requests at it. If your app file starts a server when imported, every test file competes for port 3000 and the suite becomes unreliable in ways that have nothing to do with the code you are testing.

Test two layers, and treat them as different jobs. Services are ordinary functions with no HTTP anywhere near them, so test the rules there — that is where the interesting logic lives, and those tests run instantly. Routes are tested through requests, and the four cases from lesson 12 make the checklist: the happy path, a missing resource, invalid input, and something deliberately strange.

Write a test the moment you fix a bug, and before the fix if you can manage it. It takes five minutes, it proves the fix actually works, and it is the only reliable way of stopping the same bug returning in six months when somebody refactors that file. A suite grown this way becomes a list of every mistake the project has already made once.

Example
// npm install --save-dev supertest
// package.json:  "test": "node --test"

const test = require('node:test');
const assert = require('node:assert');
const request = require('supertest');

const app = require('../src/app');   // imported, NOT started — see lesson 8

test('GET /api/books returns an array', async () => {
  const res = await request(app).get('/api/books');
  assert.strictEqual(res.status, 200);
  assert.ok(Array.isArray(res.body));
});

test('POST /api/books rejects a missing title', async () => {
  const res = await request(app)
    .post('/api/books')
    .send({ author: 'Frank Herbert' });
  assert.strictEqual(res.status, 400);      // 400, not 500
});

test('GET /api/books/9999 is a 404, not a crash', async () => {
  const res = await request(app).get('/api/books/9999');
  assert.strictEqual(res.status, 404);
});

// Services are plain functions — no HTTP, no server, nothing to mock
const { calculateTotal } = require('../src/services/orderService');

test('calculateTotal applies the discount once', () => {
  assert.strictEqual(calculateTotal([{ price: 100 }, { price: 50 }], 0.1), 135);
});
  • node --test is built in and enough for a project; Jest or Vitest add more
  • supertest makes requests against your app in memory, with no port opened
  • Exporting app without calling listen is what makes route tests possible
  • Test services as plain functions; test routes through requests
  • The same four cases: happy path, missing, invalid, deliberately strange
  • Write a test whenever you fix a bug — that is what stops it coming back
Notes
  • If you write no other tests, write the ones that assert your error paths. Everybody tests that a working request works; the bugs live in the branch that is supposed to return 400, and that is precisely the branch nobody exercises by hand, because sending deliberately bad input takes effort that feels wasted right up until it is not.

Performance and Running It in Production

Most Node performance problems are not Node being slow. They are, in order of frequency: a query without an index, a query inside a loop, and synchronous work on the request path. Measure before you optimise, because the intuitive answer is usually wrong and the real one is usually boring.

The single-thread rule from lesson 1 is what bites in production. A readFileSync, a large synchronous JSON.parse, a synchronous hash or a long loop inside a request handler blocks every other request for as long as it runs. Read configuration synchronously at startup, where blocking costs nothing at all, and keep the per-request path asynchronous without exception.

For large files, stream rather than buffer. fs.readFile loads the entire thing into memory before you can do anything with it, so a hundred-megabyte export becomes a hundred megabytes of memory per concurrent request. A stream passes it through in pieces and the memory stays flat regardless of how large the file is or how many people ask for it at once.

One Node process uses one processor core. Running several — through a process manager, or by letting your platform run several instances — is how you use the rest of the machine. The catch is the one from lesson 12: the moment there is more than one process, anything held in a module-level variable stops being shared. Sessions, caches and counters have to move into a database or a cache server, and discovering that after you scale is unpleasant.

A few small things separate an application that works from one that can be operated: a health endpoint that reports whether the database is actually reachable, graceful shutdown on SIGTERM so a deployment does not cut live requests in half, logs somewhere searchable, and something watching that tells you the service is down before a user does. None of them are large pieces of work, and together they are the difference between a project and a service.

Example
// Measure first. The answer is nearly always one of three things.
console.time('query');
const rows = await findBooks(filter);
console.timeEnd('query');

// 1. Never block the request path
// const config = JSON.parse(fs.readFileSync('config.json'));  // fine at STARTUP
app.get('/report', (req, res) => {
  // const data = fs.readFileSync(bigFile);   // blocks every other request
});

// 2. Stream large files instead of buffering them into memory
const fs = require('node:fs');
app.get('/export.csv', (req, res) => {
  res.setHeader('Content-Type', 'text/csv');
  fs.createReadStream('/data/export.csv')
    .on('error', () => res.status(500).end())
    .pipe(res);                    // memory stays flat, whatever the file size
});

// 3. A health endpoint that actually means something
app.get('/health', async (req, res) => {
  try {
    await pingDatabase();
    res.json({ status: 'ok' });
  } catch {
    res.status(503).json({ status: 'degraded' });
  }
});

// More than one process means no shared memory. Anything in a module-level
// variable is per-process from that moment on — move it to a database or cache.
  • Measure first: missing indexes, queries in loops, synchronous work — usually in that order
  • No blocking work on the request path; synchronous reads belong at startup
  • Stream large files; buffering costs memory proportional to concurrent requests
  • One process uses one core; more processes means no shared in-memory state
  • A health endpoint that checks the database, not one that answers ok unconditionally
  • Graceful shutdown on SIGTERM, so a deploy does not cut live requests in half
Notes
  • A health endpoint that always returns ok is worse than having none, because it tells your platform everything is fine while the database is unreachable and every real request is failing. Check the dependency you genuinely need and answer 503 when it is missing — that is what lets a platform take a broken instance out of rotation instead of continuing to send it traffic.
Ask AI