Lesson 24 of 25

Python Best Practices

PEP 8: Style That Removes Decisions

PEP 8 is Python's official style guide. Its value is not that its choices are objectively best — it is that they are settled. When everyone indents by four spaces and names functions in snake_case, nobody spends attention on the question, and any Python file you open anywhere looks familiar.

The naming conventions carry real information. snake_case for functions and variables, PascalCase for classes, UPPER_SNAKE_CASE for constants, a leading underscore for anything internal. Seeing Student next to student tells you which is the class and which is an instance without looking anything up.

The layout rules are about scanning. Four spaces per level, never tabs. Two blank lines before a top-level function or class, one between methods, so the eye can find where things begin. Spaces around operators and after commas, none just inside brackets. Imports at the top, grouped standard library, third-party, then your own.

Line length is the rule people argue about most. PEP 8 says 79 characters, which comes from an era of narrower screens; many modern projects use 88 or 100. Pick one number for the project and stop thinking about it — a formatter will enforce whichever you choose.

PEP 8 also says, in its own words, that consistency within a project matters more than the guide, and that readability wins over any rule. If following a rule makes a line harder to read, that is the case the guide is talking about. This is a style guide, not a set of laws.

Example
"""Marks utilities — module docstring on the first line."""

import os                       # standard library
import sys

import requests                 # third-party

from myapp import storage       # your own code


PASS_MARK = 40                  # UPPER_SNAKE_CASE for constants
MAX_MARKS = 100


class StudentRecord:            # PascalCase for classes
    """One student's marks for a term."""

    def __init__(self, name, marks):
        self.name = name
        self._marks = marks     # leading underscore: internal

    def average(self):          # one blank line between methods
        return sum(self._marks) / len(self._marks)


def has_passed(mark):           # two blank lines before top-level defs
    """Return True if the mark is a pass."""
    return mark >= PASS_MARK


# Spacing: around operators, after commas, not inside brackets
total = 10 + 5
result = has_passed(total)
marks = [78, 84, 91]

# Not this:
# total=10+5
# marks = [ 78,84 , 91 ]
# def hasPassed( mark ): ...
  • snake_case for functions, methods and variables
  • PascalCase for classes; UPPER_SNAKE_CASE for constants
  • Four spaces per indent level, never tabs
  • Two blank lines before top-level definitions, one between methods
  • Spaces around operators and after commas, none just inside brackets
  • Imports at the top: standard library, third-party, then local
Notes
  • PEP 8 governs code you write, not names Python gave you. Standard library functions such as isinstance and getattr predate the convention, and you will meet third-party libraries using camelCase throughout — follow each library's own style when calling it.

Let Tools Do It: Formatters, Linters, Type Checkers

Nobody should be enforcing style by hand or arguing about it in a code review. Three kinds of tool do the work, and setting them up takes about five minutes per project.

A formatter rewrites your file into a consistent layout. black is the well-known one and it is deliberately not configurable, which is the point — the formatting stops being anyone's opinion. Run it, and every file in the project looks the same regardless of who typed it.

A linter reads your code and reports problems. This is where the real value is, because a linter catches actual bugs and not only style: an unused import, a variable assigned and never used, a name that does not exist, a bare except, a mutable default argument. ruff is the current standard — it is very fast, and it covers what several older tools used to do separately. flake8 does the same job and is still widely used.

A type checker such as mypy reads your type hints and finds mismatches before you run anything: passing a string where an integer was expected, forgetting that a function can return None, calling a method that does not exist on that type. Python itself never checks hints at runtime, so this is the tool that makes them pay for themselves.

Once these are working, wire them into pre-commit so they run automatically before each commit. That turns "we should format our code" into something that simply happens, and it keeps style noise out of your diffs so reviews are about the logic.

Example
# Install into the project's virtual environment
python -m pip install black ruff mypy pre-commit

# Format every file in place
python -m black .

# Lint, and fix what can be fixed automatically
python -m ruff check .
python -m ruff check --fix .

# Type-check
python -m mypy src/


# What a linter catches that a formatter cannot:

import json          # F401 imported but unused

def load(path, cache=[]):     # B006 mutable default argument
    try:
        data = open(path).read()
    except:                   # E722 bare except
        pass
    return totl               # F821 undefined name 'totl'


# What a type checker catches:

def average(marks: list[int]) -> float:
    return sum(marks) / len(marks)

average("78,84")     # error: argument has incompatible type "str"


# .pre-commit-config.yaml — run them on every commit
# repos:
#   - repo: https://github.com/psf/black
#     rev: 24.4.2
#     hooks: [{id: black}]
#   - repo: https://github.com/astral-sh/ruff-pre-commit
#     rev: v0.4.4
#     hooks: [{id: ruff, args: [--fix]}]
  • black — formats automatically; not configurable, on purpose
  • ruff — fast linter; catches unused imports, undefined names, bare except, mutable defaults
  • mypy — checks type hints before the code runs
  • pre-commit — runs all of them automatically on every commit
  • Formatting in its own commit keeps real changes readable in a diff
Notes
  • Add the tools to a requirements-dev.txt rather than requirements.txt. They are needed to develop the project, not to run it, and keeping them separate keeps deployments smaller.

Names, Constants and Comments That Earn Their Place

The best readability improvement available to you is a better name, and it costs nothing. d tells the reader nothing; days_until_exam tells them everything and removes the need for a comment. Single letters are fine only where convention already gives them meaning — i for an index, x and y for coordinates, k and v for a key and value in a short loop.

Two specific characters to avoid entirely as names: lowercase l and uppercase I, which are indistinguishable from 1 in many fonts. PEP 8 mentions this explicitly, and anyone who has debugged l = 1 understands why.

A magic number is a bare literal whose meaning is not obvious — if marks >= 40, time.sleep(86400), data[7]. Give it a named constant at module level and two things improve: the reader learns what it means, and changing it later is one edit rather than a search across the file.

Comments should explain why, not what. i += 1 # add one to i is noise; the code already said that. i += 1 # the header row is not a record is worth having, because that reasoning exists nowhere else. A comment restating the code is worse than none, because it becomes wrong the moment the code changes and nobody updates it.

Docstrings are for anyone calling your function, comments are for anyone editing it. One sentence saying what a function returns and what it raises is enough, and it appears in help() and in your editor's tooltip. If a function needs a paragraph to explain, that is usually a sign it should be two functions.

Example
# Names that do the explaining
d = 30                              # what is d?
days_until_exam = 30                # no comment needed

lst = [78, 84]                      # a list of what?
marks = [78, 84]

def proc(x, y):                     # process what, using what?
    ...

def calculate_percentage(scored, total):
    ...

# Never use l, I or O as names — they look like 1 and 0
# l = 1


# Magic numbers become named constants
SECONDS_PER_DAY = 86_400
PASS_MARK = 40
MAX_RETRIES = 3
ROLL_COLUMN = 2

if marks[0] >= PASS_MARK:           # instead of >= 40
    pass


# Comments explain WHY
total = 0
for row in ["roll,marks", "101,87"][1:]:   # skip the header row
    total += int(row.split(",")[1])

# Not this:
# total += 1        # add 1 to total

# This is worth writing:
# The portal exports marks as text with a trailing space on some rows,
# so strip before converting or the int() call fails on those.


def calculate_grade(marks: int) -> str:
    """Return the letter grade for a mark out of 100.

    Raises ValueError if marks is outside 0-100.
    """
    if not 0 <= marks <= 100:
        raise ValueError(f"marks out of range: {marks}")
    return "A" if marks >= 90 else "B" if marks >= PASS_MARK else "F"
  • Descriptive names remove the need for most comments
  • Single letters only where convention already gives them meaning
  • Never name anything l, I or O
  • Replace bare literals with named constants at module level
  • Comments explain why; the code already says what
  • One-sentence docstrings on every public function, class and module
Notes
  • A comment that has to explain a confusing line is often a sign to rewrite the line. Extracting the confusing part into a well-named variable or function usually makes the comment unnecessary — and the explanation cannot then drift out of date.

Writing Python, Not Translated Java or C

Code can be correct, pass every test, and still announce that its author was thinking in another language. "Pythonic" is the word for using the idioms Python provides rather than reconstructing them, and it is one of the things interviewers notice.

The most visible tell is index-based looping. for i in range(len(items)) followed by items[i] is C written in Python. Loop over the items directly, use enumerate() when you genuinely need the position, and zip() when you are walking two collections together. The Python version is shorter and cannot go out of bounds.

The second is comparing to True and checking emptiness by length. if flag == True: is if flag:. if len(items) > 0: is if items:. Both rely on truthiness, which you now know exactly — including where the zero-versus-missing trap lies.

The rest is a short list you have already met, gathered in one place: f-strings rather than concatenation, with rather than manual close(), dict.get() where a key may be absent, unpacking rather than indexing a tuple, a comprehension where a loop only builds a list, enumerate rather than a manual counter, and try-and-handle rather than check-then-do.

One warning to balance it. Pythonic means clear, not clever. A nested comprehension with two conditions and a conditional expression uses idiomatic features and is still unreadable. If the idiomatic version is harder to follow than the plain one, write the plain one — that is what the style guide itself would tell you.

Example
items = ["Asha", "Ravi", "Meera"]
marks = [87, 92, 78]

# Index loops
for i in range(len(items)):        # translated from C
    print(items[i])

for name in items:                 # Python
    print(name)

for i, name in enumerate(items, start=1):    # when you need the position
    print(i, name)

for name, mark in zip(items, marks):         # two together
    print(name, mark)


# Truthiness
flag = True
if flag == True: pass              # noise
if flag: pass                      # Python

if len(items) > 0: pass            # translated
if items: pass                     # Python


# Manual counters and accumulators
count = 0
for m in marks:
    if m >= 40:
        count += 1
print(count)
print(sum(1 for m in marks if m >= 40))      # Python


# Strings, files, dictionaries, tuples
name, mark = "Asha", 87
print("Name: " + name + ", marks: " + str(mark))    # fragile
print(f"Name: {name}, marks: {mark}")               # Python

f = open("marks.txt", encoding="utf-8")             # may never close
try:
    data = f.read()
finally:
    f.close()

with open("marks.txt", encoding="utf-8") as f:      # Python
    data = f.read()

record = {"name": "Asha"}
if "email" in record:
    email = record["email"]
else:
    email = "not given"
email = record.get("email", "not given")            # Python

pair = ("Asha", 87)
n, m = pair[0], pair[1]
n, m = pair                                          # Python
  • for x in items, not for i in range(len(items))
  • enumerate() for positions, zip() for parallel collections
  • if flag: and if items:, not == True or len() > 0
  • f-strings over concatenation; with over manual close()
  • dict.get(key, default) where the key may be absent
  • Unpack tuples; use a comprehension when a loop only builds a list
  • Clear beats clever — an unreadable idiom is not Pythonic
Notes
  • import this at a Python prompt prints the Zen of Python, a short list of design principles. "Explicit is better than implicit", "Simple is better than complex" and "Readability counts" are the ones that settle most arguments.

Project Layout and Reproducibility

A project someone else can run in five minutes is worth far more than one that only works on your laptop. That is mostly about four files and one folder structure.

Put your code in a package folder rather than loose scripts, keep tests in a separate tests/ folder, and give the project a README.md saying what it does and exactly how to run it — the commands, in order, that take a fresh clone to working software. Write it as if for someone with your machine but not your memory, because in six months that is who you are.

Every project gets its own virtual environment, and the exact dependency versions are recorded in requirements.txt. "It works on my machine" is almost always a dependency difference. Pinning versions is what lets a project still build next year; the newer pyproject.toml does the same job and is the direction the ecosystem has moved, but requirements.txt remains perfectly respectable and is simpler to start with.

The .gitignore matters more than it looks. __pycache__/, .venv/ and *.pyc are generated files that create noisy diffs. .env is the one that actually matters: it is where your API keys live, and it must never be committed. Commit a .env.example with the variable names and empty values instead.

Two habits complete the picture. Guard the script part of any file with if __name__ == "__main__":, so importing it does not run it. And read configuration — keys, URLs, file paths — from environment variables rather than hardcoding it, so the same code runs on your laptop and on a server without editing.

Example
# marks_project/
# |- README.md              what it is, and exactly how to run it
# |- requirements.txt       pinned dependencies
# |- .gitignore             what must never be committed
# |- .env.example           variable names, no values
# |- src/
# |  |- marks/
# |     |- __init__.py
# |     |- storage.py
# |     |- report.py
# |     |- cli.py
# |- tests/
#    |- test_storage.py
#    |- test_report.py


# Setup, start to finish
python -m venv .venv
source .venv/bin/activate          # .venv\Scripts\activate on Windows
python -m pip install -r requirements.txt
python -m marks.cli

# Record what you installed
python -m pip freeze > requirements.txt


# .gitignore
# __pycache__/
# *.pyc
# .venv/
# .env
# .pytest_cache/


# src/marks/cli.py — configuration from the environment, not hardcoded
import os

DATA_PATH = os.environ.get("MARKS_DATA", "data/marks.csv")
API_KEY = os.environ.get("MARKS_API_KEY")

def main():
    if API_KEY is None:
        raise RuntimeError("set MARKS_API_KEY before running")
    print(f"reading {DATA_PATH}")

if __name__ == "__main__":
    main()
  • Code in a package folder, tests in tests/, a real README.md
  • One virtual environment per project, always
  • requirements.txt with pinned versions; pyproject.toml for larger projects
  • .gitignore must include .venv/, __pycache__/ and .env
  • Configuration from environment variables, never hardcoded
  • if __name__ == "__main__": so importing a file does not run it
Notes
  • The README is the part most students skip and the part a recruiter reads first. A project on GitHub with a clear README, a requirements file and a working run command is worth several without one.

Testing, Debugging and Logging

Tests are not an advanced topic to postpone. A test is a small function that calls your code and asserts what should come back, and having a handful means you can change your program without wondering what you broke.

pytest makes this about as easy as it can be. Put files named test_*.py in a tests/ folder, write functions named test_*, and use plain assert — no special assertion methods to learn. Running pytest finds and runs them all and prints the actual and expected values for anything that fails.

Start with the parts that are easy to test, which is another argument for the habits in this course: a function that takes values and returns a value can be tested in one line, while one that reads a file, prints and updates a global cannot. Test the normal case, the empty case, and the case that should raise — pytest.raises handles the last one.

For finding bugs, learn the debugger. breakpoint() on any line stops execution there and gives you a prompt where you can inspect every variable, step forward and continue. VS Code's debugger does the same with a click in the margin. It is a genuine step up from adding print() calls and deleting them afterwards, and it takes about twenty minutes to learn.

For programs that run unattended, replace diagnostic prints with the logging module. It gives you severity levels so you can switch detail on and off without editing code, timestamps, and the module each message came from — and logging.exception() inside an except block records the full traceback. It is the difference between a script you can diagnose from its output and one you have to re-run to understand.

Example
# src/marks/report.py
def average(marks):
    """Return the mean of a non-empty list of marks."""
    if not marks:
        raise ValueError("cannot average an empty list")
    return sum(marks) / len(marks)


# tests/test_report.py
import pytest
from marks.report import average

def test_average_of_several():
    assert average([78, 84, 91]) == pytest.approx(84.333, rel=1e-3)

def test_average_of_one():
    assert average([50]) == 50

def test_empty_list_raises():
    with pytest.raises(ValueError):
        average([])

# Run them:
#   python -m pip install pytest
#   python -m pytest


# Debugging: stop and look, instead of printing and guessing
def summarise(rows):
    total = 0
    for row in rows:
        breakpoint()          # drops into the debugger here
        total += int(row)
    return total

# At the prompt:  n = next line, c = continue, p total = print, q = quit


# Logging instead of print, for anything that runs unattended
import logging

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s %(levelname)s %(name)s: %(message)s",
)
log = logging.getLogger(__name__)

def load(path):
    log.info("loading %s", path)
    try:
        with open(path, encoding="utf-8") as f:
            return f.read()
    except FileNotFoundError:
        log.exception("could not read %s", path)   # records the traceback
        raise
  • pytest: test_*.py files, test_* functions, plain assert
  • Test the normal case, the empty case, and the case that should raise
  • pytest.raises(ValueError) — assert that something fails correctly
  • pytest.approx — compare floats with a tolerance, never with ==
  • breakpoint() — stop and inspect instead of scattering prints
  • logging over print for anything unattended; log.exception() in an except
Notes
  • Pass values to a logging call as arguments — log.info("loading %s", path) — rather than formatting them into the string yourself. The formatting is then skipped entirely when that level is switched off.
Ask AI