A Module Is Just a File
There is no special syntax for creating a module. Any file ending in .py is one. Save some functions in helpers.py and another file in the same folder can write import helpers and use them. The standard library modules you already import — math, random, json — are the same thing, written by other people and shipped with Python.
What import actually does is worth knowing, because it explains several confusing behaviours. Python finds the file, runs it from top to bottom, collects the names it defined, and binds the module object to a name in your file. Running it is the part people miss: a print() at the top level of a module fires the moment someone imports it.
It runs only once per program, though. The result is cached in sys.modules, so a second import of the same module in the same run is nearly free and does not re-execute anything. That is why editing a module while an interactive session is open does not pick up your changes — you have to restart, or use importlib.reload().
Because importing runs the file, a module that is also useful as a script needs a way to tell the difference. Python sets the variable __name__ to "__main__" when a file is run directly and to the module's own name when it is imported. Wrapping the script part in if __name__ == "__main__": means the file can be imported for its functions without also running its demo, its tests or its menu loop.
# ---- helpers.py ----
"""Small utilities for the marks project."""
PASS_MARK = 40
def average(marks):
return sum(marks) / len(marks)
def has_passed(mark):
return mark >= PASS_MARK
print("helpers.py is being run or imported") # fires on import too
if __name__ == "__main__":
# Only when run directly: python helpers.py
print(average([78, 84, 91]))
# ---- main.py, in the same folder ----
import helpers
print(helpers.PASS_MARK) # 40
print(helpers.average([80, 90])) # 85.0
print(helpers.__name__) # helpers
print(__name__) # __main__ (this file is the one running)
# Importing twice does not run it twice
import helpers
import sys
print("helpers" in sys.modules) # True — served from the cache
print(helpers.__doc__) # Small utilities for the marks project.
print(dir(helpers)) # everything the module defines dir(module)lists the names a module provides andhelp(module)prints its documentation. Both work at the interactive prompt and are faster than searching the web when you only need to remember an argument name.
The Four Import Forms
import math brings in the module as a whole; you then reach through it with math.sqrt. This is the safest form, because the module name stays attached to everything you use and a reader can see instantly where sqrt came from.
from math import sqrt copies one name into your file, so you write sqrt(16) directly. It is shorter and it is right when you use one or two things repeatedly. The cost is that the origin disappears from the call site, and if two modules both export a load function, the second import silently replaces the first with no warning at all.
import numpy as np renames on the way in. Use it for long module names and for the conventions the community has settled on — np for NumPy, pd for pandas, plt for matplotlib.pyplot. These are so standard that using different aliases makes your code harder to read, not more personal.
from math import * imports every public name at once, and you should not use it. It makes it impossible to tell which module any given name came from, and it can quietly overwrite your own variables — a file that defines e = 2 and then does from math import * has an e of 2.718. The one place you will see it legitimately is at an interactive prompt while exploring.
PEP 8 asks for imports at the top of the file, grouped: standard library first, then third-party packages, then your own modules, with a blank line between the groups. Keeping to that makes a file's dependencies readable in five seconds, which matters more than it sounds when you return to it in a month.
# 1. The whole module — origin stays visible
import math
print(math.sqrt(16)) # 4.0
print(math.ceil(4.3)) # 5
# 2. Specific names — shorter, but the origin disappears
from math import sqrt, pi
print(sqrt(25), round(pi, 4)) # 5.0 3.1416
# 3. An alias — for long names and community conventions
import datetime as dt
print(dt.datetime.now().year)
# import numpy as np
# import pandas as pd
# import matplotlib.pyplot as plt
# 4. Everything — do not do this
e = 2
from math import *
print(e) # 2.718281828459045 — your variable is gone
# The collision problem, in miniature
# from json import load
# from pickle import load # silently replaces the first one
# PEP 8 import order
import os # standard library
import sys
import requests # third-party
import helpers # your own modules import module— safest; every use says where it came fromfrom module import name— shorter; risks silent name collisionsimport module as alias— for long names and the standardnp/pd/pltconventionsfrom module import *— avoid; hides origins and overwrites your own names- Imports go at the top, grouped: standard library, third-party, then your own
- An import inside a function is legal and occasionally useful — for a slow optional dependency, or to break a circular import. It is an exception, not a style; the default remains the top of the file.
Where Python Looks — and the Bug Everyone Hits
When you write import random, Python searches a list of folders held in sys.path. The first entry is the folder containing the script you ran, followed by the installed library locations. First match wins, and that ordering is the source of the most common import bug in a beginner's first month.
If you save your own practice file as random.py, your folder is searched first, so import random finds your file. The symptom is bizarre: AttributeError: module 'random' has no attribute 'randint', or a program that imports fine and then behaves as though the standard library vanished. The same happens with math.py, json.py, string.py, email.py and test.py. The fix is simply to rename your file — and to delete the __pycache__ folder next to it, because a stale compiled copy can keep the problem alive after the rename.
The other frequent failure is ModuleNotFoundError for something you are certain you installed. Nine times out of ten there are two Pythons on the machine and pip installed into the one you are not running. Installing with python -m pip install X removes the ambiguity, because -m uses the pip belonging to the interpreter you just named. import sys; print(sys.executable) tells you exactly which interpreter is running.
A third case appears once your project has several files: a circular import, where a.py imports b.py and b.py imports a.py. Python starts running one, reaches the import of the other, starts running that, and comes back to a module that is only half-defined. The error mentions a partially initialised module. The real fix is structural — move the shared piece into a third module that both can import.
import sys
print(sys.executable) # which interpreter is running this
print(sys.path[:3]) # where imports are searched, in order
# The shadowing bug, step by step:
# 1. You save a practice file as random.py
# 2. In the same folder you run another script containing:
# import random
# print(random.randint(1, 10))
# 3. AttributeError: module 'random' has no attribute 'randint'
#
# Python found YOUR random.py first, because sys.path[0] is your folder.
#
# Fix: rename the file (random_practice.py) and delete __pycache__/
# Which file did an import actually load?
import random
print(random.__file__) # ...\Lib\random.py if it is the real one
# Installing into the interpreter you are actually running
# python -m pip install requests <- unambiguous
# pip install requests <- may target a different Python ModuleNotFoundErroris a subclass ofImportError. It means the module could not be found at all; a plainImportErrorusually means the module was found but the specific name you asked for is not in it.
Your Own Modules and Packages
Splitting a program into files is the point at which it starts being a project rather than a script. Group related functions into a module named for what it does — storage.py, report.py, validation.py — and let the entry-point file import them. A file that is more than a few hundred lines is usually two files in disguise.
A package is a folder of modules. Import from it with a dot: from myapp.storage import load_tasks. Traditionally the folder needed an __init__.py file to be recognised as a package; since Python 3.3 a plain folder often works, but including an empty __init__.py is still the clearer and safer choice, and it gives you somewhere to put the package's public interface if you want one.
Imports inside a package come in two forms. An absolute import names the full path from the project root: from myapp.storage import load_tasks. A relative import uses dots for "this package": from .storage import load_tasks. Absolute imports are clearer and are what PEP 8 recommends; relative ones are more compact inside a large package.
Relative imports have one sharp edge worth naming, because it wastes an afternoon the first time. They only work when the file is being run as part of a package. Run that same file directly with python myapp/storage.py and you get ImportError: attempted relative import with no known parent package. The fix is to run the package instead — python -m myapp.storage from the project root — rather than to change the import.
# A small project laid out as a package
#
# project/
# main.py
# myapp/
# __init__.py
# storage.py
# report.py
# ---- myapp/storage.py ----
import json
def load_tasks(path="tasks.json"):
with open(path) as f:
return json.load(f)
def save_tasks(tasks, path="tasks.json"):
with open(path, "w") as f:
json.dump(tasks, f, indent=2)
# ---- myapp/report.py ----
from myapp.storage import load_tasks # absolute — clear from anywhere
# from .storage import load_tasks # relative — only inside the package
def pending_count():
return sum(1 for t in load_tasks() if not t["done"])
# ---- main.py ----
from myapp.report import pending_count
if __name__ == "__main__":
print(f"{pending_count()} tasks pending")
# Run it from the project folder:
# python main.py
# Run a module inside the package directly:
# python -m myapp.report (works)
# python myapp/report.py (breaks relative imports) - Name your modules in lowercase with underscores, and avoid names that clash with the standard library. If a module name feels generic —
utils.py,helpers.py,misc.py— it usually means the file has no single job and will grow into a dumping ground.
The Standard Library Worth Knowing Now
Python ships with roughly two hundred modules, and the phrase people use for it is "batteries included". You do not need to learn them, but knowing that about a dozen exist saves you from writing code that already exists — which is also what an interviewer is checking when they ask how you would shuffle a list or count word frequencies.
math and statistics cover numbers: square roots, constants, factorials, and mean, median and mode without needing NumPy. random covers picking, shuffling and generating; random.seed() makes a run reproducible, which matters when you are testing something that uses randomness.
datetime handles dates and times, with timedelta for arithmetic on them and strftime for formatting. pathlib is the modern way to work with file paths — Path("data") / "marks.csv" builds a correct path on Windows and Linux alike, which string concatenation does not. os and sys cover the rest of the environment: working directory, environment variables, command-line arguments.
json and csv read and write the two formats you will actually meet. collections holds Counter, defaultdict and deque, all covered earlier in this course. itertools has combinatorial tools such as combinations and permutations that turn several nested loops into one line, and re handles pattern matching once simple string methods run out.
The habit to build is to check the standard library before installing anything. It is already on every machine, it needs no version management, and it will still work in five years.
import math, random, statistics
from datetime import datetime, timedelta
from pathlib import Path
from collections import Counter
from itertools import combinations
marks = [78, 84, 91, 84, 65]
print(math.sqrt(144), math.floor(4.7), math.factorial(5)) # 12.0 4 120
print(statistics.mean(marks), statistics.median(marks))
print(statistics.mode(marks)) # 84
random.seed(42) # reproducible for testing
print(random.randint(1, 100))
print(random.choice(["Heads", "Tails"]))
print(random.sample(marks, 2)) # 2 distinct items
random.shuffle(marks) # in place, returns None
now = datetime.now()
print(now.strftime("%d %B %Y")) # 02 August 2026
print((now + timedelta(days=30)).strftime("%Y-%m-%d"))
print((datetime(2026, 12, 31) - now).days, "days left in the year")
path = Path("data") / "marks.csv" # correct on every OS
print(path, path.suffix, path.exists())
print(Counter("mississippi").most_common(2)) # [('i', 4), ('s', 4)]
print(list(combinations(["A", "B", "C"], 2))) # every pair math,statistics— arithmetic, constants, mean/median/moderandom— choice, shuffle, sample;seed()makes runs repeatabledatetime— dates, differences withtimedelta, formatting withstrftimepathlib,os,sys— paths, environment, command-line argumentsjson,csv— the two data formats you will meet mostcollections,itertools,re— counting, combinatorics, pattern matching
- Anything using
randomfrom therandommodule is pseudo-random and predictable from its seed. That is exactly right for simulations and games and wrong for anything security-related; use thesecretsmodule for tokens and passwords.
pip, Virtual Environments and requirements.txt
Everything outside the standard library comes from the Python Package Index, and pip is the tool that fetches it. Install with python -m pip install requests. The -m form is worth making a habit, because it always installs into the interpreter you named rather than whichever pip happens to be first on your PATH.
Installing packages system-wide causes a specific problem, and it appears the moment you have two projects. One needs an older version of a library and one needs the newer, and only one of them can win. A virtual environment solves it by giving each project its own private copy of the installed packages. Create one with python -m venv .venv, activate it, and every pip install after that affects only that project.
Activation differs by platform: .venv\Scripts\activate on Windows, source .venv/bin/activate on macOS and Linux. Once active, your prompt shows the environment's name and python means the environment's Python. deactivate returns you to normal. The .venv folder is generated, not source code — add it to .gitignore and never commit it.
What you commit instead is the list of what you installed. python -m pip freeze prints every package with its exact version, and redirecting that into requirements.txt records it. Anyone who clones your project then runs python -m pip install -r requirements.txt and gets the same set. Pinning exact versions this way is what makes a project reproducible six months later, when the libraries have moved on and yours has not.
One last piece of advice for a shared or college machine: never run pip with administrator rights to fix a permissions error. Doing so writes into the system Python that the operating system itself depends on. Create a virtual environment, or add --user, and the problem disappears without risk.
# Install into the interpreter you are actually running
python -m pip install requests
python -m pip install "pandas==2.2.0" # a specific version
python -m pip list # what is installed
python -m pip show requests # version, location, dependencies
# One environment per project
python -m venv .venv
# Activate it
.venv\Scripts\activate # Windows
source .venv/bin/activate # macOS / Linux
# Now pip installs only into this project
python -m pip install requests pandas
# Record exactly what this project needs
python -m pip freeze > requirements.txt
# Someone else reproduces it
python -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
deactivate # leave the environment
# .gitignore should contain:
# .venv/
# __pycache__/
# *.pyc python -m pip install X— install into the interpreter you namedpython -m venv .venv— a private package set for this project- Activate:
.venv\Scripts\activate(Windows) orsource .venv/bin/activate python -m pip freeze— the exact versions installed, forrequirements.txtpython -m pip install -r requirements.txt— recreate the same set elsewhere- Commit
requirements.txt; never commit.venv/or__pycache__/
- If
activatefails on Windows PowerShell with a message about execution policies, that is Windows blocking scripts rather than a Python problem. Running the same command from Command Prompt, or from the terminal inside VS Code with the interpreter selected, is the quickest way past it.
