open(), and Why with Is Not Optional
open() asks the operating system for access to a file and hands back a file object that you read from or write to. When you are finished, that access has to be released — the operating system limits how many files a program may hold open, and, more importantly, written data sits in a buffer until the file is closed or the buffer fills.
You can call f.close() yourself, and you will get it wrong eventually. Any exception between open() and close() skips the close entirely, and then the buffered text is never written. The file exists, it is empty or truncated, and nothing told you.
The with statement fixes this permanently. with open(...) as f: closes the file when the block ends — normally, on a return, or when an exception is thrown out of it. It is the same as a try/finally pair, written in one line, and it is why you will almost never see a bare open() in decent Python.
The mechanism has a name worth knowing: with works with any context manager, an object that defines what should happen on entry and on exit. Files are the first one you meet; database connections, network sockets and locks all use the same pattern. Once the idea clicks for files, it transfers everywhere.
One clarification that saves confusion later: the file object is only valid inside the with block. Reading from f after the block raises ValueError: I/O operation on closed file. If you need the contents afterwards, store them in a variable inside the block.
# The right way
with open("notes.txt", "w", encoding="utf-8") as f:
f.write("first line\n")
f.write("second line\n")
# closed here, automatically, whatever happened inside
# What with replaces
f = open("notes.txt", "r", encoding="utf-8")
try:
data = f.read()
finally:
f.close() # runs even if read() raised
# What goes wrong without either
f = open("report.txt", "w", encoding="utf-8")
f.write("important results")
# an exception here means close() never runs and the text is lost
# The file object dies with the block
with open("notes.txt", encoding="utf-8") as f:
contents = f.read() # keep what you need
print(f.closed) # True
# f.read() # ValueError: I/O operation on closed file
print(contents)
# Several files in one with
with open("in.txt", encoding="utf-8") as src, \
open("out.txt", "w", encoding="utf-8") as dst:
for line in src:
dst.write(line.upper()) - Notebooks and interactive sessions hide this problem, because closing the session flushes everything. A script that works in Jupyter and produces an empty file when run from the terminal is very often a missing
with.
Modes, and the One That Deletes Your Data
The second argument to open() is the mode, and choosing wrongly is the most expensive mistake in this lesson. "w" means write, and it empties the file the moment it is opened — before you have written anything, before any error, irreversibly. Opening an existing file with "w" to "update" it destroys what was there.
"a" is append: the file is created if missing and new text goes to the end, leaving existing content alone. This is what you want for logs and for adding records. "r" is read, is the default, and raises FileNotFoundError if the file does not exist — which is usually the behaviour you want, since silently creating an empty input file would hide the real problem.
"x" is exclusive creation: it creates the file and raises FileExistsError if something is already there. Reach for it when overwriting would be a bug, such as when exporting a report you must not clobber.
Adding "b" gives binary mode, and it is not optional for images, PDFs, zip archives or anything that is not text. In binary mode you read and write bytes rather than str, and no encoding or newline translation is applied. Opening a JPEG in text mode either raises UnicodeDecodeError or silently corrupts it.
There is also "r+" and friends for reading and writing the same handle. They exist, they involve thinking about the file position, and for almost everything a beginner needs, the safer pattern is: read the whole file, work on the data in memory, then write it back.
# "w" truncates immediately — before anything is written
with open("data.txt", "w", encoding="utf-8") as f:
pass # the file is now empty
# "a" adds to the end and creates the file if needed
with open("log.txt", "a", encoding="utf-8") as f:
f.write("run finished\n")
# "r" is the default and complains about a missing file
try:
with open("missing.txt", encoding="utf-8") as f:
print(f.read())
except FileNotFoundError:
print("no such file — create it or check the path")
# "x" refuses to overwrite
try:
with open("log.txt", "x", encoding="utf-8") as f:
f.write("fresh")
except FileExistsError:
print("already exists — not touching it")
# Binary mode for anything that is not text
with open("logo.png", "rb") as f:
header = f.read(8)
print(type(header)) # <class 'bytes'>
# The safe update pattern: read, change, write back
with open("data.txt", encoding="utf-8") as f:
lines = f.readlines()
lines = [line for line in lines if line.strip()]
with open("data.txt", "w", encoding="utf-8") as f:
f.writelines(lines) "r"— read; the default;FileNotFoundErrorif absent"w"— write; empties an existing file immediately"a"— append; creates if missing, never truncates"x"— create only;FileExistsErrorif it is already there"rb"/"wb"— binary; works inbytes, no encoding, no newline translation"r+","w+","a+"— read and write together; think about the file position first
- Before running a script that writes with
"w", check the filename twice. There is no undo, no recycle bin and no confirmation — the operating system has already been told to empty the file.
Reading Without Running Out of Memory
There are four ways to read, and the difference between them is memory. f.read() loads the entire file into one string. That is fine for a configuration file and disastrous for a two-gigabyte log, where it will either fill your RAM or fail outright.
f.readlines() has the same problem in a different shape: it builds a list holding every line, so the whole file is still in memory, plus the overhead of a list. Its useful property is that you get a real list you can index and re-read.
Looping over the file object directly — for line in f: — reads one line at a time and keeps only that line in memory. It is the shortest to write and the one that scales, so make it your default. The file object is an iterator, which also means it is consumed once: a second loop over the same open file yields nothing until you rewind with f.seek(0).
Every line you read keeps its trailing newline, because that is what is actually in the file. print(line) then produces double spacing, which is the first thing everyone notices. line.rstrip("\n") removes exactly the newline; line.strip() also removes leading and trailing spaces, which is usually what you want for data but not when the spacing is meaningful.
The last line of a file may or may not end with a newline, depending on who wrote it. Code that assumes it does will produce one wrong record; stripping each line and skipping empty ones handles both cases without special-casing.
# Everything at once — only for small files
with open("notes.txt", encoding="utf-8") as f:
text = f.read()
print(len(text), "characters")
# Every line as a list — still holds the whole file
with open("notes.txt", encoding="utf-8") as f:
lines = f.readlines()
print(lines[:2]) # ['first line\n', 'second line\n']
# One line at a time — the default choice
with open("notes.txt", encoding="utf-8") as f:
for line in f:
print(line.rstrip("\n"))
# The file object is an iterator: one pass only
with open("notes.txt", encoding="utf-8") as f:
first = list(f)
second = list(f) # [] — already exhausted
f.seek(0) # rewind to the start
third = list(f)
print(len(first), len(second), len(third))
# A realistic read: skip blanks and comments, count what matters
total = 0
with open("marks.txt", encoding="utf-8") as f:
for line_no, line in enumerate(f, start=1):
line = line.strip()
if not line or line.startswith("#"):
continue
try:
total += int(line)
except ValueError:
print(f"line {line_no}: not a number -> {line!r}")
print("total:", total) {line!r}inside an f-string appliesrepr(), so the value prints with its quotes and escape characters visible. When a file line is misbehaving, that instantly shows whether there is a hidden space or a stray carriage return.
Writing: Newlines and print(file=)
f.write(text) writes exactly the string you give it and adds nothing. No newline, no separator. Forgetting that is why a first attempt at writing records produces one enormous line. Add \n yourself at the end of each piece.
f.writelines(lines) is named misleadingly: it writes a sequence of strings and, again, adds no newlines between them. It is exactly equivalent to a loop of write() calls. It is genuinely useful for writing back lines you read with readlines(), since those already carry their newlines.
There is a friendlier alternative you already know. print(value, file=f) writes to the file instead of the screen and behaves exactly like print otherwise — it converts non-strings with str(), joins several values with a space, and adds the newline for you. For anything with formatting, this is usually the more readable choice.
write() returns the number of characters written, which is occasionally useful and mostly ignored. What is worth remembering is that writing is buffered: text may sit in memory until the file is closed. Inside a with block this is invisible and safe. If you need output on disk immediately — for a log another program is watching — call f.flush().
students = [("Asha", 87), ("Ravi", 92), ("Meera", 78)]
# write() adds nothing — the \n is yours to supply
with open("marks.txt", "w", encoding="utf-8") as f:
for name, mark in students:
f.write(f"{name},{mark}\n")
# Forgetting it gives one long line
with open("broken.txt", "w", encoding="utf-8") as f:
for name, mark in students:
f.write(f"{name},{mark}") # Asha,87Ravi,92Meera,78
# writelines adds nothing either
with open("marks.txt", encoding="utf-8") as f:
lines = f.readlines() # these still end in \n
with open("copy.txt", "w", encoding="utf-8") as f:
f.writelines(lines) # so this is correct
# print(file=) converts and adds the newline for you
with open("report.txt", "w", encoding="utf-8") as f:
print("Marks report", file=f)
print("=" * 12, file=f)
for name, mark in students:
print(f"{name:<8}{mark:>4}", file=f)
print("Average:", sum(m for _, m in students) / len(students), file=f)
# write() reports how much it wrote
with open("log.txt", "a", encoding="utf-8") as f:
n = f.write("done\n")
print(n, "characters") # 5
f.flush() # force it to disk now f.write(s)— writes exactlys; you supply the\nf.writelines(seq)— a loop of writes; still adds no newlinesprint(x, file=f)— converts withstr()and adds the newline- Writing is buffered;
withflushes on exit, or callf.flush()yourself f.write()returns the number of characters written
- For structured data, do not invent your own text format.
csvhandles fields containing commas and quotes correctly, andjsonpreserves types and nesting. Both are covered in the last section, and both save you writing a parser.
Encoding and Paths: Two Sources of Mysterious Errors
A text file on disk is bytes. Turning those bytes into characters requires knowing the encoding, and if you do not say which one, Python uses a platform default that differs between machines. That is why a script that works on your laptop can fail on a classmate's with UnicodeDecodeError: 'charmap' codec can't decode byte .... The file did not change; the assumed encoding did.
The fix is one argument: always pass encoding="utf-8" when opening a text file. UTF-8 handles every script — English, Hindi, Odia, emoji — and it is what nearly everything on the internet produces. Being explicit also makes your code behave identically on Windows, macOS and Linux, which is worth more than the six characters it costs.
If you are handed a file in some other encoding and cannot change it, either name that encoding, or pass errors="replace" to substitute a marker for the bytes that will not decode. Use the second one for a quick look at unknown data, not for anything you will store.
The other mysterious failure is FileNotFoundError for a file that is plainly sitting next to your script. A relative path such as "marks.txt" is resolved against the current working directory — the folder your terminal was in when it started Python — not the folder containing the script. Run the same script from a different directory and it breaks.
Two habits solve it. Path(__file__).parent / "marks.txt" builds a path relative to the script itself, wherever it is run from. And using pathlib generally is better than gluing strings together: Path("data") / "marks.csv" puts the right separator in for the platform, and .exists(), .stem, .suffix and .mkdir(parents=True, exist_ok=True) cover most of what you need.
from pathlib import Path
# Always say the encoding for text files
with open("hindi.txt", "w", encoding="utf-8") as f:
f.write("नमस्ते\n")
with open("hindi.txt", encoding="utf-8") as f:
print(f.read().strip()) # नमस्ते
# Without it, another machine may raise:
# UnicodeDecodeError: 'charmap' codec can't decode byte 0x8f ...
# A file you cannot re-encode
with open("legacy.txt", encoding="latin-1") as f:
pass
with open("legacy.txt", encoding="utf-8", errors="replace") as f:
pass # unreadable bytes become a marker character
# Relative paths follow the TERMINAL's folder, not the script's
print(Path.cwd()) # where relative paths start from
# Anchor to the script instead
DATA = Path(__file__).parent / "marks.txt"
print(DATA)
# pathlib handles separators and the common questions
p = Path("data") / "reports" / "august.csv"
print(p.name, p.stem, p.suffix) # august.csv august .csv
print(p.exists())
p.parent.mkdir(parents=True, exist_ok=True) # create the folders if needed
# Short forms for small files
p.write_text("a,b\n1,2\n", encoding="utf-8")
print(p.read_text(encoding="utf-8")) - Never delete or overwrite a file to "fix" an encoding error. The bytes are fine; only your assumption about them is wrong. Try
utf-8first, then the encoding the source system used, and keep the original untouched until something reads cleanly.
CSV and JSON: Structured Files
CSV is the format your data will arrive in — exported from a spreadsheet, a college portal or a database. It looks trivially easy to parse with split(","), and that works right up until a field contains a comma. Real CSV puts such fields in quotes, and a quoted field can itself contain quotes, and at that point your hand-written parser is wrong in ways you will not notice until the numbers are off. Use the csv module.
csv.reader gives each row as a list of strings; csv.DictReader uses the header row to give each row as a dictionary, so you write row["marks"] instead of row[2] and your code survives a column being moved. Both give you strings — a CSV file has no types, so int(row["marks"]) is your job.
When writing, pass newline="" to open(). Without it, Windows translates the module's line endings a second time and every row is followed by a blank line. It is an odd requirement and it is in the documentation for a reason.
JSON is the other format, and it is what APIs speak. Unlike CSV it keeps types and nesting, so a list of dictionaries survives a round trip intact. json.dump(data, f) writes and json.load(f) reads; the s variants — dumps and loads — work with strings instead of files. Pass indent=2 when a human will read the file, and ensure_ascii=False to keep non-English text readable rather than escaped.
Choosing between them is simple. CSV for flat tables that a spreadsheet must open. JSON for anything nested, or anything where types matter. If you find yourself writing a comma inside a CSV cell or storing a list in one, the data has outgrown the format.
import csv, json
rows = [
{"name": "Asha", "branch": "CSE", "marks": 87},
{"name": "Ravi, Kumar", "branch": "ECE", "marks": 92}, # note the comma
]
# Writing CSV — newline="" is required
with open("students.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["name", "branch", "marks"])
writer.writeheader()
writer.writerows(rows)
# The comma-containing field is quoted automatically:
# name,branch,marks
# Asha,CSE,87
# "Ravi, Kumar",ECE,92
# Reading it back with named columns
with open("students.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
print(row["name"], int(row["marks"]) + 1) # values are strings
# Why split(",") is not enough
line = '"Ravi, Kumar",ECE,92'
print(line.split(",")) # 4 pieces — the name was torn in half
print(next(csv.reader([line]))) # ['Ravi, Kumar', 'ECE', '92']
# JSON keeps types and nesting
with open("students.json", "w", encoding="utf-8") as f:
json.dump(rows, f, indent=2, ensure_ascii=False)
with open("students.json", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded[0]["marks"] + 1) # 88 — still an int, no conversion needed
# The string versions
text = json.dumps({"ok": True})
print(text, json.loads(text)["ok"]) csv.reader— rows as lists;csv.DictReader— rows as dictionariescsv.writer/csv.DictWriter— always open withnewline=""- Every CSV value is a string; convert with
int()orfloat()yourself json.dump/json.loadfor files;dumps/loadsfor stringsindent=2for readability,ensure_ascii=Falsefor non-English text- CSV for flat tables; JSON for nested data and preserved types
- For serious tabular work,
pandas.read_csv()reads a file into a DataFrame in one line and handles types, missing values and large files far better than hand-written loops. Thecsvmodule is still worth knowing — it is always available, and interviewers ask about it.
