One Shape, Four Results
A comprehension is a compact way to say "build a collection from another collection". The shape is always the same — an output expression, a for clause, and an optional if — and only the brackets change what you get back.
Square brackets give a list. Curly braces with a colon give a dictionary. Curly braces without a colon give a set, so duplicates in the result collapse. Round brackets do something different in kind: they give a generator, which produces values one at a time rather than building anything, and the second half of this lesson is about that.
There is no tuple comprehension. (x for x in items) is the generator form, so if you want a tuple you wrap the generator: tuple(x for x in items). That trips people up exactly once.
Why prefer a comprehension to a loop at all? Not brevity for its own sake. A comprehension announces its purpose in its first two characters: the reader knows a new collection is being built and that nothing else is happening. A for loop could be doing anything until you have read all of it. Comprehensions are also somewhat faster, because the appending happens inside the interpreter rather than as a method call for every item — but readability is the reason that matters.
The corresponding rule is that a comprehension should only build a collection. If the body needs to print, log, write to a file or update something else, write a loop. A comprehension whose result you throw away is a loop wearing the wrong clothes.
marks = [78, 45, 91, 32, 65]
# List: square brackets
print([m + 5 for m in marks]) # [83, 50, 96, 37, 70]
# Dictionary: braces with a colon
names = ["Asha", "Ravi", "Meera"]
print({n: len(n) for n in names}) # {'Asha': 4, 'Ravi': 4, 'Meera': 5}
# Set: braces without a colon — duplicates collapse
print({m % 10 for m in marks}) # {8, 1, 2, 5}
# Generator: round brackets — nothing is built yet
gen = (m + 5 for m in marks)
print(gen) # <generator object <genexpr> at 0x...>
print(list(gen)) # [83, 50, 96, 37, 70]
# There is no tuple comprehension — wrap the generator
print(tuple(m for m in marks if m >= 40)) # (78, 45, 91, 65)
# Building a lookup table from two lists
branches = ["CSE", "ECE", "CSE"]
print({n: b for n, b in zip(names, branches)})
# {'Asha': 'CSE', 'Ravi': 'ECE', 'Meera': 'CSE'}
# Inverting a dictionary (duplicate values collapse)
scores = {"Maths": 92, "Physics": 88}
print({v: k for k, v in scores.items()}) # {92: 'Maths', 88: 'Physics'} [expr for item in items]— a list{key: value for item in items}— a dictionary{expr for item in items}— a set; duplicates disappear(expr for item in items)— a generator, not a tupletuple(expr for item in items)— how to actually get a tuple- Use a comprehension to build a collection; use a loop for anything else
- The loop variable in a comprehension is local to it and does not leak into the surrounding code — this differs from Python 2, where it did. Naming it
iorncannot clobber a variable of the same name outside.
Filtering and Transforming: Where the if Goes
Two different things in a comprehension both use the word if, they sit in different places, and they do different jobs. Getting them straight removes most of the confusion beginners have with the syntax.
An if at the end, after the for, is a filter. It decides whether each item is included at all, so the result may be shorter than the input. [m for m in marks if m >= 40] keeps the passes and drops the rest.
An if ... else at the front, before the for, is a conditional expression — the ternary from Lesson 7 — and it is part of the output. Every item is included; the condition only decides what value is produced. ["pass" if m >= 40 else "fail" for m in marks] gives one label per student, so the result is always the same length as the input.
The clue that tells them apart is the else. A filter cannot have one, because there is nothing to put in the output when an item is excluded. A conditional expression must have one, because it has to produce something either way.
You can use both at once, and this is where readability starts to suffer: ["A" if m >= 90 else "B" for m in marks if m >= 40] filters to the passes and then labels them. It is correct and it takes a moment to parse. Two clauses is about the ceiling; beyond that, use a loop or split it into two steps.
marks = [78, 45, 91, 32, 65]
# Filter: if at the END. Result may be shorter.
print([m for m in marks if m >= 40]) # [78, 45, 91, 65]
print(len(marks), len([m for m in marks if m >= 40])) # 5 4
# Transform: if/else at the FRONT. Result is always the same length.
print(["pass" if m >= 40 else "fail" for m in marks])
# ['pass', 'pass', 'pass', 'fail', 'pass']
# The else is the giveaway
# [m for m in marks if m >= 40 else 0] # SyntaxError — a filter has no else
# ["pass" if m >= 40 for m in marks] # SyntaxError — an expression needs one
# Both together: filter, then label
print(["A" if m >= 90 else "B" for m in marks if m >= 40])
# ['B', 'B', 'A', 'B']
# Two filters chain as 'and'
print([m for m in marks if m >= 40 if m < 90]) # [78, 45, 65]
print([m for m in marks if 40 <= m < 90]) # same, and clearer
# Beyond two clauses, write the loop
labels = []
for m in marks:
if m < 40:
continue
if m >= 90:
labels.append("distinction")
elif m >= 60:
labels.append("first class")
else:
labels.append("pass")
print(labels) - Calling a function twice in one comprehension — once in the filter and once in the output — runs it twice per item. From Python 3.8 the walrus operator avoids that:
[y for x in data if (y := f(x)) > 0]computesf(x)once.
Nesting: Grids and Flattening
There are two ways a comprehension can nest, they look similar, and they do opposite things.
A comprehension inside the output expression builds a nested structure. [[i * j for j in range(1, 4)] for i in range(1, 4)] produces a list of lists, because the inner comprehension runs once for each i and its result becomes one element of the outer list. This is the correct way to build a grid, and it fixes the [[0] * 3] * 2 trap from Lesson 9 — the inner expression is evaluated afresh each time, so no row is shared.
Two for clauses in the same comprehension flatten instead. [n for row in matrix for n in row] produces one flat list. The order to remember is that the clauses read left to right in exactly the order you would write the nested loops: outer loop first, inner loop second. Getting them the wrong way round gives a NameError, because the second clause is using a name the first has not defined yet.
Each for clause can have its own filter, and later clauses can use the values from earlier ones. That is genuinely useful for pulling values out of nested data — every mark of every student, only the passing ones — and it is also where a comprehension stops being readable. Two for clauses is the practical limit.
When you do exceed it, the fix is rarely to persevere. Either extract a helper function and call it from a simple comprehension, or write the nested loops out. Nobody has ever been thanked for a four-clause comprehension.
# Nested OUTPUT: builds a grid, one fresh row per pass
table = [[i * j for j in range(1, 5)] for i in range(1, 4)]
for row in table:
print(row)
# [1, 2, 3, 4]
# [2, 4, 6, 8]
# [3, 6, 9, 12]
# And no shared rows, unlike [[0] * 3] * 2
grid = [[0] * 3 for _ in range(2)]
grid[0][0] = 9
print(grid) # [[9, 0, 0], [0, 0, 0]]
# Two for CLAUSES: flattens. Read them outer-first, like nested loops.
matrix = [[1, 2], [3, 4], [5, 6]]
print([n for row in matrix for n in row]) # [1, 2, 3, 4, 5, 6]
# The equivalent loops, in the same order
flat = []
for row in matrix:
for n in row:
flat.append(n)
print(flat)
# Wrong order fails
# [n for n in row for row in matrix] # NameError: name 'row' is not defined
# Later clauses can use earlier values, and each can filter
students = [
{"name": "Asha", "marks": [78, 91]},
{"name": "Ravi", "marks": [32, 65]},
]
print([(s["name"], m) for s in students for m in s["marks"] if m >= 40])
# [('Asha', 78), ('Asha', 91), ('Ravi', 65)]
# Past two clauses, use a helper or a loop
def passing(student):
return [(student["name"], m) for m in student["marks"] if m >= 40]
print([pair for s in students for pair in passing(s)]) [[...] for ...]— a comprehension in the output builds a nested structure[x for a in outer for x in a]— two clauses flatten- Clauses read left to right in the order the nested loops would be written
[[0] * 3 for _ in range(2)]— the safe way to build a grid- Each
forclause may have its ownif; later clauses see earlier values - Two clauses is the readability limit — extract a function beyond that
- For flattening one level,
itertools.chain.from_iterable(matrix)says the same thing without nesting and returns a lazy iterator, which matters when the data is large.
Generators: Values on Demand
A comprehension builds the whole result before you use any of it. For a million rows that means a million objects in memory, even if you only wanted their sum. A generator produces values one at a time, on request, and remembers nothing it has already handed over.
You write one by using yield instead of return. That single keyword changes what the function is: calling it does not run the body at all, it creates a generator object. The body runs only when values are requested, and yield pauses the function — local variables, the loop position, everything — until the next value is asked for.
That pausing is what makes generators more than a memory trick. State is kept between values without you managing it, so a Fibonacci sequence is four lines, and an infinite sequence is perfectly reasonable because nothing is ever built. You take what you need and stop.
next(gen) asks for one value; a for loop asks repeatedly until the generator is finished, at which point it raises StopIteration internally and the loop ends. You will rarely write next() yourself, but knowing it is there explains what a for loop has been doing all along.
The property to keep in mind is that a generator is single-use. Once consumed, it is empty — a second loop over the same generator produces nothing at all, with no error to tell you. If you need the values twice, either build a list or call the generator function again to get a fresh one.
def count_up(limit):
n = 1
while n <= limit:
yield n # produce a value and PAUSE here
n += 1
# Calling it runs nothing yet
g = count_up(3)
print(g) # <generator object count_up at 0x...>
print(next(g)) # 1 — runs up to the first yield
print(next(g)) # 2 — resumes where it paused
print(next(g)) # 3
# print(next(g)) # StopIteration — nothing left
# A for loop does all of that for you
for n in count_up(3):
print(n, end=" ") # 1 2 3
print()
# State is kept for you, so infinite sequences are fine
def fibonacci():
a, b = 0, 1
while True: # never terminates — and never needs to
yield a
a, b = b, a + b
fib = fibonacci()
print([next(fib) for _ in range(10)])
# [0, 1, 1, 2, 3, 5, 8, 13, 21, 34]
from itertools import islice
print(list(islice(fibonacci(), 5))) # [0, 1, 1, 2, 3] — take the first n
# Single use: the second pass finds nothing
g = count_up(3)
print(list(g)) # [1, 2, 3]
print(list(g)) # [] — exhausted, and no warning
# yield from delegates to another iterable
def both():
yield from count_up(2)
yield from ["a", "b"]
print(list(both())) # [1, 2, 'a', 'b'] yield— produce a value and pause; the function becomes a generator- Calling a generator function runs nothing; it returns a generator object
next(gen)— resume and get one value;StopIterationwhen finished- A
forloop callsnext()for you and stops onStopIteration - Single-use: once consumed, a generator is empty for good
yield from— hand over to another iterable
- A
returninside a generator ends it early rather than producing a value. Any value returned is attached to theStopIterationand is invisible to an ordinaryforloop, so treatreturnin a generator as "stop here".
Generator Expressions and Pipelines
A generator expression is a comprehension written with round brackets. It is the lazy version of the same idea: no list is built, and values appear as they are asked for. The memory difference is not incremental — a list comprehension over a million items holds a million objects, while the generator expression is a few hundred bytes whatever the size of the source.
When a generator expression is the only argument to a function, the brackets can be dropped: sum(m for m in marks) rather than sum((m for m in marks)). This is why sum, max, any and all so often appear with what looks like a comprehension missing its brackets, and it is the most common way you will meet generator expressions in real code.
The real power is pipelines. Because each stage is lazy, you can chain them — read lines, strip them, drop the blanks, parse the numbers — and nothing is stored at any stage. One value travels the whole chain, then the next. A pipeline over a five-gigabyte log file runs in constant memory, which is not something a list-based version can do at all.
Laziness also means work is skipped. any(is_valid(x) for x in rows) stops at the first row that matches; the rest of the file is never read. With a list comprehension, every row is processed first and then the answer is computed.
The trade-offs are real and worth naming. You cannot take len() of a generator, you cannot index into it, and you cannot loop over it twice. If any of those is what you need, build the list. The choice is not "generators are better" — it is "lazy when the data is large or the work might be skipped, eager when you need the collection itself".
import sys
big_list = [x * x for x in range(100_000)]
big_gen = (x * x for x in range(100_000))
print(sys.getsizeof(big_list)) # hundreds of thousands of bytes
print(sys.getsizeof(big_gen)) # a couple of hundred, whatever the size
# Sole argument: drop the extra brackets
marks = [78, 45, 91, 32, 65]
print(sum(m for m in marks if m >= 40)) # 279
print(max(len(w) for w in ["a", "abc"])) # 3
print(any(m >= 90 for m in marks)) # True
print(all(m >= 40 for m in marks)) # False
# A pipeline: nothing is stored at any stage
def read_lines(path):
with open(path, encoding="utf-8") as f:
for line in f:
yield line.rstrip("\n")
lines = read_lines("marks.txt")
stripped = (l.strip() for l in lines)
useful = (l for l in stripped if l and not l.startswith("#"))
numbers = (int(l) for l in useful)
print(sum(numbers)) # one value flows through at a time
# Laziness skips work: this stops at the first match
rows = (n for n in range(1, 10_000_000))
print(any(n > 5 for n in rows)) # True — nowhere near ten million checks
# What you give up
g = (x for x in range(5))
# print(len(g)) # TypeError: object of type 'generator' has no len()
# print(g[0]) # TypeError: 'generator' object is not subscriptable
print(len(list(g))) # build the list when you need those things sys.getsizeof()reports the size of the container object itself, not of everything inside it. It is still the right tool here, because the whole point is that the generator has nothing inside it.
Choosing Between Them, and the Traps
The decision is usually easy once you ask two questions. Do you need the collection itself — its length, its items by position, more than one pass over it? Then build a list. Are you going to consume the values once, and might the data be large or the work skippable? Then use a generator.
The first trap is silent exhaustion. A generator that has already been consumed returns nothing, and nothing warns you. The symptom is a total of zero or an empty report from code that clearly worked a moment ago. If a value is used twice, materialise it once with list() and use that.
The second is that laziness moves when things happen. A generator's body does not run until values are pulled, so exceptions surface at the point of consumption, not at the point of creation — which can be in a completely different function. Similarly, a generator that opens a file keeps it open until it is exhausted or discarded, which is why a with block inside the generator function is the right place for it.
The third is debugging. You cannot print a generator to see what is in it, because printing shows only <generator object ...>, and wrapping it in list() to look consumes it. During development, keep a list; convert to a generator once the logic is settled and the size demands it.
The last trap is over-enthusiasm. A generator for ten items saves nothing and costs clarity. Comprehensions and generators are tools for specific problems — building a collection, and streaming one — not a style to apply everywhere. When either makes a line harder to read than the loop it replaced, the loop was the right answer.
def marks_over(limit, marks):
for m in marks:
if m >= limit:
yield m
marks = [78, 45, 91, 32, 65]
# Trap 1: used twice, empty the second time
passed = marks_over(40, marks)
print(sum(passed)) # 279
print(len(list(passed))) # 0 — already exhausted
passed = list(marks_over(40, marks)) # materialise once
print(sum(passed), len(passed)) # 279 4
# Trap 2: the body does not run until values are pulled
def risky():
print("starting")
raise ValueError("boom")
yield 1
g = risky() # nothing printed, nothing raised
print("created")
# next(g) # only NOW does it print and raise
# Which is why the file handling belongs inside the generator
def read_lines(path):
with open(path, encoding="utf-8") as f: # closed when exhausted
for line in f:
yield line.rstrip("\n")
# Trap 3: you cannot inspect it without consuming it
g = (x for x in range(3))
print(g) # <generator object <genexpr> at 0x...>
print(list(g)) # [0, 1, 2] — and now g is empty
# Right tool, right size
names = ["Asha", "Ravi"]
print([n.upper() for n in names]) # small and needed as a list
print(sum(len(n) for n in names)) # consumed once — generator is fine - Need
len(), indexing or a second pass — build a list - Consumed once, data large or work skippable — use a generator
- A consumed generator is silently empty; materialise with
list()if reused - Laziness delays exceptions and side effects to the point of consumption
- Keep the
withblock inside the generator so the file closes correctly - For small collections, prefer whichever version reads better
itertoolsis the standard toolbox for generators:isliceto take the first n,chainto join several,teeto duplicate one, andgroupbyto batch consecutive items. All of them are lazy, so they compose into pipelines without giving up the memory benefit.
