Lesson 9 of 25

Lists

Ordered, Mutable, and Happy to Hold Anything

A list is an ordered collection written in square brackets. Three words describe it completely. Ordered means the items stay in the sequence you put them in and each has a position. Mutable means you can change it after it is created — replace an item, add one, remove one — which is the single biggest difference from strings and tuples. And it is heterogeneous: one list can hold numbers, strings, other lists, or a mixture, because Python stores references to objects rather than the objects themselves.

Indexing and slicing work exactly as they do on strings, and for the same reason: both are sequences. Position 0 is the first item, negative numbers count from the end, and a slice takes start up to but not including stop. Reading past the end with a plain index raises IndexError; a slice that runs past the end just gives you what exists.

Because lists are mutable, assignment into a position works: fruits[1] = "blueberry" replaces one item in place. So does assignment into a slice, which can replace a run of items with a different number of items. That is a genuinely useful trick and also a fine way to confuse a reader, so keep it for the obvious cases.

Lists can nest, and a list of lists is how you represent a table or a grid before you reach for a library. grid[1][2] reads as "row 1, column 2". Nesting is also where the copying rules in a later section start to matter, so keep it in mind.

Example
fruits = ["apple", "banana", "cherry"]
numbers = [1, 2, 3, 4, 5]
mixed = [1, "hello", 3.14, True, [7, 8]]
empty = []

print(fruits[0])      # apple
print(fruits[-1])     # cherry
print(fruits[1:3])    # ['banana', 'cherry']
print(fruits[:2])     # ['apple', 'banana']
# print(fruits[9])    # IndexError: list index out of range
print(fruits[5:9])    # []  — slices do not complain

# Mutable: you can change an item in place
fruits[1] = "blueberry"
print(fruits)         # ['apple', 'blueberry', 'cherry']

# A slice can be replaced by a different number of items
numbers[1:3] = [20, 30, 40]
print(numbers)        # [1, 20, 30, 40, 4, 5]

# Nested lists model a table
grid = [[1, 2, 3],
        [4, 5, 6]]
print(grid[1][2])     # 6
print(len(grid))      # 2  rows
print(len(grid[0]))   # 3  columns
  • [] — ordered, mutable, may hold any mix of types
  • items[i] — one item; items[-1] is the last
  • items[a:b] — a new list from a up to but not including b
  • items[i] = x — replace in place; strings do not allow this
  • len(items) — how many items; on a nested list it counts rows
Notes
  • A single-item list still needs to look like a list: [5]. Writing (5) gives you the number 5 in brackets, not a container — that is the tuple trap covered in the next lesson.

Adding and Removing Items

There are three ways to add and four ways to remove, and picking the wrong one is one of the most common early mistakes. append(x) adds one item to the end. extend(other) adds every item of another collection. insert(i, x) puts an item at a given position and shifts the rest along.

The confusion is between append and extend. Calling append with a list puts that whole list inside as a single item, so your list of three names becomes a list of four things, one of which is a list. Calling extend with a string is worse: a string is iterable, so it is unpacked into individual characters. The rule to remember is that append takes one object and extend takes a collection to unpack.

For removal, remove(value) deletes the first item equal to that value and raises ValueError if it is not there — so check with in first, or be ready to handle the exception. pop() removes and returns the last item, or the item at an index you give, which makes it the tool for using a list as a stack. del items[i] deletes by position without returning anything and also works on a slice. clear() empties the list while keeping the same object, which matters when other names point at it.

One practical note: + and += look interchangeable and are not. a = a + [4] builds a new list and points a at it. a += [4] modifies the existing list in place, exactly like extend. If a second name refers to that list, the two spellings give different results — the next section explains why.

Example
names = ["Asha", "Ravi"]

names.append("Meera")             # one item
print(names)                      # ['Asha', 'Ravi', 'Meera']

names.extend(["Karan", "Divya"])  # every item of another collection
print(names)                      # [..., 'Karan', 'Divya']

names.insert(0, "Zoya")           # at the front; everything shifts right
print(names[0])                   # Zoya

# append vs extend — the classic mix-up
a = [1, 2]
a.append([3, 4])
print(a)          # [1, 2, [3, 4]]   one new item, which is a list

b = [1, 2]
b.extend([3, 4])
print(b)          # [1, 2, 3, 4]

c = [1, 2]
c.extend("hi")    # a string is iterable...
print(c)          # [1, 2, 'h', 'i']  ...so it unpacks into characters

# Removing
items = ["a", "b", "c", "b"]
items.remove("b")        # first match only
print(items)             # ['a', 'c', 'b']
# items.remove("z")      # ValueError: list.remove(x): x not in list

last = items.pop()       # removes AND returns
print(last, items)       # b ['a', 'c']

del items[0]
print(items)             # ['c']
items.clear()
print(items)             # []
  • .append(x) — add one object to the end
  • .extend(iterable) — add each item of another collection
  • .insert(i, x) — add at a position; everything after shifts
  • .remove(value) — delete the first match; ValueError if absent
  • .pop() / .pop(i) — remove and return; IndexError on an empty list
  • del items[i], del items[a:b] — delete by position, return nothing
  • .clear() — empty the list but keep the same object
Notes
  • .index(value) finds a position and .count(value) counts occurrences; both scan from the front, and .index() raises ValueError when the value is absent. If you only want to know whether it is there, value in items is clearer.

Sorting: In Place or a New List

Python gives you two ways to sort and the difference matters. items.sort() reorders the list in place and returns None. sorted(items) leaves the original alone and hands back a new sorted list. Use sort() when you own the list and want it ordered from now on; use sorted() when you need an ordering for one purpose and the original order still means something.

The trap follows directly from "returns None". Writing items = items.sort() looks reasonable and silently replaces your list with None, after which the next line fails with a confusing message about NoneType. This is a general Python convention worth internalising now: methods that change an object in place return None, so any line of the form x = x.something() deserves a second look.

Both accept reverse=True for descending order and a key function that says what to sort by. key is the feature that makes sorting genuinely useful: key=len sorts by length, key=str.lower sorts text the way a human expects rather than putting every capital letter first, and a small lambda sorts a list of tuples or dictionaries by whichever field you choose.

Two facts round it off. Python's sort is stable, meaning items that compare equal keep their original relative order — which lets you sort by one field, then by another, and have the first ordering survive as a tie-breaker. And sorting a list of mixed types raises TypeError, because Python refuses to guess whether a string comes before a number.

Example
marks = [78, 45, 91, 65]

print(sorted(marks))          # [45, 65, 78, 91]  — a new list
print(marks)                  # [78, 45, 91, 65]  — untouched

marks.sort()                  # reorders in place
print(marks)                  # [45, 65, 78, 91]
print(marks.sort())           # None  <- it returns nothing

# The trap
values = [3, 1, 2]
values = values.sort()        # values is now None
print(values)                 # None
# print(len(values))          # TypeError: object of type 'NoneType' has no len()

# key= decides what to sort by
names = ["Meera", "asha", "Ravi"]
print(sorted(names))              # ['Meera', 'Ravi', 'asha']  capitals first
print(sorted(names, key=str.lower))  # ['asha', 'Meera', 'Ravi']
print(sorted(names, key=len))     # shortest first

# Sorting records by a chosen field
students = [("Asha", 87), ("Ravi", 92), ("Meera", 87)]
print(sorted(students, key=lambda s: s[1], reverse=True))
# [('Ravi', 92), ('Asha', 87), ('Meera', 87)]  — the 87s keep their order

# Mixed types will not sort
# print(sorted([3, "a"]))     # TypeError: '<' not supported between instances
Notes
  • The same in-place-returns-None rule covers .append(), .extend(), .insert(), .remove(), .reverse() and .clear(). .pop() is the deliberate exception: it returns the item it removed.

Two Names, One List: Aliasing and Copying

This is the most important section in the lesson, and the source of bugs that survive into professional code. In Python a variable is a name attached to an object, not a box holding a value. Writing b = a does not copy the list; it attaches a second name to the same list. Change it through either name and both see the change, because there is only one list.

With numbers and strings this never bites, because they are immutable — you can never change them, only point a name at a different value. With lists it bites constantly. The same rule explains why passing a list into a function lets the function change your data: the parameter is one more name for the same object. That is often exactly what you want, and it must be a decision rather than an accident.

To get an independent list, copy it explicitly. a.copy(), list(a) and a[:] all do the same thing and all produce a shallow copy: a new outer list containing the same inner objects. For a flat list of numbers or strings that is a complete copy and you need nothing more.

For a nested list it is not. The outer list is new, but the inner lists are shared, so appending to copy_of[0] also changes original[0]. When you need every level duplicated, copy.deepcopy() from the standard library walks the whole structure. It costs more time and memory, so reach for it when you actually have nesting rather than out of habit.

The quickest way to tell aliasing from copying while debugging is is. a == b asks whether the contents match; a is b asks whether they are the same object. Two lists that are equal but not identical are the sign that your copy worked.

Example
a = [1, 2, 3]
b = a                # NOT a copy — a second name for one list
b.append(4)
print(a)             # [1, 2, 3, 4]   a changed too
print(a is b)        # True

# A real (shallow) copy — three equivalent spellings
a = [1, 2, 3]
b = a.copy()         # or list(a), or a[:]
b.append(4)
print(a, b)          # [1, 2, 3] [1, 2, 3, 4]
print(a is b)        # False

# The same rule applies to function arguments
def add_bonus(scores):
    scores.append(5)         # changes the caller's list

results = [10, 20]
add_bonus(results)
print(results)       # [10, 20, 5]

# Shallow copies are only one level deep
marks = [[87, 92], [78, 65]]
shallow = marks.copy()
shallow[0].append(100)       # the inner lists are still shared
print(marks)         # [[87, 92, 100], [78, 65]]  — original changed

import copy
deep = copy.deepcopy(marks)
deep[0].append(50)
print(marks)         # [[87, 92, 100], [78, 65]]  — untouched
print(deep)          # [[87, 92, 100, 50], [78, 65]]
Notes
  • If a function should not modify what it is given, copy the argument at the top of the function or return a new list instead. Silently mutating a caller's data is legal, occasionally useful, and a habit that makes larger programs hard to reason about.

Building Lists: Loops, Comprehensions and the Repetition Trap

A large share of list code follows one shape: start with an empty list, loop over something, and append a transformed or filtered version of each item. A list comprehension writes that shape as a single expression: [expression for item in iterable if condition], where the condition is optional. Read it left to right as "give me this, for each of these, where that holds".

Comprehensions are not just shorter. They tell the reader immediately that a new list is being built and nothing else is happening, whereas a loop with append could be doing anything until you have read all of it. They are also modestly faster, since the appending happens inside the interpreter rather than as a method call per item.

They have a readability ceiling, and it arrives sooner than people expect. One for and one if is comfortable. Two for clauses — used for flattening a nested list, where you read them outer loop first, exactly as you would write the nested loops — is about the limit. A conditional expression inside the output, plus a filter, plus nesting, is a line nobody wants to debug. When it stops fitting on one line comfortably, write the loop.

Finally, the repetition trap. [0] * 5 is a fine way to build a list of five zeros, because integers are immutable and sharing them is harmless. [[0] * 3] * 2 is not a grid: it creates one inner list and puts two references to it in the outer list, so setting grid[0][0] changes what looks like a separate row. The fix is a comprehension, which evaluates the inner expression afresh on every pass.

Example
# The loop shape
squares = []
for x in range(1, 6):
    squares.append(x ** 2)

# The same thing as one expression
squares = [x ** 2 for x in range(1, 6)]
print(squares)                 # [1, 4, 9, 16, 25]

# With a filter
marks = [78, 45, 91, 32, 65]
print([m for m in marks if m >= 40])       # [78, 45, 91, 65]

# Transform and filter together
names = ["asha", "ravi", "meera"]
print([n.capitalize() for n in names if len(n) > 4])   # ['Meera']

# Flattening: the for clauses read outer first
matrix = [[1, 2], [3, 4], [5, 6]]
print([n for row in matrix for n in row])  # [1, 2, 3, 4, 5, 6]

# Repetition is safe for immutable items
row = [0] * 5
print(row)                     # [0, 0, 0, 0, 0]

# ...and a trap for mutable ones
grid = [[0] * 3] * 2           # TWO references to ONE inner list
grid[0][0] = 9
print(grid)                    # [[9, 0, 0], [9, 0, 0]]  — both rows changed

grid = [[0] * 3 for _ in range(2)]   # a fresh inner list each pass
grid[0][0] = 9
print(grid)                    # [[9, 0, 0], [0, 0, 0]]
Notes
  • The underscore in for _ in range(2) is an ordinary variable name, used by convention when the loop value is not needed. It carries no special meaning to Python — it is a signal to the reader.

What Lists Cost, and When to Use Something Else

A list is the default container and usually the right one, but it has a shape, and knowing the shape tells you when to reach for something else. Items are stored one after another with their positions known, so reading items[i] and calling len(items) are instant regardless of size, and appending to the end is cheap.

Two operations are not cheap. Searching with in, .index(), .remove() or .count() has to walk the list from the front, so the cost grows with its length. If you are checking membership inside a loop, that turns into a slow program on real data — convert the collection to a set once and check against that instead, since a set lookup does not depend on size. Inserting or deleting near the front is the other one: every item after the change has to shift along. Using pop(0) to consume a queue is a classic slow spot, and collections.deque is the container built for that job.

Choosing a container is worth two seconds of thought before you start. A list is for an ordered collection you will change. A tuple is for a fixed group of values, and being immutable lets it act as a dictionary key. A set is for membership and uniqueness, with no order and no duplicates. A dictionary is for looking a value up by a key. The next three lessons take each of them in turn; the point here is only that reaching for a list every time is a habit, not a decision.

One last practical note about writing lists: a trailing comma after the final item is legal and encouraged in multi-line lists. It keeps the diff clean when someone adds an entry later, and it prevents the mistake of gluing two strings together by leaving out a comma — ["a" "b"] is a one-item list containing "ab", with no error at all.

Example
valid_ids = ["A101", "A102", "A103"]
print("A102" in valid_ids)     # True — but Python scans from the front

valid_ids = {"A101", "A102", "A103"}   # a set
print("A102" in valid_ids)     # True — found directly, whatever the size

# Front operations shift everything after them
queue = [1, 2, 3, 4]
queue.pop(0)                   # every remaining item moves down one place

from collections import deque
queue = deque([1, 2, 3, 4])
queue.popleft()                # cheap at any size
queue.appendleft(0)
print(list(queue))             # [0, 2, 3, 4]

# Trailing commas keep multi-line lists safe to edit
subjects = [
    "Maths",
    "Physics",
    "Chemistry",
]

# A missing comma is not an error — it is silent string joining
oops = ["Maths" "Physics"]
print(oops)                    # ['MathsPhysics']  — one item, not two
  • Instant: items[i], len(items), .append(), .pop() from the end
  • Grows with length: in, .index(), .remove(), .count()
  • Grows with length: .insert(0, x) and .pop(0) — use deque for queues
  • Repeated membership tests: convert to a set first
  • Fixed group of values: a tuple; unique items: a set; lookup by key: a dictionary
Notes
  • Lists are not arrays in the numerical sense — they hold references to objects, so arithmetic on a whole list does not work: [1, 2] * 2 repeats the list rather than doubling its numbers. For element-wise maths, use NumPy.
Ask AI