Lesson 20 of 25

OOP: Encapsulation & Abstraction

Encapsulation Is About Change, Not Secrecy

Encapsulation means separating what a class promises from how it delivers it. The promise — the method names, what they take and what they return — is the public interface. Everything else is implementation, and you should be able to change it without breaking anyone.

The point is not security. Nothing here stops a determined programmer, and it is not meant to. The point is that if a class exposes twenty attributes, all twenty become things other code depends on, and every one is now frozen. Expose three and you can rebuild the other seventeen next month without touching a line elsewhere.

Python enforces almost none of this, and that is a deliberate design decision. The community's phrase for it is "we are all consenting adults": the language marks intent with naming conventions and trusts you to respect them. Coming from Java or C++, the absence of private feels like something missing. In practice, the convention does the same job for the same reason — it tells the reader which parts are safe to rely on.

The practical rule is to make the interface as small as the class allows. Ask what a caller needs to do, expose exactly that, and mark everything else as internal. A BankAccount needs deposit, withdraw and a way to read the balance. The list of transactions, the overdraft rule and the internal balance field are its own business.

Example
class BankAccount:
    def __init__(self, owner, balance=0):
        self.owner = owner          # public: callers legitimately need it
        self._balance = balance     # internal: the leading _ says so
        self._history = []          # internal

    # ---- the public interface ----
    def deposit(self, amount):
        if amount <= 0:
            raise ValueError("deposit must be positive")
        self._apply("deposit", amount)
        return self._balance

    def withdraw(self, amount):
        if amount > self._balance:
            raise ValueError(f"balance is only {self._balance}")
        self._apply("withdraw", -amount)
        return self._balance

    def statement(self):
        return [f"{kind}: {amt}" for kind, amt in self._history]

    # ---- internal helper: free to change ----
    def _apply(self, kind, delta):
        self._balance += delta
        self._history.append((kind, delta))


acc = BankAccount("Asha", 1000)
acc.deposit(500)
acc.withdraw(200)
print(acc.statement())      # ['deposit: 500', 'withdraw: -200']

# Nothing stops this — the underscore is a message, not a wall
print(acc._balance)         # 1300
acc._balance = 999999       # legal, and clearly the caller's own fault
Notes
  • A good test of an interface: could you rewrite the whole class to store its data in a database instead of in memory, without changing a single line of calling code? If not, something internal has leaked into the public surface.

_single, __double, and Name Mangling

One leading underscore — _balance — is pure convention. It means "internal; do not rely on this". Python does nothing at all about it, apart from excluding such names from from module import *. Everyone in the Python world reads it the same way, which is what makes it work.

Two leading underscores do something real. Inside a class body, self.__balance is rewritten by the interpreter to self._BankAccount__balance. This is name mangling, and the effect is that a subclass defining its own __balance gets a completely separate attribute rather than silently overwriting the parent's.

Notice what that is for: it prevents accidental collisions in a hierarchy, not access from outside. The mangled name is perfectly reachable — acc._BankAccount__balance works fine. Treating the double underscore as "private" in the Java sense will mislead you and, worse, will make subclassing and debugging your class unnecessarily awkward.

So the guidance is simple. Use one underscore for anything internal; that covers ninety-nine per cent of cases. Use two only when you have a real reason to fear a name clash in a class designed for inheritance — a base class in a library, for instance. Do not use two out of a general wish for privacy.

One more convention completes the set: names with two underscores at both ends, such as __init__ and __str__, are reserved for Python's own protocol. They are not mangled and you should not invent new ones.

Example
class Account:
    def __init__(self, balance):
        self.owner = "Asha"        # public
        self._internal = "note"    # convention: internal
        self.__secret = balance    # mangled to _Account__secret

    def show(self):
        return self.__secret       # inside the class, the short name works


a = Account(1000)
print(a.owner, a._internal)     # both readable
print(a.show())                 # 1000

# print(a.__secret)             # AttributeError: 'Account' object has no
                                # attribute '__secret'
print(a._Account__secret)       # 1000 — not private, just renamed
print([n for n in vars(a) if "secret" in n])   # ['_Account__secret']


# What mangling is actually for: avoiding collisions in a hierarchy
class Base:
    def __init__(self):
        self.__value = "base"      # _Base__value
    def base_value(self):
        return self.__value

class Child(Base):
    def __init__(self):
        super().__init__()
        self.__value = "child"     # _Child__value — a DIFFERENT attribute
    def child_value(self):
        return self.__value

c = Child()
print(c.base_value(), c.child_value())   # base child — neither overwrote
print(sorted(vars(c)))                   # ['_Base__value', '_Child__value']
  • name — public; part of the promise you are making
  • _name — internal by convention; nothing is enforced
  • __name — mangled to _Class__name; prevents subclass collisions
  • __name__ — reserved for Python's own protocol; do not invent your own
  • Prefer one underscore; reach for two only in classes designed to be subclassed
Notes
  • Name mangling happens at compile time, based on the class the code is written in — not on the object's runtime type. That is exactly why a subclass's __value and a parent's __value can coexist without either knowing about the other.

@property: Attributes with Logic Behind Them

In Java you write getters and setters for every field from the start, because adding one later changes how callers write the code. Python does not have that problem, and @property is the reason.

Start with a plain attribute. If you later need validation, a computed value or a log line whenever it is read, convert it into a property. @property turns a method into something accessed like an attribute, so account.balance still works exactly as before while now running your code. No caller changes. This is why writing getters "just in case" is un-Pythonic — you can always add one when it is actually needed.

A getter alone gives a read-only attribute: assigning to it raises AttributeError. Adding a matching @balance.setter lets assignment work and gives you somewhere to validate, so account.balance = -100 can raise ValueError instead of quietly recording an impossible balance. The internal value conventionally lives in an attribute of the same name with a leading underscore.

The other main use is a computed or derived value — an area from width and height, a full name from two parts, an average from a list of marks. Expressing it as a property rather than a method says it is a fact about the object rather than an action, and it cannot drift out of date the way a stored copy can.

The rule of thumb for choosing between a property and a method: use a property when the caller would think of it as data and getting it is cheap. Use a method when it does real work, has side effects, or takes arguments. Nobody expects obj.total to open a network connection.

Example
class Account:
    def __init__(self, balance=0):
        self._balance = balance     # the storage behind the property

    @property
    def balance(self):              # read: acc.balance
        return self._balance

    @balance.setter
    def balance(self, value):       # write: acc.balance = 500
        if value < 0:
            raise ValueError("balance cannot be negative")
        self._balance = value


acc = Account(1000)
print(acc.balance)        # 1000 — looks like an attribute
acc.balance = 1500        # runs the setter, with validation
print(acc.balance)        # 1500

try:
    acc.balance = -50
except ValueError as e:
    print(e)              # balance cannot be negative


# Read-only: define only the getter
class Circle:
    def __init__(self, radius):
        self.radius = radius        # a plain attribute is fine

    @property
    def area(self):                 # computed, always current
        return 3.14159 * self.radius ** 2

    @property
    def diameter(self):
        return self.radius * 2


c = Circle(5)
print(round(c.area, 2))     # 78.54
c.radius = 10
print(round(c.area, 2))     # 314.16 — no stale copy to update
# c.area = 50               # AttributeError: property 'area' has no setter


# Start simple; convert only when a rule appears
class Student:
    def __init__(self, first, last):
        self.first = first          # plain attributes, no ceremony
        self.last = last

    @property
    def full_name(self):
        return f"{self.first} {self.last}"

print(Student("Asha", "Rao").full_name)   # Asha Rao
Notes
  • A property must be defined on the class, not assigned to an instance. Attaching one to self inside __init__ does not work — Python looks for properties on the type.

Property Traps

The first trap catches nearly everyone once. Inside the setter, writing self.balance = value calls the setter again, which calls it again, until Python gives up with RecursionError: maximum recursion depth exceeded. The setter must assign to the storage attribute — self._balance — not to the property's own name. The same applies to the getter: return self.balance inside the getter is infinite too.

The second is the decorator name. The setter is written @balance.setter, using the name of the property, and it must come after the getter, and the method underneath must have that same name. Writing @property.setter, or giving the setter a different method name, produces errors that do not obviously point at the cause.

The third is a design trap rather than a syntax one. Property access looks free, so putting slow work behind one is misleading: report.total that runs a database query surprises everyone who reads the calling code, and doubly so inside a loop. If it is expensive, make it a method so the parentheses warn the reader — or cache the result, which functools.cached_property does for you in one line.

The fourth is subtle and worth knowing: a property that validates on assignment cannot protect a mutable value it hands out. If marks returns the internal list, a caller can append to it and skip every check. Return a copy, or expose methods for the operations you want to allow.

Finally, do not convert every attribute into a property out of habit. A property with a getter that only returns the value and a setter that only stores it is four lines that do exactly what a plain attribute does. Add the property when there is a rule to enforce or a value to compute.

Example
# 1. The recursion trap
class Broken:
    def __init__(self, v):
        self.value = v

    @property
    def value(self):
        return self.value          # calls itself forever

    @value.setter
    def value(self, v):
        self.value = v             # calls itself forever

# Broken(5)   # RecursionError: maximum recursion depth exceeded

class Fixed:
    def __init__(self, v):
        self.value = v             # goes through the setter — that is fine

    @property
    def value(self):
        return self._value         # the STORAGE attribute

    @value.setter
    def value(self, v):
        if v < 0:
            raise ValueError("must not be negative")
        self._value = v

print(Fixed(5).value)              # 5


# 2. Expensive work hidden behind attribute syntax
from functools import cached_property

class Report:
    def __init__(self, rows):
        self.rows = rows

    @cached_property
    def total(self):               # computed once, then remembered
        print("(computing...)")
        return sum(self.rows)

r = Report([1, 2, 3])
print(r.total)     # (computing...) then 6
print(r.total)     # 6 — no recomputation


# 3. A property cannot protect a mutable value it hands out
class Student:
    def __init__(self):
        self._marks = []

    @property
    def marks(self):
        return list(self._marks)   # a copy: callers cannot edit the original

    def add_mark(self, m):
        if not 0 <= m <= 100:
            raise ValueError(f"out of range: {m}")
        self._marks.append(m)

s = Student()
s.add_mark(87)
s.marks.append(999)                # edits the copy, not the student
print(s.marks)                     # [87]
  • Setter must assign to the storage attribute (self._x), never to itself
  • The decorator is @x.setter, defined after the getter, on a method of the same name
  • Keep properties cheap; use a method when the work is real
  • functools.cached_property computes once and remembers
  • Return a copy of any mutable value a property exposes
  • Do not add a property that only gets and sets — a plain attribute already does that
Notes
  • @property also supports a deleter, written @x.deleter, which runs on del obj.x. It is rarely needed; most classes never define one.

Abstraction and Abstract Base Classes

Abstraction is stating what something does without saying how. An abstract base class names the methods every implementation must provide and leaves them empty, so the base defines a contract while the subclasses supply the behaviour.

The abc module makes that contract enforceable. Inherit from ABC, decorate the required methods with @abstractmethod, and Python refuses to create an instance of any class that has not implemented all of them: TypeError: Can't instantiate abstract class ... with abstract method ....

Compare that with raising NotImplementedError in the base method, which you saw in the previous lesson. Both work; they differ in when they fail. NotImplementedError fires when the missing method is finally called, possibly deep inside a long run. An abstract base class fails at the moment the object is created — before any damage is done — and it lists every method that is missing. For anything that is genuinely a contract, that is the better trade.

An abstract class is not limited to empty methods. It can have a normal __init__, ordinary concrete methods that all subclasses inherit, and abstract properties (write @property above @abstractmethod). That mixture is the usual shape: the base supplies the parts that are the same everywhere and demands the parts that differ.

Python also offers a duck-typed alternative. typing.Protocol describes a set of methods that a class must have, without requiring it to inherit from anything at all — so third-party classes can satisfy your interface without knowing you exist. Use an abstract base class when you also want to share code with the subclasses, and a Protocol when you only want to state the shape.

Example
from abc import ABC, abstractmethod

class Storage(ABC):
    """Every storage backend must provide load() and save()."""

    def __init__(self, name):
        self.name = name            # abstract classes can have __init__

    @abstractmethod
    def load(self):
        ...

    @abstractmethod
    def save(self, data):
        ...

    @property
    @abstractmethod
    def location(self):             # an abstract property
        ...

    def describe(self):             # concrete: inherited by everyone
        return f"{self.name} at {self.location}"


class MemoryStorage(Storage):
    def __init__(self):
        super().__init__("memory")
        self._data = []

    def load(self):
        return list(self._data)

    def save(self, data):
        self._data = list(data)

    @property
    def location(self):
        return "RAM"


class Incomplete(Storage):
    def load(self):
        return []


m = MemoryStorage()
m.save([1, 2, 3])
print(m.load(), m.describe())    # [1, 2, 3] memory at RAM

# Storage("x")      # TypeError: Can't instantiate abstract class Storage
# Incomplete("y")   # TypeError: ... with abstract methods location, save


# The duck-typed alternative: no inheritance required
from typing import Protocol, runtime_checkable

@runtime_checkable
class Readable(Protocol):
    def load(self): ...

def show(source: Readable):
    print(source.load())

class Unrelated:                 # inherits from nothing
    def load(self): return "data"

show(Unrelated())                # data
print(isinstance(Unrelated(), Readable))   # True — it has the right shape
Notes
  • An abstract base class checks that the method names exist; it does not check their parameters. A subclass whose save(self) takes no data still counts as implemented and will fail later at the call. A type checker catches that; Python at runtime does not.

Designing a Surface People Can Use

Everything in this lesson serves one goal: a class that is easy to use correctly and hard to use wrongly. A few habits get you most of the way there.

Start with the smallest public interface that does the job. Every public name is a promise, and promises are easier to add than to withdraw. If you are unsure whether something should be public, make it internal — nobody will complain later that you exposed too little.

Use plain attributes until a rule appears. Python's flexibility means you lose nothing by starting simple: the day validation is needed, @property adds it without changing a single caller. Writing accessors in advance is work with no payoff.

Validate at the boundary, and refuse invalid values rather than storing them. A class that can never hold impossible data is far easier to reason about than one where every method has to re-check. This is the same idea as validating user input once, at the point where it enters the program.

Say what the class is for. A short docstring on the class and on each public method — one sentence on what it returns and what it raises — is worth more than a page of comments inside. And keep the class small enough that the docstring is honest: if you cannot describe it without "and", it is doing two jobs.

Finally, remember the Python attitude behind all of this. The conventions are not enforcement, and they are not weakness — they are a shared language for saying "this part is stable, that part is mine". Code that uses them consistently is readable by every Python programmer who comes after you.

Example
class TaskList:
    """An ordered list of tasks, each with a title and a done flag."""

    def __init__(self, owner):
        self.owner = owner        # public: a caller legitimately needs it
        self._tasks = []          # internal: shape may change

    @property
    def pending(self):
        """How many tasks are not yet done."""
        return sum(1 for t in self._tasks if not t["done"])

    def add(self, title):
        """Add a task. Raises ValueError if the title is blank."""
        title = title.strip()
        if not title:
            raise ValueError("title must not be empty")   # refuse, do not store
        self._tasks.append({"title": title, "done": False})

    def complete(self, title):
        """Mark a task done. Raises KeyError if there is no such task."""
        for t in self._tasks:
            if t["title"] == title:
                t["done"] = True
                return
        raise KeyError(title)

    def titles(self):
        """Return the task titles, newest last."""
        return [t["title"] for t in self._tasks]     # a new list, not ours

    def __len__(self):
        return len(self._tasks)

    def __repr__(self):
        return f"TaskList({self.owner!r}, {len(self._tasks)} tasks)"


tl = TaskList("Asha")
tl.add("revise DSA")
tl.add("submit lab record")
tl.complete("revise DSA")

print(tl)               # TaskList('Asha', 2 tasks)
print(len(tl), tl.pending)   # 2 1
print(tl.titles())

tl.titles().append("sneaky")     # edits the returned copy only
print(len(tl))                   # 2
  • Expose the smallest interface that does the job; everything else gets a leading underscore
  • Plain attributes first; add @property when a rule or a computation appears
  • Validate on the way in and refuse bad values rather than storing them
  • Return copies of internal collections so callers cannot bypass your rules
  • One-sentence docstrings on the class and every public method
  • If the description needs "and", the class is doing two jobs
Notes
  • Implementing __len__ and __repr__ costs four lines and makes your class behave like the built-ins in loops, logs and the debugger. It is the cheapest usability improvement available to a Python class.
Ask AI