Lesson 18 of 25

Object-Oriented Programming: Classes & Objects

What a Class Is Actually For

A class bundles data and the behaviour that belongs to that data into one named thing. Without classes you keep them apart: a dictionary holding a student's details, and separate functions that each take that dictionary as their first argument. That works, and at some size it stops being pleasant, because nothing links the functions to the data they were written for and nothing stops someone passing the wrong dictionary.

A class is the blueprint; an object or instance is one thing built from it. Student is the class, and Asha's record is an instance. Every instance has its own copy of the data and shares the same set of methods, so you write the behaviour once and it applies to every student you create.

The honest test for whether you need one is whether the data and the operations travel together. If you find yourself passing the same dictionary into six functions, or writing if record["type"] == ... in several places, a class will simplify things. If you have one function that transforms some input into some output, it should stay a function — wrapping it in a class with no state adds ceremony and nothing else.

Two conventions before the syntax. Class names use PascalCase — Student, BankAccount, TaskManager — which is how you tell a class from a function at a glance. And a class should represent one thing you could describe in a sentence; a class called Manager that does storage, validation and printing is three classes waiting to be separated.

Example
# Without a class: data here, behaviour somewhere else
student = {"name": "Asha", "marks": [78, 84, 91]}

def average(record):
    return sum(record["marks"]) / len(record["marks"])

def has_passed(record):
    return average(record) >= 40

print(average(student))          # nothing links these functions to the data


# With a class: one thing, with its own behaviour
class Student:
    def __init__(self, name, marks):
        self.name = name
        self.marks = marks

    def average(self):
        return sum(self.marks) / len(self.marks)

    def has_passed(self):
        return self.average() >= 40


asha = Student("Asha", [78, 84, 91])
ravi = Student("Ravi", [30, 35, 28])

print(asha.name, round(asha.average(), 2))   # Asha 84.33
print(ravi.has_passed())                     # False

print(type(asha))                # <class '__main__.Student'>
print(isinstance(asha, Student)) # True
Notes
  • You have been using objects since Lesson 1. "text".upper() is a method on a str object and [1, 2].append(3) is a method on a list object. Writing your own class is doing the same thing for your own data.

__init__ and the self You Have to Type

__init__ runs automatically when you create an instance, and its job is to set up that instance's data. The double underscores mark it as a method Python calls for you rather than one you call by name — you write Student("Asha", [78]) and Python calls __init__ behind the scenes. It is not the object's creator in the strict sense, but treating it as "the constructor" is close enough for now.

Every method's first parameter is self, and it is the instance the method was called on. This is not a keyword — you could name it anything and Python would not object — but every Python programmer expects self, so using anything else makes your code strange to read.

What is really happening is simple once you see it: asha.average() is shorthand for Student.average(asha). The object before the dot is passed in as the first argument. That is why self must appear in the definition and must not appear at the call site, and it is why forgetting it gives the confusing message takes 0 positional arguments but 1 was given — Python passed the instance and your method had nowhere to put it.

Anything you attach with self.something = ... becomes an instance attribute, stored on that object and independent of every other instance. Reading an attribute that was never set raises AttributeError, so define everything an object needs inside __init__, even the fields that start as None or an empty list. It makes the class self-documenting — one place shows every piece of data an instance carries.

Example
class BankAccount:
    def __init__(self, holder, balance=0):
        self.holder = holder        # instance attributes
        self.balance = balance
        self.history = []           # define it even though it starts empty

    def deposit(self, amount):
        if amount <= 0:
            raise ValueError("deposit must be positive")
        self.balance += amount
        self.history.append(("deposit", amount))
        return self.balance

    def withdraw(self, amount):
        if amount > self.balance:
            raise ValueError(f"balance is only {self.balance}")
        self.balance -= amount
        self.history.append(("withdraw", amount))
        return self.balance


acc = BankAccount("Asha", 1000)
acc.deposit(500)
acc.withdraw(200)
print(acc.balance, acc.history)     # 1300 [('deposit', 500), ('withdraw', 200)]

# Each instance has its own data
other = BankAccount("Ravi")
print(other.balance, other.history)  # 0 []

# What the dot really does
print(BankAccount.deposit(acc, 100) == acc.balance)   # True — same call

# Forgetting self
class Broken:
    def greet():                    # no self
        return "hi"

# Broken().greet()   # TypeError: greet() takes 0 positional arguments
                     # but 1 was given
  • class Name: — PascalCase; the blueprint
  • __init__(self, ...) — runs when an instance is created
  • self — the instance; first parameter of every instance method, by convention
  • self.x = value — an instance attribute, private to that object
  • obj.method() is Class.method(obj) — that is where self comes from
  • Set every attribute in __init__, even the ones that start empty
Notes
  • Python lets you attach new attributes to an object at any time: acc.branch = "Kolkata" works even though branch is nowhere in the class. That flexibility means a typo — acc.blance = 0 — creates a second attribute instead of raising an error, and the real one stays unchanged.

Class Attributes, and the Shared-List Trap

An assignment in the class body, outside any method, creates a class attribute: one value shared by the class and every instance of it. Instance attributes, created with self.x = ..., belong to a single object. Both are read with the same dot syntax, which is exactly why the difference catches people out.

Class attributes are right for things that genuinely belong to the type rather than to an individual: a constant such as a pass mark, a piece of configuration, or a counter of how many instances have been made. Reading works through either the class or an instance, since Python looks on the instance first and falls back to the class.

The first trap is shadowing. Writing emp.company = "NewCorp" does not change the class attribute — it creates an instance attribute with the same name that hides it for that one object. Every other instance still sees the original. To change it for everyone, assign through the class: Employee.company = "NewCorp".

The second trap is far worse, because it produces wrong data rather than surprising data. A mutable class attribute — a list, a dictionary, a set — is one object shared by every instance. Appending to it through one object appends for all of them. It is exactly the mutable-default-argument bug from Lesson 13, wearing a different hat: one object created once, shared forever.

The fix is the same in spirit. Anything that should be per-instance must be created inside __init__, where the line runs afresh for every object. Keep class attributes for immutable values — numbers, strings, tuples — and you will never meet this problem.

Example
class Employee:
    company = "Priodemy"      # class attribute: shared
    count = 0                 # a counter across all instances

    def __init__(self, name, salary):
        self.name = name      # instance attributes: per object
        self.salary = salary
        Employee.count += 1   # note: through the CLASS, not self


a = Employee("Asha", 95000)
b = Employee("Ravi", 85000)
print(Employee.count, a.company, b.company)   # 2 Priodemy Priodemy

# Shadowing: this changes ONE object, not the class
a.company = "NewCorp"
print(a.company, b.company, Employee.company)  # NewCorp Priodemy Priodemy

# To change it everywhere, go through the class
Employee.company = "Priodemy Labs"
print(b.company)                               # Priodemy Labs
print(a.company)                               # NewCorp — still shadowed


# THE TRAP: a mutable class attribute is one shared object
class StudentBad:
    subjects = []             # ONE list for every student ever created

    def __init__(self, name):
        self.name = name

asha = StudentBad("Asha")
asha.subjects.append("Maths")

ravi = StudentBad("Ravi")
print(ravi.subjects)          # ['Maths'] — Ravi never chose anything
print(asha.subjects is ravi.subjects)   # True — literally the same list


# The fix: create it per instance, inside __init__
class Student:
    pass_mark = 40            # immutable class attribute is fine

    def __init__(self, name):
        self.name = name
        self.subjects = []    # a NEW list for every student

asha, ravi = Student("Asha"), Student("Ravi")
asha.subjects.append("Maths")
print(ravi.subjects)          # []
Notes
  • obj.__dict__ shows only the instance's own attributes, so it is the quickest way to check whether a value is genuinely on the object or coming from the class behind it.

Dunder Methods: Making Objects Behave Like Python

Print an object of your own class and you get something like <__main__.Student object at 0x000001A2>. That is Python's default, and it is useless in a log or a debugger. Dunder methods — short for double underscore — let you hook your class into the language's own operations so that print(), ==, len() and the rest do something sensible.

Two of them handle display and they are for different audiences. __str__ is the friendly form, used by print() and str(): what a user should see. __repr__ is the unambiguous form, used by repr(), by the interactive prompt, and — importantly — when your object appears inside a list or dictionary. Its convention is to look like the code that would recreate the object. If you write only one, write __repr__, because Python falls back to it when __str__ is missing but not the other way round.

__eq__ controls ==. By default, two objects are equal only if they are literally the same object, so two Student instances holding identical data compare as different. Defining __eq__ to compare the fields that matter fixes that, and it also fixes in, .count() and .index(), which all use == underneath.

There is a catch attached: defining __eq__ makes your class unhashable, so instances can no longer go in a set or be used as dictionary keys. Python does this deliberately, because equal objects must hash the same, and it cannot guess your rule. If you need hashability, define __hash__ to return a hash of the same fields you compared.

Beyond those, __len__ makes len(obj) work — and, as a side effect, makes an empty object falsy in an if. There are dozens more for arithmetic, indexing and iteration. You do not need them all; you need to know they exist, so that when you want your object to work with a built-in, you can look up which method to write.

Example
class Student:
    def __init__(self, name, marks):
        self.name = name
        self.marks = marks

    def __repr__(self):        # for developers, logs, containers
        return f"Student({self.name!r}, {self.marks!r})"

    def __str__(self):         # for users
        return f"{self.name}: {sum(self.marks)} marks"

    def __eq__(self, other):
        if not isinstance(other, Student):
            return NotImplemented
        return (self.name, self.marks) == (other.name, other.marks)

    def __hash__(self):        # needed because we defined __eq__
        return hash((self.name, tuple(self.marks)))

    def __len__(self):
        return len(self.marks)


a = Student("Asha", [78, 84])
b = Student("Asha", [78, 84])

print(a)               # Asha: 162 marks          <- __str__
print(repr(a))         # Student('Asha', [78, 84]) <- __repr__
print([a])             # [Student('Asha', [78, 84])] — containers use repr

print(a == b)          # True  — same data, thanks to __eq__
print(a is b)          # False — still two different objects
print(a in [b])        # True  — `in` uses ==

print(len(a))          # 2     — __len__
print(bool(Student("Ravi", [])))   # False — no marks, so __len__ is 0

print({a, b})          # one element — equal and hashing the same
  • __init__ — set up a new instance
  • __repr__ — unambiguous form; used by the prompt and inside containers
  • __str__ — friendly form; used by print(); falls back to __repr__
  • __eq__ — defines ==, and therefore in, .count(), .index()
  • __hash__ — required alongside __eq__ if instances go in sets or dict keys
  • __len__ — makes len() work and decides truthiness
Notes
  • Returning NotImplemented from __eq__ for an unrelated type is the correct signal, not an error. Python then tries the other object's comparison before concluding the two are unequal.

Methods That Are Not About One Instance

Most methods work on a single object and take self. Two decorators change that, and both solve real problems.

@classmethod receives the class itself as its first parameter, conventionally named cls, instead of an instance. Its main use is an alternative constructor: a second way of creating an object. __init__ takes the fields directly, while a classmethod can take a CSV line, a dictionary from an API, or a database row, do the parsing, and then call cls(...). The caller gets a clearly named entry point — Student.from_csv(line) — instead of a pile of parsing code before every construction.

@staticmethod receives neither. It is a plain function that lives inside the class because it belongs there conceptually — a validation helper, a formatting rule. If it never touches self or cls, making it static says so, and it can be called on the class without creating an instance.

Be honest about the third case, though: if a function neither uses the instance nor conceptually belongs to the class, it should be a module-level function. Classes are not folders, and a class full of static methods is usually a module wearing a costume.

Example
class Student:
    pass_mark = 40

    def __init__(self, name, marks):
        self.name = name
        self.marks = marks

    # Alternative constructor: takes cls, returns an instance
    @classmethod
    def from_csv(cls, line):
        """Build a Student from 'Asha,78,84,91'."""
        name, *scores = line.strip().split(",")
        return cls(name, [int(s) for s in scores])

    @classmethod
    def from_dict(cls, data):
        return cls(data["name"], data["marks"])

    # No self, no cls — a helper that belongs here conceptually
    @staticmethod
    def is_valid_mark(value):
        return isinstance(value, int) and 0 <= value <= 100

    def average(self):
        return sum(self.marks) / len(self.marks)


a = Student("Asha", [78, 84, 91])
b = Student.from_csv("Ravi,65,72,80")
c = Student.from_dict({"name": "Meera", "marks": [90, 95]})

print(b.name, round(b.average(), 2))     # Ravi 72.33

# A static method needs no instance
print(Student.is_valid_mark(105))        # False
print(Student.is_valid_mark(87))         # True

# Reading several rows becomes one line
lines = ["Asha,78,84", "Ravi,65,72"]
students = [Student.from_csv(l) for l in lines]
print([s.name for s in students])        # ['Asha', 'Ravi']
Notes
  • Using cls(...) rather than the class name by hand matters once inheritance appears in the next lesson: a subclass calling the inherited from_csv then gets an instance of the subclass, not of the parent.

Dataclasses, and When Not to Write a Class

A great many classes exist only to hold a few fields. Writing __init__, __repr__ and __eq__ for them by hand is tedious and easy to get subtly wrong. The @dataclass decorator generates all three from the field declarations, so a twenty-line class becomes five.

You declare each field with a name and a type hint. The decorator writes an __init__ taking them in order, a __repr__ showing them, and an __eq__ comparing them. Defaults work as usual, and you can still add your own methods — a dataclass is an ordinary class with some boilerplate filled in.

It also protects you from the shared-mutable trap in a way plain Python does not. Writing marks: list = [] in a dataclass raises ValueError at definition time, telling you to use field(default_factory=list) instead. The factory is called once per instance, which is exactly the fix from Lesson 13, made compulsory. Adding frozen=True goes further and makes instances immutable and hashable.

Knowing when not to write a class is as valuable. A small fixed record that never changes is often better as a namedtuple. Data coming straight from JSON that you only read once is fine as a dictionary. A calculation with no state should stay a function. Reach for a class when there is data and behaviour and the two belong together — and when several instances will exist independently.

Two habits to carry into the next two lessons. Keep classes small enough to describe in a sentence, and prefer holding another object over inheriting from one unless the relationship is genuinely "is a kind of". Both save far more trouble than they cost.

Example
from dataclasses import dataclass, field

@dataclass
class Student:
    name: str
    branch: str = "CSE"
    marks: list = field(default_factory=list)   # a NEW list per instance

    def average(self):
        return sum(self.marks) / len(self.marks) if self.marks else 0


a = Student("Asha", "CSE", [78, 84])
b = Student("Asha", "CSE", [78, 84])

print(a)          # Student(name='Asha', branch='CSE', marks=[78, 84])
print(a == b)     # True  — __eq__ was generated
print(round(a.average(), 2))    # 81.0

c = Student("Ravi")
c.marks.append(90)
print(Student("Meera").marks)   # [] — not shared

# A plain mutable default is refused outright
# @dataclass
# class Bad:
#     marks: list = []    # ValueError: mutable default ... use default_factory

# frozen=True gives immutable, hashable instances
@dataclass(frozen=True)
class Point:
    x: int
    y: int

p = Point(1, 2)
print({p: "origin-ish"})        # usable as a dictionary key
# p.x = 5                       # FrozenInstanceError

# When a class is overkill
from collections import namedtuple
Colour = namedtuple("Colour", "r g b")   # fixed record, no behaviour
print(Colour(255, 0, 0).r)               # 255
  • @dataclass generates __init__, __repr__ and __eq__ from the fields
  • field(default_factory=list) — a fresh mutable default per instance
  • @dataclass(frozen=True) — immutable and hashable instances
  • You can still add your own methods; it is an ordinary class underneath
  • Fixed record, no behaviour: prefer a namedtuple or a dictionary
  • Stateless calculation: prefer a plain function
Notes
  • The type hints in a dataclass are not enforced — Student("Asha", 5, "oops") runs happily. They exist for readers and for your editor, and the dataclass machinery uses them only to know which names are fields.
Ask AI