Class 11Computer Science · Computer SystemsFull chapter

Encoding Schemes and Number System

The whole chapter in one place — read it, then test yourself. Clear notes, a reference sheet, a practice quiz, and worked NCERT solutions & PYQs.

Why Binary, and How Positional Systems Work

Quick answer A computer stores numbers in base 2 because a transistor reliably has only two states, and every base - 2, 8, 10 or 16 - works by the same rule: each digit is worth the digit times the base raised to its place number.

A processor has no idea what a 5 is. Inside the chip there are only transistors, and a transistor is dependably in one of two conditions: current flowing or not flowing, high voltage or low voltage. Two states can be told apart even when the chip is hot, the power supply sags a little, and there is electrical noise on the line. Ten different voltage levels cannot be told apart that reliably. That one engineering fact is the whole reason your Aadhaar number, a UPI payment of two hundred rupees, and this sentence are all finally stored in base 2.

One two-state cell is a bit. Eight bits grouped together are a byte. With n bits you can build 2n different patterns, so a byte has 28 = 256 patterns, which hold the values 0 to 255. This is not a rule someone invented; it is just counting. Each extra bit doubles the number of patterns because that bit can be 0 or 1 on top of everything the earlier bits could already do.

Positional notation. Decimal already works the way binary does, you have simply never had to think about it. In 4507 the 5 is not "five", it is "five hundred", and the only reason is where it sits. Every positional system has a base (also called the radix) and a set of legal digits running from 0 up to base minus 1. The value of the whole number is:

value = d(k) x b^k + ... + d(2) x b^2 + d(1) x b^1 + d(0) x b^0

Place numbers start at 0 at the rightmost digit and count leftwards. So (4507)10 = 4x1000 + 5x100 + 0x10 + 7x1. Change the base and nothing about the method changes - only the number you raise to the power.

SystemBaseLegal digitsPython prefixExample
Binary20 10b0b1010
Octal80 1 2 3 4 5 6 70o0o12
Decimal100 to 9none10
Hexadecimal160 to 9, then A B C D E F0x0xA

Base 16 runs out of ordinary digits after 9, so the letters A to F are pressed into service for the values 10 to 15. A is not a letter here, it is the number ten.

Why octal and hexadecimal exist at all. Machines are happy with binary; humans are not. Writing 2026 as 11111101010 and copying it without dropping a bit is painful. But 8 = 23 and 16 = 24, so exactly three bits collapse into one octal digit and exactly four bits collapse into one hex digit, with no arithmetic and no carrying. Decimal gets no such shortcut, because 10 is not a power of 2. That is why memory addresses, MAC addresses and web colour codes are written in hex - it is compressed binary that a human can read aloud.

Worked example. The three ideas above, checked by running them:

# A decimal number is really a sum: digit x base ** place
print(4 * 10**3 + 5 * 10**2 + 0 * 10**1 + 7 * 10**0)

# One byte = 8 bits. How many patterns does that give?
print(2**8, "patterns, values 0 to", 2**8 - 1)

# Python reads all four bases directly. All four below are the same number.
print(0b1010, 0o12, 10, 0xA)

Output:

4507
256 patterns, values 0 to 255
10 10 10 10

The last line is the point worth remembering: 0b1010, 0o12, 10 and 0xA are four ways of typing one single number. The prefix affects nothing inside the machine - once stored, all four are the same pattern of bits, and Python prints them the same way.

Second worked example. A web colour such as #FF9933 is nothing but three two-digit hex numbers stuck together - red, green, blue:

# A web colour like #FF9933 is just three hex numbers: red, green, blue
print(int("FF", 16), int("99", 16), int("33", 16))
print("largest value one pair of hex digits can hold:", int("FF", 16))

Output:

255 153 51
largest value one pair of hex digits can hold: 255

Two hex digits cover exactly 0 to 255, which is exactly one byte. That is not a coincidence - 2 hex digits are 8 bits.

Positional value value = d(k) x b^k + ... + d(1) x b^1 + d(0) x b^0 b is the base. Place numbers start at 0 on the right of the number.
Legal digits in base b 0, 1, 2, ... , (b - 1) Base 16 runs out at 9, so A-F stand for 10-15. Seeing an 8 in an octal number means the number is invalid.
Patterns from n bits n bits -> 2**n values, range 0 to 2**n - 1 bits · 8 bits = 256 values = 0 to 255. This is why one byte tops out at 255.
Python base literals 0b1010 == 0o12 == 10 == 0xA The prefix is only about how you TYPE the number. All four are stored identically.
Why 8 and 16 were chosen 8 = 2**3 and 16 = 2**4 One octal digit is exactly 3 bits, one hex digit exactly 4 bits. 10 is not a power of 2, so decimal gets no grouping trick.
Remember
  • Computers use base 2 because a transistor can hold two states reliably; ten voltage levels cannot be distinguished dependably, so base 10 hardware is impractical.
  • Every positional system follows one rule: value = sum of (digit x base ** place), with places numbered from 0 at the right.
  • A base b uses digits 0 to b-1; hexadecimal borrows A to F for the values 10 to 15.
  • n bits give 2**n patterns, holding values 0 to 2**n - 1; a byte therefore holds 0 to 255.
  • Octal and hex exist for human convenience, because 8 = 2**3 and 16 = 2**4 let bits be grouped exactly; decimal has no such shortcut.

Decimal to Binary, Octal and Hexadecimal

Quick answer Divide the whole part repeatedly by the target base and read the remainders bottom-up; multiply the fractional part repeatedly by the base and read the integer parts top-down.

Going from decimal into any other base uses two separate procedures - one for the part before the decimal point, one for the part after it. Do them separately, then join the two answers with a point.

Whole part: repeated division. Divide the number by the target base. Write down the remainder. Divide the quotient by the base again. Keep going until the quotient becomes 0. Then read the remainders bottom to top.

The reason this works is worth knowing rather than memorising. Any division gives n = quotient x b + remainder, and the remainder is always smaller than b - so the remainder is exactly the units digit of n in base b. Dividing away that remainder shifts everything one place right, and the next remainder is the next digit up. The digits therefore come out in reverse order, which is why you read them upwards.

Worked example: (173)10 to binary. Each new remainder is glued to the front of the answer string, which is exactly what "read bottom-up" means:

n = 173
bits = ""
while n > 0:
    print(n, "/ 2  ->  quotient", n // 2, ", remainder", n % 2)
    bits = str(n % 2) + bits
    n = n // 2

print("Remainders read bottom-up:", bits)
print("Cross-check with Python  :", bin(173), oct(173), hex(173))

Output:

173 / 2  ->  quotient 86 , remainder 1
86 / 2  ->  quotient 43 , remainder 0
43 / 2  ->  quotient 21 , remainder 1
21 / 2  ->  quotient 10 , remainder 1
10 / 2  ->  quotient 5 , remainder 0
5 / 2  ->  quotient 2 , remainder 1
2 / 2  ->  quotient 1 , remainder 0
1 / 2  ->  quotient 0 , remainder 1
Remainders read bottom-up: 10101101
Cross-check with Python  : 0b10101101 0o255 0xad

So (173)10 = (10101101)2 = (255)8 = (AD)16. Note that (255)8 has nothing to do with the decimal number 255 - the same three symbols mean 173 when read in base 8.

The identical method for base 8 and base 16. Only the divisor changes. For (2026)10 into hexadecimal, divide by 16 each time and convert any remainder above 9 into a letter:

StepQuotientRemainderHex digit
2026 / 1612610A
126 / 16714E
7 / 16077

Reading the hex digits upwards gives 7EA, and Python confirms it: hex(2026) returns '0x7ea'. Similarly (173)10 in octal is 173/8 = 21 rem 5, 21/8 = 2 rem 5, 2/8 = 0 rem 2, read upwards as 255.

Fractional part: repeated multiplication. Multiply the fraction by the base. Whatever integer part pops out on the left is the next digit. Drop that integer part and multiply the leftover fraction again. Read the digits top to bottom - the opposite direction from the whole part, because multiplying by b shifts the point one place to the right, pushing out the digit that sits closest to the point first.

Worked example: (45.375)10 to binary.

frac = 0.375
for step in range(4):
    frac = frac * 2
    print("x 2  ->", frac, " carry out the integer part:", int(frac))
    frac = frac - int(frac)
    if frac == 0:
        print("remainder is 0, stop")
        break

# 45 is 101101, and .375 is .011, so 45.375 should be 101101.011
print("check integer part:", bin(45))
print("check fraction    :", 0 * 2**-1 + 1 * 2**-2 + 1 * 2**-3)

Output:

x 2  -> 0.75  carry out the integer part: 0
x 2  -> 1.5  carry out the integer part: 1
x 2  -> 1.0  carry out the integer part: 1
remainder is 0, stop
check integer part: 0b101101
check fraction    : 0.375

Digits read top-down are 0, 1, 1. So (45.375)10 = (101101.011)2.

The gotcha you must know. A decimal fraction only terminates in binary if its denominator is a power of 2. 0.375 is 3/8, so it stops after three bits. But 0.7 is 7/10, and 10 has a factor of 5, so it never stops - its bits are 1011001100110011... repeating forever. A computer must cut it off somewhere, which stores a value very slightly off. That is the honest reason 0.1 + 0.2 prints 0.30000000000000004 in Python rather than 0.3. It is not a bug in Python; it is base 2 meeting a fraction it cannot write down.

Decimal to base b (whole part) divide by b repeatedly, collect remainders, read BOTTOM-UP Works because n = quotient x b + remainder, and that remainder is the units digit in base b.
Decimal to base b (fraction) multiply the fraction by b repeatedly, collect integer parts, read TOP-DOWN Multiplying by b shifts the point one place right, so the digit nearest the point comes out first.
bin() bin(n) -> str, e.g. bin(173) is '0b10101101' Takes an int only. Remove the prefix with bin(n)[2:].
oct() oct(n) -> str, e.g. oct(173) is '0o255' (255) in base 8 is decimal 173, not decimal 255. Never read an octal answer as decimal.
hex() hex(n) -> str, e.g. hex(2026) is '0x7ea' Letters come out lowercase. Boards usually expect uppercase, so write 7EA.
format() with a width format(45, '08b') -> '00101101' Pads with leading zeros to a fixed width. Base codes are 'b', 'o', 'x' and 'X'.
Remember
  • Whole part: divide repeatedly by the target base, collect remainders, read them BOTTOM-UP.
  • Fractional part: multiply repeatedly by the base, collect integer parts, read them TOP-DOWN.
  • Remainders come out in reverse because each division peels off the current units digit first.
  • For hexadecimal, convert any remainder of 10 to 15 into A to F before writing it down.
  • A decimal fraction terminates in binary only if its denominator is a power of 2; 0.375 stops, 0.7 repeats forever, which is why 0.1 + 0.2 gives 0.30000000000000004.

Binary, Octal and Hexadecimal Back to Decimal

Quick answer Every base converts to decimal by one method - multiply each digit by the base raised to its place number and add, with negative powers used to the right of the point.

Going the other way is easier, because there is only one method and it works for every base. Number the places: 0 for the digit immediately left of the point, then 1, 2, 3 going left; and -1, -2, -3 going right from the point. Multiply each digit by base raised to its place number, then add everything.

Worked example: (110101101)2 to decimal. Written out in full so you can see every place value. The slice b[::-1] walks the string right to left, which is the direction the place numbers grow:

# (110101101) base 2 -> decimal, by expanding place values
b = "110101101"
total = 0
place = 0
for bit in b[::-1]:
    value = int(bit) * 2**place
    total = total + value
    print("bit", bit, "at place", place, "->", bit, "x", 2**place, "=", value)
    place = place + 1

print("Total:", total)
print("Python's int() agrees:", int(b, 2))

Output:

bit 1 at place 0 -> 1 x 1 = 1
bit 0 at place 1 -> 0 x 2 = 0
bit 1 at place 2 -> 1 x 4 = 4
bit 1 at place 3 -> 1 x 8 = 8
bit 0 at place 4 -> 0 x 16 = 0
bit 1 at place 5 -> 1 x 32 = 32
bit 0 at place 6 -> 0 x 64 = 0
bit 1 at place 7 -> 1 x 128 = 128
bit 1 at place 8 -> 1 x 256 = 256
Total: 429
Python's int() agrees: 429

In an exam you do not write nine lines. Write the place values 256 128 64 32 16 8 4 2 1 above the bits, and add only the columns where the bit is 1: 256 + 128 + 32 + 8 + 4 + 1 = 429. For binary the digit is only ever 0 or 1, so there is no multiplication at all - it is pure addition of the place values you keep.

Octal and hexadecimal. Same expansion, different powers. Powers of 8 are 1, 8, 64, 512. Powers of 16 are 1, 16, 256, 4096. For hex you must first replace letters by their values: A=10, B=11, C=12, D=13, E=14, F=15.

print("(725) base 8  =", 7 * 8**2 + 2 * 8**1 + 5 * 8**0, "  int() says", int("725", 8))
print("(3AF) base 16 =", 3 * 16**2 + 10 * 16**1 + 15 * 16**0, "  int() says", int("3AF", 16))
print("(1011.101) base 2 =", int("1011", 2) + 1*2**-1 + 0*2**-2 + 1*2**-3)

Output:

(725) base 8  = 469   int() says 469
(3AF) base 16 = 943   int() says 943
(1011.101) base 2 = 11.625

So (725)8 = 7x64 + 2x8 + 5x1 = 448 + 16 + 5 = 469, and (3AF)16 = 3x256 + 10x16 + 15x1 = 768 + 160 + 15 = 943.

Digits after the point. The place values to the right of the point are b-1, b-2, b-3. In binary those are 0.5, 0.25 and 0.125. So .101 means 0.5 + 0 + 0.125 = 0.625, giving (1011.101)2 = 11 + 0.625 = 11.625. Students lose marks here by numbering the first place after the point as 0 instead of -1 - the place immediately right of the point is always -1.

Python's shortcut, and a deliberate error. int(s, base) converts a string written in any base from 2 to 36 into a decimal int. It also checks the number is legal, which is genuinely useful:

print(int("129", 8))

Output:

Traceback (most recent call last):
  File "demo.py", line 1, in 
    print(int("129", 8))
          ~~~^^^^^^^^^^
ValueError: invalid literal for int() with base 8: '129'

The error is the teaching point, not a mistake. Base 8 has no digit 9, so 129 is not an octal number at all - the same way 1G is not a hexadecimal number. Before you convert anything, check that every digit is legal for that base.

Base b to decimal value = sum of (digit x b ** place) The single method that works for every base and every direction from the point.
Places after the point b**-1, b**-2, b**-3 ... In binary these are 0.5, 0.25, 0.125. The first place after the point is -1, never 0.
int() with a base int(s, base) -> int, e.g. int('3AF', 16) is 943 s must be a STRING. int(3AF, 16) is a syntax error. Base may be 2 to 36.
Hex letter values A=10, B=11, C=12, D=13, E=14, F=15 Case does not matter to int(): int('3af', 16) == int('3AF', 16) is True.
Illegal digit check int('129', 8) -> ValueError Octal has no 8 or 9; hex has no G. Validate the digits before you start converting.
Binary shortcut write 128 64 32 16 8 4 2 1 above the bits and add where the bit is 1 Fastest exam method for 8-bit numbers, with no multiplication needed.
Remember
  • One method converts any base to decimal: multiply each digit by base raised to its place number, then add.
  • Places are numbered 0, 1, 2 going left from the point and -1, -2, -3 going right from it.
  • For binary just add the place values where the bit is 1: 256+128+32+8+4+1 = 429 for (110101101) base 2.
  • Replace hex letters with values before expanding: A=10, B=11, C=12, D=13, E=14, F=15.
  • int(s, base) does the whole job and raises ValueError if a digit is illegal for that base, e.g. int('129', 8).

Grouping Shortcuts Between Binary, Octal and Hex

Quick answer Because 8 = 2 cubed and 16 = 2 to the fourth, three bits collapse into one octal digit and four bits into one hex digit, so these conversions need grouping rather than arithmetic.

Converting binary to octal or hex through decimal is slow and invites mistakes. There is a shortcut, and it exists for a precise reason: 8 = 23, so every possible group of 3 bits maps onto exactly one octal digit (000 to 111 covers 0 to 7 with nothing left over). And 16 = 24, so every group of 4 bits maps onto exactly one hex digit (0000 to 1111 covers 0 to F). No arithmetic, no carrying - it is pure substitution.

This is also why there is no such trick for decimal. There is no whole number k with 2k = 10, so bits never group evenly into decimal digits.

The rule. Group from the point outwards: leftwards in 3s (octal) or 4s (hex) for the whole part, padding the left with zeros; rightwards for the fractional part, padding the right with zeros. Padding on the wrong side changes the value, so this is where marks are lost.

BitsOctal digitBitsHex digitBitsHex digit
00000000010008
00110001110019
0102001021010A
0113001131011B
1004010041100C
1015010151101D
1106011061110E
1117011171111F

Worked example: (101110110)2 to octal and to hex. Nine bits split cleanly into three groups of 3, because 9 is a multiple of 3. They do not split cleanly into 4s: 9 is not a multiple of 4, so three zeros must be padded on the left to reach twelve bits, giving 0001 0111 0110.

b = "101110110"

# Octal: group the 9 bits in 3s from the right
for g in ["101", "110", "110"]:
    print(g, "->", int(g, 2))
print("octal answer:", oct(int(b, 2)))

# Hex: the same 9 bits need 3 leading zeros to make 12
for g in ["0001", "0111", "0110"]:
    print(g, "->", format(int(g, 2), "X"))
print("hex answer  :", hex(int(b, 2)))

# And backwards: each hex digit expands to exactly 4 bits
for d in "1B2":
    print(d, "->", format(int(d, 16), "04b"))
print("so (1B2) base 16 =", bin(int("1B2", 16)))

Output:

101 -> 5
110 -> 6
110 -> 6
octal answer: 0o566
0001 -> 1
0111 -> 7
0110 -> 6
hex answer  : 0x176
1 -> 0001
B -> 1011
2 -> 0010
so (1B2) base 16 = 0b110110010

So (101110110)2 = (566)8 = (176)16, and running the process backwards, (1B2)16 = (110110010)2. Note the leading zero of 0001 is simply dropped in the final binary answer, exactly as you would drop it in decimal.

Octal to hexadecimal has no direct route. 8 and 16 are both powers of 2, but neither is a power of the other, so 3-bit groups do not line up with 4-bit groups. You must expand to binary and regroup. This is the standard three-step answer expected in the exam:

# Octal -> hexadecimal always goes through binary. (3752) base 8 is the target.
bits = ""
for d in "3752":
    bits = bits + format(int(d, 8), "03b")
print("each octal digit as 3 bits:", bits)

# Regroup the SAME 12 bits in 4s
hexdigits = ""
i = 0
while i < len(bits):
    group = bits[i:i + 4]
    print("group", group, "->", format(int(group, 2), "X"))
    hexdigits = hexdigits + format(int(group, 2), "X")
    i = i + 4

print("as hex digits:", hexdigits)
print("cross-check  :", hex(int("3752", 8)), "and decimal", int("3752", 8))

Output:

each octal digit as 3 bits: 011111101010
group 0111 -> 7
group 1110 -> E
group 1010 -> A
as hex digits: 7EA
cross-check  : 0x7ea and decimal 2026

The bits never changed - 011111101010 is the same twelve bits both times. Only the grouping changed, and that is the whole idea: 3 at a time you read it as (3752)8, 4 at a time you read it as (7EA)16, and both are the decimal number 2026. Here the twelve bits happened to divide evenly by 4, so no padding was needed; when they do not, pad on the left.

Binary to octal group bits in 3s from the point outwards, pad left with 0s 3 bits per digit because 8 = 2**3. (101)(110)(110) -> 566.
Binary to hexadecimal group bits in 4s from the point outwards, pad left with 0s 4 bits per digit because 16 = 2**4. (0001)(0111)(0110) -> 176.
Octal digit to bits format(int(d, 8), '03b') Always exactly 3 bits, leading zeros kept. Digit 3 becomes 011, not 11.
Hex digit to bits format(int(d, 16), '04b') Always exactly 4 bits. B becomes 1011, F becomes 1111.
Octal to hexadecimal octal -> binary (3s) -> regroup (4s) -> hexadecimal No direct shortcut exists; the exam expects the binary middle step to be shown.
Padding side whole part pads LEFT, fraction pads RIGHT Padding the wrong side silently changes the value. Most common mistake in this topic.
Remember
  • One octal digit is exactly 3 bits because 8 = 2**3; one hex digit is exactly 4 bits because 16 = 2**4.
  • Group from the point outwards: pad the LEFT with zeros for the whole part, the RIGHT with zeros for the fraction.
  • Binary to octal or hex is substitution, not arithmetic, so it is far faster and safer than routing through decimal.
  • Octal and hexadecimal cannot be converted directly into each other; expand to binary and regroup.
  • Decimal gets no grouping shortcut because 10 is not a power of 2.

ASCII, ISCII and Unicode

Quick answer An encoding scheme is an agreed number for every character: ASCII covers 128 English characters in 7 bits, ISCII added Indian scripts in an 8-bit code, and Unicode gives every character in every script one code point, stored as UTF-8 or UTF-32 bytes.

Memory holds only numbers, so text has to be turned into numbers. An encoding scheme is simply a public agreement about which number stands for which character. The agreement is the entire point - if you save a file with one scheme and open it with another, the receiver decodes different characters and you get garbage on screen.

ASCII. American Standard Code for Information Interchange. It is a 7-bit code, so it has 27 = 128 codes, numbered 0 to 127. Its layout is deliberate, not random:

CodesCharactersNote
0 to 31, and 127Control charactersNot printable, e.g. code 10 is newline
32SpaceThe first printable code
48 to 57'0' to '9'In order, so ord(c) - 48 gives the value
65 to 90'A' to 'Z'In alphabetical order
97 to 122'a' to 'z'Exactly 32 above the uppercase letters

The gap of 32 between the cases is not an accident. 32 is 100000 in binary, so switching between upper and lower case flips one single bit - cheap for 1960s hardware to do. Extended ASCII uses the eighth bit to reach 256 codes, but there was never one agreed meaning for codes 128 to 255, and that disagreement is exactly the problem the later schemes had to solve.

Worked example: ord() and chr(). These are the two built-in functions the syllabus wants. ord() takes one character and returns its code; chr() does the reverse.

print(ord("A"), ord("Z"), ord("a"), ord("z"), ord("0"), ord("9"), ord(" "))
print(chr(65), chr(97), chr(48))

# Lowercase is exactly 32 more than uppercase, and 32 is 100000 in binary
print(ord("a") - ord("A"), bin(32))
print(chr(ord("P") + 32))

# The digit character "7" is not the number 7
print(ord("7"), ord("7") - ord("0"))

Output:

65 90 97 122 48 57 32
A a 0
32 0b100000
p
55 7

That last line matters. The character '7' has code 55; it is a picture of a digit, not the quantity seven. Subtracting ord('0') converts the picture into the number.

ISCII. Indian Script Code for Information Interchange, standardised by the Bureau of Indian Standards as IS 13194:1991. ASCII has no place for क or ম or த, so India needed its own scheme. ISCII is an 8-bit code: codes 0 to 127 are identical to ASCII, so plain English text keeps working unchanged, and the upper half, codes 128 to 255, holds Indian script characters.

Its clever idea is that a single common chart serves many Indian scripts. Devanagari, Bengali, Gurmukhi, Gujarati, Odia, Tamil, Telugu, Kannada and Malayalam all descend from Brahmi and share the same underlying phonetic structure - the same consonants and vowels in the same order. So ISCII assigns one code per sound, and a separate script-selection code tells the display which script to draw it in. The same stored byte can appear as क or as ক depending on the script currently selected.

Why Unicode superseded it. ISCII was a sound design for 8-bit machines, but it hit hard limits:

  • Only the 128 codes of the upper half were left for all of India's characters.
  • Because script selection is a mode rather than part of each character's identity, a document mixing Hindi and Tamil freely is awkward, and one lost mode-switch code corrupts everything after it.
  • Scripts not derived from Brahmi, such as the Perso-Arabic script used for Urdu, do not fit the shared phonetic chart.
  • It solved nothing outside India. A file with Hindi and Chinese and Arabic in it was still impossible, and international software had no reason to support it.

Unicode. Unicode gives every character in every script of the world its own permanent number, called a code point, written as U+ followed by hex digits. The range runs from U+0000 to U+10FFFF, which is 1,114,112 possible code points - far more than enough. ASCII was folded in unchanged, so U+0041 is still 'A' with value 65. ISCII's work was not thrown away either: Unicode's Devanagari block, U+0900 to U+097F, was based on the ISCII layout, so the Indian standard's design survives inside the global one.

for ch in ["A", "अ", "क", "ব", "ம", "₹"]:
    print(ch, "  code point", ord(ch), " written as U+%04X" % ord(ch))

Output:

A   code point 65  written as U+0041
अ   code point 2309  written as U+0905
क   code point 2325  written as U+0915
ব   code point 2476  written as U+09AC
ம   code point 2990  written as U+0BAE
₹   code point 8377  written as U+20B9

Note that क and ব now have different code points, unlike ISCII where they could share one. Each character carries its own script identity, so one file can hold Hindi, Bengali, Tamil and English at once with no mode switching. The rupee sign ₹ is U+20B9 - which is why a UPI app or an IRCTC ticket can print it at all.

Code point versus encoding. U+0915 is क's identity. It does not say how many bytes to write to disk. That is the job of a UTF (Unicode Transformation Format), and the syllabus names two:

  • UTF-8 is variable-length, using 1 to 4 bytes per character. U+0000 to U+007F take 1 byte, U+0080 to U+07FF take 2, U+0800 to U+FFFF take 3, and U+10000 upwards take 4. Its great advantage is that the 1-byte range is byte-for-byte identical to ASCII, so every old English file is already valid UTF-8. Devanagari falls in the 3-byte band. UTF-8 is the default on the web.
  • UTF-32 is fixed-length: every character takes exactly 4 bytes, holding the code point directly. Finding the 500th character is instant arithmetic, but plain English text becomes four times larger than it needs to be.
for ch in ["A", "क", "₹"]:
    u8 = ch.encode("utf-8")
    u32 = ch.encode("utf-32-be")
    print(ch, "| UTF-8:", len(u8), "bytes", u8.hex(" "), "| UTF-32:", len(u32), "bytes", u32.hex(" "))

msg = "Namaste नमस्ते"
print(msg, "has", len(msg), "characters")
print("UTF-8 needs ", len(msg.encode("utf-8")), "bytes")
print("UTF-32 needs", len(msg.encode("utf-32-be")), "bytes")

Output:

A | UTF-8: 1 bytes 41 | UTF-32: 4 bytes 00 00 00 41
क | UTF-8: 3 bytes e0 a4 95 | UTF-32: 4 bytes 00 00 09 15
₹ | UTF-8: 3 bytes e2 82 b9 | UTF-32: 4 bytes 00 00 20 b9
Namaste नमस्ते has 14 characters
UTF-8 needs  26 bytes
UTF-32 needs 56 bytes

Read the UTF-32 bytes for क: 00 00 09 15. That is U+0915 written out directly, which is exactly what fixed-length means. The UTF-8 bytes e0 a4 95 look nothing like 0915 because UTF-8 spreads the code point across three bytes with marker bits. And the count at the end shows the trade-off plainly: 14 characters cost 26 bytes in UTF-8 (8 English characters at 1 byte each, 6 Devanagari at 3 bytes each) but a flat 14 x 4 = 56 bytes in UTF-32.

ord() ord(ch) -> int, e.g. ord('A') is 65 Takes exactly ONE character. ord('AB') raises TypeError.
chr() chr(n) -> str, e.g. chr(97) is 'a' The inverse of ord(). chr(ord(x)) gives back x.
ASCII case gap ord('a') - ord('A') == 32 32 is 100000 in binary, so changing case flips a single bit. chr(ord('P') + 32) gives 'p'.
Digit character to number ord(c) - ord('0') ord('7') is 55, not 7. The character is a picture of the digit, not its value.
Unicode code point range U+0000 to U+10FFFF (1,114,112 code points) Written in hexadecimal after U+. Devanagari occupies U+0900 to U+097F.
UTF-8 byte count 1 byte: U+0000-U+007F ; 2: U+0080-U+07FF ; 3: U+0800-U+FFFF ; 4: U+10000-U+10FFFF bytes · ASCII stays 1 byte, so old English files are already valid UTF-8. Devanagari needs 3.
UTF-32 byte count always 4 bytes per character bytes · len(s.encode('utf-32-be')) == 4 * len(s). Fixed width means instant indexing but 4x the space for English.
Remember
  • An encoding scheme is an agreed number per character; mismatched schemes on save and open produce garbage text.
  • ASCII is 7-bit with 128 codes: 'A'=65, 'a'=97, '0'=48, space=32, and codes 0-31 plus 127 are control characters.
  • ISCII (IS 13194:1991, Bureau of Indian Standards) is 8-bit: 0-127 same as ASCII, 128-255 for Indian scripts, with one shared chart for the Brahmi-derived scripts plus a script-selection mechanism.
  • Unicode replaced ISCII because 128 slots were too few, mixing scripts needed mode switching, non-Brahmi scripts like Urdu did not fit, and it had no international reach; Unicode's Devanagari block was itself based on ISCII.
  • UTF-8 is variable-length (1-4 bytes, ASCII-compatible, Devanagari takes 3); UTF-32 is fixed at 4 bytes per character, simpler to index but wasteful for English.

The formula sheet

Every formula in this chapter, in one place — screenshot it before your exam.

value = d(k) x b^k + ... + d(1) x b^1 + d(0) x b^0
Positional value
0, 1, 2, ... , (b - 1)
Legal digits in base b
n bits -> 2**n values, range 0 to 2**n - 1
Patterns from n bitsbits
0b1010 == 0o12 == 10 == 0xA
Python base literals
8 = 2**3 and 16 = 2**4
Why 8 and 16 were chosen
divide by b repeatedly, collect remainders, read BOTTOM-UP
Decimal to base b (whole part)
multiply the fraction by b repeatedly, collect integer parts, read TOP-DOWN
Decimal to base b (fraction)
bin(n) -> str, e.g. bin(173) is '0b10101101'
bin()
oct(n) -> str, e.g. oct(173) is '0o255'
oct()
hex(n) -> str, e.g. hex(2026) is '0x7ea'
hex()
format(45, '08b') -> '00101101'
format() with a width
value = sum of (digit x b ** place)
Base b to decimal
b**-1, b**-2, b**-3 ...
Places after the point
int(s, base) -> int, e.g. int('3AF', 16) is 943
int() with a base
A=10, B=11, C=12, D=13, E=14, F=15
Hex letter values
int('129', 8) -> ValueError
Illegal digit check
write 128 64 32 16 8 4 2 1 above the bits and add where the bit is 1
Binary shortcut
group bits in 3s from the point outwards, pad left with 0s
Binary to octal
group bits in 4s from the point outwards, pad left with 0s
Binary to hexadecimal
format(int(d, 8), '03b')
Octal digit to bits
format(int(d, 16), '04b')
Hex digit to bits
octal -> binary (3s) -> regroup (4s) -> hexadecimal
Octal to hexadecimal
whole part pads LEFT, fraction pads RIGHT
Padding side
ord(ch) -> int, e.g. ord('A') is 65
ord()
chr(n) -> str, e.g. chr(97) is 'a'
chr()
ord('a') - ord('A') == 32
ASCII case gap
ord(c) - ord('0')
Digit character to number
U+0000 to U+10FFFF (1,114,112 code points)
Unicode code point range
1 byte: U+0000-U+007F ; 2: U+0080-U+07FF ; 3: U+0800-U+FFFF ; 4: U+10000-U+10FFFF
UTF-8 byte countbytes
always 4 bytes per character
UTF-32 byte countbytes

Test yourself

Tap an answer to check it instantly — you'll see why it's right, and what to revise if it isn't.

0 correct · 0/12 answered
Q1

What is the exact output of this code?print(bin(29), oct(29), hex(29))

Q2

What is the exact output of this code?print(int('101', 2) + int('101', 8) + int('101', 16))

Q3

What is the exact output of this code?print(0b1010 + 0o12 + 0xA)

Q4

What is the exact output of this code?print(hex(int('11011', 2)))

Q5

A file contains the single Devanagari character क. What is the output ofprint(len('क'.encode('utf-8')), len('क'.encode('utf-32-be')))

Q6

Convert (110101101)₂ to decimal.

Q7

(3752)₈ is equal to which hexadecimal number?

Q8

Why can a group of exactly 4 bits always be replaced by a single hexadecimal digit?

Q9

If ord('A') is 65 and ord('a') is 97, what does chr(ord('K') + 32) return?

Q10

Which statement about ISCII is correct?

Q11

Which pair of statements about UTF-8 and UTF-32 is correct?

Q12

Which of these decimal fractions has a TERMINATING representation in binary?

NCERT solutions & previous-year questions

Step-by-step model answers — tap a question to reveal the full solution.

NCERT questions 6

1 Convert the following binary numbers into their decimal equivalents: (a) (1010)₂ (b) (101101)₂ (c) (111111)₂ (d) (10110.101)₂Binary to decimal conversion

Method: write the place values above the bits and add the columns where the bit is 1. To the right of the point the place values are 2-1 = 0.5, 2-2 = 0.25, 2-3 = 0.125.

(a) (1010)2
Place values: 8 4 2 1, so 8 + 0 + 2 + 0 = 10

(b) (101101)2
Place values: 32 16 8 4 2 1, so 32 + 0 + 8 + 4 + 0 + 1 = 45

(c) (111111)2
All six bits are 1: 32 + 16 + 8 + 4 + 2 + 1 = 63. Shortcut worth remembering: n bits all set to 1 give 2n - 1, so 26 - 1 = 63.

(d) (10110.101)2
Whole part 10110: 16 + 0 + 4 + 2 + 0 = 22
Fraction .101: 1x0.5 + 0x0.25 + 1x0.125 = 0.625
Total = 22.625

Verified in Python: int('1010',2) gives 10, int('101101',2) gives 45, int('111111',2) gives 63, and 22 + 0.625 gives 22.625.

2 Convert the decimal numbers 45, 128 and 1000 into their binary, octal and hexadecimal equivalents.Decimal to binary, octal and hexadecimal

Use repeated division by the target base and read the remainders bottom-up. Shown fully for 45 into binary and 1000 into hexadecimal, then tabulated.

45 into binary:
45 / 2 = 22 rem 1
22 / 2 = 11 rem 0
11 / 2 = 5 rem 1
5 / 2 = 2 rem 1
2 / 2 = 1 rem 0
1 / 2 = 0 rem 1
Reading upwards: (101101)2

1000 into hexadecimal:
1000 / 16 = 62 rem 8
62 / 16 = 3 rem 14 = E
3 / 16 = 0 rem 3
Reading upwards: (3E8)16

DecimalBinaryOctalHexadecimal
45101101552D
1281000000020080
1000111110100017503E8

Notice 128 is 27, so its binary form is a single 1 followed by seven 0s - a useful sanity check whenever the number is a power of 2.

Verified with Python: bin(45), oct(45), hex(45) give 0b101101 0o55 0x2d; bin(128), oct(128), hex(128) give 0b10000000 0o200 0x80; bin(1000), oct(1000), hex(1000) give 0b1111101000 0o1750 0x3e8.

3 Convert the octal numbers (236)₈ and (705)₈ into their binary and hexadecimal equivalents.Octal to binary and hexadecimal

Octal to hexadecimal has no direct route, so go through binary. Expand each octal digit into exactly 3 bits, then regroup the same bits in 4s from the right.

(236)8
2 becomes 010, 3 becomes 011, 6 becomes 110
Binary: 010011110, that is (10011110)2 after dropping the leading zero.
Regroup in 4s: 1001 1110, giving 9 and E, so (9E)16
Check in decimal: 9x16 + 14 = 158, and 2x64 + 3x8 + 6 = 158. Match.

(705)8
7 becomes 111, 0 becomes 000, 5 becomes 101
Binary: (111000101)2
Nine bits are not a multiple of 4, so pad three zeros on the left to reach 12 bits and regroup in 4s: 0001 1100 0101, giving 1, C, 5, so (1C5)16
Check in decimal: 1x256 + 12x16 + 5 = 453, and 7x64 + 0x8 + 5 = 453. Match.

Verified in Python: int('236',8) = 158 = int('9E',16), and int('705',8) = 453 = int('1C5',16).

4 Convert the hexadecimal numbers (2A)₁₆, (1FF)₁₆ and (ABC)₁₆ into decimal and binary.Hexadecimal to decimal and binary

For decimal, expand with powers of 16 after replacing letters by values (A=10, B=11, C=12, D=13, E=14, F=15). For binary, replace each hex digit by exactly 4 bits.

(2A)16
Decimal: 2x16 + 10x1 = 32 + 10 = 42
Binary: 2 becomes 0010, A becomes 1010, so (101010)2 after dropping leading zeros.

(1FF)16
Decimal: 1x256 + 15x16 + 15x1 = 256 + 240 + 15 = 511
Binary: 1 becomes 0001, F becomes 1111, F becomes 1111, so (111111111)2 - nine 1s, which is 29 - 1 = 511. Consistent.

(ABC)16
Decimal: 10x256 + 11x16 + 12x1 = 2560 + 176 + 12 = 2748
Binary: A becomes 1010, B becomes 1011, C becomes 1100, so (101010111100)2

Verified in Python: int('2A',16) = 42, int('1FF',16) = 511, int('ABC',16) = 2748, and int('101010111100',2) = 2748.

5 What is the difference between ASCII and Unicode?Encoding schemes

Both are encoding schemes - agreed tables of which number stands for which character - but they were built for different sized problems.

PointASCIIUnicode
Size of the code7-bit, 128 codes (0 to 127)Code points U+0000 to U+10FFFF, 1,114,112 possible
CoverageEnglish letters, digits, punctuation, control characters onlyEvery script in the world - Devanagari, Bengali, Tamil, Arabic, Chinese, plus symbols such as ₹
Storage7 bits, normally held in one byte per character with the eighth bit spareDepends on the encoding: UTF-8 uses 1 to 4 bytes, UTF-32 uses a fixed 4
Mixing languagesImpossible - there is no code for क at allNatural, since every character has its own permanent code point

The key relationship: Unicode did not discard ASCII. The first 128 Unicode code points are exactly the ASCII codes, so 'A' is 65 in both. This is why a plain English text file written decades ago is still readable today as UTF-8 without any conversion.

Demonstration: ord('A') returns 65, the ASCII value, while ord('क') returns 2325, which is U+0915 - a code point that simply does not exist in ASCII.

6 What is ISCII? Why was it created, and why has Unicode replaced it?ISCII and Unicode

What it is. ISCII stands for Indian Script Code for Information Interchange. It was standardised by the Bureau of Indian Standards as IS 13194:1991. It is an 8-bit encoding scheme: codes 0 to 127 are identical to ASCII, so English text works unchanged, and codes 128 to 255 are used for Indian script characters.

Why it was created. ASCII has 128 codes and every one of them is spoken for by English. There was no code for क, ম, த or ಕ, so an Indian-language document simply could not be stored or exchanged. India needed a national standard that fitted the 8-bit hardware of the time.

Its clever design. Devanagari, Bengali, Gurmukhi, Gujarati, Odia, Tamil, Telugu, Kannada and Malayalam all descend from Brahmi and share the same phonetic structure - the same consonants and vowels in the same order. ISCII therefore uses a single common code chart based on Devanagari, and a separate script-selection mechanism tells the display which script to render. The same stored code can appear as क or as ক depending on the selected script, which let one set of codes serve many languages.

Why Unicode replaced it.

  • Too few codes. Only the upper half of an 8-bit code, 128 slots, was available for the whole of India's writing.
  • Script is a mode, not an identity. Because the script is chosen separately, freely mixing Hindi and Tamil in one document is awkward, and losing one mode-switch code corrupts every character after it.
  • Non-Brahmi scripts do not fit. The Perso-Arabic script used for Urdu does not share the common phonetic chart.
  • No international reach. A file needing Hindi and Chinese together was still impossible, and software outside India had no reason to implement ISCII.

Unicode fixes all four by giving every character in every script its own permanent code point, so क is U+0915 and ব is U+09AC - different characters, no mode switching, and one file can hold any mixture of scripts. Importantly, ISCII's work was not wasted: Unicode's Devanagari block, U+0900 to U+097F, was based on the ISCII layout, so the Indian standard's design lives on inside the global one.

Previous-year board questions 4

Q1 Convert (1011.011)₂ into its decimal equivalent, showing the working. Board pattern - 2 marks

Number the places from the point: 3, 2, 1, 0 going left and -1, -2, -3 going right.

Digit1011.011
Place3210-1-2-3
Place value84210.50.250.125

Whole part: 1x8 + 0x4 + 1x2 + 1x1 = 8 + 0 + 2 + 1 = 11

Fractional part: 0x0.5 + 1x0.25 + 1x0.125 = 0 + 0.25 + 0.125 = 0.375

Answer: (1011.011)2 = (11.375)10

Marking note: the commonest error is numbering the first place after the point as 0 instead of -1, which shifts the whole fraction and gives 0.75 instead of 0.375. Always start at -1.

Verified in Python: int('1011', 2) + 0*2**-1 + 1*2**-2 + 1*2**-3 evaluates to 11.375.

Q2 Convert (256)₁₀ into its hexadecimal and octal equivalents. Board pattern - 2 marks

Into hexadecimal, divide repeatedly by 16:

DivisionQuotientRemainder
256 / 16160
16 / 1610
1 / 1601

Reading the remainders bottom-up: (100)16

Into octal, divide repeatedly by 8:

DivisionQuotientRemainder
256 / 8320
32 / 840
4 / 804

Reading bottom-up: (400)8

Check: (100)16 = 1x162 = 256. (400)8 = 4x82 = 4x64 = 256. Both correct.

Warning: (100)16 is not one hundred. The same three symbols mean 256 in base 16 and 64 in base 8. Always write the base as a subscript in your answer, or you lose the mark.

Verified in Python: hex(256) gives '0x100' and oct(256) gives '0o400'.

Q3 Convert (A2B)₁₆ into its binary and octal equivalents. Board pattern - 3 marks

Step 1, hexadecimal to binary. Replace each hex digit with exactly 4 bits (16 = 24):

  • A = 10 gives 1010
  • 2 = 2 gives 0010
  • B = 11 gives 1011

Joining them: (101000101011)2

Step 2, regroup the same bits in 3s from the right for octal (8 = 23). Twelve bits divide evenly into four groups:

101 | 000 | 101 | 011 gives 5, 0, 5, 3

(5053)8

Check via decimal:
(A2B)16 = 10x256 + 2x16 + 11 = 2560 + 32 + 11 = 2603
(5053)8 = 5x512 + 0x64 + 5x8 + 3 = 2560 + 0 + 40 + 3 = 2603. Match.

Marking note: full marks require the binary middle step to be shown. There is no direct hexadecimal-to-octal conversion, because although 8 and 16 are both powers of 2, neither is a power of the other, so 4-bit groups never line up with 3-bit groups.

Verified in Python: int('A2B',16) = 2603 = int('5053',8) = int('101000101011',2).

Q4 Differentiate between UTF-8 and UTF-32 with a suitable example. Which one would you choose for a website that shows content in both English and Hindi, and why? Board pattern - 3 marks

Difference. Both are ways of storing Unicode code points as bytes. A code point such as U+0915 is only the character's identity; the UTF decides the actual bytes written to disk.

PointUTF-8UTF-32
LengthVariable: 1 to 4 bytes per characterFixed: always 4 bytes per character
ASCII characters1 byte, byte-identical to ASCII4 bytes, with three leading zero bytes
Devanagari3 bytes (U+0800 to U+FFFF band)4 bytes
AdvantageCompact; backward compatible with ASCIIFinding the nth character is instant arithmetic
DisadvantageCharacter positions must be scanned, not calculatedWastes space, 4x for plain English

Example. The string "Namaste नमस्ते" has 14 characters, 8 English (including the space) and 6 Devanagari.

  • UTF-8: 8x1 + 6x3 = 26 bytes
  • UTF-32: 14x4 = 56 bytes

Confirmed by running len(msg.encode('utf-8')) which gives 26, and len(msg.encode('utf-32-be')) which gives 56.

The individual bytes make the difference visible. For क: UTF-32 stores 00 00 09 15, which is U+0915 written out directly. UTF-8 stores e0 a4 95, spreading the same code point across three bytes with marker bits.

Choice for a bilingual website: UTF-8. Two reasons. First, size - every page is sent over the network on every visit, and UTF-8 is well under half the size of UTF-32 for mixed English and Hindi text, so pages load faster on a slow mobile connection. Second, compatibility - HTML tags, CSS, JavaScript and URLs are all ASCII, and UTF-8 stores those in 1 byte each, byte-identical to ASCII, so older tools and servers handle the file correctly. UTF-32's only real advantage, instant indexing by character position, does not matter to a web page that is simply displayed.

Part of Priodemy for School

Interactive CBSE lessons, Class 8–12 — free with every school on Priodemy EduSuite. Explore more chapters and labs on the Priodemy for School hub.

Ask AI