The Problem Git Solves
Before you learn a single command, be clear about the problem. You are building a college project and it finally works. You try a new feature, and by midnight nothing works and you cannot remember which of the eleven files you touched. So you do what almost everybody does the first time: you copy the whole folder and call it project-backup, then project-final, then project-final-REAL. A week later you have six folders, no idea which one actually runs, and no way to tell what is different between any two of them.
Now add a teammate. You send them a ZIP. They edit style.css; so do you. When their copy comes back, one of you has to merge the two versions by hand, line by line. With three people this stops being annoying and becomes impossible — and it is exactly the point in a group project where someone's work quietly gets overwritten and nobody notices until the demo.
Version control is the category of software built to solve this. It records every change you make, who made it, when, and why. It lets you return to any earlier state, work on two ideas at once without them colliding, and combine several people's work with the machine doing the tedious comparison instead of you. Git is the version control system essentially the whole industry has settled on, which is why almost every internship description mentions it.
- A dated history of every change, so "it worked yesterday" is something you can go back to
- The ability to try a risky idea in isolation and throw it away cleanly if it fails
- Real answers to "who wrote this line and why", months later
- Automatic merging of several people's edits, with a clear report when two genuinely clash
- One shared source of truth, so nobody has to ask which copy is current
What Git Actually Is
Git is a free, open-source program that runs on your own computer. Linus Torvalds wrote it in 2005 because the Linux kernel project needed something faster and more reliable than what existed at the time. You install it once, and from then on it is available in your terminal as the command git.
The word that matters in "distributed version control system" is distributed. In older systems the entire history lived on one central server; if the server was down or your Wi-Fi died, you could not commit, see history or compare versions. Git works the other way round. When you copy a project from GitHub, you receive the whole history onto your own disk — every commit, every branch, from the first day of the project.
That one design choice explains most of what feels strange about Git at the start. Committing, branching and inspecting history are all local, so they are instant and work with no internet. But it also means committing changes nothing on GitHub. Sending work to a server is a separate step called pushing, and students lose weeks to this misunderstanding: they commit faithfully for a fortnight, never push, and then the laptop dies.
- Local and offline — commit, branch and inspect history on a train with no network
- Distributed — every clone contains the full history, so every teammate's copy is effectively a backup
- Fast — because nothing has to travel over a network for the everyday commands
- Content-addressed — every commit is identified by a hash of what it contains, so history cannot be quietly altered
- Free and open source, and available for Windows, macOS and Linux
- Git does one job well: tracking text that changes. It is a poor fit for very large binary files such as video or design files, because it cannot compare two versions meaningfully and stores each version in full. Keep those outside the repository, or use Git LFS, which is built for that case.
Snapshots, Not a List of Edits
Some version control systems store history as a list of changes per file: this line was added to index.html, that line was removed from app.js. Git takes a different view. Each time you commit, it records a snapshot of the entire project as it looked at that instant.
That sounds wasteful, and it would be if Git stored every file again each time. It does not — a file that has not changed is stored as a reference to content Git already has. So a snapshot of a two-hundred-file project where you edited one file is barely larger than that one file. This model is why switching branches is close to instant, and why every commit is a complete working state on its own rather than a fragment that only makes sense after the twelve before it.
Every commit is identified by a long hexadecimal hash such as 9f4c2a1b8e..., computed from the commit's content, its author and its parent. In practice you type only the first seven or so characters. Because the hash depends on the parent, changing an old commit changes the identity of every commit after it. Hold on to that fact — it is why rewriting history that other people already have causes so much pain later in this course.
# Your project history is a chain of snapshots.
# Each commit points back at the one before it.
# 9f4c2a1 3b7e0d4 a81c9f2
# (oldest) <-- (middle) <-- (newest)
# "Initial "Add login "Fix header
# commit" form" spacing"
#
# ^
# |
# main branch
# Ask Git to show you exactly this, in one line per commit:
git log --oneline
# a81c9f2 Fix header spacing
# 3b7e0d4 Add login form
# 9f4c2a1 Initial commit The Three Areas, and Why Staging Exists
Git keeps your work in three places, and almost every confusing error message becomes readable once you can say which of the three it is about. The working directory is the ordinary folder on your disk. The staging area, also called the index, is a holding zone where you assemble the exact set of changes your next commit will contain. The repository is the hidden .git folder, where committed snapshots live permanently.
Beginners nearly always find the staging area pointless. Why not just commit and be done? Because a good commit describes one logical change, and real work is never that tidy. You sat down to fix a broken login button; along the way you corrected a spelling mistake in the README and reformatted a function that was bothering you. Staging lets you commit the login fix by itself, with a message that is actually true, and commit the README fix separately.
The second reason is that staging gives you a review step. Once changes are staged, git diff --staged shows precisely what is about to be written into history. That is the moment you notice the console.log(password) you left in, or the .env file full of API keys you were about to publish to the internet. Reading that diff before every commit will save you more trouble than any other habit in this course.
# The journey of one change:
#
# working directory --git add--> staging area --git commit--> repository
# (you edit files) (chosen changes) (permanent snapshot)
#
# And the commands that show you each gap:
git status # a summary of all three at once
git diff # working directory vs staging area (not yet added)
git diff --staged # staging area vs last commit (what a commit would record) - A file in a Git repository is always in one of four states. Untracked means Git has never seen it. Modified means it is tracked and you have changed it since the last commit. Staged means the current version of it is queued for the next commit. Committed means it is safely stored in the repository.
git statusexists to tell you which files are in which state, and it is the command you should run most often.
Commits, HEAD and Branches
A commit stores five things: the snapshot, the author's name and email, a timestamp, the message you wrote, and a pointer to the commit before it. Follow those pointers backwards and you have the project's entire history. That is genuinely all a Git history is.
A branch sounds like a heavy concept, but in Git it is astonishingly light: a movable label pointing at one commit. Commit while on a branch and the label slides forward. Creating a branch writes one small file containing a hash, which is why it is instant even on a huge project — and why you should feel free to branch for the smallest experiment. HEAD is a second label, pointing at the branch you currently have checked out, and almost every command acts relative to it.
New repositories name their first branch main. Older repositories, and Git installations that were never configured, use master. They are ordinary names with no special powers; you will meet both, and a project can rename one to the other at any time. This course uses main throughout.
git init # start tracking the current folder
git status # what has changed?
git add . # stage everything that has changed
git commit -m "Add homepage layout"
git log --oneline # see the history you have built
# Where am I?
git branch --show-current # main
# HEAD-relative shorthand you will use constantly:
# HEAD the commit you are sitting on
# HEAD~1 one commit before that
# HEAD~3 three commits before that Git Is Not GitHub
These two names get used interchangeably in conversation, and keeping them apart will save you real confusion. Git is the program on your computer that tracks your project's history. GitHub is a website that stores copies of Git repositories and wraps them in collaboration features — code review, issue tracking, automated testing. GitLab, Bitbucket and Codeberg do the same job; GitHub is simply the most widely used.
You can use Git for years without ever creating a GitHub account, and everything in the first half of this course works entirely on your own machine. GitHub becomes valuable when you want an off-site copy of your work, want to collaborate with people who are not sitting next to you, or want a public record of what you have built.
The practical consequence is worth stating plainly one more time: git commit writes to your laptop. Only git push sends anything to GitHub. If you have committed a hundred times and never pushed, GitHub knows nothing about your project and your work exists in exactly one place.
- Git — the version control program, installed locally, works offline
- GitHub — a hosting service for Git repositories, with review and automation on top
git commit— saves a snapshot on your machinegit push— copies your commits to GitHubgit pull— brings other people's commits into your copy
- Work along in a real folder rather than only reading. Git is a set of habits more than a set of facts, and the commands stop feeling arbitrary once you have used them on a project you actually care about.
