Lesson 10 of 15

Cloning & Forking

Cloning: Getting a Repository onto Your Machine

git clone copies a repository from a remote onto your disk. The word "copy" is doing more work than it looks: you do not receive the current files, you receive the whole repository — every commit, every branch, the entire history from the first day. That is the distributed model in action, and it is why you can clone a project on college Wi-Fi and then read its complete history on a train with no signal.

Cloning also does three setup jobs for you, which is why it is easier than git init plus manual configuration. It creates the folder, registers the source as a remote called origin, and checks out the default branch with tracking already configured. From the moment the clone finishes, git push and git pull know where to go.

Pick the URL form that matches how you authenticate. HTTPS URLs start with https:// and need a personal access token when you push. SSH URLs start with git@github.com: and use the key pair you set up earlier. Cloning a public repository over HTTPS needs no authentication at all — you only need credentials to push. If you clone with the wrong form, you do not have to start again; git remote set-url changes it in place.

One habit worth building: never download a project as a ZIP from GitHub when you intend to work on it. A ZIP contains the files with no .git folder, which means no history, no branches, and no way to push. It looks like the same thing and is not.

Example
# HTTPS — no setup needed to read a public repository
git clone https://github.com/ananya/fest-website.git

# SSH — uses your key, best if you will be pushing
git clone git@github.com:ananya/fest-website.git

# Clone into a folder with a different name
git clone git@github.com:ananya/fest-website.git my-copy

# Only the latest commit — much faster on a huge project,
# but you get no history and limited branch operations
git clone --depth 1 https://github.com/some/large-project.git

# What clone set up for you
cd fest-website
git remote -v          # origin, already configured
git branch -vv         # main is tracking origin/main
git log --oneline      # the entire history is already here
Notes
  • Do not clone one repository inside another. If you clone into a folder that is already tracked by Git, the outer repository sees a strange half-tracked directory and behaves confusingly. Keep a plain projects folder that is not itself a repository, and clone into that.

Fork Versus Clone: Two Different Things

These two get confused constantly, and the difference is simple once stated: a clone is a copy on your computer; a fork is a copy on GitHub, under your account. They happen in different places and solve different problems.

Forking exists because of permissions. You cannot push to a repository you do not have write access to — you cannot push to React's repository, and you cannot push to a senior's project you found interesting. Forking gives you a repository you do own, containing the same history, which you can push to freely. You then ask the original owner to take your changes with a pull request.

The usual sequence is therefore fork then clone: fork on the GitHub website to get your own copy, clone your fork to get it onto your machine, work locally, push to your fork, and open a pull request against the original. Your fork is the middle step that makes the whole thing possible.

The mistake to avoid is forking when you do not need to. If you are on a four-person team project and everyone has write access to the same repository, forking is the wrong tool — you should all clone the one repository and use branches. Four forks means four separate repositories and a great deal of unnecessary synchronising. Fork when you lack write access; branch when you have it.

  • Clone — GitHub to your laptop; you can push only if you have write access
  • Fork — GitHub to GitHub, into your own account; you always have write access to your fork
  • Contributing to a project you do not own → fork, then clone your fork
  • Working on your own team's repository → clone it directly and use branches
  • A fork keeps a link to the original, so GitHub can offer to sync it and can open pull requests back to it
  • Forks of public repositories are public, and the fact that you forked is visible on the original repository
Notes
  • Forking is also a reasonable way to keep a snapshot of someone's public project you find useful, or to experiment with a library. Nobody is notified in an intrusive way, and nothing about the original changes.

origin and upstream: Keeping a Fork Current

A fork is a snapshot, and the moment you make it, it starts going out of date. The original project keeps receiving commits; your fork does not receive them automatically. Work for two weeks on a stale fork and your pull request will conflict with changes you never saw.

The standard solution is a second remote. By convention origin is your fork — the one you can push to — and upstream is the original project, which you can only read from. You add upstream yourself after cloning, because cloning only sets up origin.

Syncing then has a clear shape. Fetch from upstream to learn what the original project has done, merge upstream/main into your local main, and push the result to origin so your fork on GitHub is current too. That last push is the step people forget, which is why their fork's front page still says it is behind by twenty commits.

Do this before starting each new piece of work, and always branch off an updated main rather than working on main itself. Keeping your fork's main as a clean mirror of the original makes syncing a fast-forward every time, which means it never conflicts. GitHub also offers a Sync fork button on the web interface that does the same job for simple cases.

Example
# 1. Fork on GitHub (the Fork button), then clone YOUR fork
git clone git@github.com:ananya/awesome-project.git
cd awesome-project

# 2. Add the original project as 'upstream' (read-only for you)
git remote add upstream https://github.com/original-owner/awesome-project.git
git remote -v
# origin    git@github.com:ananya/awesome-project.git (fetch/push)   <- yours
# upstream  https://github.com/original-owner/... (fetch/push)       <- theirs

# 3. Sync before starting anything new
git fetch upstream
git switch main
git merge upstream/main        # usually a clean fast-forward
git push origin main           # update your fork on GitHub too

# 4. Branch, work, push to YOUR fork
git switch -c fix/typo-in-docs
git commit -am "Fix typo in installation guide"
git push -u origin fix/typo-in-docs

# 5. Open a pull request from your branch to the original project
Notes
  • If git push on a fork fails with a permission error mentioning the original owner's name, you cloned the original repository instead of your fork. Check git remote -v — origin should contain your username. Fix it with git remote set-url origin rather than starting over.

Your First Open Source Contribution

The fork workflow is how essentially all open source contribution happens, and it is more approachable than students expect. You do not need to be an expert to be useful. Documentation that is out of date, a broken link, an error message that does not explain itself, a missing example — these are real problems, and fixing one teaches you the entire workflow with low stakes.

Before you start, read the project's CONTRIBUTING.md if it has one. It will tell you the branch naming they expect, how they want commit messages written, and whether they want an issue opened before a pull request. Ignoring it is the fastest way to have an otherwise good contribution rejected. Many projects also label beginner-friendly issues with tags such as good first issue.

Keep the first contribution small and single-purpose. A pull request that fixes one typo gets merged in a day. A pull request that fixes a typo, refactors three files and changes the build configuration will sit unreviewed for weeks, because reviewing it is a large job for a volunteer.

Expect review comments and do not read them as rejection. Being asked to rename a variable or add a test is the normal shape of a contribution, and responding to that feedback well is exactly the skill the exercise is meant to build.

  • Find a project you actually use, and an issue labelled for newcomers
  • Read CONTRIBUTING.md and the code of conduct before writing anything
  • Fork, clone your fork, add upstream, and sync
  • Branch with a descriptive name — never work directly on main in your fork
  • Make one focused change, and run the project's tests if it has them
  • Push to your fork and open a pull request describing what and why
  • Respond to review comments by pushing more commits to the same branch — the pull request updates itself
Notes
  • Contributions show on your GitHub profile, and a merged pull request into a project other people use is far more convincing to a recruiter than another tutorial project. One real, small, merged fix is worth more than ten half-finished clones of a to-do app.
Ask AI