Lesson 12 of 20

File Handling

Reading and Writing Whole Files

For small files, PHP gives you two functions that do the whole job in one call. file_get_contents() reads a file into a string, and file_put_contents() writes a string to a file, creating it if it does not exist and replacing it entirely if it does. That last part is worth pausing on: writing without the append flag destroys the previous contents, with no warning and no undo.

Adding FILE_APPEND makes it add to the end instead, which is what a log file wants. Adding LOCK_EX asks PHP to take an exclusive lock while writing, which matters as soon as two requests might write at the same moment — without it, two simultaneous appends can interleave and produce a corrupted line. On a web server, "two requests at the same moment" is the normal case, not an edge case.

file() is a variation that returns an array of lines rather than one string, which is convenient for anything line-oriented. Pass FILE_IGNORE_NEW_LINES to strip the line endings and FILE_SKIP_EMPTY_LINES to drop blank ones, or you will be trimming every element yourself.

All of these return false on failure, and failure is common — a missing file, a folder you cannot write to, a typo in the path. Check the return value. Ignoring it produces the classic symptom of a script that reports success while nothing was actually saved.

Example
<?php
$path = __DIR__ . '/data/notes.txt';

// Write - REPLACES everything already there
file_put_contents($path, "First line\n");

// Append, with a lock so concurrent requests do not interleave
file_put_contents($path, "Second line\n", FILE_APPEND | LOCK_EX);

// Read the whole file
$content = file_get_contents($path);
if ($content === false) {
    exit('Could not read the notes file.');
}
echo nl2br(htmlspecialchars($content, ENT_QUOTES, 'UTF-8'));

// Read as an array of lines
$lines = file($path, FILE_IGNORE_NEW_LINES | FILE_SKIP_EMPTY_LINES);
foreach ($lines as $i => $line) {
    echo ($i + 1) . ': ' . htmlspecialchars($line) . '<br>';
}

// Always check before you assume
if (!is_readable($path)) {
    exit('File is missing or not readable.');
}
echo filesize($path) . ' bytes';

// Other everyday operations
// copy($path, $path . '.bak');
// rename($path, __DIR__ . '/data/archive.txt');
// unlink($path);          // delete - there is no recycle bin
Notes
  • file_get_contents() loads the entire file into memory at once. A 500 MB log file will exhaust memory_limit and kill the request. Use it for configuration, small data files and templates; use the streaming approach in the next section for anything that could grow.

Line by Line, and CSV

When a file is too large to hold in memory, open it and read a piece at a time. fopen() returns a handle, fgets() reads one line, and fclose() releases the handle when you are done. Only one line is ever in memory, so the file's size stops mattering.

The second argument to fopen() is the mode, and choosing the wrong one is destructive. 'r' is read-only. 'a' appends, creating the file if needed. 'w' opens for writing and truncates the file to zero bytes immediately — before you have written anything. Reaching for 'w' when you meant 'a' is how a log file disappears.

CSV deserves its own mention because it is the format everyone actually exchanges data in — attendance sheets, marks lists, exports from a college ERP. fgetcsv() reads one row and hands you an array of fields, correctly handling quoted values that contain commas. Writing your own splitter with explode(',', $line) looks fine until the first address field contains a comma, and then it silently shifts every column after it.

fputcsv() is the counterpart for generating a download. Combined with the right headers, you can stream a CSV straight to the browser using the special handle php://output, without ever creating a file on disk.

One practical detail when importing a CSV that a spreadsheet produced: the first row is usually a header. Read it separately and use it to build associative rows with array_combine(), so the rest of your code refers to $row['email'] rather than $row[3]. Column positions change; names usually do not.

Example
<?php
// Streaming: memory usage stays flat regardless of file size
$handle = fopen(__DIR__ . '/logs/access.log', 'r');
if ($handle === false) {
    exit('Cannot open the log file.');
}
while (($line = fgets($handle)) !== false) {
    if (str_contains($line, 'ERROR')) {
        echo htmlspecialchars($line) . '<br>';
    }
}
fclose($handle);

// Importing a CSV with a header row
$in = fopen(__DIR__ . '/uploads/students.csv', 'r');
$header = fgetcsv($in);          // ['roll', 'name', 'email']

$stmt = $pdo->prepare('INSERT INTO students (roll, name, email) VALUES (?, ?, ?)');
while (($cells = fgetcsv($in)) !== false) {
    if (count($cells) !== count($header)) {
        continue;                // skip malformed rows
    }
    $row = array_combine($header, $cells);
    $stmt->execute([$row['roll'], trim($row['name']), trim($row['email'])]);
}
fclose($in);

// Generating a CSV download without writing to disk
header('Content-Type: text/csv; charset=utf-8');
header('Content-Disposition: attachment; filename="marks.csv"');

$out = fopen('php://output', 'w');
fputcsv($out, ['Roll', 'Name', 'Marks']);
foreach ($rows as $r) {
    fputcsv($out, [$r['roll'], $r['name'], $r['marks']]);
}
fclose($out);
exit;
Notes
  • Importing thousands of rows one execute() at a time is slow because each insert is its own transaction. Wrap the whole loop in $pdo->beginTransaction() and $pdo->commit() and the same import can finish many times faster — and if a row fails halfway, you can roll the whole thing back instead of leaving a half-imported mess.

Paths and Permissions

A relative path like 'data/notes.txt' is resolved against the current working directory, which is the directory of the script the browser requested — not the directory of the file the code happens to live in. So a helper in lib/storage.php that opens 'data/notes.txt' works when called from index.php in the project root and breaks when called from admin/users.php. This is why paths mysteriously stop working after you move a file.

Always build paths from __DIR__, the magic constant holding the folder of the file it appears in. __DIR__ . '/../data/notes.txt' means the same thing no matter which script was requested. This is the same reasoning that applies to require, and it is one of the highest-value habits in this whole lesson.

Permission errors are the other common cause of failure, especially on Linux hosting. The web server runs as its own user, and if that user cannot write to your folder you get "failed to open stream: Permission denied". Create your writable folders deliberately and check with is_writable() rather than guessing. mkdir() with the third argument set to true creates parent folders as needed.

One structural decision worth making early: keep uploaded and generated files outside your public folder wherever you can, and serve them through a PHP script that checks permissions first. A file sitting in the public folder is downloadable by anyone who guesses its URL, and "nobody will guess it" is not access control.

Example
<?php
// Fragile - depends on which script was requested
// $data = file_get_contents('data/settings.json');

// Reliable - relative to THIS file, always
$dataDir = __DIR__ . '/../storage';
$file    = $dataDir . '/settings.json';

// Create the folder if it is missing (with parents)
if (!is_dir($dataDir) && !mkdir($dataDir, 0775, true) && !is_dir($dataDir)) {
    exit('Could not create the storage folder.');
}

if (!is_writable($dataDir)) {
    exit('The storage folder is not writable by the web server.');
}

// Inspecting a path
echo pathinfo($file, PATHINFO_EXTENSION);   // json
echo basename($file);                        // settings.json
echo dirname($file);                         // the folder

// Listing files
foreach (glob($dataDir . '/*.json') as $found) {
    echo basename($found) . '<br>';
}

// A recommended layout
// project/
//   public/        <- the web root; only this is reachable by URL
//     index.php
//   src/
//   storage/       <- uploads, logs, caches: NOT reachable by URL
//   .env           <- credentials, never in public/
Notes
  • Use forward slashes in paths even on Windows — PHP accepts them everywhere, and code written with backslashes breaks the moment it is deployed to a Linux host, which is where almost all PHP eventually runs.

Never Build a Path From User Input

A page that serves a file named in the URL — download.php?file=notes.pdf — looks harmless and is one of the classic ways sites get broken into. The attack is called path traversal, and it works by sending ../ segments to climb out of your folder: ?file=../../config.php hands over your database password.

Stripping ../ from the string is not a fix. Attackers encode it, double it, or nest it so that removing one copy reveals another. Any defence that consists of deleting dangerous-looking characters from a path is a defence you will eventually lose.

There are two approaches that do work. The stronger one is to never accept a path at all: accept an id, look the record up in your database, and read the stored filename from there. The user never influences the path, so there is nothing to traverse. This is the right design for anything with access rules attached.

Where you genuinely must accept a name, combine two checks. Run it through basename(), which discards every directory component and leaves only the final name. Then resolve the full path with realpath() and confirm the result is still inside your intended folder — realpath() collapses all the .. segments and symbolic links, so the comparison is on the real location rather than on the text of the request.

The same discipline applies to include and require. Including a filename derived from user input is even more serious than reading one, because the file's contents get executed. If you are choosing which page to load from a URL parameter, map the parameter through an allow-list of known page names, never straight into a path.

Example
<?php
// DANGEROUS
// $file = $_GET['file'];
// readfile(__DIR__ . '/uploads/' . $file);
//   ?file=../../.env  ->  your credentials

// SAFEST: accept an id, look up the real path yourself
$id = filter_input(INPUT_GET, 'id', FILTER_VALIDATE_INT);
if ($id === false || $id === null) {
    http_response_code(400);
    exit('Invalid request.');
}

$stmt = $pdo->prepare('SELECT stored_name, original_name FROM documents WHERE id = ? AND user_id = ?');
$stmt->execute([$id, $_SESSION['user_id'] ?? 0]);
$doc = $stmt->fetch();

if (!$doc) {
    http_response_code(404);
    exit('Not found.');
}

$path = __DIR__ . '/../storage/uploads/' . $doc['stored_name'];

// If you must accept a name: basename + realpath containment
$baseDir = realpath(__DIR__ . '/../storage/uploads');
$wanted  = realpath($baseDir . '/' . basename($_GET['file'] ?? ''));

if ($wanted === false || !str_starts_with($wanted, $baseDir . DIRECTORY_SEPARATOR)) {
    http_response_code(403);
    exit('Forbidden.');
}

// Choosing a page: allow-list, never a raw path
$pages = ['home' => 'home.php', 'about' => 'about.php'];
$key   = $_GET['page'] ?? 'home';
require __DIR__ . '/views/' . ($pages[$key] ?? $pages['home']);
Notes
  • Notice the AND user_id = ? in that query. Checking that the document belongs to the person asking for it is a separate control from path safety, and forgetting it produces a different bug: anyone who increments the id in the URL can read everyone else's files.

File Uploads: Validate the Content, Not the Name

An upload arrives in $_FILES, not $_POST, and the form must carry enctype="multipart/form-data" or the file simply will not be sent. Each entry gives you the original name, a client-supplied type, a temporary path in tmp_name, a size, and an error code.

Check error first, and check it against UPLOAD_ERR_OK rather than assuming success. The other codes tell you useful things — that the file exceeded the server limit, that only part of it arrived, or that no file was chosen. Reporting the right one saves your users from guessing.

Now the central point of this section. The filename and the reported type both come from the client and mean nothing. Checking that a name ends in .jpg proves nothing, because anyone can rename a file. Checking $_FILES['photo']['type'] proves even less, because that string is sent by the browser and can be set to whatever the attacker likes. A PHP script renamed to photo.jpg passes both checks.

Inspect the actual bytes instead. The finfo extension examines the file's content and reports its real type, and getimagesize() returns false for anything that is not a genuine image. Use those to decide, and then choose the extension yourself from an allow-list based on what you detected — never from what the user sent.

Two more rules complete the picture. Move the file with move_uploaded_file(), which verifies that the source really was an upload for this request rather than an arbitrary path. And give it a new random name in a folder where scripts cannot be executed, ideally outside the web root entirely. The worst-case upload bug is one where an attacker uploads a PHP file and then simply requests its URL; if uploaded files can never be executed, that whole class of attack disappears.

Example
<?php
// <form method="post" enctype="multipart/form-data">
//   <input type="file" name="photo" accept="image/*">

$file = $_FILES['photo'] ?? null;

if ($file === null || $file['error'] === UPLOAD_ERR_NO_FILE) {
    exit('Please choose a file.');
}

if ($file['error'] !== UPLOAD_ERR_OK) {
    $messages = [
        UPLOAD_ERR_INI_SIZE   => 'The file is larger than the server allows.',
        UPLOAD_ERR_FORM_SIZE  => 'The file is larger than the form allows.',
        UPLOAD_ERR_PARTIAL    => 'The upload was interrupted. Please try again.',
        UPLOAD_ERR_NO_TMP_DIR => 'Server misconfiguration: no temp folder.',
        UPLOAD_ERR_CANT_WRITE => 'Server could not write the file to disk.',
    ];
    exit($messages[$file['error']] ?? 'Upload failed.');
}

// Size limit enforced in PHP too, not only in php.ini
if ($file['size'] > 2 * 1024 * 1024) {
    exit('Images must be 2 MB or smaller.');
}

// Detect the REAL type from the file's bytes
$finfo = new finfo(FILEINFO_MIME_TYPE);
$mime  = $finfo->file($file['tmp_name']);

$allowed = [
    'image/jpeg' => 'jpg',
    'image/png'  => 'png',
    'image/webp' => 'webp',
];

if (!isset($allowed[$mime]) || getimagesize($file['tmp_name']) === false) {
    exit('Only JPG, PNG or WebP images are accepted.');
}

// New random name; extension chosen by US, from the detected type
$newName = bin2hex(random_bytes(16)) . '.' . $allowed[$mime];
$target  = __DIR__ . '/../storage/uploads/' . $newName;

if (!move_uploaded_file($file['tmp_name'], $target)) {
    exit('Could not save the uploaded file.');
}

// Store the original name for display, the new name for retrieval
$stmt = $pdo->prepare(
    'INSERT INTO documents (user_id, original_name, stored_name) VALUES (?, ?, ?)'
);
$stmt->execute([$_SESSION['user_id'], $file['name'], $newName]);
Notes
  • A file can be a perfectly valid image and contain PHP code in its metadata. That is why passing getimagesize() is not sufficient on its own: the real protection is that the uploads folder is outside the web root, or that script execution is disabled there in the server configuration. Detection narrows the risk; not executing uploads removes it.

JSON Files, and When a File Beats a Database

JSON is the easiest format for storing structured data in a file. json_encode() turns a PHP array into a JSON string and json_decode() turns it back. Pass true as the second argument to json_decode() to get an associative array rather than an object — for configuration and data files that is nearly always what you want.

Both functions fail silently by default, returning false or null and leaving you to check an error function afterwards. Since PHP 7.3 you can pass JSON_THROW_ON_ERROR instead, which turns a malformed file into a catchable exception with a real message. Use it — a truncated JSON file that quietly decodes to null is genuinely hard to debug.

Two other flags earn their place. JSON_PRETTY_PRINT formats the output over multiple lines, which matters if a human will ever read or edit the file. JSON_UNESCAPED_UNICODE keeps non-English characters as themselves instead of escape sequences, so a file containing Hindi or Tamil text stays readable.

So when is a file the right choice at all? Files work well for configuration, for append-only logs, for caching an expensive result, and for import and export. They work badly the moment you need to search, sort, update one record among many, or handle several users writing at once — because every one of those means reading the entire file, changing it in memory, and writing the whole thing back, with a race condition in the middle.

The honest rule: if the data has records that get created, updated and queried, it belongs in a database. Databases exist precisely to solve concurrent access and indexed lookup, and the next lessons show how little code it takes to use one.

Example
<?php
$settings = [
    'site_name' => 'Campus Portal',
    'per_page'  => 20,
    'tagline'   => 'सीखते रहिए',
];

$path = __DIR__ . '/../storage/settings.json';

// Write, readable by humans and preserving non-English text
file_put_contents(
    $path,
    json_encode($settings, JSON_PRETTY_PRINT | JSON_UNESCAPED_UNICODE | JSON_THROW_ON_ERROR),
    LOCK_EX
);

// Read, with real errors instead of a silent null
try {
    $raw    = file_get_contents($path);
    $loaded = json_decode($raw, true, 512, JSON_THROW_ON_ERROR);
} catch (JsonException $e) {
    exit('settings.json is corrupted: ' . $e->getMessage());
}

echo $loaded['site_name'];

// A simple file cache: recompute only when older than an hour
$cache = __DIR__ . '/../storage/cache/stats.json';
if (is_file($cache) && time() - filemtime($cache) < 3600) {
    $stats = json_decode(file_get_contents($cache), true);
} else {
    $stats = expensiveCalculation();
    file_put_contents($cache, json_encode($stats, JSON_THROW_ON_ERROR), LOCK_EX);
}
Notes
  • Storing a list of users or orders in a JSON file will work for your first demo and fail on the day two people submit at once: both requests read the same file, both add their own record, and the second write erases the first. That is not a bug you can fix with more careful code — it is the problem databases were invented to solve.
Ask AI