Skip to content

Stream a file line by line (without loading it) snippet

Reading a file line by line without loading it is the difference between constant and O(file) memory — the convenient APIs (read(), file_get_contents(), ReadFile, Files.readAllLines) slurp everything, and a multi-gigabyte log becomes an OOM.

Reading a file line by line without loading it is the difference between constant and O(file) memory — the convenient APIs (read(), file_get_contents(), ReadFile, Files.readAllLines) slurp everything, and a multi-gigabyte log becomes an OOM. Every language ships a scanner that yields lines from a small buffer; the two remaining traps are the final line without a trailing newline (naive loops drop it) and silent decoding (bytes vs str boundary).

Runnable recipe · 12 languages
Files & Streamsfilesiostreamingmemoryreadlinesbuffered-io

Every language

12 implementations, copy-ready. One at a time with syntax highlighting, or all inline.

JSJavaScript
import { createReadStream } from 'node:fs';
import { createInterface } from 'node:readline/promises';

const rl = createInterface({
  input: createReadStream('access.log', { encoding: 'utf8' }),
  crlfDelay: Infinity, // treat \r\n and \n as one break
});

for await (const line of rl) {
  console.log(line); // ONE line at a time — the buffer is KBs, never the file
}

rl is the async iterable itself — `for await (const line of rl)` pulls lines as chunks arrive, and the interface yields the final line even when the file has no trailing \n.

TSTypeScript
import { createReadStream } from 'node:fs';
import { createInterface } from 'node:readline/promises';

export async function* lines(path: string): AsyncGenerator<string> {
  const rl = createInterface({
    input: createReadStream(path, { encoding: 'utf8' }), // encoding lives HERE
  });
  for await (const line of rl) yield line; // final unterminated line included
}

for await (const line of lines('app.log')) {
  // line: string — already decoded
}

createInterface has no encoding option of its own — the decode setting lives on the stream (createReadStream's option). Skip it and you are iterating Buffers at the bytes/str boundary.

GoGo
import (
	"bufio"
	"log"
	"os"
)

func main() {
	f, err := os.Open("access.log")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()

	scanner := bufio.NewScanner(f)
	scanner.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // allow 4MB lines
	for scanner.Scan() { // Scan strips \n and yields the last unterminated line
		log.Println(scanner.Text())
	}
	if err := scanner.Err(); err != nil {
		log.Fatal(err) // bufio.ErrTooLong lands HERE, not in the loop
	}
}

Scanner's default max token is 64KB — one long minified line ends the loop silently with bufio.ErrTooLong unless Buffer() raised the ceiling. Always check scanner.Err() after the loop.

RsRust
use std::fs::File;
use std::io::{BufRead, BufReader};

fn main() -> std::io::Result<()> {
    for line in BufReader::new(File::open("access.log")?).lines() {
        let line = line?; // io::Result per line — a read can fail mid-file
        println!("{line}");
    }
    Ok(())
}

lines() iterates io::Result items — `?` per item, or .collect::<Result<Vec<_>, _>>() to fold the whole pass into one Result. The trailing chunk after the last \n is still yielded as a final item.

PHPPHP
function lines(string $path): \Generator
{
    $h = fopen($path, 'r');
    if ($h === false) {
        throw new \RuntimeException("cannot open {$path}");
    }
    try {
        while (($line = fgets($h)) !== false) { // false is EOF; '' never happens
            yield rtrim($line, "\r\n");          // fgets KEEPS the newline
        }
    } finally {
        fclose($h);
    }
}

foreach (lines('access.log') as $line) {
    // one line in memory at a time — file('access.log') would load it all
}

file($path) returns the entire file as an array — the O(file) memory trap; the generator over fgets yields lazily. fgets keeps the trailing \n, hence the rtrim.

PyPython
from pathlib import Path

with Path('access.log').open(encoding='utf-8') as f:
    for line in f:  # the file object IS the line iterator
        process(line.rstrip('\n'))

readlines() slurps every line into a list — the memory trap; plain iteration yields from the file's internal buffer, and a last line without a trailing \n still arrives.

C#C#
foreach (var line in File.ReadLines("access.log")) // lazy IEnumerable<string>
{
    Process(line);
}

// File.ReadAllLines("access.log") — string[] of EVERY line, loaded eagerly.
// One word apart in the editor, one order of magnitude apart in memory.

ReadLines yields from a buffered read; ReadAllLines materializes a string[] of the whole file — THE trap pair in C#. For explicit encoding and error control, loop StreamReader.ReadLine yourself.

JvJava
import java.io.BufferedReader;
import java.nio.file.Files;
import java.nio.file.Path;

try (BufferedReader reader = Files.newBufferedReader(Path.of("access.log"))) {
    reader.lines().forEach(line -> {
        // pulled lazily, one buffered read at a time
    });
}

Files.lines(path) is the one-liner twin but the Stream does NOT close the handle itself — it must be the resource of a try-with-resources or the fd leaks until GC. Owning the BufferedReader makes that explicit.

SwSwift
import Foundation

// Swift's stdlib has no line iterator, and String(contentsOf:) loads the whole
// file — the eager trap. The honest pattern is chunked reads + a carry buffer:

struct ChunkedLineReader {
    private let handle: FileHandle
    private var carry = Data() // bytes of a partial line, carried across chunks

    init(_ path: String) throws {
        handle = try FileHandle(forReadingFrom: URL(fileURLWithPath: path))
    }
    deinit { try? handle.close() }

    mutating func nextLine() throws -> String? {
        while true {
            if let nl = carry.firstIndex(of: 0x0A) {         // a full line arrived
                let line = carry.subdata(in: carry.startIndex..<nl)
                carry.removeSubrange(carry.startIndex...nl)  // drop line + \n
                return String(decoding: line, as: UTF8.self)
            }
            guard let chunk = try handle.read(upToCount: 65_536),
                  !chunk.isEmpty
            else { // EOF: the final line may lack the trailing \n — flush it
                let rest = carry
                carry = Data()
                return rest.isEmpty ? nil : String(decoding: rest, as: UTF8.self)
            }
            carry.append(chunk)
        }
    }
}

var reader = try ChunkedLineReader("access.log")
while let line = try reader.nextLine() { print(line) } // constant memory

Honest omission: Swift ships no stdlib line iterator — String(contentsOf:) is eager, so ~20 lines of buffered reading are yours to write. String(decoding:as:) replaces invalid UTF-8 (lossy) rather than erroring at the bytes/str boundary.

KtKotlin
import java.io.File

File("access.log").bufferedReader()
    .useLines { lines -> // closes the reader however the block exits
        lines.forEach { line -> /* lazy Sequence, one line at a time */ }
    }

val total = File("access.log").useLines { it.count() } // nothing materialized

useLines closes the reader in a finally — the Kotlin fix for Java's Files.lines leak; the block's last expression is its return value, so counting or folding stays allocation-free.

RbRuby
File.foreach('access.log', chomp: true) do |line|
  puts line # yielded lazily, one buffered read at a time
end

# File.read(path).each_line { } loads the whole file into a String first —
# fine for KBs, the trap this recipe exists to avoid for GBs

foreach on a File yields lazily; each_line is the same iterator over a String already held in memory — reaching for it means you already slurped the file. chomp: true strips the separator foreach keeps.

ZigZig
const std = @import("std");

pub fn main() !void {
    var file = try std.fs.cwd().openFile("access.log", .{});
    defer file.close();

    var ring: [64 * 1024]u8 = undefined; // cap on ONE line — StreamTooLong past it
    var buffered = std.io.bufferedReader(file.reader());
    const r = buffered.reader();

    while (try r.readUntilDelimiterOrEof(&ring, '\n')) |line| {
        const clean = std.mem.trimRight(u8, line, "\r"); // tolerate CRLF
        _ = clean; // process(clean)
    }
}

readUntilDelimiterOrEof returns the trailing bytes when EOF lands mid-line — the unterminated last line comes through as the final iteration. The buffer size is the line cap: anything longer errors with StreamTooLong.