Skip to content

Lire un fichier en flux ligne par ligne (sans le charger) snippet

Lire un fichier ligne par ligne sans le charger, c'est la différence entre mémoire constante et O(taille du fichier) — les APIs confortables (read(), file_get_contents(), ReadFile, Files.readAllLines) avalent tout, et un log de plusieurs gigaoctets devient un OOM.

Lire un fichier ligne par ligne sans le charger, c'est la différence entre mémoire constante et O(taille du fichier) — les APIs confortables (read(), file_get_contents(), ReadFile, Files.readAllLines) avalent tout, et un log de plusieurs gigaoctets devient un OOM. Chaque langage embarque un scanner qui produit les lignes depuis un petit buffer ; les deux pièges restants sont la dernière ligne sans saut de ligne final (les boucles naïves la perdent) et le décodage silencieux (frontière bytes contre str).

Recette exécutable · 12 langages
Files & Streamsfilesiostreamingmemoryreadlinesbuffered-io

Every language

12 langages, copy-ready. One at a time with syntax highlighting, or all inline.

JSJavaScript
import { createReadStream } from 'node:fs';
import { createInterface } from 'node:readline/promises';

const rl = createInterface({
  input: createReadStream('access.log', { encoding: 'utf8' }),
  crlfDelay: Infinity, // treat \r\n and \n as one break
});

for await (const line of rl) {
  console.log(line); // ONE line at a time — the buffer is KBs, never the file
}

rl is the async iterable itself — `for await (const line of rl)` pulls lines as chunks arrive, and the interface yields the final line even when the file has no trailing \n.

TSTypeScript
import { createReadStream } from 'node:fs';
import { createInterface } from 'node:readline/promises';

export async function* lines(path: string): AsyncGenerator<string> {
  const rl = createInterface({
    input: createReadStream(path, { encoding: 'utf8' }), // encoding lives HERE
  });
  for await (const line of rl) yield line; // final unterminated line included
}

for await (const line of lines('app.log')) {
  // line: string — already decoded
}

createInterface has no encoding option of its own — the decode setting lives on the stream (createReadStream's option). Skip it and you are iterating Buffers at the bytes/str boundary.

GoGo
import (
	"bufio"
	"log"
	"os"
)

func main() {
	f, err := os.Open("access.log")
	if err != nil {
		log.Fatal(err)
	}
	defer f.Close()

	scanner := bufio.NewScanner(f)
	scanner.Buffer(make([]byte, 0, 64*1024), 4*1024*1024) // allow 4MB lines
	for scanner.Scan() { // Scan strips \n and yields the last unterminated line
		log.Println(scanner.Text())
	}
	if err := scanner.Err(); err != nil {
		log.Fatal(err) // bufio.ErrTooLong lands HERE, not in the loop
	}
}

Scanner's default max token is 64KB — one long minified line ends the loop silently with bufio.ErrTooLong unless Buffer() raised the ceiling. Always check scanner.Err() after the loop.

RsRust
use std::fs::File;
use std::io::{BufRead, BufReader};

fn main() -> std::io::Result<()> {
    for line in BufReader::new(File::open("access.log")?).lines() {
        let line = line?; // io::Result per line — a read can fail mid-file
        println!("{line}");
    }
    Ok(())
}

lines() iterates io::Result items — `?` per item, or .collect::<Result<Vec<_>, _>>() to fold the whole pass into one Result. The trailing chunk after the last \n is still yielded as a final item.

PHPPHP
function lines(string $path): \Generator
{
    $h = fopen($path, 'r');
    if ($h === false) {
        throw new \RuntimeException("cannot open {$path}");
    }
    try {
        while (($line = fgets($h)) !== false) { // false is EOF; '' never happens
            yield rtrim($line, "\r\n");          // fgets KEEPS the newline
        }
    } finally {
        fclose($h);
    }
}

foreach (lines('access.log') as $line) {
    // one line in memory at a time — file('access.log') would load it all
}

file($path) returns the entire file as an array — the O(file) memory trap; the generator over fgets yields lazily. fgets keeps the trailing \n, hence the rtrim.

PyPython
from pathlib import Path

with Path('access.log').open(encoding='utf-8') as f:
    for line in f:  # the file object IS the line iterator
        process(line.rstrip('\n'))

readlines() slurps every line into a list — the memory trap; plain iteration yields from the file's internal buffer, and a last line without a trailing \n still arrives.

C#C#
foreach (var line in File.ReadLines("access.log")) // lazy IEnumerable<string>
{
    Process(line);
}

// File.ReadAllLines("access.log") — string[] of EVERY line, loaded eagerly.
// One word apart in the editor, one order of magnitude apart in memory.

ReadLines yields from a buffered read; ReadAllLines materializes a string[] of the whole file — THE trap pair in C#. For explicit encoding and error control, loop StreamReader.ReadLine yourself.

JvJava
import java.io.BufferedReader;
import java.nio.file.Files;
import java.nio.file.Path;

try (BufferedReader reader = Files.newBufferedReader(Path.of("access.log"))) {
    reader.lines().forEach(line -> {
        // pulled lazily, one buffered read at a time
    });
}

Files.lines(path) is the one-liner twin but the Stream does NOT close the handle itself — it must be the resource of a try-with-resources or the fd leaks until GC. Owning the BufferedReader makes that explicit.

SwSwift
import Foundation

// Swift's stdlib has no line iterator, and String(contentsOf:) loads the whole
// file — the eager trap. The honest pattern is chunked reads + a carry buffer:

struct ChunkedLineReader {
    private let handle: FileHandle
    private var carry = Data() // bytes of a partial line, carried across chunks

    init(_ path: String) throws {
        handle = try FileHandle(forReadingFrom: URL(fileURLWithPath: path))
    }
    deinit { try? handle.close() }

    mutating func nextLine() throws -> String? {
        while true {
            if let nl = carry.firstIndex(of: 0x0A) {         // a full line arrived
                let line = carry.subdata(in: carry.startIndex..<nl)
                carry.removeSubrange(carry.startIndex...nl)  // drop line + \n
                return String(decoding: line, as: UTF8.self)
            }
            guard let chunk = try handle.read(upToCount: 65_536),
                  !chunk.isEmpty
            else { // EOF: the final line may lack the trailing \n — flush it
                let rest = carry
                carry = Data()
                return rest.isEmpty ? nil : String(decoding: rest, as: UTF8.self)
            }
            carry.append(chunk)
        }
    }
}

var reader = try ChunkedLineReader("access.log")
while let line = try reader.nextLine() { print(line) } // constant memory

Honest omission: Swift ships no stdlib line iterator — String(contentsOf:) is eager, so ~20 lines of buffered reading are yours to write. String(decoding:as:) replaces invalid UTF-8 (lossy) rather than erroring at the bytes/str boundary.

KtKotlin
import java.io.File

File("access.log").bufferedReader()
    .useLines { lines -> // closes the reader however the block exits
        lines.forEach { line -> /* lazy Sequence, one line at a time */ }
    }

val total = File("access.log").useLines { it.count() } // nothing materialized

useLines closes the reader in a finally — the Kotlin fix for Java's Files.lines leak; the block's last expression is its return value, so counting or folding stays allocation-free.

RbRuby
File.foreach('access.log', chomp: true) do |line|
  puts line # yielded lazily, one buffered read at a time
end

# File.read(path).each_line { } loads the whole file into a String first —
# fine for KBs, the trap this recipe exists to avoid for GBs

foreach on a File yields lazily; each_line is the same iterator over a String already held in memory — reaching for it means you already slurped the file. chomp: true strips the separator foreach keeps.

ZigZig
const std = @import("std");

pub fn main() !void {
    var file = try std.fs.cwd().openFile("access.log", .{});
    defer file.close();

    var ring: [64 * 1024]u8 = undefined; // cap on ONE line — StreamTooLong past it
    var buffered = std.io.bufferedReader(file.reader());
    const r = buffered.reader();

    while (try r.readUntilDelimiterOrEof(&ring, '\n')) |line| {
        const clean = std.mem.trimRight(u8, line, "\r"); // tolerate CRLF
        _ = clean; // process(clean)
    }
}

readUntilDelimiterOrEof returns the trailing bytes when EOF lands mid-line — the unterminated last line comes through as the final iteration. The buffer size is the line cap: anything longer errors with StreamTooLong.