Skip to content

Markdown Table Generator — PHP source

Turn pipe, CSV, tab, semicolon, or space-separated data into a clean GitHub-Flavored Markdown table. Auto-detects the delimiter, pads columns, escapes pipes, and supports per-column alignment - all in your browser.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * Markdown Table Generator — pure logic, PHP polyglot showcase port.
 *
 * Language:    PHP (8.0+; uses match, str_starts_with, str_ends_with)
 * Origin:      CosmoDev polyglot showcase port of the `markdown-table` tool.
 * Ported from: src/lib/markdown-table.ts (the canonical, live TypeScript lib).
 *
 * Parsing and rendering are deterministic and depend only on their inputs.
 * This file is display source — part of CosmoDev's polyglot tool pages, where
 * the same pure logic is shown side-by-side across languages.
 */

declare(strict_types=1);

// Delimiter candidates considered during auto-detection, in priority order.
// Structural delimiters (pipe, tab) outrank punctuation (`,` / `;`) outrank space.
const CANDIDATES = ['|', "\t", ';', ',', ' '];

// Structural weight per delimiter. Higher wins ties.
const WEIGHT = ['|' => 3, "\t" => 3, ';' => 2, ',' => 2, ' ' => 1];

/**
 * Split a single line by delimiter, trimming each resulting cell.
 *
 * - `|`  strips one leading/trailing pipe (so `| a | b |` works) then splits.
 * - ` `  splits on runs of whitespace.
 * - `,` `\t` `;` split on the literal character.
 *
 * @param string $line
 * @param string $delimiter One of CANDIDATES.
 * @return list<string>
 */
function split_line(string $line, string $delimiter): array
{
    if ($delimiter === '|') {
        $l = trim($line);
        if (str_starts_with($l, '|')) {
            $l = substr($l, 1);
        }
        if (str_ends_with($l, '|')) {
            $l = substr($l, 0, -1);
        }
        // A fully-empty line collapses to a single empty cell, not zero cells.
        if ($l === '') {
            return [''];
        }
        return array_map('trim', explode('|', $l));
    }
    if ($delimiter === ' ') {
        // parse_table only passes non-empty (post-trim) lines here, so this
        // always yields >= 1 token — no empty-line guard needed.
        $tokens = preg_split('/\s+/', trim($line));
        return $tokens === false ? [''] : $tokens;
    }
    return array_map('trim', explode($delimiter, $line));
}

/**
 * Parse input into a 2-D grid of trimmed cells. Blank lines are skipped;
 * each remaining line is split by delimiter.
 *
 * @param string $input
 * @param string $delimiter One of CANDIDATES.
 * @return list<list<string>>
 */
function parse_table(string $input, string $delimiter): array
{
    $lines = preg_split('/\r?\n/', $input);
    if ($lines === false) {
        return [];
    }
    $grid = [];
    foreach ($lines as $line) {
        $trimmed = trim($line);
        if ($trimmed !== '') {
            $grid[] = split_line($trimmed, $delimiter);
        }
    }
    return $grid;
}

/**
 * Count the delimiter occurrences in a line. Whitespace counts *runs* of
 * whitespace, not individual space characters.
 *
 * @param string $line
 * @param string $delimiter One of CANDIDATES.
 * @return int
 */
function count_occurrences(string $line, string $delimiter): int
{
    if ($delimiter === ' ') {
        $tokens = preg_split('/\s+/', trim($line));
        return ($tokens === false ? 0 : count($tokens)) - 1;
    }
    return substr_count($line, $delimiter);
}

/**
 * Heuristic delimiter detection. Each candidate is scored by
 *   frequency x cross-line consistency x structural weight,
 * and the best wins. Falls back to `,` when nothing scores (single column or
 * empty input).
 *
 * @param string $sample
 * @return string One of CANDIDATES.
 */
function detect_delimiter(string $sample): string
{
    $lines = preg_split('/\r?\n/', $sample);
    if ($lines === false) {
        return ',';
    }
    $lines = array_values(array_filter(array_map('trim', $lines), fn ($l) => $l !== ''));
    if (count($lines) === 0) {
        return ',';
    }

    $best = ',';
    $bestScore = 0.0;
    foreach (CANDIDATES as $d) {
        $counts = array_map(fn ($l) => (float) count_occurrences($l, $d), $lines);
        $avg = array_sum($counts) / count($counts);
        if ($avg == 0.0) {
            continue;
        }
        // population variance across lines -> lower is more consistent
        $variance = array_sum(array_map(fn ($c) => ($c - $avg) ** 2, $counts)) / count($counts);
        $consistency = 1.0 / (1.0 + $variance);
        $score = $avg * $consistency * (float) WEIGHT[$d];
        if ($score > $bestScore) {
            $bestScore = $score;
            $best = $d;
        }
    }
    return $best;
}

/**
 * Escape a cell for GFM: collapse newlines (CRLF or LF) to a single space,
 * then escape literal `|` so it does not terminate the cell.
 *
 * @param string $cell
 * @return string
 */
function escape_cell(string $cell): string
{
    $collapsed = preg_replace('/\r?\n/', ' ', $cell);
    return str_replace('|', '\\|', $collapsed === null ? $cell : $collapsed);
}

/**
 * Pad a cell to width honoring alignment. Center splits the slack with the
 * floor on the left. Uses mb_strlen so multibyte cells line up correctly.
 *
 * @param string $cell
 * @param int    $width
 * @param string $align One of 'left' | 'center' | 'right' | 'none'.
 * @return string
 */
function pad_cell(string $cell, int $width, string $align): string
{
    $diff = $width - mb_strlen($cell, 'UTF-8');
    if ($diff <= 0) {
        return $cell;
    }
    if ($align === 'right') {
        return str_repeat(' ', $diff) . $cell;
    }
    if ($align === 'center') {
        $left = intdiv($diff, 2);
        return str_repeat(' ', $left) . $cell . str_repeat(' ', $diff - $left);
    }
    return $cell . str_repeat(' ', $diff); // 'left' | 'none'
}

/**
 * Render a separator cell (`---`, `:--`, `--:`, `:-:`) of at least 3 dashes.
 *
 * @param string $align One of 'left' | 'center' | 'right' | 'none'.
 * @param int    $width
 * @return string
 */
function sep_cell(string $align, int $width): string
{
    $w = max(3, $width);
    return match ($align) {
        'center' => ':' . str_repeat('-', $w - 2) . ':',
        'right'  => str_repeat('-', $w - 1) . ':',
        'left'   => ':' . str_repeat('-', $w - 1),
        default  => str_repeat('-', $w),
    };
}

/**
 * Render a 2-D grid as a GitHub-Flavored Markdown table. Cells are padded to
 * equal column widths (computed from the escaped text), literal `|` is
 * escaped, and the separator row carries the per-column alignment.
 *
 * Returns '' for an empty grid.
 *
 * @param list<list<string>> $rows
 * @param array{header: bool, align?: list<string>} $opts
 * @return string
 */
function to_markdown(array $rows, array $opts): string
{
    if (count($rows) === 0) {
        return '';
    }

    // Column count is the longest row.
    $cols = 0;
    foreach ($rows as $r) {
        $cols = max($cols, count($r));
    }

    // Escape every cell and normalize each row to the column count.
    $grid = [];
    foreach ($rows as $r) {
        $out = array_map('escape_cell', $r);
        while (count($out) < $cols) {
            $out[] = '';
        }
        $grid[] = $out;
    }

    // Per-column alignment: missing entries default to 'none'; surplus ignored.
    $aligns = [];
    for ($i = 0; $i < $cols; $i++) {
        $aligns[] = $opts['align'][$i] ?? 'none';
    }

    // Per-column width: at least 3 (GFM separator minimum), grown to fit the
    // widest escaped cell in the column.
    $widths = [];
    for ($c = 0; $c < $cols; $c++) {
        $max = 3;
        foreach ($grid as $r) {
            $len = mb_strlen($r[$c], 'UTF-8');
            if ($len > $max) {
                $max = $len;
            }
        }
        $widths[] = $max;
    }

    // Frame one row of cells with the GFM pipe scaffolding.
    $renderLine = function (array $cells) use ($widths, $aligns): string {
        $padded = [];
        foreach ($cells as $i => $c) {
            $padded[] = pad_cell($c, $widths[$i], $aligns[$i]);
        }
        return '| ' . implode(' | ', $padded) . ' |';
    };

    $sepParts = [];
    foreach ($aligns as $i => $a) {
        $sepParts[] = sep_cell($a, $widths[$i]);
    }
    $separator = '| ' . implode(' | ', $sepParts) . ' |';

    // When there is no header, synthesize a blank header row so the table is
    // still valid GFM.
    if ($opts['header']) {
        $header = $renderLine($grid[0]);
        $dataStart = 1;
    } else {
        $header = $renderLine(array_fill(0, $cols, ''));
        $dataStart = 0;
    }

    $lines = [$header, $separator];
    for ($i = $dataStart; $i < count($grid); $i++) {
        $lines[] = $renderLine($grid[$i]);
    }
    return implode("\n", $lines);
}

/**
 * Transpose a grid (rows <-> columns). Jagged grids are filled with ''.
 *
 * @param list<list<string>> $rows
 * @return list<list<string>>
 */
function transpose(array $rows): array
{
    if (count($rows) === 0) {
        return [];
    }
    $cols = 0;
    foreach ($rows as $r) {
        $cols = max($cols, count($r));
    }
    $out = [];
    for ($c = 0; $c < $cols; $c++) {
        $column = [];
        foreach ($rows as $r) {
            $column[] = $r[$c] ?? '';
        }
        $out[] = $column;
    }
    return $out;
}

/*
 * Example usage (this file is a library; uncomment to run as a script):
 *
 * $sample = "name,role,team\nAda,engineer,platform\nLin,designer,brand";
 * $d = detect_delimiter($sample);
 * $grid = parse_table($sample, $d);
 * echo to_markdown($grid, ['header' => true, 'align' => ['left', 'left', 'center']]) . "\n";
 */

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →