Skip to content

System Prompt Builder — PHP source

Assemble a system prompt from ordered blocks — role, context, constraints, output format — with a live token count, soft-limit warnings, and a shareable URL. 100% client-side.

This is the PHP implementation — the same logic the interactive tool runs, in a shareable, citable form.

<?php
/**
 * System Prompt Builder — assemble an ordered list of prompt blocks into a
 * markdown-structured system prompt, with pure list operations, presets,
 * warnings, and a compact URL codec for shareable state.
 *
 * Language: PHP (8.1+, standard library only)
 * Source:   CosmoDev polyglot showcase port of the System Prompt Builder
 *           tool, ported from src/lib/systemPromptBuilder.ts (the canonical
 *           TypeScript implementation).
 * Tool page: https://dev.cosmolabs.org/tools/system-prompt-builder
 * License:  display source — part of CosmoDev's polyglot tool pages.
 *
 * Token counting inlines the chars-per-token heuristic from
 * src/lib/tokenEstimator.ts (the original imports it). All list operations
 * are pure — they return new arrays and never mutate their input.
 */

declare(strict_types=1);

/**
 * One editable section of the system prompt.
 */
final class PromptBlock
{
    public function __construct(
        public readonly string $id,
        public string $title,
        public string $content = '',
        public bool $enabled = true,
    ) {
    }
}

/**
 * Assemble + count + lint in one pass — the island's live report.
 */
final class PromptReport
{
    public function __construct(
        public string $assembled,
        public int $tokens,
        public array $warnings = [],
    ) {
    }
}

/**
 * Blocks whose assembled size starts crowding the context on most models.
 */
const SYSTEM_PROMPT_SOFT_LIMIT_TOKENS = 2000;

/**
 * Ordered starter templates — the recommended skeleton of a system prompt.
 * Each entry: {id, title, description, content}.
 */
const SYSTEM_PROMPT_PRESETS = [
    [
        'id' => 'role',
        'title' => 'Role',
        'description' => 'Who the model is and what it optimizes for.',
        'content' => 'You are a senior software engineer. You give correct, concise answers and say so plainly when you are unsure.',
    ],
    [
        'id' => 'context',
        'title' => 'Context',
        'description' => 'The situation the model is working in.',
        'content' => 'The user is a developer working in a TypeScript codebase. Prefer runnable examples over prose when both work.',
    ],
    [
        'id' => 'constraints',
        'title' => 'Constraints',
        'description' => 'Hard rules the model must not break.',
        'content' => "- Never invent library APIs; use only the ones in the provided code.\n- Keep answers under 300 words unless asked for more.",
    ],
    [
        'id' => 'output-format',
        'title' => 'Output format',
        'description' => 'The exact shape of the answer.',
        'content' => 'Respond with: 1) a one-line summary, 2) a fenced code block, 3) any caveats as bullet points.',
    ],
    [
        'id' => 'examples',
        'title' => 'Examples',
        'description' => 'Few-shot demonstrations of the desired behavior.',
        'content' => "Input: reverse \"abc\"\nOutput: \"cba\"",
    ],
    [
        'id' => 'tone',
        'title' => 'Tone',
        'description' => 'Voice and register.',
        'content' => 'Direct and friendly. No filler openers, no apologies.',
    ],
    [
        'id' => 'refusal',
        'title' => 'Refusal policy',
        'description' => 'How to handle out-of-scope requests.',
        'content' => 'If a request is outside your scope, say so in one sentence and suggest the closest thing you can do.',
    ],
    [
        'id' => 'safety',
        'title' => 'Safety',
        'description' => 'Guardrails for sensitive content.',
        'content' => 'Refuse requests that could cause harm, and never echo secrets, keys, or credentials back in full.',
    ],
];

// ---- token estimate (tokens figure only, from tokenEstimator.ts) ------------

/**
 * Average characters per token, by content type.
 */
const SPB_CHARS_PER_TOKEN = ['prose' => 4.0, 'code' => 3.5, 'json' => 3.0, 'cjk' => 1.5];

/**
 * Classify a single line by its shape. Order: json, cjk, code, prose.
 */
function spb_detect_line_type(string $line): string
{
    $trimmed = trim($line);
    // JSON-ish: opens like a JSON fragment AND carries a separator.
    $first = $trimmed !== '' ? $trimmed[0] : '';
    if (($first === '{' || $first === '}' || $first === '[' || $first === '"')
        && (str_contains($line, ':') || str_contains($line, ','))) {
        return 'json';
    }
    // CJK ideographs / kana / Hangul pack roughly one token per 1.5 chars.
    // (The Hangul arm stops at U+D7FF — PCRE rejects ranges that cross the
    // surrogate block, which cannot appear in valid UTF-8 anyway.)
    if (preg_match('/[\x{4E00}-\x{9FFF}\x{3040}-\x{30FF}\x{AC00}-\x{D7FF}]/u', $line) === 1) {
        return 'cjk';
    }
    // Code: symbol-dense, or a statement terminator / block opener at EOL.
    $len = mb_strlen($line, 'UTF-8');
    $density = $len > 0
        ? preg_match_all('/[{}();=<>\\[\\]#]/', $line) / $len
        : 0.0;
    if ($density > 0.08 || str_ends_with($trimmed, ';') || str_ends_with($trimmed, '{') || str_ends_with($trimmed, '}')) {
        return 'code';
    }
    return 'prose';
}

/**
 * Sum of per-line token estimates (excludes chat framing). $contentType may
 * be 'prose' | 'code' | 'json' | 'cjk' | 'auto'; 'auto' classifies per line,
 * with a document that parses as JSON counted as json throughout.
 */
function spb_estimate_tokens(string $text, string $contentType = 'prose'): int
{
    $forced = ($contentType !== '' && $contentType !== 'auto') ? $contentType : null;
    $wholeTextJson = false;
    if ($forced === null && trim($text) !== '') {
        json_decode($text);
        $wholeTextJson = json_last_error() === JSON_ERROR_NONE;
    }
    $tokens = 0;
    foreach (preg_split('/\r?\n/', $text) ?: [] as $line) {
        if (trim($line) === '') {
            continue;
        }
        $type = $forced ?? ($wholeTextJson ? 'json' : spb_detect_line_type($line));
        // max(1, round(len / rate)) — floor(x + 0.5) is JS-style rounding.
        $tokens += max(1, (int) floor(mb_strlen($line, 'UTF-8') / SPB_CHARS_PER_TOKEN[$type] + 0.5));
    }
    return $tokens;
}

/**
 * Render enabled, non-empty blocks (in order) as one markdown-structured
 * prompt (drop the "## Title" headers by passing headers=false).
 */
function assemble_prompt(array $blocks, bool $headers = true): string
{
    $rendered = [];
    foreach ($blocks as $b) {
        if (!($b instanceof PromptBlock) || !$b->enabled || trim($b->content) === '') {
            continue;
        }
        $rendered[] = $headers
            ? '## ' . (trim($b->title) === '' ? 'Untitled' : trim($b->title)) . "\n" . trim($b->content)
            : trim($b->content);
    }
    return trim(implode("\n\n", $rendered));
}

/**
 * Append a block (caller supplies the id so the lib stays pure).
 */
function add_block(
    array $blocks,
    string $id,
    string $title,
    string $content = '',
    bool $enabled = true,
): array {
    $blocks[] = new PromptBlock($id, $title, $content, $enabled);
    return $blocks;
}

/**
 * Patch one block by id (patch keys: title/content/enabled); unknown ids
 * leave the list unchanged.
 */
function update_block(array $blocks, string $id, array $patch = []): array
{
    return array_map(
        static fn (PromptBlock $b): PromptBlock => $b->id === $id
            ? new PromptBlock(
                $b->id,
                array_key_exists('title', $patch) ? $patch['title'] : $b->title,
                array_key_exists('content', $patch) ? $patch['content'] : $b->content,
                array_key_exists('enabled', $patch) ? $patch['enabled'] : $b->enabled,
            )
            : $b,
        $blocks,
    );
}

/**
 * Flip one block's enabled flag by id.
 */
function toggle_block(array $blocks, string $id): array
{
    return array_map(
        static fn (PromptBlock $b): PromptBlock => $b->id === $id
            ? new PromptBlock($b->id, $b->title, $b->content, !$b->enabled)
            : $b,
        $blocks,
    );
}

/**
 * Remove one block by id.
 */
function remove_block(array $blocks, string $id): array
{
    return array_values(array_filter(
        $blocks,
        static fn (PromptBlock $b): bool => $b->id !== $id,
    ));
}

/**
 * Move a block (clamped; no-op when indexes are out of range or equal).
 */
function move_block(array $blocks, int $from, int $to): array
{
    $n = count($blocks);
    if ($from < 0 || $from >= $n || $to < 0 || $to >= $n || $from === $to) {
        return $blocks;
    }
    $moved = array_splice($blocks, $from, 1);
    array_splice($blocks, $to, 0, $moved);
    return $blocks;
}

/**
 * Assemble + count + lint in one pass — the island's live report.
 */
function build_report(array $blocks, string $contentType = 'prose'): PromptReport
{
    $assembled = assemble_prompt($blocks);
    $tokens = $assembled !== '' ? spb_estimate_tokens($assembled, $contentType) : 0;
    $warnings = [];
    if ($tokens > SYSTEM_PROMPT_SOFT_LIMIT_TOKENS) {
        $warnings[] = sprintf(
            'Assembled prompt is ~%s tokens — beyond %s it starts crowding the context window on most models.',
            number_format($tokens),
            number_format(SYSTEM_PROMPT_SOFT_LIMIT_TOKENS),
        );
    }
    $hasRole = false;
    foreach ($blocks as $b) {
        if ($b instanceof PromptBlock && $b->enabled && mb_strtolower(trim($b->title), 'UTF-8') === 'role') {
            $hasRole = true;
            break;
        }
    }
    if (count($blocks) > 0 && !$hasRole) {
        $warnings[] = 'No enabled "Role" block — stating who the model is tends to anchor every following instruction.';
    }
    if (count($blocks) > 0 && $assembled === '') {
        $warnings[] = 'Every block is disabled or empty — the assembled prompt is empty.';
    }
    return new PromptReport($assembled, $tokens, $warnings);
}

// ---- shareable state codec (URL-safe, compact) ------------------------------
// Triples of [enabled(0/1), title, content] keep URLs far smaller than the
// full object shape; ids are regenerated on decode (they are UI-local).

const SPB_MAX_ENCODED_LENGTH = 4000;

/**
 * Encode blocks to a compact base64url string; '' when blocks are empty.
 */
function encode_blocks(array $blocks): string
{
    if (count($blocks) === 0) {
        return '';
    }
    $compact = array_map(
        static fn (PromptBlock $b): array => [$b->enabled ? 1 : 0, $b->title, $b->content],
        $blocks,
    );
    // JSON_UNESCAPED_* matches JSON.stringify output (no \/ or \uXXXX noise).
    $json = json_encode($compact, JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE);
    return rtrim(strtr(base64_encode($json), '+/', '-_'), '=');
}

/**
 * True when the encoded form would make an uncomfortably long URL.
 */
function encoded_too_long(string $encoded): bool
{
    return strlen($encoded) > SPB_MAX_ENCODED_LENGTH;
}

/**
 * Decode `encode_blocks` output; regenerates ids (b1, b2, …). Returns null
 * on malformed input — never throws; '' decodes to [].
 */
function decode_blocks(string $encoded): ?array
{
    if ($encoded === '') {
        return [];
    }
    $b64 = strtr($encoded, '-_', '+/') . str_repeat('=', (4 - strlen($encoded) % 4) % 4);
    $json = base64_decode($b64, true);
    if ($json === false) {
        return null;
    }
    $raw = json_decode($json);
    if (!is_array($raw)) {
        return null;
    }
    $blocks = [];
    foreach ($raw as $i => $entry) {
        if (!is_array($entry) || count($entry) !== 3) {
            return null;
        }
        [$enabled, $title, $content] = $entry;
        if (!is_int($enabled) || !is_string($title) || !is_string($content)) {
            return null;
        }
        $blocks[] = new PromptBlock('b' . ($i + 1), $title, $content, $enabled === 1);
    }
    return $blocks;
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →