Skip to content

Context Window Planner — Java source

Paste your system prompt, docs, and history — see how they fill any model's context window, with overflow warnings and output headroom.

This is the Java implementation — the same logic the interactive tool runs, in a shareable, citable form.

// Context Window Planner — plan labeled prompt sections against a model's
// context window.
//
// Language: Java (Java 17, standard library only)
// Source:   CosmoDev polyglot showcase port of the Context Window Planner
//           tool, ported from src/lib/contextPlanner.ts (the canonical
//           TypeScript implementation).
// Live at:  https://dev.cosmolabs.org/tools/context-window-planner
// License:  display source — part of CosmoDev's polyglot tool pages.
//
// Design goals:
//   - Pure + deterministic; never throws (public API returns plain values).
//   - Functionally equivalent to the TS reference: same inputs -> same outputs.
//   - Self-contained: JDK only (no Jackson — the stdlib equivalent would be
//     javax.json; this port instead ships the same small strict
//     recursive-descent validator the other polyglot siblings use, so the
//     JSON grammar matches JSON.parse exactly rather than approximately).
//
// Port notes: the TS lib delegates to two siblings — `estimateTokens` from
// src/lib/tokenEstimator.ts and `fitsWindow` from src/lib/ai/models.ts (which
// defaults to the bundled pricing snapshot, src/data/ai-models.json). A
// dependency-free port cannot load that file, so the estimator is inlined
// below in the exact form the planner uses it (`estimateTokens(text).tokens`,
// auto content type — the full heuristic lives in the token-estimator port),
// window math is inlined from `fitsWindow` and `models` is an explicit
// parameter, never re-derived.
//
// Faithfulness notes (the places Java's stdlib silently differs from JS):
//   - Length: TS's `String.length` counts UTF-16 code units — and so does
//     Java's `String.length()`, so no helper is needed; astral-plane
//     characters (emoji, rare CJK ext-B ideographs) already count as 2 on
//     both sides.
//   - Rounding: `Math.round(double)` is floor(x + 0.5) — exactly JS
//     `Math.round` (halfway cases up). No banker's-rounding trap here.

import java.util.ArrayList;
import java.util.List;
import java.util.Optional;

/**
 * Context Window Planner — a faithful, stdlib-only port of the tool's pure
 * logic. Every public member mirrors the TS lib one-for-one.
 */
final class ContextWindowPlanner {

    private ContextWindowPlanner() {}

    /** One labeled block of the prompt (system / docs / history / ...).
     *  Mirrors the TS {@code PlanSection} interface. */
    record PlanSection(String label, String text) {}

    /** Convenience factory mirroring the TS object literal
     *  {@code { label, text }}. */
    static PlanSection sec(String label, String text) {
        return new PlanSection(label, text);
    }

    /** The subset of the TS {@code AiModel} record the planner reads.
     *  Production code passes the full snapshot entry; only these fields
     *  influence the plan. */
    record Model(String id, long contextWindow, long maxOutput) {}

    /** Sample table for standalone use (mirrors the shared test fixtures).
     *  Production code passes the model snapshot instead. */
    static final List<Model> SAMPLE_MODELS = List.of(
            new Model("alpha-mini", 200_000L, 10_000L),
            new Model("beta-pro", 1_000_000L, 10_000L),
            new Model("gamma-open", 100_000L, 10_000L));

    /** Result of {@link #planWindow}. Field-for-field twin of the TS
     *  {@code WindowPlan} interface. */
    record WindowPlan(
            String id,            // the model id planned against
            long inputTokens,     // sum of per-section token estimates
            long contextWindow,   // the model's context window
            long free,            // window - input; negative on overflow
            boolean fits,         // raw fit: free >= 0
            boolean outputReserveOk, // room for the output reserve
            long maxOutput) {}    // the model's output cap (informational)

    /** Content classification of a single line. The planner only needs each
     *  type's chars-per-token rate (mirrors {@code CHARS_PER_TOKEN} in
     *  src/lib/tokenEstimator.ts: prose 4, code 3.5, json 3, cjk 1.5). */
    private enum ContentType { PROSE, CODE, JSON, CJK }

    private static double charsPerToken(ContentType t) {
        return switch (t) {
            case PROSE -> 4.0;
            case CODE -> 3.5;
            case JSON -> 3.0;
            case CJK -> 1.5;
        };
    }

    /** Reports whether {@code s} contains a CJK ideograph (U+4E00–U+9FFF),
     *  kana (U+3040–U+30FF), or a Hangul syllable (U+AC00–U+D7AF).
     *  Mirrors {@code CJK_RE} in the TS lib. */
    private static boolean hasCjk(String s) {
        for (int i = 0; i < s.length(); i++) {
            char c = s.charAt(i);
            if ((c >= 0x4E00 && c <= 0x9FFF) ||
                (c >= 0x3040 && c <= 0x30FF) ||
                (c >= 0xAC00 && c <= 0xD7AF))
                return true;
        }
        return false;
    }

    /** Reports whether {@code c} is one of the code-flavored symbols counted
     *  by {@code CODE_SYMBOL_RE} ({@code {}();=<>[]#}). */
    private static boolean isCodeSymbol(char c) {
        return c == '{' || c == '}' || c == '(' || c == ')' || c == ';' ||
               c == '=' || c == '<' || c == '>' || c == '[' || c == ']' || c == '#';
    }

    /** Splits {@code text} on LF or CRLF, mirroring
     *  {@code text.split(/\r?\n/)}: strip the optional CR that belongs to the
     *  newline, then split on LF. A lone CR is NOT a line break. */
    private static String[] splitLines(String text) {
        List<String> lines = new ArrayList<>();
        int start = 0;
        while (start <= text.length()) {
            int nl = text.indexOf('\n', start);
            String line = nl < 0 ? text.substring(start)
                                 : text.substring(start, nl);
            if (line.endsWith("\r")) line = line.substring(0, line.length() - 1);
            lines.add(line);
            if (nl < 0) break;
            start = nl + 1;
        }
        return lines.toArray(new String[0]);
    }

    /** Classifies a single line by its shape. Order: json, cjk, code, prose.
     *  Inlined from {@code detectLineType()} in src/lib/tokenEstimator.ts. */
    private static ContentType detectLineType(String line) {
        String trimmed = line.strip();
        // JSON-ish: opens like a JSON fragment AND carries a separator.
        boolean startsJsonish = !trimmed.isEmpty() &&
                (trimmed.charAt(0) == '{' || trimmed.charAt(0) == '}' ||
                 trimmed.charAt(0) == '[' || trimmed.charAt(0) == '"');
        if (startsJsonish && (line.contains(":") || line.contains(",")))
            return ContentType.JSON;
        // CJK ideographs / kana / Hangul pack roughly one token per 1.5 chars.
        if (hasCjk(line))
            return ContentType.CJK;
        // Code: symbol-dense, or a statement terminator / block opener at EOL.
        int length = line.length(); // UTF-16 code units, same unit as TS
        int symbols = 0;
        for (int i = 0; i < length; i++)
            if (isCodeSymbol(line.charAt(i))) symbols++;
        double density = length > 0 ? (double) symbols / length : 0.0;
        if (density > 0.08 || trimmed.endsWith(";") || trimmed.endsWith("{") ||
            trimmed.endsWith("}"))
            return ContentType.CODE;
        return ContentType.PROSE;
    }

    /** A strict JSON syntax validator — the exact grammar {@code JSON.parse}
     *  accepts, walked with a cursor. Same shape as the Rust/C/C++/C#
     *  siblings' validators, so whole-text JSON detection behaves
     *  identically across every port. */
    private static final class JsonParser {
        private final String s;
        private int pos;

        JsonParser(String text) { this.s = text; }

        /** value := ws* (object | array | string | number | 'true' | 'false' | 'null') ws* */
        boolean value() {
            skipWs();
            char c = peek();
            if (pos >= s.length()) return false;
            switch (c) {
                case '{': return object();
                case '[': return array();
                case '"': return string();
                case '-': case '0': case '1': case '2': case '3': case '4':
                case '5': case '6': case '7': case '8': case '9':
                    return number();
                case 't': return literal("true");
                case 'f': return literal("false");
                case 'n': return literal("null");
                default: return false;
            }
        }

        boolean atEnd() {
            skipWs();
            return pos == s.length(); // reject trailing garbage
        }

        private void skipWs() {
            while (pos < s.length()) {
                char c = s.charAt(pos);
                if (c == ' ' || c == '\t' || c == '\n' || c == '\r') pos++;
                else break;
            }
        }

        private char peek() { return pos < s.length() ? s.charAt(pos) : '\0'; }

        private boolean eat(char b) {
            if (pos < s.length() && s.charAt(pos) == b) { pos++; return true; }
            return false;
        }

        private boolean literal(String lit) {
            if (s.regionMatches(pos, lit, 0, lit.length())) {
                pos += lit.length();
                return true;
            }
            return false;
        }

        /** object := '{' ws* (string ws* ':' value (ws* ',' ...)*)? ws* '}' */
        private boolean object() {
            if (!eat('{')) return false;
            skipWs();
            if (eat('}')) return true;
            while (true) {
                if (!string()) return false;
                skipWs();
                if (!eat(':')) return false;
                if (!value()) return false;
                skipWs();
                if (eat(',')) skipWs();
                else return eat('}');
            }
        }

        /** array := '[' ws* (value (ws* ',' ws* value)*)? ws* ']' */
        private boolean array() {
            if (!eat('[')) return false;
            skipWs();
            if (eat(']')) return true;
            while (true) {
                if (!value()) return false;
                skipWs();
                if (eat(',')) skipWs();
                else return eat(']');
            }
        }

        /** string := '"' (escape | any char >= 0x20)* '"'
         *  escape := '\' ('"' | '/' | '\' | 'b' | 'f' | 'n' | 'r' | 't' | 'u' hex4) */
        private boolean string() {
            if (!eat('"')) return false;
            while (pos < s.length()) {
                char c = s.charAt(pos);
                if (c == '"') { pos++; return true; }
                if (c == '\\') {
                    pos++;
                    if (pos >= s.length()) return false;
                    char esc = s.charAt(pos++);
                    switch (esc) {
                        case '"': case '/': case '\\': case 'b':
                        case 'f': case 'n': case 'r': case 't':
                            break;
                        case 'u':
                            for (int i = 0; i < 4; i++) {
                                char h = peek();
                                boolean hex = (h >= '0' && h <= '9') ||
                                              (h >= 'a' && h <= 'f') ||
                                              (h >= 'A' && h <= 'F');
                                if (!hex) return false;
                                pos++;
                            }
                            break;
                        default:
                            return false;
                    }
                } else if (c < 0x20) {
                    return false; // raw control characters not allowed in strings
                } else {
                    pos++;
                }
            }
            return false; // unterminated string
        }

        /** number := '-'? int frac? exp? — no leading zeros, like JSON.parse. */
        private boolean number() {
            eat('-');
            if (peek() == '0') pos++;
            else if (peek() >= '1' && peek() <= '9') {
                while (peek() >= '0' && peek() <= '9') pos++;
            } else {
                return false;
            }
            if (peek() == '.') {
                pos++;
                int digits = 0;
                while (peek() >= '0' && peek() <= '9') { pos++; digits++; }
                if (digits == 0) return false;
            }
            if (peek() == 'e' || peek() == 'E') {
                pos++;
                if (peek() == '+' || peek() == '-') pos++;
                int digits = 0;
                while (peek() >= '0' && peek() <= '9') { pos++; digits++; }
                if (digits == 0) return false;
            }
            return true;
        }
    }

    /** Whole-text JSON gate: a document that parses as JSON is json all the
     *  way down. Mirrors {@code isValidJson()} ({@code JSON.parse} in a
     *  try/catch); empty/whitespace text is not. */
    static boolean isValidJson(String text) {
        if (text == null || text.strip().isEmpty()) return false;
        JsonParser p = new JsonParser(text);
        return p.value() && p.atEnd();
    }

    /** Token count of {@code text} under auto content detection — exactly the
     *  slice of {@code estimateTokens()} the planner consumes
     *  ({@code .tokens}): per non-empty line,
     *  {@code max(1, round(len / charsPerToken))}. Framing tokens are the
     *  caller's job. */
    static long estimateTokens(String text) {
        // AUTO + whole-text JSON: json's 3 chars/token rate applies to every
        // line, not just the reported content type.
        boolean wholeTextJson = isValidJson(text);
        long tokens = 0;
        for (String line : splitLines(text)) {
            if (line.strip().isEmpty()) continue;
            ContentType t = wholeTextJson ? ContentType.JSON : detectLineType(line);
            // Math.round is floor(x + 0.5) — JS Math.round, exactly.
            long lineTokens = Math.round(line.length() / charsPerToken(t));
            tokens += Math.max(1, lineTokens);
        }
        return tokens;
    }

    /** Sum of per-section token estimates (framing tokens are the caller's
     *  job). Mirrors {@code inputTokenTotal()} in the TS lib. */
    static long inputTokenTotal(List<PlanSection> sections) {
        long total = 0;
        for (PlanSection s : sections) total += estimateTokens(s.text());
        return total;
    }

    /** Plan one section set against one model's context window. Returns
     *  {@code Optional.empty()} for an unknown model id (window math is
     *  {@code fitsWindow}'s, never re-derived). Mirrors {@code planWindow()}
     *  in the TS lib. */
    static Optional<WindowPlan> planWindow(List<PlanSection> sections,
                                           String modelId, long outputReserve,
                                           List<Model> models) {
        long inputTokens = inputTokenTotal(sections);
        // Fit check inlined from fitsWindow() in src/lib/ai/models.ts.
        Model m = null;
        for (Model cand : models)
            if (cand.id().equals(modelId)) { m = cand; break; }
        if (m == null) return Optional.empty();
        long free = m.contextWindow() - inputTokens;
        return Optional.of(new WindowPlan(modelId, inputTokens, m.contextWindow(),
                free, free >= 0, free >= outputReserve, m.maxOutput()));
    }

    /** Plan against several models; unknown ids are dropped from the result.
     *  Mirrors {@code planAll()} in the TS lib. */
    static List<WindowPlan> planAll(List<PlanSection> sections,
                                    List<String> modelIds, long outputReserve,
                                    List<Model> models) {
        List<WindowPlan> plans = new ArrayList<>();
        for (String id : modelIds)
            planWindow(sections, id, outputReserve, models).ifPresent(plans::add);
        return plans;
    }

    // ---------- tests (showcase-only; the canonical suite lives in src/lib) ----------

    /** A 1600-char single line of 'a' is pure prose: 1600 / 4 = 400 tokens. */
    private static List<PlanSection> twoSections() {
        String lineA = "a".repeat(1600);
        return List.of(sec("sys", lineA), sec("docs", lineA));
    }

    public static void main(String[] args) {
        List<PlanSection> two = twoSections();

        // input totals
        check(inputTokenTotal(two) == 800);
        check(inputTokenTotal(List.of()) == 0);
        check(inputTokenTotal(List.of(sec("sys", ""))) == 0);

        // plans two 400-token sections against beta-pro
        WindowPlan p = planWindow(two, "beta-pro", 0, SAMPLE_MODELS).orElseThrow();
        check(p.id().equals("beta-pro"));
        check(p.inputTokens() == 800);
        check(p.contextWindow() == 1_000_000);
        check(p.free() == 999_200);
        check(p.fits());
        check(p.outputReserveOk());
        check(p.maxOutput() == 10_000);

        // reserve larger than free leaves raw fit true
        p = planWindow(two, "beta-pro", 1_000_000, SAMPLE_MODELS).orElseThrow();
        check(p.fits());
        check(!p.outputReserveOk());

        // reserve exactly equal to free is ok
        p = planWindow(two, "beta-pro", 999_200, SAMPLE_MODELS).orElseThrow();
        check(p.outputReserveOk());

        // smaller window leaves 199,200 free
        p = planWindow(two, "alpha-mini", 0, SAMPLE_MODELS).orElseThrow();
        check(p.contextWindow() == 200_000);
        check(p.free() == 199_200);
        check(p.fits());

        // unknown model id returns empty
        check(planWindow(two, "ghost", 0, SAMPLE_MODELS).isEmpty());

        // no sections: full window free
        p = planWindow(List.of(), "beta-pro", 0, SAMPLE_MODELS).orElseThrow();
        check(p.inputTokens() == 0);
        check(p.free() == 1_000_000);
        check(p.fits());

        // overflow: fits false, reserve false
        p = planWindow(List.of(sec("big", "z".repeat(4_400_000))), "beta-pro", 0,
                       SAMPLE_MODELS).orElseThrow();
        check(p.inputTokens() == 1_100_000);
        check(p.free() == -100_000);
        check(!p.fits());
        check(!p.outputReserveOk());

        // plan_all drops unknown ids and keeps order
        List<WindowPlan> plans = planAll(two,
                List.of("beta-pro", "alpha-mini", "ghost"), 0, SAMPLE_MODELS);
        check(plans.size() == 2);
        check(plans.get(0).id().equals("beta-pro"));
        check(plans.get(1).id().equals("alpha-mini"));
        check(plans.get(1).free() == 199_200);
        check(planAll(two, List.of(), 0, SAMPLE_MODELS).isEmpty());

        System.out.println("context-window-planner (Java): all tests passed");
    }

    private static void check(boolean cond) {
        if (!cond) throw new AssertionError("context-window-planner test failed");
    }
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →