Skip to content

LLM Cost Calculator — C++ source

Estimate LLM costs per request or per month at billion-token scale — with realistic prompt-cache hit rates, four-lane pricing, and side-by-side model comparison from a dated pricing snapshot.

This is the C++ implementation — the same logic the interactive tool runs, in a shareable, citable form.

// llm-cost-calculator — C++ port: per-million-token cost math for LLM workloads.
#include <algorithm>
#include <iomanip>
#include <iostream>
#include <optional>
#include <string>
#include <vector>
// Multiplier applied to Batch API pricing (the standard 50% discount).
inline constexpr double kBatchDiscount = 0.5;
// Price pair for a model or a custom rate card. nullopt = unpriced.
struct CostRates {
    std::optional<double> input_per_m;  // USD per 1M input tokens.
    std::optional<double> output_per_m; // USD per 1M output tokens.
};
// One workload to price.
struct CostInput {
    long input_tokens;   // Input tokens per request.
    long output_tokens;  // Output tokens per request.
    long requests = 1;   // Request count.
    bool batch = false;  // Apply the batch discount multiplier.
};
// Pricing projection of a model — the only three fields cost math needs.
struct Model {
    std::string id;
    std::optional<double> input_per_m;
    std::optional<double> output_per_m;
};
struct ModelCost {
    std::string id;
    std::optional<double> cost;              // CostFor with the model's rates.
    std::optional<double> tokens_per_dollar; // 1e6 / output_per_m.
};
// Cost in USD, or nullopt when either rate is unpriced:
// ((in/1e6)*input + (out/1e6)*output) * requests * (batch ? 0.5 : 1).
std::optional<double> CostFor(const CostRates &r, const CostInput &w) {
    if (!r.input_per_m || !r.output_per_m) return std::nullopt;
    double base = w.input_tokens / 1e6 * *r.input_per_m
                + w.output_tokens / 1e6 * *r.output_per_m;
    return base * w.requests * (w.batch ? kBatchDiscount : 1.0);
}
// Output tokens per USD: 1e6 / output_per_m, or nullopt when unpriced.
std::optional<double> TokensPerDollar(const Model &m) {
    if (!m.output_per_m) return std::nullopt;
    return 1e6 / *m.output_per_m;
}
// Cost every model for one workload, sorted: cost asc, unpriced last, id asc.
std::vector<ModelCost> CompareModels(std::vector<Model> models, const CostInput &w) {
    std::vector<ModelCost> rows;
    for (const auto &m : models)
        rows.push_back({m.id, CostFor({m.input_per_m, m.output_per_m}, w),
                        TokensPerDollar(m)});
    std::sort(rows.begin(), rows.end(), [](const ModelCost &a, const ModelCost &b) {
        if (a.cost && b.cost) {
            if (*a.cost != *b.cost) return *a.cost < *b.cost;
            return a.id < b.id;
        }
        if (a.cost.has_value() != b.cost.has_value()) return a.cost.has_value();
        return a.id < b.id;
    });
    return rows;
}
int main() {
    // Rate cards flow from the model snapshot in the TS lib; one is unpriced.
    const std::vector<Model> models = {
        {"sonnet-class", 3.0, 15.0},
        {"budget-class", 0.5, 2.0},
        {"preview", std::nullopt, std::nullopt}, // unpriced model
    };
    const CostInput work{.input_tokens = 12000, .output_tokens = 3000,
                         .requests = 250, .batch = true};
    for (const auto &row : CompareModels(models, work)) {
        std::cout << std::left << std::setw(14) << row.id;
        if (row.cost)
            std::cout << "$" << *row.cost << "  (" << *row.tokens_per_dollar
                      << " output tokens/USD)\n";
        else
            std::cout << "unpriced\n";
    }
    return 0;
}

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →