Skip to content

Embedding Chunk Planner — Ruby source

Plan document chunking for RAG — chunk counts with overlap math, vector counts, and embedding costs per model.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# Embedding Chunk Planner — pure chunking math for RAG pipelines.
#
# Language: Ruby (Ruby 3.2, standard library only)
# Source:   CosmoDev polyglot showcase port of the Embedding Chunk Planner
#           tool, ported from src/lib/embeddingPlanner.ts (the canonical
#           TypeScript implementation).
# Tool page: https://dev.cosmolabs.org/tools/embedding-chunk-planner
# License:  display source — part of CosmoDev's polyglot tool pages.
#
# Design goals:
#   - Pure + deterministic; never raises (nil instead of an error).
#   - Functionally equivalent to the TS reference: same inputs -> same outputs.
#   - Self-contained: stdlib only (no gems). The model price table is inlined
#     below, mirrored from src/lib/ai/embeddings.ts — prices NEVER live in the
#     planner itself.
#
# Behavior (mirrors the TS source exactly):
#   - chunk_size <= 0 or total_tokens <= 0 -> ZERO_PLAN (nothing to embed).
#   - Negative overlap is treated as 0; overlap then clamps to at most
#     chunk_size / 2 so consecutive chunks always advance.
#   - chunks = max(1, ceil((total_tokens - overlap) / (chunk_size - overlap)))
#     — a tiny document still yields one chunk.

module EmbeddingChunkPlanner
  # One embedding model's offered dimensions (ascending, Matryoshka shortening
  # included) and pricing: USD per 1M input tokens.
  EmbeddingModel = Struct.new(:id, :vendor, :dims, :input_per_m, keyword_init: true)

  # How a document splits into overlapping chunks.
  ChunkPlan = Struct.new(:chunks, :total_tokens_with_overlap, :overhead_tokens, keyword_init: true)

  # Chunk plan plus pricing for one embedding call. The three chunk fields are
  # flattened in (the TS `...plan` spread) so the struct reads like the TS
  # `EmbeddingPlan extends ChunkPlan`. `vectors` is one per chunk; `cost` is
  # USD: total_tokens_with_overlap / 1e6 * model.input_per_m.
  EmbeddingPlan = Struct.new(
    :chunks, :total_tokens_with_overlap, :overhead_tokens, :model, :vectors, :cost,
    keyword_init: true
  )

  # Chunking knobs, in tokens. Mirrors the TS `Partial<ChunkOptions>`: each
  # field is independently optional — nil falls back to the 512 / 64 default,
  # and setting one leaves the other at its default.
  ChunkOptions = Struct.new(:chunk_size, :overlap, keyword_init: true)

  # Embedding model price table — the SSOT for pricing, mirrored from
  # src/lib/ai/embeddings.ts. Refresh both files together.
  EMBEDDING_MODELS = [
    EmbeddingModel.new(id: 'text-embedding-3-small', vendor: 'OpenAI',    dims: [512, 1536],           input_per_m: 0.02),
    EmbeddingModel.new(id: 'text-embedding-3-large', vendor: 'OpenAI',    dims: [256, 1024, 3072],     input_per_m: 0.13),
    EmbeddingModel.new(id: 'embed-english-v3.0',     vendor: 'Cohere',    dims: [512, 1024, 1536],     input_per_m: 0.1),
    EmbeddingModel.new(id: 'voyage-3-lite',          vendor: 'Voyage AI', dims: [512, 1024],           input_per_m: 0.02)
  ].freeze

  # Default knobs: 512-token chunks, 64-token overlap (TS DEFAULT_CHUNK_OPTIONS).
  DEFAULT_CHUNK_SIZE = 512
  DEFAULT_OVERLAP = 64

  # The "nothing to embed" plan the TS source returns for zero/negative input
  # or a non-positive chunk size.
  ZERO_PLAN = ChunkPlan.new(chunks: 0, total_tokens_with_overlap: 0, overhead_tokens: 0).freeze

  module_function

  # Look up an embedding model by id. Returns nil for unknown ids.
  def get_embedding_model(id)
    EMBEDDING_MODELS.find { |model| model.id == id }
  end

  # Plan how +total_tokens+ split into overlapping chunks. +opts+ may be nil
  # (both defaults), mirroring the TS optional parameter; each nil field falls
  # back to the 512 / 64 default independently. (0 is truthy in Ruby, so an
  # explicit chunk_size of 0 reaches the degenerate-input guard as-is.)
  def plan_chunks(total_tokens, opts = nil)
    chunk_size = opts&.chunk_size || DEFAULT_CHUNK_SIZE
    overlap_raw = opts&.overlap || DEFAULT_OVERLAP

    return ZERO_PLAN if chunk_size <= 0 || total_tokens <= 0

    # [ [overlap_raw, 0].max, chunk_size / 2 ].min — the TS clamp. Overlap that
    # large would never advance, so consecutive chunks always gain at least
    # half a chunk. (chunk_size >= 1 here, so chunk_size - overlap never hits
    # zero — Integer#/ truncates like TS's floor of a positive quotient.)
    overlap = [[overlap_raw, 0].max, chunk_size / 2].min

    # Float division + ceil mirrors TS's Math.ceil exactly (a tiny document
    # lands the quotient just below zero; ceil brings it to 0 and the max(1, ..)
    # lifts it back to one chunk).
    chunks = [1, (total_tokens - overlap).fdiv(chunk_size - overlap).ceil].max

    total_with_overlap = total_tokens + (chunks - 1) * overlap
    ChunkPlan.new(
      chunks: chunks,
      total_tokens_with_overlap: total_with_overlap,
      overhead_tokens: total_with_overlap - total_tokens
    )
  end

  # Chunk a document AND price its embedding for +model_id+ at +dims+
  # dimensions. Unknown model, or dims the model does not offer -> nil.
  def plan_embedding(total_tokens, model_id, dims, opts = nil)
    model = get_embedding_model(model_id)
    return nil if model.nil? || !model.dims.include?(dims)

    plan = plan_chunks(total_tokens, opts)
    EmbeddingPlan.new(
      chunks: plan.chunks,
      total_tokens_with_overlap: plan.total_tokens_with_overlap,
      overhead_tokens: plan.overhead_tokens,
      model: model,
      vectors: plan.chunks,
      cost: plan.total_tokens_with_overlap / 1e6 * model.input_per_m
    )
  end
end

# ---------- showcase examples (the canonical suite lives in src/lib) ----------
if __FILE__ == $PROGRAM_NAME
  # 1,000 tokens: ceil((1000-64)/(512-64)) = 3 chunks, 2 seams x 64.
  a = EmbeddingChunkPlanner.plan_chunks(1000)
  raise 'defaults example' unless a == EmbeddingChunkPlanner::ChunkPlan.new(
    chunks: 3, total_tokens_with_overlap: 1128, overhead_tokens: 128
  )

  # A document that fits one chunk has no seam overhead.
  b = EmbeddingChunkPlanner.plan_chunks(512)
  raise 'single chunk example' unless b.chunks == 1 && b.total_tokens_with_overlap == 512 && b.overhead_tokens.zero?

  # Zero/negative input or non-positive chunk size -> the zero plan.
  zero = EmbeddingChunkPlanner::ZERO_PLAN
  raise 'degenerate example' unless EmbeddingChunkPlanner.plan_chunks(0) == zero &&
                                    EmbeddingChunkPlanner.plan_chunks(-100) == zero &&
                                    EmbeddingChunkPlanner.plan_chunks(1000, EmbeddingChunkPlanner::ChunkOptions.new(chunk_size: 0)) == zero &&
                                    EmbeddingChunkPlanner.plan_chunks(1000, EmbeddingChunkPlanner::ChunkOptions.new(chunk_size: -8)) == zero

  # overlap 600 > floor(512/2) = 256 -> clamped to 256.
  clamped = EmbeddingChunkPlanner.plan_chunks(1000, EmbeddingChunkPlanner::ChunkOptions.new(overlap: 600))
  raise 'clamp example' unless clamped.chunks == 3 && clamped.total_tokens_with_overlap == 1512 && clamped.overhead_tokens == 512

  # Negative overlap clamps to 0: 1000 tokens -> ceil(1000/512) = 2 chunks.
  negative = EmbeddingChunkPlanner.plan_chunks(1000, EmbeddingChunkPlanner::ChunkOptions.new(overlap: -5))
  raise 'negative overlap example' unless negative.chunks == 2 && negative.total_tokens_with_overlap == 1000 && negative.overhead_tokens.zero?

  # chunk_size without overlap: ceil((1000-64)/192) = 5 chunks.
  partial = EmbeddingChunkPlanner.plan_chunks(1000, EmbeddingChunkPlanner::ChunkOptions.new(chunk_size: 256))
  raise 'partial options example' unless partial.chunks == 5 && partial.total_tokens_with_overlap == 1256 && partial.overhead_tokens == 256

  # Shorter than the overlap still yields one chunk.
  tiny = EmbeddingChunkPlanner.plan_chunks(50, EmbeddingChunkPlanner::ChunkOptions.new(overlap: 64))
  raise 'tiny document example' unless tiny.chunks == 1 && tiny.total_tokens_with_overlap == 50 && tiny.overhead_tokens.zero?

  # chunk_size of 1 clamps overlap to 0: ceil(3/1) = 3 chunks.
  ones = EmbeddingChunkPlanner.plan_chunks(3, EmbeddingChunkPlanner::ChunkOptions.new(chunk_size: 1))
  raise 'chunk size one example' unless ones.chunks == 3 && ones.total_tokens_with_overlap == 3 && ones.overhead_tokens.zero?

  # Pricing: 1,000 tokens on text-embedding-3-small @ 1536 dims.
  priced = EmbeddingChunkPlanner.plan_embedding(1000, 'text-embedding-3-small', 1536)
  raise 'pricing example' if priced.nil? || priced.chunks != 3 ||
                              priced.total_tokens_with_overlap != 1128 || priced.vectors != 3
  raise 'pricing cost example' unless (priced.cost - 0.00002256).abs < 1e-12 # 1128 / 1e6 * $0.02

  # A single-chunk document on voyage-3-lite @ 512 dims.
  single = EmbeddingChunkPlanner.plan_embedding(512, 'voyage-3-lite', 512)
  raise 'single pricing example' if single.nil? || single.vectors != 1
  raise 'single cost example' unless (single.cost - 0.00001024).abs < 1e-12 # 512 / 1e6 * $0.02

  # Unknown model or unoffered dims -> nil.
  raise 'unknown dims example' unless EmbeddingChunkPlanner.plan_embedding(1000, 'text-embedding-3-small', 999).nil?
  raise 'unknown model example' unless EmbeddingChunkPlanner.plan_embedding(1000, 'ghost', 1536).nil?

  # Zero tokens price out to a zero-cost plan.
  zero_priced = EmbeddingChunkPlanner.plan_embedding(0, 'text-embedding-3-small', 1536)
  raise 'zero pricing example' if zero_priced.nil? || zero_priced.chunks != 0 ||
                                   zero_priced.vectors != 0 || zero_priced.cost != 0.0

  puts 'all showcase examples passed'
end

Also available in 12 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →