Skip to content

Mock LLM Responder — Ruby source

Generate deterministic mock LLM API responses - chat completion JSON, SSE event streams with chunk timing, and a replay curl - for testing clients without an API key.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# mock-llm-responder — Ruby port: deterministic mock LLM responses (seeded PRNG + token math).

POEM_WORDS = %w[
  cosmos nebula quantum signal photon drift
  orbit vector cipher lumen aurora echo
  helix nova pulse tide vertex zenith
  quasar ion halo flux prism comet
].freeze

# FNV-1a 32-bit hash — turns the spec into a deterministic seed / id.
def hash_string(s)
  h = 0x811c9dc5
  s.each_char { |c| h = ((h ^ c.ord) * 0x01000193) & 0xffffffff }
  h
end

# mulberry32 — tiny seeded PRNG; same seed, same sequence, forever.
# 32-bit wrapping is explicit: Ruby ints are arbitrary precision.
def mulberry32(seed)
  a = seed & 0xffffffff
  lambda do
    a = (a + 0x6d2b79f5) & 0xffffffff
    t = ((a ^ (a >> 15)) * (a | 1)) & 0xffffffff
    t = ((t + (((t ^ (t >> 7)) * (t | 61)) & 0xffffffff)) & 0xffffffff) ^ t
    (t ^ (t >> 14)) / 4294967296.0
  end
end

# ~4 chars per token, floor of 1 — deterministic, no tokenizer needed.
def token_count(text)
  return 0 if text.empty?

  [(text.length / 4.0).ceil, 1].max
end

# Cut text so it fits in max_tokens tokens (4 chars each).
def truncate_to_tokens(text, max_tokens)
  return text if token_count(text) <= max_tokens

  text[0, max_tokens * 4].rstrip
end

# One poem line of 5-7 vocabulary words.
def make_line(rng)
  n = 5 + (rng.call * 3).floor
  Array.new(n) { POEM_WORDS[(rng.call * POEM_WORDS.length).floor] }.join(' ')
end

# Poem-ish lorem, grown line by line until the token budget is full.
def build_poem(seed, max_tokens)
  rng = mulberry32(seed)
  text = ''
  loop do
    line = make_line(rng)
    candidate = text.empty? ? line : "#{text}\n#{line}"
    break if !text.empty? && token_count(candidate) > max_tokens

    text = candidate
  end
  truncate_to_tokens(text, max_tokens)
end

# Split content into stream chunks. Chunks reassemble to the exact content.
def chunk_content(content, per_chunk)
  content.scan(/\S+\s*/).each_slice(per_chunk).map(&:join)
end

spec = 'streamed-lorem|mock-gpt-4o-mini|24'
puts format('id=chatcmpl-mock-%08x', hash_string("#{spec}|id"))
poem = build_poem(hash_string("#{spec}|poem"), 24)
puts poem
puts "tokens=#{token_count(poem)} chunks=#{chunk_content(poem, 4).size}"

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →