Skip to content

Rate Limit Planner — Ruby source

Turn RPM/TPM limits into a concrete request schedule — batch size, spacing, binding limit, and total run time, with a safety factor for retries. 100% client-side.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# Rate Limit Planner — turn provider rate limits plus a workload into a
# concrete schedule: batch size, spacing, what caps it, and finish time.
# Deterministic — no clock reads.
#
# Language: Ruby (3.x, zero dependencies)
# Port of src/lib/rateLimitPlanner.ts (the canonical TypeScript
# implementation). Method names are snake_case per Ruby convention.
# Tool page: https://dev.cosmolabs.org/tools/rate-limit-planner

module RateLimitPlanner
  WINDOW_MS = 60_000
  DEFAULT_SAFETY = 0.8

  # Requests/tokens per minute; nil = not limited.
  Limits = Struct.new(:rpm, :tpm, keyword_init: true)

  Workload = Struct.new(:requests, :avg_tokens_per_request, keyword_init: true)

  Options = Struct.new(:safety_factor, keyword_init: true)

  BatchSlice = Struct.new(:batch, :at_ms, :requests, :tokens, keyword_init: true)

  Plan = Struct.new(
    :batch_size, :interval_ms, :max_concurrent, :bounded_by,
    :timeline, :total_ms, :warnings, keyword_init: true
  )

  module_function

  def fmt(v)
    v.to_s.reverse.gsub(/(\d{3})(?=\d)/, '\1,').reverse
  end

  # Plan a schedule. Raises ArgumentError where TS raises RangeError.
  def plan_rate_limit(limits, workload, opts = Options.new)
    sf = opts.safety_factor || DEFAULT_SAFETY
    warnings = []
    if workload.requests.negative? || workload.avg_tokens_per_request.negative?
      raise ArgumentError, 'requests and avgTokensPerRequest must be >= 0'
    end
    raise ArgumentError, 'safetyFactor must be in (0, 1]' if sf <= 0 || sf > 1

    rpm_eff = limits.rpm ? limits.rpm * sf : nil
    tpm_eff = limits.tpm ? limits.tpm * sf : nil

    # Impossible: one request alone exceeds the token budget.
    if tpm_eff && workload.avg_tokens_per_request > tpm_eff && workload.requests.positive?
      return Plan.new(
        batch_size: 0, interval_ms: 0, max_concurrent: 0, bounded_by: 'tpm',
        timeline: [], total_ms: Float::INFINITY,
        warnings: [
          "A single request averages #{fmt(workload.avg_tokens_per_request)} tokens but " \
          "the effective token limit is #{fmt(tpm_eff.floor)}/min — no schedule can run " \
          'this. Shrink requests or raise the tier.'
        ]
      )
    end

    by_rpm = rpm_eff || Float::INFINITY
    by_tokens = if tpm_eff.nil? || workload.avg_tokens_per_request.zero?
                  Float::INFINITY
                else
                  tpm_eff / workload.avg_tokens_per_request
                end

    if by_rpm.infinite? && by_tokens.infinite?
      warnings << 'No limits set — the plan assumes an unbounded endpoint. ' \
                  'Add RPM or TPM for a real schedule.'
    end

    steady = [1, [by_rpm, by_tokens].min.floor].max
    bounded_by =
      if by_rpm.infinite? && by_tokens.infinite? then 'none'
      elsif by_rpm.floor == by_tokens.floor then 'both'
      elsif by_rpm < by_tokens then 'rpm'
      else 'tpm'
      end

    # Even pacing inside the window: batch_size requests spread over 60s.
    interval_ms = (WINDOW_MS.to_f / steady).round
    # Concurrency >1 only helps sub-interval latencies; the safe published
    # floor is 1 — batch bursts raise it to batch_size/4.
    max_concurrent = steady == 1 ? 1 : [steady, (steady / 4.0).ceil].min

    timeline = []
    remaining = workload.requests
    batch = 0
    while remaining.positive? && batch < 10
      take = [steady, remaining].min
      timeline << BatchSlice.new(
        batch: batch + 1,
        at_ms: batch * WINDOW_MS,
        requests: take,
        tokens: take * workload.avg_tokens_per_request
      )
      remaining -= take
      batch += 1
    end

    windows_needed = workload.requests.positive? ? (workload.requests.to_f / steady).ceil : 0
    last_window_requests = windows_needed.positive? ? workload.requests - (windows_needed - 1) * steady : 0
    total_ms = windows_needed.positive? ? (windows_needed - 1) * WINDOW_MS + interval_ms * last_window_requests : 0

    if rpm_eff && workload.requests.positive? && steady > by_rpm
      warnings << 'Rounded up to at least one request per window — even a single ' \
                  'request per minute keeps the schedule honest.'
    end

    Plan.new(
      batch_size: steady, interval_ms: interval_ms, max_concurrent: max_concurrent,
      bounded_by: bounded_by, timeline: timeline, total_ms: total_ms, warnings: warnings
    )
  end

  # Human summary line for the plan.
  def describe_plan(plan)
    return 'No viable schedule.' if plan.batch_size.zero?
    if plan.bounded_by == 'none'
      return "#{plan.batch_size}+ requests per window — endpoint treated as unbounded."
    end

    limiter = if plan.bounded_by == 'both'
                'both limits bind together'
              else
                "the #{plan.bounded_by.upcase} limit binds first"
              end
    "#{plan.batch_size} requests per 60s window (one every #{plan.interval_ms}ms) — #{limiter}."
  end
end

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →