Skip to content

robots.txt Generator — Ruby source

Build a standards-compliant robots.txt with per-user-agent allow/disallow rules, crawl-delay, and sitemap entries.

This is the Ruby implementation — the same logic the interactive tool runs, in a shareable, citable form.

# robots-txt-generator — Ruby port: standards-compliant robots.txt generator + parser.
#
# Display port of the CosmoDev Robots.txt Generator tool — same contract as
# cli/robots-txt-generator/robots-txt-generator.go (the live Go twin) and
# src/lib/robotsTxt.ts (canonical TypeScript). Pure + deterministic; stdlib
# only. A group carries a LIST of user-agents (the REP spec's stacked
# User-agent lines); parsing aggregates repeated blocks per agent, then
# re-merges agents whose accumulated rules are identical into one stacked
# group, exactly like the TS lib.

module RobotsTxt
  RuleGroup = Struct.new(:user_agents, :disallow, :allow, :crawl_delay, keyword_init: true)
  RobotsConfig = Struct.new(:groups, :sitemaps, keyword_init: true)

  module_function

  # cleanAgents in the TS lib: trim + drop blanks. A non-empty list that trims
  # away entirely degrades to ['*']; an empty list means "no group yet".
  # (Struct fields default to nil, so coerce to [] first.)
  def clean_agents(list)
    list = Array(list)
    trimmed = list.map(&:strip).reject(&:empty?)
    trimmed.empty? && !list.empty? ? ['*'] : trimmed
  end

  # JS-style number rendering for Crawl-delay: 10.0 -> '10', 1.5 -> '1.5'.
  def num_str(f)
    f == f.truncate ? f.truncate.to_s : f.to_s
  end

  def generate_robots(config)
    out = []
    config.groups.each do |g|
      uas = clean_agents(g.user_agents)
      next if uas.empty?

      uas.each { |ua| out << "User-agent: #{ua}" }
      Array(g.allow).each { |a| out << "Allow: #{a.strip}" unless a.strip.empty? }
      disallow = Array(g.disallow)
      if disallow.empty?
        out << 'Disallow:' # no entries -> the allow-all marker
      else
        disallow.each { |d| out << "Disallow: #{d}" } # '' kept verbatim
      end
      out << "Crawl-delay: #{num_str(g.crawl_delay)}" if g.crawl_delay&.finite?
      out << '' # blank line separates groups
    end
    config.sitemaps.each { |s| out << "Sitemap: #{s.strip}" unless s.strip.empty? }

    out.join("\n").gsub(/\n{3,}/, "\n\n").rstrip + "\n"
  end

  # Move the pending User-agent stack into per-agent rule buckets (Google's
  # merge semantics: a rule line applies to every agent of the immediately
  # preceding consecutive User-agent stack). Returns the active bucket list.
  def flush_stack(stack, by_ua, groups)
    current = stack.map do |ua|
      by_ua[ua] ||= RuleGroup.new(user_agents: [ua], disallow: [], allow: []).tap { |g| groups << g }
    end
    stack.clear
    current
  end

  # ParseRobots in the Go twin: strip '#'-comments, split on the FIRST ':'
  # (values may contain colons), unknown directives ignored. to_f coercion
  # differs from JS Number() only for garbage values (0.0 vs NaN) — both are
  # dropped or merge as unset in practice.
  def parse_robots(text)
    groups = []
    by_ua = {}
    sitemaps = []
    stack = []
    current = []

    (text || '').split("\n").each do |raw|
      line = raw.sub(/#.*$/, '').strip
      next if line.empty?
      field, value = line.split(':', 2)
      next if value.nil?
      value = value.strip

      case field.strip.downcase
      when 'user-agent'
        ua = value.empty? ? '*' : value
        stack << ua unless stack.include?(ua)
      when 'disallow'
        current = flush_stack(stack, by_ua, groups) unless stack.empty?
        current.each { |g| g.disallow << value }
      when 'allow'
        current = flush_stack(stack, by_ua, groups) unless stack.empty?
        current.each { |g| g.allow << value }
      when 'crawl-delay'
        current = flush_stack(stack, by_ua, groups) unless stack.empty?
        current.each { |g| g.crawl_delay = value.to_f }
      when 'sitemap'
        current = flush_stack(stack, by_ua, groups) unless stack.empty?
        sitemaps << value
      end
    end

    # Merge agents with identical rule sets into one stacked group.
    merged = []
    groups.each do |g|
      twin = merged.find { |m| [m.disallow, m.allow, m.crawl_delay] == [g.disallow, g.allow, g.crawl_delay] }
      if twin
        twin.user_agents.concat(g.user_agents)
      else
        merged << RuleGroup.new(user_agents: g.user_agents.dup, disallow: g.disallow,
                                allow: g.allow, crawl_delay: g.crawl_delay)
      end
    end
    RobotsConfig.new(groups: merged, sitemaps: sitemaps)
  end
end

# Sample config -> generated body -> parse round-trip check (tool page behavior).
config = RobotsTxt::RobotsConfig.new(
  groups: [
    RobotsTxt::RuleGroup.new(user_agents: ['*'], allow: ['/admin/public/'],
                             disallow: ['/private/', '/admin/']),
    RobotsTxt::RuleGroup.new(user_agents: ['GPTBot', 'CCBot'], disallow: ['/'], crawl_delay: 10)
  ],
  sitemaps: ['https://example.com/sitemap.xml']
)
body = RobotsTxt.generate_robots(config)
puts body
puts "round-trip stable: #{RobotsTxt.generate_robots(RobotsTxt.parse_robots(body)) == body}"

Also available in 13 other languages

Every CosmoDev tool ships its pure logic in TypeScript (web) and Go (CLI), with authored implementations in a dozen-plus languages — the same contract, ported. Compare all languages side by side →