robots-txt-gen
Generate, validate, and test robots.txt files from the command line.
Quick Start
# Generate a robots.txt for a platform
python3 scripts/robots_txt_gen.py generate --preset nextjs --sitemap https://example.com/sitemap.xml
# Validate an existing robots.txt
python3 scripts/robots_txt_gen.py validate --file robots.txt
# Validate a remote robots.txt
python3 scripts/robots_txt_gen.py validate --url https://example.com/robots.txt
# Test if a URL is allowed for a user-agent
python3 scripts/robots_txt_gen.py test --file robots.txt --url /admin/dashboard --agent Googlebot
# Generate with custom rules
python3 scripts/robots_txt_gen.py generate --allow "/" --disallow "/admin" --disallow "/api" --disallow "/private" --sitemap https://example.com/sitemap.xml --agent "*"
Commands
generate
Create a robots.txt file with custom rules or platform presets.
Options:
--preset <name> — Use a platform preset: wordpress, nextjs, django, rails, laravel, static, spa, ecommerce
--agent <name> — User-agent (default: *). Repeat for multiple agents.
--allow <path> — Allow path. Repeatable.
--disallow <path> — Disallow path. Repeatable.
--sitemap <url> — Sitemap URL. Repeatable.
--crawl-delay <seconds> — Crawl delay directive.
--block-ai — Add rules to block common AI crawlers (GPTBot, ChatGPT-User, CCBot, Google-Extended, anthropic-ai, etc.)
--output <file> — Write to file instead of stdout.
validate
Check a robots.txt file for syntax errors and best-practice warnings.
Options:
--file <path> — Local file to validate.
--url <url> — Remote robots.txt URL to fetch and validate.
test
Test whether a specific URL path is allowed or disallowed for a given user-agent.
Options:
--file <path> — robots.txt file to test against.
--url <path> — URL path to test (e.g., /admin/login).
--agent <name> — User-agent to test as (default: Googlebot).
Platform Presets
| Preset |
What it blocks |
Notes |
wordpress |
/wp-admin/, /wp-includes/, query params |
Allows /wp-admin/admin-ajax.php |
nextjs |
/_next/static/, /api/, /.next/ |
Standard Next.js paths |
django |
/admin/, /static/admin/, /media/private/ |
Django admin and private media |
rails |
/admin/, /assets/, /tmp/ |
Rails conventions |
laravel |
/admin/, /storage/, /vendor/ |
Laravel conventions |
static |
Nothing blocked |
Simple allow-all with sitemap |
spa |
/api/, /assets/ |
Single-page app pattern |
ecommerce |
/cart/, /checkout/, /account/, /search? |
Prevents crawling user sessions |
AI Crawler Blocking
The --block-ai flag adds disallow rules for known AI training crawlers:
- GPTBot, ChatGPT-User (OpenAI)
- Google-Extended (Google AI)
- CCBot (Common Crawl)
- anthropic-ai (Anthropic)
- Bytespider (ByteDance)
- ClaudeBot (Anthropic)
- FacebookBot (Meta)
1---2name: robots-txt-gen3description: Generate, validate, and analyze robots.txt files for websites. Use when creating robots.txt from scratch, validating existing robots.txt syntax, checking if a URL is allowed/blocked by robots.txt rules, or generating robots.txt for common platforms (WordPress, Next.js, Django, Rails). Also use when auditing crawl directives or debugging search engine indexing issues.4---56# robots-txt-gen78Generate, validate, and test robots.txt files from the command line.910## Quick Start1112```bash13# Generate a robots.txt for a platform14python3 scripts/robots_txt_gen.py generate --preset nextjs --sitemap https://example.com/sitemap.xml1516# Validate an existing robots.txt17python3 scripts/robots_txt_gen.py validate --file robots.txt1819# Validate a remote robots.txt20python3 scripts/robots_txt_gen.py validate --url https://example.com/robots.txt2122# Test if a URL is allowed for a user-agent23python3 scripts/robots_txt_gen.py test --file robots.txt --url /admin/dashboard --agent Googlebot2425# Generate with custom rules26python3 scripts/robots_txt_gen.py generate --allow "/" --disallow "/admin" --disallow "/api" --disallow "/private" --sitemap https://example.com/sitemap.xml --agent "*"27```2829## Commands3031### `generate`32Create a robots.txt file with custom rules or platform presets.3334Options:35- `--preset <name>` — Use a platform preset: `wordpress`, `nextjs`, `django`, `rails`, `laravel`, `static`, `spa`, `ecommerce`36- `--agent <name>` — User-agent (default: `*`). Repeat for multiple agents.37- `--allow <path>` — Allow path. Repeatable.38- `--disallow <path>` — Disallow path. Repeatable.39- `--sitemap <url>` — Sitemap URL. Repeatable.40- `--crawl-delay <seconds>` — Crawl delay directive.41- `--block-ai` — Add rules to block common AI crawlers (GPTBot, ChatGPT-User, CCBot, Google-Extended, anthropic-ai, etc.)42- `--output <file>` — Write to file instead of stdout.4344### `validate`45Check a robots.txt file for syntax errors and best-practice warnings.4647Options:48- `--file <path>` — Local file to validate.49- `--url <url>` — Remote robots.txt URL to fetch and validate.5051### `test`52Test whether a specific URL path is allowed or disallowed for a given user-agent.5354Options:55- `--file <path>` — robots.txt file to test against.56- `--url <path>` — URL path to test (e.g., `/admin/login`).57- `--agent <name>` — User-agent to test as (default: `Googlebot`).5859## Platform Presets6061| Preset | What it blocks | Notes |62|--------|---------------|-------|63| `wordpress` | `/wp-admin/`, `/wp-includes/`, query params | Allows `/wp-admin/admin-ajax.php` |64| `nextjs` | `/_next/static/`, `/api/`, `/.next/` | Standard Next.js paths |65| `django` | `/admin/`, `/static/admin/`, `/media/private/` | Django admin and private media |66| `rails` | `/admin/`, `/assets/`, `/tmp/` | Rails conventions |67| `laravel` | `/admin/`, `/storage/`, `/vendor/` | Laravel conventions |68| `static` | Nothing blocked | Simple allow-all with sitemap |69| `spa` | `/api/`, `/assets/` | Single-page app pattern |70| `ecommerce` | `/cart/`, `/checkout/`, `/account/`, `/search?` | Prevents crawling user sessions |7172## AI Crawler Blocking7374The `--block-ai` flag adds disallow rules for known AI training crawlers:75- GPTBot, ChatGPT-User (OpenAI)76- Google-Extended (Google AI)77- CCBot (Common Crawl)78- anthropic-ai (Anthropic)79- Bytespider (ByteDance)80- ClaudeBot (Anthropic)81- FacebookBot (Meta)