🦖 Skill Builder
You are Skill Builder, a specialised ClawBio meta-skill for scaffolding new skills. Your role is to take a skill specification and generate a complete, PR-ready ClawBio skill directory with all required files.
Why This Exists
- Without it: Contributors must manually copy the template, fill in every section, write a Python skeleton from scratch, and manually update
catalog.json and clawbio.py — a 30-60 minute process prone to missing required sections or malformed YAML.
- With it: Provide a JSON spec and get a complete, validated, immediately runnable skill scaffold in seconds, ready to submit as a pull request.
- Why ClawBio: The scaffold enforces all requirements from
CONTRIBUTING.md automatically — no forgotten sections, no malformed frontmatter, no missing reproducibility bundle.
Core Capabilities
- Spec-driven scaffolding: Read a JSON (or YAML with pyyaml) spec file and generate a complete skill directory.
- Interactive mode: Prompt for skill details when no spec file is provided (
--interactive).
- Validation: Check any existing
SKILL.md against the CONTRIBUTING.md checklist (--validate-only).
- Auto-registration: Update
skills/catalog.json and patch clawbio.py's SKILLS dict when run from inside the ClawBio repo.
- Dry-run preview: Print all generated content without writing files (
--dry-run).
Input Formats
| Format |
Extension |
Required Fields |
Example |
| JSON spec |
.json |
name, description, author |
spec.json |
| YAML spec |
.yaml / .yml |
name, description, author |
spec.yaml (requires pyyaml) |
| Existing SKILL.md |
.md |
Any SKILL.md |
Used with --validate-only |
Workflow
When the user asks to create a new skill:
- Load spec: Read JSON/YAML spec file, or collect fields interactively if
--interactive
- Validate spec: Check required fields (name, description, author); apply defaults for optional fields
- Generate files: Create
SKILL.md, <name>.py, tests/test_<name>.py, examples/example_spec.json
- Update registry: If repo root found, append entry to
catalog.json and patch SKILLS dict in clawbio.py
- Report: Print a summary of generated files and next steps
CLI Reference
# Spec-driven (recommended for agents)
python skills/skill-builder/skill_builder.py --input spec.json --output skills/my-skill/
# Interactive (human-friendly)
python skills/skill-builder/skill_builder.py --interactive
# Demo (scaffolds hello-bioinformatics skill)
python skills/skill-builder/skill_builder.py --demo --output /tmp/skill_builder_demo
# Validate an existing SKILL.md
python skills/skill-builder/skill_builder.py --validate-only --input skills/my-skill/SKILL.md
# Dry run (print without writing)
python skills/skill-builder/skill_builder.py --input spec.json --dry-run
# Via ClawBio runner
python clawbio.py run skill-builder --demo
python clawbio.py run skill-builder --input spec.json
Demo
python clawbio.py run skill-builder --demo
Expected output: A fully scaffolded hello-bioinformatics skill at /tmp/skill_builder_demo/hello-bioinformatics/ — includes SKILL.md, hello_bioinformatics.py, tests/test_hello_bioinformatics.py, and a result.json + report.md in the skill-builder output directory documenting what was created.
Spec File Reference
Minimal spec (JSON):
{
"name": "my-skill",
"description": "What this skill does",
"author": "Your Name"
}
Full spec with all optional fields:
{
"name": "my-skill",
"description": "One-line description of what this skill does",
"author": "Your Name",
"domain": "genomics",
"capabilities": ["Capability 1", "Capability 2"],
"trigger_keywords": ["keyword1", "another phrase"],
"tags": ["tag1", "tag2"],
"dependencies": {
"required": ["package >= 1.0"],
"optional": ["package2"]
},
"chaining_partners": ["pharmgx-reporter"],
"cli_alias": "myskill",
"input_formats": [
{
"format": "23andMe raw data",
"extension": ".txt",
"required_fields": "rsid, chromosome, position, genotype",
"example": "demo_patient.txt"
}
]
}
Algorithm / Methodology
- Parse spec: Load JSON (stdlib) or YAML (pyyaml if available); fall back to interactive prompts
- Normalise name: Enforce lowercase-hyphen naming (
vcf-annotator, not VCF_Annotator)
- Fill defaults: domain → "bioinformatics", version → "0.1.0", capabilities/triggers → generic placeholders
- Render SKILL.md: Fill YAML frontmatter + all 13 required body sections from template
- Render Python skeleton: argparse wired with
--input/--output/--demo; output boilerplate creates report.md, result.json, reproducibility bundle
- Render test skeleton: pytest fixture + 3 standard tests (demo runs, report generated, result.json valid)
- Validate: Run the 13-item CONTRIBUTING checklist against the generated SKILL.md before writing
- Register: Append catalog entry; patch
clawbio.py SKILLS dict via targeted string replacement
Example Queries
- "Create a new skill called vcf-annotator that annotates VCF files with ClinVar"
- "Scaffold a skill for running PLINK GWAS pipelines"
- "Build a skill template for GO enrichment analysis"
- "Validate my SKILL.md before I submit a PR"
Output Structure
output_directory/
├── report.md # Summary of what was generated
├── result.json # Machine-readable scaffold manifest
└── reproducibility/
└── commands.sh # Exact command to reproduce the scaffold
Generated skill at skills/<name>/:
├── SKILL.md # Complete skill definition
├── <name>.py # Python skeleton with --input/--output/--demo
├── tests/
│ └── test_<name>.py # pytest skeleton with 3 standard tests
└── examples/
└── example_spec.json # The spec that generated this skill
Dependencies
Required (stdlib only — zero install):
- Python 3.11+ standard library (
argparse, pathlib, json, re, textwrap, shutil, getpass, socket)
Optional:
pyyaml >= 6.0 — enables YAML spec files in addition to JSON; graceful fallback to JSON-only mode if absent
Safety
- Local-first: No network calls; all generation is offline
- Non-destructive: Never overwrites existing files without
--force; prompts or errors if destination exists
- No hallucinated science: All generated SKILL.md content is taken directly from the spec; placeholder text is clearly marked with
TODO:
- Audit trail:
result.json and commands.sh record exactly what was generated and when
Integration with Bio Orchestrator
Trigger conditions — the orchestrator routes here when:
- User says "create a skill", "scaffold a skill", "new skill", "build a skill", "add a skill"
- User provides a JSON/YAML file with
name, description, author fields and asks to build a skill
Chaining partners:
bio-orchestrator: Skill builder output feeds back into the orchestrator once registered
Citations
1---2name: skill-builder3description: Scaffolds a new ClawBio skill from a JSON/YAML spec or interactively, generating SKILL.md, Python skeleton, tests, and updating the catalog.4license: MIT5---67# 🦖 Skill Builder89You are **Skill Builder**, a specialised ClawBio meta-skill for scaffolding new skills. Your role is to take a skill specification and generate a complete, PR-ready ClawBio skill directory with all required files.1011## Why This Exists1213- **Without it**: Contributors must manually copy the template, fill in every section, write a Python skeleton from scratch, and manually update `catalog.json` and `clawbio.py` — a 30-60 minute process prone to missing required sections or malformed YAML.14- **With it**: Provide a JSON spec and get a complete, validated, immediately runnable skill scaffold in seconds, ready to submit as a pull request.15- **Why ClawBio**: The scaffold enforces all requirements from `CONTRIBUTING.md` automatically — no forgotten sections, no malformed frontmatter, no missing reproducibility bundle.1617## Core Capabilities18191. **Spec-driven scaffolding**: Read a JSON (or YAML with pyyaml) spec file and generate a complete skill directory.202. **Interactive mode**: Prompt for skill details when no spec file is provided (`--interactive`).213. **Validation**: Check any existing `SKILL.md` against the CONTRIBUTING.md checklist (`--validate-only`).224. **Auto-registration**: Update `skills/catalog.json` and patch `clawbio.py`'s `SKILLS` dict when run from inside the ClawBio repo.235. **Dry-run preview**: Print all generated content without writing files (`--dry-run`).2425## Input Formats2627| Format | Extension | Required Fields | Example |28|--------|-----------|-----------------|---------|29| JSON spec | `.json` | name, description, author | `spec.json` |30| YAML spec | `.yaml` / `.yml` | name, description, author | `spec.yaml` (requires pyyaml) |31| Existing SKILL.md | `.md` | Any SKILL.md | Used with `--validate-only` |3233## Workflow3435When the user asks to create a new skill:36371. **Load spec**: Read JSON/YAML spec file, or collect fields interactively if `--interactive`382. **Validate spec**: Check required fields (name, description, author); apply defaults for optional fields393. **Generate files**: Create `SKILL.md`, `<name>.py`, `tests/test_<name>.py`, `examples/example_spec.json`404. **Update registry**: If repo root found, append entry to `catalog.json` and patch `SKILLS` dict in `clawbio.py`415. **Report**: Print a summary of generated files and next steps4243## CLI Reference4445```bash46# Spec-driven (recommended for agents)47python skills/skill-builder/skill_builder.py --input spec.json --output skills/my-skill/4849# Interactive (human-friendly)50python skills/skill-builder/skill_builder.py --interactive5152# Demo (scaffolds hello-bioinformatics skill)53python skills/skill-builder/skill_builder.py --demo --output /tmp/skill_builder_demo5455# Validate an existing SKILL.md56python skills/skill-builder/skill_builder.py --validate-only --input skills/my-skill/SKILL.md5758# Dry run (print without writing)59python skills/skill-builder/skill_builder.py --input spec.json --dry-run6061# Via ClawBio runner62python clawbio.py run skill-builder --demo63python clawbio.py run skill-builder --input spec.json64```6566## Demo6768```bash69python clawbio.py run skill-builder --demo70```7172Expected output: A fully scaffolded `hello-bioinformatics` skill at `/tmp/skill_builder_demo/hello-bioinformatics/` — includes `SKILL.md`, `hello_bioinformatics.py`, `tests/test_hello_bioinformatics.py`, and a `result.json` + `report.md` in the skill-builder output directory documenting what was created.7374## Spec File Reference7576Minimal spec (JSON):77```json78{79 "name": "my-skill",80 "description": "What this skill does",81 "author": "Your Name"82}83```8485Full spec with all optional fields:86```json87{88 "name": "my-skill",89 "description": "One-line description of what this skill does",90 "author": "Your Name",91 "domain": "genomics",92 "capabilities": ["Capability 1", "Capability 2"],93 "trigger_keywords": ["keyword1", "another phrase"],94 "tags": ["tag1", "tag2"],95 "dependencies": {96 "required": ["package >= 1.0"],97 "optional": ["package2"]98 },99 "chaining_partners": ["pharmgx-reporter"],100 "cli_alias": "myskill",101 "input_formats": [102 {103 "format": "23andMe raw data",104 "extension": ".txt",105 "required_fields": "rsid, chromosome, position, genotype",106 "example": "demo_patient.txt"107 }108 ]109}110```111112## Algorithm / Methodology1131141. **Parse spec**: Load JSON (stdlib) or YAML (pyyaml if available); fall back to interactive prompts1152. **Normalise name**: Enforce lowercase-hyphen naming (`vcf-annotator`, not `VCF_Annotator`)1163. **Fill defaults**: domain → "bioinformatics", version → "0.1.0", capabilities/triggers → generic placeholders1174. **Render SKILL.md**: Fill YAML frontmatter + all 13 required body sections from template1185. **Render Python skeleton**: argparse wired with `--input`/`--output`/`--demo`; output boilerplate creates `report.md`, `result.json`, reproducibility bundle1196. **Render test skeleton**: pytest fixture + 3 standard tests (demo runs, report generated, result.json valid)1207. **Validate**: Run the 13-item CONTRIBUTING checklist against the generated SKILL.md before writing1218. **Register**: Append catalog entry; patch `clawbio.py` SKILLS dict via targeted string replacement122123## Example Queries124125- "Create a new skill called vcf-annotator that annotates VCF files with ClinVar"126- "Scaffold a skill for running PLINK GWAS pipelines"127- "Build a skill template for GO enrichment analysis"128- "Validate my SKILL.md before I submit a PR"129130## Output Structure131132```133output_directory/134├── report.md # Summary of what was generated135├── result.json # Machine-readable scaffold manifest136└── reproducibility/137 └── commands.sh # Exact command to reproduce the scaffold138139Generated skill at skills/<name>/:140├── SKILL.md # Complete skill definition141├── <name>.py # Python skeleton with --input/--output/--demo142├── tests/143│ └── test_<name>.py # pytest skeleton with 3 standard tests144└── examples/145 └── example_spec.json # The spec that generated this skill146```147148## Dependencies149150**Required** (stdlib only — zero install):151- Python 3.11+ standard library (`argparse`, `pathlib`, `json`, `re`, `textwrap`, `shutil`, `getpass`, `socket`)152153**Optional**:154- `pyyaml` >= 6.0 — enables YAML spec files in addition to JSON; graceful fallback to JSON-only mode if absent155156## Safety157158- **Local-first**: No network calls; all generation is offline159- **Non-destructive**: Never overwrites existing files without `--force`; prompts or errors if destination exists160- **No hallucinated science**: All generated SKILL.md content is taken directly from the spec; placeholder text is clearly marked with `TODO:`161- **Audit trail**: `result.json` and `commands.sh` record exactly what was generated and when162163## Integration with Bio Orchestrator164165**Trigger conditions** — the orchestrator routes here when:166- User says "create a skill", "scaffold a skill", "new skill", "build a skill", "add a skill"167- User provides a JSON/YAML file with `name`, `description`, `author` fields and asks to build a skill168169**Chaining partners**:170- `bio-orchestrator`: Skill builder output feeds back into the orchestrator once registered171172## Citations173174- [CONTRIBUTING.md](https://github.com/ClawBio/ClawBio/blob/main/CONTRIBUTING.md) — skill submission guidelines and checklist175- [templates/SKILL-TEMPLATE.md](https://github.com/ClawBio/ClawBio/blob/main/templates/SKILL-TEMPLATE.md) — canonical SKILL.md template