nfcore-scrnaseq-wrapper
You are nfcore-scrnaseq-wrapper, a specialised ClawBio agent for upstream single-cell RNA-seq preprocessing from FASTQ using the nf-core/scrnaseq Nextflow pipeline.
Trigger
Fire when:
- User wants to run
scrnaseq from raw FASTQ files
- User asks to preprocess 10x Chromium single-cell data
- User wants to execute
nf-core/scrnaseq
- User wants to generate
.h5ad from raw single-cell FASTQs
- User asks for primary scRNA preprocessing (FASTQ → h5ad)
- User mentions
simpleaf, STARsolo, alevin-fry, or kb-python for upstream processing
Do NOT fire when:
- User already has an
.h5ad and wants clustering, UMAP, or markers → route to scrna-orchestrator
- User asks for scVI, scANVI, batch correction, or dimensionality reduction → route to
scrna-embedding
- User asks about bulk RNA-seq, differential expression, or pseudo-bulk analysis → route to
rnaseq-de
- Input is an already-processed count matrix, not raw FASTQs
Scope
One skill, one task: run upstream scRNA preprocessing from FASTQ using nf-core/scrnaseq and produce canonical outputs for downstream ClawBio skills.
This skill does NOT perform clustering, normalization, marker detection, dimensionality reduction, or any analysis on the .h5ad it produces.
Why This Exists
- Without it: Users hand-build samplesheets, guess reference combinations, miss backend issues, and struggle to locate the correct
.h5ad for downstream analysis.
- With it: One validated command runs the pipeline, captures provenance, writes a reproducibility bundle, and points directly to the best downstream handoff artifact.
- Why ClawBio: The wrapper keeps execution local-first, validates before launching Nextflow, and makes the run chainable into
scrna and scrna-embedding.
Core Capabilities
- Strict Preflight: Validate Java, Nextflow, backend, samplesheet, FASTQs, and references before execution.
- Curated Presets: Expose all six pipeline modes (
standard, star, kallisto, cellranger, cellrangerarc, cellrangermulti).
- Controlled Execution: Always run with
-params-file, a fixed pipeline source, and explicit reproducibility artifacts.
- Output Resolution: Detect MultiQC, pipeline_info,
.h5ad, .rds, and select a canonical preferred_h5ad when possible.
- Downstream Handoff: Recommend the next command for
scrna-orchestrator (automatic via --run-downstream); scrna-embedding can follow as a second step.
Input Formats
| Format |
Extension |
Required columns |
Optional columns |
| Samplesheet |
.csv |
sample, fastq_1, fastq_2 |
expected_cells, seq_center, sample_type, fastq_barcode, feature_type |
| Demo mode |
n/a |
none — test profile provides its own data |
— |
Workflow
- Validate: Check the selected preset, samplesheet structure, FASTQ accessibility, references, Java, Nextflow, and backend.
- Normalize: Write a validated samplesheet copy with absolute POSIX paths into the reproducibility bundle.
- Configure: Build one effective
params.yaml and a fixed Nextflow command.
- Execute: Run
nf-core/scrnaseq using the local sibling checkout when available, or the pinned remote tag.
- Parse: Detect MultiQC, pipeline_info,
.h5ad, .rds, and CellBender-derived outputs.
- Generate: Write
report.md, result.json, provenance JSON files, and reproducibility artifacts.
- Hand off: Recommend the next ClawBio command using the
preferred_h5ad when handoff_available = true.
Algorithm / Methodology
The wrapper executes a strictly ordered 7-step pipeline. A failure at any step raises a structured SkillError with an error_code and a fix hint; no subsequent step runs.
Pipeline source resolution (pipeline_source.py): Prefer a local sibling scrnaseq/ checkout (pinned commit, audit-safe). Fall back to the remote pipeline tag when no checkout is found or the checkout path contains whitespace (macOS Docker restriction).
Samplesheet validation (samplesheet_builder.py): Parse the CSV, resolve FASTQ paths relative to the CSV parent directory, normalize sample-name whitespace to underscores, verify readability and FASTQ extensions, reject FASTQ basenames with whitespace, enforce consistent expected_cells (≥1) and seq_center for repeated sample rows, reject exact duplicate FASTQ rows, and write a normalized copy with absolute POSIX paths to reproducibility/samplesheet.valid.csv.
Preflight (preflight.py): Verify Java (≥17) and Nextflow (≥25.04.0). Compare version tuples after zero-padding to 3 elements (avoids false negatives such as (24, 4) < (24, 4, 0)). For docker, run docker info and gate on exit code. For conda/mamba, locate the binary. For singularity/apptainer, accept either binary interchangeably. For wave and gpu, no binary check is needed (Nextflow-native features). All subprocess calls have a 30-second timeout.
Params construction (params_builder.py): Translate the preset + CLI flags into a params.yaml consumed by Nextflow via -params-file. All file paths use .as_posix() for forward-slash consistency across platforms. igenomes_ignore is automatically set to true whenever any explicit reference path is provided (suppresses nf-schema DNS validation of the default iGenomes S3 URL). Skip flags are only written when true, keeping params.yaml minimal.
Command build + execution (command_builder.py, executor.py): Construct the nextflow run command with -params-file and -work-dir <output>/upstream/work, then launch via subprocess.Popen with stdout and stderr piped to log files on disk — never buffered in RAM. On TimeoutExpired, the process is killed and EXECUTION_FAILED is raised.
Output parsing (outputs_parser.py): Scan the upstream results tree for MultiQC HTML, pipeline_info/, .h5ad (combined matrix preferred over per-sample, filtered preferred over raw), .rds, and CellBender-derived files. handoff_available is set to true only when a preferred_h5ad is confirmed on disk.
Provenance + reporting (provenance.py, reporting.py): Write JSON provenance bundles, a SHA-256 checksum manifest (files only — never directories), environment.yml, a portable commands.sh, report.md, and result.json.
Presets
| Preset |
Aligner |
Use case |
standard |
simpleaf (alevin-fry) |
Default for 10x GEX; fast, memory-efficient |
star |
STARsolo |
Best FASTQ QC metrics; supports RNA velocity (--star-feature "Gene Velocyto") |
kallisto |
kb-python / BUStools |
Pseudo-alignment; fastest; lamanno/nac RNA velocity via --kb-workflow |
cellranger |
CellRanger |
CellRanger v2/v3 compatibility; requires CellRanger binary in PATH |
cellrangerarc |
CellRanger ARC |
Multiome (GEX + ATAC); accepts prebuilt --cellranger-index or reference-build inputs |
cellrangermulti |
CellRanger Multi |
GEX + VDJ + feature barcoding; --cellranger-multi-barcodes required for CMO/FFPE multiplexing |
Each preset requires at least one reference option: --genome <iGenomes_shortcut> OR a pre-built index (--star-index, --simpleaf-index, etc.) OR --fasta + --gtf.
CLI Reference
# Standard usage
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./scrnaseq_run
# Preflight check only (no Nextflow execution)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./scrnaseq_run --check
# Demo mode (runs the upstream nf-core test profile; forces star preset; Docker must be running)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--demo --output ./scrnaseq_demo
# Via ClawBio runner
python clawbio.py run scrnaseq-pipeline --input samplesheet.csv --output ./scrnaseq_run
python clawbio.py run scrnaseq-pipeline --demo --output ./scrnaseq_demo
# STARsolo with local FASTA+GTF (STAR index built by the pipeline)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# STARsolo with prebuilt STAR index
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--star-index /refs/star_index
# STARsolo RNA velocity
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star \
--star-feature "Gene Velocyto" --star-ignore-sjdbgtf \
--fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# Simpleaf (standard) with UMI resolution override
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset standard --protocol 10XV3 \
--simpleaf-umi-resolution cr-like-em --genome GRCh38
# Kallisto RNA velocity (NAC workflow)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset kallisto \
--kb-workflow nac --fasta /refs/hg38.fa --gtf /refs/hg38.gtf
# Air-gapped cluster: local iGenomes mirror
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset star --protocol 10XV3 \
--genome GRCh38 --igenomes-base /mnt/local_igenomes
# CellRanger Multi (CMO multiplexing)
python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \
--input samplesheet.csv --output ./run --preset cellrangermulti \
--cellranger-index /refs/refdata-gex-GRCh38 \
--gex-cmo-set /refs/cmo_set.csv \
--cellranger-multi-barcodes /refs/multi_barcodes.csv
Key flags
| Flag |
Type |
Default |
Description |
--preset |
string |
standard |
Aligner preset |
--profile |
string |
docker |
Execution backend: docker, conda, mamba, singularity, apptainer, podman, shifter, charliecloud, wave, gpu |
--pipeline-version |
string |
4.1.0 |
Remote nf-core/scrnaseq tag or commit (used when no local sibling checkout is found) |
--protocol |
string |
None |
Chemistry/protocol forwarded to the aligner. Required and non-auto for standard, star, and kallisto. standard/simpleaf additionally rejects smartseq. For cellranger, cellrangerarc, and cellrangermulti any value is accepted — auto is the typical upstream recommendation |
--genome |
string |
— |
iGenomes shortcut (GRCh38, mm10, etc.) — mutually exclusive with --fasta/--gtf and all index flags |
--igenomes-base |
string |
— |
Base URL or local path for iGenomes (default s3://ngi-igenomes/igenomes/). Use for local mirrors or air-gapped clusters |
--fasta |
path |
— |
Genome FASTA (.fa, .fna, .fasta, .gz variants; no whitespace in path) |
--gtf |
path |
— |
Gene annotation GTF |
--star-index |
path |
— |
Prebuilt STAR genome index directory |
--simpleaf-index |
path |
— |
Prebuilt simpleaf/alevin-fry index |
--kallisto-index |
path |
— |
Prebuilt kallisto index |
--cellranger-index |
path |
— |
Prebuilt CellRanger or CellRanger ARC reference |
--transcript-fasta |
path |
— |
Transcriptome FASTA for simpleaf |
--txp2gene |
path |
— |
Transcript-to-gene mapping for simpleaf |
--barcode-whitelist |
path |
— |
Custom barcode whitelist (per-aligner format) |
--star-feature |
enum |
— |
STARsolo feature type: Gene, GeneFull, Gene Velocyto |
--star-ignore-sjdbgtf |
flag |
— |
Do not use GTF for SJDB construction (required for Gene Velocyto) |
--seq-center |
string |
— |
Sequencing center name for BAM read group tag |
--simpleaf-umi-resolution |
enum |
— |
UMI resolution strategy for alevin-fry: cr-like, cr-like-em, parsimony, parsimony-em, parsimony-gene, parsimony-gene-em |
--kb-workflow |
enum |
— |
Kallisto workflow: standard, lamanno, nac |
--kb-t1c |
path |
— |
cDNA transcripts-to-capture file for RNA velocity (lamanno/nac) |
--kb-t2c |
path |
— |
Intron transcripts-to-capture file for RNA velocity (lamanno/nac) |
--skip-cellbender |
flag |
— |
Disable the CellBender ambient RNA removal subworkflow |
--skip-emptydrops |
flag |
— |
Skip emptyDrops cell calling (independent of --skip-cellbender) |
--skip-fastqc |
flag |
— |
Skip FastQC quality control |
--skip-multiqc |
flag |
— |
Skip MultiQC report generation |
--skip-cellranger-renaming |
flag |
— |
Skip automatic sample renaming in CellRanger modules |
--skip-cellrangermulti-vdjref |
flag |
— |
Skip mkvdjref in cellrangermulti (when VDJ data is absent or a prebuilt --cellranger-vdj-index is supplied) |
--save-reference |
flag |
— |
Save the built reference index for future reuse |
--save-align-intermeds |
flag |
— |
Save alignment intermediate BAM files (disabled by default) |
--expected-cells |
int |
— |
Override expected cell count (≥1) for all samples |
--email |
string |
— |
Email address for pipeline completion notification |
--multiqc-title |
string |
— |
Custom title for the MultiQC report |
--resume |
flag |
— |
Nextflow resume (checksum-verified against prior manifest) |
--run-downstream |
flag |
— |
Opt in to scrna_orchestrator handoff after pipeline completion |
--cellrangerarc-config |
path |
— |
Config JSON for CellRanger ARC index construction |
--cellrangerarc-reference |
string |
— |
Reference genome name used inside the CellRanger ARC config |
--motifs |
path |
— |
Motif file (e.g. JASPAR) for CellRanger ARC |
--cellranger-vdj-index |
path |
— |
Prebuilt CellRanger VDJ reference |
--gex-frna-probe-set |
path |
— |
Probe set CSV for FFPE fixed RNA profiling (cellrangermulti) |
--gex-target-panel |
path |
— |
Target panel CSV for targeted GEX (cellrangermulti) |
--gex-cmo-set |
path |
— |
CMO reference CSV for multiplexed samples (cellrangermulti) |
--gex-barcode-sample-assignment |
path |
— |
Barcode-to-sample assignment CSV for OCM multiplexing (cellrangermulti) |
--fb-reference |
path |
— |
Feature-barcode reference CSV for antibody capture (cellrangermulti) |
--vdj-inner-enrichment-primers |
path |
— |
V(D)J cDNA enrichment primer sequences (cellrangermulti) |
--cellranger-multi-barcodes |
path |
— |
Multiplexed sample samplesheet for CMO/FFPE demultiplexing (cellrangermulti) |
Output Structure
output_directory/
├── report.md # Wrapper run summary
├── result.json # Structured result payload
├── logs/
│ ├── stdout.txt # Nextflow stdout
│ └── stderr.txt # Nextflow stderr
├── upstream/
│ └── results/ # nf-core/scrnaseq output tree
│ ├── fastqc/ # Per-read FastQC reports
│ ├── multiqc/ # MultiQC HTML and data
│ │ └── multiqc_report.html
│ ├── pipeline_info/ # Execution report, timeline, trace, DAG
│ └── <aligner>/ # Aligner-specific outputs
│ ├── <sample>/ # Per-sample STAR/simpleaf/kallisto outputs
│ └── mtx_conversions/ # AnnData (.h5ad), SCE (.rds), Seurat (.rds)
│ ├── <sample>_filtered_matrix.h5ad
│ ├── <sample>_raw_matrix.h5ad
│ ├── combined_filtered_matrix.h5ad ← preferred_h5ad (when present)
│ └── combined_raw_matrix.h5ad
├── reproducibility/
│ ├── samplesheet.valid.csv # Normalized samplesheet (absolute POSIX paths)
│ ├── params.yaml # Effective Nextflow parameters
│ ├── commands.sh # Portable replay script
│ ├── environment.yml # Conda environment spec (for reference)
│ ├── checksums.sha256 # SHA-256 for samplesheet, params, refs, h5ad, MultiQC
│ ├── manifest.json # Run metadata: preset, profile, versions, checksums
│ ├── macos_docker.config # macOS+Docker workarounds (VirtioFS, ARM64, STAR FIFOs)
│ └── remap_paths.py # Helper for replaying on a different machine
└── provenance/
├── inputs.json # Samplesheet and reference paths + checksums
├── invocation.json # Timestamp, preset, profile, pipeline version
├── preflight.json # Java/Nextflow/backend info
├── upstream.json # Pipeline source resolution details
├── outputs.json # Detected artifacts
├── runtime.json # Execution timing
└── skill.json # Skill name and version
Example Output
result.json (abbreviated):
{
"skill": "scrnaseq-pipeline",
"version": "0.1.0",
"summary": {
"preset": "star",
"aligner_effective": "star",
"pipeline_source_kind": "remote_repo",
"pipeline_version_or_commit": "4.1.0",
"profile": "docker",
"preferred_h5ad": "<output>/upstream/results/star/mtx_conversions/combined_filtered_matrix.h5ad",
"handoff_available": true,
"samples_detected": 2,
"cellbender_used": false
}
}
report.md closes with:
## Next Steps
- python clawbio.py run scrna --input <preferred_h5ad> --output <dir>
- python clawbio.py run scrna-embedding --input <preferred_h5ad> --output <dir>
Gotchas
- Preflight runs before any Nextflow call. If Java, Nextflow, or the backend are missing or too old, the pipeline never starts and you get a structured JSON error with
error_code and a fix hint. Nextflow ≥25.04.0 is required.
--genome conflicts with any explicit reference flag. Providing --genome alongside --fasta, --gtf, or any index flag raises CONFLICTING_REFERENCES in preflight. Use either --genome <shortcut> or explicit flags — never both.
igenomes_ignore is set automatically. Whenever any explicit reference path (fasta, gtf, any index) is provided, the wrapper writes igenomes_ignore: true to suppress nf-schema DNS validation of the default iGenomes S3 URL. You do not need to set this manually. Use --igenomes-base only for local iGenomes mirrors.
- Protocol compatibility is enforced before Nextflow starts.
standard, star, and kallisto presets require an explicit non-auto protocol (e.g., 10XV3, dropseq, or a supported custom chemistry string). standard/simpleaf additionally rejects smartseq — use star or kallisto for Smart-seq data. cellranger, cellrangerarc, and cellrangermulti accept any protocol value; auto is the typical upstream recommendation but is not enforced.
--skip-cellbender and --skip-emptydrops are independent flags. --skip-cellbender disables the CellBender ambient RNA removal subworkflow; --skip-emptydrops disables the emptyDrops cell calling step. Both are written as separate parameters in params.yaml and each replays as its own flag in commands.sh. Setting one does not imply the other.
--demo forces preset=star and skip_cellbender=true. The nf-core upstream test profile ships STAR-compatible data and explicitly disables CellBender (which does not work on small test datasets). If a different preset is requested with --demo, the wrapper warns and overrides it.
preferred_h5ad may be absent. If no combined matrix is produced and there are multiple per-sample files, handoff_available is false. Always check result.json before chaining to scrna-orchestrator or scrna-embedding.
- No arbitrary Nextflow passthrough. All pipeline configuration flows through the preset system and
params.yaml. Direct -c, --outdir, or custom Nextflow flags cannot be injected.
--resume enforces strict compatibility. The wrapper checks that the stored manifest matches the current preset, profile, and pipeline source. Mismatches raise INVALID_RESUME_STATE.
- RNA velocity requires two coordinated flags. For STARsolo:
--star-feature "Gene Velocyto" AND --star-ignore-sjdbgtf must be passed together. For Kallisto: --kb-workflow lamanno or nac with --kb-t1c and --kb-t2c.
cellrangerarc config and reference are paired. Providing --cellrangerarc-config without --cellrangerarc-reference (or vice versa) raises INVALID_PRESET_CONFIGURATION. --motifs is optional and independent.
cellrangermulti validates against the samplesheet. feature_type=ab requires --fb-reference; feature_type=cmo requires --cellranger-multi-barcodes; feature_type=vdj with --skip-cellrangermulti-vdjref requires --cellranger-vdj-index. CMO, FFPE probe-set, and OCM multiplexing modes are mutually exclusive.
- FASTA schema validation. The FASTA path must match
^\S+\.fn?a(sta)?(\.gz)?$ (the nf-core/scrnaseq 4.1.0 schema). Paths with whitespace or non-standard extensions are rejected in preflight.
- Local checkout must be a sibling directory. The wrapper looks for
../scrnaseq relative to the ClawBio repo root. If the checkout path contains whitespace (common on macOS), the wrapper warns and falls back to the remote pipeline. Ensure ClawBio/ and scrnaseq/ share the same parent folder.
- macOS Docker workaround is applied automatically. On macOS with Docker,
macos_docker.config is written to the reproducibility bundle and passed to Nextflow. It sets stageInMode = "copy" (avoids VirtioFS EDEADLK), --platform linux/amd64 (Rosetta emulation), and routes STAR _STARtmp to the container's /tmp (avoids VirtioFS FIFO limitation). Output directories under /tmp emit a WARNING — use a path under HOME.
Safety
- Local-first: User FASTQs and outputs remain on the local filesystem.
- Strict preflight: Nextflow is never invoked if validation fails.
- No hallucinated outputs: Only artifacts confirmed on disk are reported.
- Disclaimer: Every report includes the ClawBio medical disclaimer.
Agent Boundary
The agent dispatches and explains; this skill executes.
Agent: Interpret the user's preprocessing intent, choose the preset, and verify that handoff_available is true in result.json before routing to downstream skills.
Skill: Validate environment and inputs, run the pipeline with controlled parameters, write all provenance and reproducibility artifacts, and report the detected preferred_h5ad.
Chaining Partners
| Skill |
When to chain |
scrna-orchestrator |
After a successful run, pass preferred_h5ad for clustering, QC, and markers |
scrna-embedding |
Pass preferred_h5ad for scVI/scANVI batch integration and latent embeddings |
multiqc-reporter |
Re-aggregate QC across multiple wrapper runs |
Maintenance
Review cadence: After each nf-core/scrnaseq major release. Check NEXTFLOW_MIN_VERSION (schemas.py), SUPPORTED_PRESETS, SUPPORTED_PROFILES, and this SKILL.md for accuracy.
Staleness signals:
- Preflight rejects a Nextflow version that the current pipeline supports → update
NEXTFLOW_MIN_VERSION in schemas.py and reproducibility/pinned_versions.json.
- New aligners appear upstream but are absent from
PRESET_ALIGNERS → add to schemas.py and update tests.
- The VirtioFS macOS workaround (
stageInMode = "copy") is only necessary while Apple Silicon runs Docker via QEMU. Remove _write_macos_docker_config when a native arm64 Docker runtime eliminates VirtioFS deadlocks.
Deprecation criteria: Deprecate if nf-core/scrnaseq releases a Python SDK with equivalent preflight, params, and provenance APIs.
Citations
1---2name: nfcore-scrnaseq-wrapper3description: Wrapper skill for running nf-core/scrnaseq upstream single-cell RNA-seq preprocessing from FASTQ with strict preflight, reproducibility outputs, and downstream handoff to ClawBio scRNA skills.4license: MIT5---67# nfcore-scrnaseq-wrapper89You are **nfcore-scrnaseq-wrapper**, a specialised ClawBio agent for upstream single-cell RNA-seq preprocessing from FASTQ using the `nf-core/scrnaseq` Nextflow pipeline.1011## Trigger1213**Fire when:**14- User wants to run `scrnaseq` from raw FASTQ files15- User asks to preprocess 10x Chromium single-cell data16- User wants to execute `nf-core/scrnaseq`17- User wants to generate `.h5ad` from raw single-cell FASTQs18- User asks for primary scRNA preprocessing (FASTQ → h5ad)19- User mentions `simpleaf`, `STARsolo`, `alevin-fry`, or `kb-python` for upstream processing2021**Do NOT fire when:**22- User already has an `.h5ad` and wants clustering, UMAP, or markers → route to `scrna-orchestrator`23- User asks for scVI, scANVI, batch correction, or dimensionality reduction → route to `scrna-embedding`24- User asks about bulk RNA-seq, differential expression, or pseudo-bulk analysis → route to `rnaseq-de`25- Input is an already-processed count matrix, not raw FASTQs2627## Scope2829One skill, one task: run upstream scRNA preprocessing from FASTQ using `nf-core/scrnaseq` and produce canonical outputs for downstream ClawBio skills.3031This skill does NOT perform clustering, normalization, marker detection, dimensionality reduction, or any analysis on the `.h5ad` it produces.3233## Why This Exists3435- **Without it**: Users hand-build samplesheets, guess reference combinations, miss backend issues, and struggle to locate the correct `.h5ad` for downstream analysis.36- **With it**: One validated command runs the pipeline, captures provenance, writes a reproducibility bundle, and points directly to the best downstream handoff artifact.37- **Why ClawBio**: The wrapper keeps execution local-first, validates before launching Nextflow, and makes the run chainable into `scrna` and `scrna-embedding`.3839## Core Capabilities40411. **Strict Preflight**: Validate Java, Nextflow, backend, samplesheet, FASTQs, and references before execution.422. **Curated Presets**: Expose all six pipeline modes (`standard`, `star`, `kallisto`, `cellranger`, `cellrangerarc`, `cellrangermulti`).433. **Controlled Execution**: Always run with `-params-file`, a fixed pipeline source, and explicit reproducibility artifacts.444. **Output Resolution**: Detect MultiQC, pipeline_info, `.h5ad`, `.rds`, and select a canonical `preferred_h5ad` when possible.455. **Downstream Handoff**: Recommend the next command for `scrna-orchestrator` (automatic via `--run-downstream`); `scrna-embedding` can follow as a second step.4647## Input Formats4849| Format | Extension | Required columns | Optional columns |50|--------|-----------|------------------|------------------|51| Samplesheet | `.csv` | `sample`, `fastq_1`, `fastq_2` | `expected_cells`, `seq_center`, `sample_type`, `fastq_barcode`, `feature_type` |52| Demo mode | n/a | none — test profile provides its own data | — |5354## Workflow55561. **Validate**: Check the selected preset, samplesheet structure, FASTQ accessibility, references, Java, Nextflow, and backend.572. **Normalize**: Write a validated samplesheet copy with absolute POSIX paths into the reproducibility bundle.583. **Configure**: Build one effective `params.yaml` and a fixed Nextflow command.594. **Execute**: Run `nf-core/scrnaseq` using the local sibling checkout when available, or the pinned remote tag.605. **Parse**: Detect MultiQC, pipeline_info, `.h5ad`, `.rds`, and CellBender-derived outputs.616. **Generate**: Write `report.md`, `result.json`, provenance JSON files, and reproducibility artifacts.627. **Hand off**: Recommend the next ClawBio command using the `preferred_h5ad` when `handoff_available = true`.6364## Algorithm / Methodology6566The wrapper executes a strictly ordered 7-step pipeline. A failure at any step raises a structured `SkillError` with an `error_code` and a `fix` hint; no subsequent step runs.67681. **Pipeline source resolution** (`pipeline_source.py`): Prefer a local sibling `scrnaseq/` checkout (pinned commit, audit-safe). Fall back to the remote pipeline tag when no checkout is found or the checkout path contains whitespace (macOS Docker restriction).69702. **Samplesheet validation** (`samplesheet_builder.py`): Parse the CSV, resolve FASTQ paths relative to the CSV parent directory, normalize sample-name whitespace to underscores, verify readability and FASTQ extensions, reject FASTQ basenames with whitespace, enforce consistent `expected_cells` (≥1) and `seq_center` for repeated sample rows, reject exact duplicate FASTQ rows, and write a normalized copy with absolute POSIX paths to `reproducibility/samplesheet.valid.csv`.71723. **Preflight** (`preflight.py`): Verify Java (≥17) and Nextflow (≥25.04.0). Compare version tuples after zero-padding to 3 elements (avoids false negatives such as `(24, 4) < (24, 4, 0)`). For `docker`, run `docker info` and gate on exit code. For `conda`/`mamba`, locate the binary. For `singularity`/`apptainer`, accept either binary interchangeably. For `wave` and `gpu`, no binary check is needed (Nextflow-native features). All subprocess calls have a 30-second timeout.73744. **Params construction** (`params_builder.py`): Translate the preset + CLI flags into a `params.yaml` consumed by Nextflow via `-params-file`. All file paths use `.as_posix()` for forward-slash consistency across platforms. `igenomes_ignore` is automatically set to `true` whenever any explicit reference path is provided (suppresses nf-schema DNS validation of the default iGenomes S3 URL). Skip flags are only written when `true`, keeping `params.yaml` minimal.75765. **Command build + execution** (`command_builder.py`, `executor.py`): Construct the `nextflow run` command with `-params-file` and `-work-dir <output>/upstream/work`, then launch via `subprocess.Popen` with stdout and stderr piped to log files on disk — never buffered in RAM. On `TimeoutExpired`, the process is killed and `EXECUTION_FAILED` is raised.77786. **Output parsing** (`outputs_parser.py`): Scan the upstream results tree for MultiQC HTML, `pipeline_info/`, `.h5ad` (combined matrix preferred over per-sample, filtered preferred over raw), `.rds`, and CellBender-derived files. `handoff_available` is set to `true` only when a `preferred_h5ad` is confirmed on disk.79807. **Provenance + reporting** (`provenance.py`, `reporting.py`): Write JSON provenance bundles, a SHA-256 checksum manifest (files only — never directories), `environment.yml`, a portable `commands.sh`, `report.md`, and `result.json`.8182## Presets8384| Preset | Aligner | Use case |85|--------|---------|---------|86| `standard` | simpleaf (alevin-fry) | Default for 10x GEX; fast, memory-efficient |87| `star` | STARsolo | Best FASTQ QC metrics; supports RNA velocity (`--star-feature "Gene Velocyto"`) |88| `kallisto` | kb-python / BUStools | Pseudo-alignment; fastest; lamanno/nac RNA velocity via `--kb-workflow` |89| `cellranger` | CellRanger | CellRanger v2/v3 compatibility; requires CellRanger binary in PATH |90| `cellrangerarc` | CellRanger ARC | Multiome (GEX + ATAC); accepts prebuilt `--cellranger-index` or reference-build inputs |91| `cellrangermulti` | CellRanger Multi | GEX + VDJ + feature barcoding; `--cellranger-multi-barcodes` required for CMO/FFPE multiplexing |9293Each preset requires at least one reference option: `--genome <iGenomes_shortcut>` OR a pre-built index (`--star-index`, `--simpleaf-index`, etc.) OR `--fasta` + `--gtf`.9495## CLI Reference9697```bash98# Standard usage99python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \100 --input samplesheet.csv --output ./scrnaseq_run101102# Preflight check only (no Nextflow execution)103python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \104 --input samplesheet.csv --output ./scrnaseq_run --check105106# Demo mode (runs the upstream nf-core test profile; forces star preset; Docker must be running)107python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \108 --demo --output ./scrnaseq_demo109110# Via ClawBio runner111python clawbio.py run scrnaseq-pipeline --input samplesheet.csv --output ./scrnaseq_run112python clawbio.py run scrnaseq-pipeline --demo --output ./scrnaseq_demo113114# STARsolo with local FASTA+GTF (STAR index built by the pipeline)115python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \116 --input samplesheet.csv --output ./run --preset star --protocol 10XV3 \117 --fasta /refs/hg38.fa --gtf /refs/hg38.gtf118119# STARsolo with prebuilt STAR index120python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \121 --input samplesheet.csv --output ./run --preset star --protocol 10XV3 \122 --star-index /refs/star_index123124# STARsolo RNA velocity125python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \126 --input samplesheet.csv --output ./run --preset star \127 --star-feature "Gene Velocyto" --star-ignore-sjdbgtf \128 --fasta /refs/hg38.fa --gtf /refs/hg38.gtf129130# Simpleaf (standard) with UMI resolution override131python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \132 --input samplesheet.csv --output ./run --preset standard --protocol 10XV3 \133 --simpleaf-umi-resolution cr-like-em --genome GRCh38134135# Kallisto RNA velocity (NAC workflow)136python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \137 --input samplesheet.csv --output ./run --preset kallisto \138 --kb-workflow nac --fasta /refs/hg38.fa --gtf /refs/hg38.gtf139140# Air-gapped cluster: local iGenomes mirror141python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \142 --input samplesheet.csv --output ./run --preset star --protocol 10XV3 \143 --genome GRCh38 --igenomes-base /mnt/local_igenomes144145# CellRanger Multi (CMO multiplexing)146python skills/nfcore-scrnaseq-wrapper/nfcore_scrnaseq_wrapper.py \147 --input samplesheet.csv --output ./run --preset cellrangermulti \148 --cellranger-index /refs/refdata-gex-GRCh38 \149 --gex-cmo-set /refs/cmo_set.csv \150 --cellranger-multi-barcodes /refs/multi_barcodes.csv151```152153### Key flags154155| Flag | Type | Default | Description |156|------|------|---------|-------------|157| `--preset` | string | `standard` | Aligner preset |158| `--profile` | string | `docker` | Execution backend: `docker`, `conda`, `mamba`, `singularity`, `apptainer`, `podman`, `shifter`, `charliecloud`, `wave`, `gpu` |159| `--pipeline-version` | string | `4.1.0` | Remote `nf-core/scrnaseq` tag or commit (used when no local sibling checkout is found) |160| `--protocol` | string | `None` | Chemistry/protocol forwarded to the aligner. Required and non-`auto` for `standard`, `star`, and `kallisto`. `standard`/simpleaf additionally rejects `smartseq`. For `cellranger`, `cellrangerarc`, and `cellrangermulti` any value is accepted — `auto` is the typical upstream recommendation |161| `--genome` | string | — | iGenomes shortcut (`GRCh38`, `mm10`, etc.) — mutually exclusive with `--fasta`/`--gtf` and all index flags |162| `--igenomes-base` | string | — | Base URL or local path for iGenomes (default `s3://ngi-igenomes/igenomes/`). Use for local mirrors or air-gapped clusters |163| `--fasta` | path | — | Genome FASTA (`.fa`, `.fna`, `.fasta`, `.gz` variants; no whitespace in path) |164| `--gtf` | path | — | Gene annotation GTF |165| `--star-index` | path | — | Prebuilt STAR genome index directory |166| `--simpleaf-index` | path | — | Prebuilt simpleaf/alevin-fry index |167| `--kallisto-index` | path | — | Prebuilt kallisto index |168| `--cellranger-index` | path | — | Prebuilt CellRanger or CellRanger ARC reference |169| `--transcript-fasta` | path | — | Transcriptome FASTA for simpleaf |170| `--txp2gene` | path | — | Transcript-to-gene mapping for simpleaf |171| `--barcode-whitelist` | path | — | Custom barcode whitelist (per-aligner format) |172| `--star-feature` | enum | — | STARsolo feature type: `Gene`, `GeneFull`, `Gene Velocyto` |173| `--star-ignore-sjdbgtf` | flag | — | Do not use GTF for SJDB construction (required for `Gene Velocyto`) |174| `--seq-center` | string | — | Sequencing center name for BAM read group tag |175| `--simpleaf-umi-resolution` | enum | — | UMI resolution strategy for alevin-fry: `cr-like`, `cr-like-em`, `parsimony`, `parsimony-em`, `parsimony-gene`, `parsimony-gene-em` |176| `--kb-workflow` | enum | — | Kallisto workflow: `standard`, `lamanno`, `nac` |177| `--kb-t1c` | path | — | cDNA transcripts-to-capture file for RNA velocity (lamanno/nac) |178| `--kb-t2c` | path | — | Intron transcripts-to-capture file for RNA velocity (lamanno/nac) |179| `--skip-cellbender` | flag | — | Disable the CellBender ambient RNA removal subworkflow |180| `--skip-emptydrops` | flag | — | Skip emptyDrops cell calling (independent of `--skip-cellbender`) |181| `--skip-fastqc` | flag | — | Skip FastQC quality control |182| `--skip-multiqc` | flag | — | Skip MultiQC report generation |183| `--skip-cellranger-renaming` | flag | — | Skip automatic sample renaming in CellRanger modules |184| `--skip-cellrangermulti-vdjref` | flag | — | Skip mkvdjref in cellrangermulti (when VDJ data is absent or a prebuilt `--cellranger-vdj-index` is supplied) |185| `--save-reference` | flag | — | Save the built reference index for future reuse |186| `--save-align-intermeds` | flag | — | Save alignment intermediate BAM files (disabled by default) |187| `--expected-cells` | int | — | Override expected cell count (≥1) for all samples |188| `--email` | string | — | Email address for pipeline completion notification |189| `--multiqc-title` | string | — | Custom title for the MultiQC report |190| `--resume` | flag | — | Nextflow resume (checksum-verified against prior manifest) |191| `--run-downstream` | flag | — | Opt in to `scrna_orchestrator` handoff after pipeline completion |192| `--cellrangerarc-config` | path | — | Config JSON for CellRanger ARC index construction |193| `--cellrangerarc-reference` | string | — | Reference genome name used inside the CellRanger ARC config |194| `--motifs` | path | — | Motif file (e.g. JASPAR) for CellRanger ARC |195| `--cellranger-vdj-index` | path | — | Prebuilt CellRanger VDJ reference |196| `--gex-frna-probe-set` | path | — | Probe set CSV for FFPE fixed RNA profiling (`cellrangermulti`) |197| `--gex-target-panel` | path | — | Target panel CSV for targeted GEX (`cellrangermulti`) |198| `--gex-cmo-set` | path | — | CMO reference CSV for multiplexed samples (`cellrangermulti`) |199| `--gex-barcode-sample-assignment` | path | — | Barcode-to-sample assignment CSV for OCM multiplexing (`cellrangermulti`) |200| `--fb-reference` | path | — | Feature-barcode reference CSV for antibody capture (`cellrangermulti`) |201| `--vdj-inner-enrichment-primers` | path | — | V(D)J cDNA enrichment primer sequences (`cellrangermulti`) |202| `--cellranger-multi-barcodes` | path | — | Multiplexed sample samplesheet for CMO/FFPE demultiplexing (`cellrangermulti`) |203204## Output Structure205206```text207output_directory/208├── report.md # Wrapper run summary209├── result.json # Structured result payload210├── logs/211│ ├── stdout.txt # Nextflow stdout212│ └── stderr.txt # Nextflow stderr213├── upstream/214│ └── results/ # nf-core/scrnaseq output tree215│ ├── fastqc/ # Per-read FastQC reports216│ ├── multiqc/ # MultiQC HTML and data217│ │ └── multiqc_report.html218│ ├── pipeline_info/ # Execution report, timeline, trace, DAG219│ └── <aligner>/ # Aligner-specific outputs220│ ├── <sample>/ # Per-sample STAR/simpleaf/kallisto outputs221│ └── mtx_conversions/ # AnnData (.h5ad), SCE (.rds), Seurat (.rds)222│ ├── <sample>_filtered_matrix.h5ad223│ ├── <sample>_raw_matrix.h5ad224│ ├── combined_filtered_matrix.h5ad ← preferred_h5ad (when present)225│ └── combined_raw_matrix.h5ad226├── reproducibility/227│ ├── samplesheet.valid.csv # Normalized samplesheet (absolute POSIX paths)228│ ├── params.yaml # Effective Nextflow parameters229│ ├── commands.sh # Portable replay script230│ ├── environment.yml # Conda environment spec (for reference)231│ ├── checksums.sha256 # SHA-256 for samplesheet, params, refs, h5ad, MultiQC232│ ├── manifest.json # Run metadata: preset, profile, versions, checksums233│ ├── macos_docker.config # macOS+Docker workarounds (VirtioFS, ARM64, STAR FIFOs)234│ └── remap_paths.py # Helper for replaying on a different machine235└── provenance/236 ├── inputs.json # Samplesheet and reference paths + checksums237 ├── invocation.json # Timestamp, preset, profile, pipeline version238 ├── preflight.json # Java/Nextflow/backend info239 ├── upstream.json # Pipeline source resolution details240 ├── outputs.json # Detected artifacts241 ├── runtime.json # Execution timing242 └── skill.json # Skill name and version243```244245## Example Output246247`result.json` (abbreviated):248```json249{250 "skill": "scrnaseq-pipeline",251 "version": "0.1.0",252 "summary": {253 "preset": "star",254 "aligner_effective": "star",255 "pipeline_source_kind": "remote_repo",256 "pipeline_version_or_commit": "4.1.0",257 "profile": "docker",258 "preferred_h5ad": "<output>/upstream/results/star/mtx_conversions/combined_filtered_matrix.h5ad",259 "handoff_available": true,260 "samples_detected": 2,261 "cellbender_used": false262 }263}264```265266`report.md` closes with:267```268## Next Steps269- python clawbio.py run scrna --input <preferred_h5ad> --output <dir>270- python clawbio.py run scrna-embedding --input <preferred_h5ad> --output <dir>271```272273## Gotchas274275- **Preflight runs before any Nextflow call.** If Java, Nextflow, or the backend are missing or too old, the pipeline never starts and you get a structured JSON error with `error_code` and a `fix` hint. Nextflow ≥25.04.0 is required.276- **`--genome` conflicts with any explicit reference flag.** Providing `--genome` alongside `--fasta`, `--gtf`, or any index flag raises `CONFLICTING_REFERENCES` in preflight. Use either `--genome <shortcut>` or explicit flags — never both.277- **`igenomes_ignore` is set automatically.** Whenever any explicit reference path (fasta, gtf, any index) is provided, the wrapper writes `igenomes_ignore: true` to suppress nf-schema DNS validation of the default iGenomes S3 URL. You do not need to set this manually. Use `--igenomes-base` only for local iGenomes mirrors.278- **Protocol compatibility is enforced before Nextflow starts.** `standard`, `star`, and `kallisto` presets require an explicit non-`auto` protocol (e.g., `10XV3`, `dropseq`, or a supported custom chemistry string). `standard`/simpleaf additionally rejects `smartseq` — use `star` or `kallisto` for Smart-seq data. `cellranger`, `cellrangerarc`, and `cellrangermulti` accept any protocol value; `auto` is the typical upstream recommendation but is not enforced.279- **`--skip-cellbender` and `--skip-emptydrops` are independent flags.** `--skip-cellbender` disables the CellBender ambient RNA removal subworkflow; `--skip-emptydrops` disables the emptyDrops cell calling step. Both are written as separate parameters in `params.yaml` and each replays as its own flag in `commands.sh`. Setting one does not imply the other.280- **`--demo` forces preset=star and skip_cellbender=true.** The nf-core upstream `test` profile ships STAR-compatible data and explicitly disables CellBender (which does not work on small test datasets). If a different preset is requested with `--demo`, the wrapper warns and overrides it.281- **`preferred_h5ad` may be absent.** If no combined matrix is produced and there are multiple per-sample files, `handoff_available` is `false`. Always check `result.json` before chaining to `scrna-orchestrator` or `scrna-embedding`.282- **No arbitrary Nextflow passthrough.** All pipeline configuration flows through the preset system and `params.yaml`. Direct `-c`, `--outdir`, or custom Nextflow flags cannot be injected.283- **`--resume` enforces strict compatibility.** The wrapper checks that the stored manifest matches the current preset, profile, and pipeline source. Mismatches raise `INVALID_RESUME_STATE`.284- **RNA velocity requires two coordinated flags.** For STARsolo: `--star-feature "Gene Velocyto"` AND `--star-ignore-sjdbgtf` must be passed together. For Kallisto: `--kb-workflow lamanno` or `nac` with `--kb-t1c` and `--kb-t2c`.285- **`cellrangerarc` config and reference are paired.** Providing `--cellrangerarc-config` without `--cellrangerarc-reference` (or vice versa) raises `INVALID_PRESET_CONFIGURATION`. `--motifs` is optional and independent.286- **`cellrangermulti` validates against the samplesheet.** `feature_type=ab` requires `--fb-reference`; `feature_type=cmo` requires `--cellranger-multi-barcodes`; `feature_type=vdj` with `--skip-cellrangermulti-vdjref` requires `--cellranger-vdj-index`. CMO, FFPE probe-set, and OCM multiplexing modes are mutually exclusive.287- **FASTA schema validation.** The FASTA path must match `^\S+\.fn?a(sta)?(\.gz)?$` (the nf-core/scrnaseq 4.1.0 schema). Paths with whitespace or non-standard extensions are rejected in preflight.288- **Local checkout must be a sibling directory.** The wrapper looks for `../scrnaseq` relative to the ClawBio repo root. If the checkout path contains whitespace (common on macOS), the wrapper warns and falls back to the remote pipeline. Ensure `ClawBio/` and `scrnaseq/` share the same parent folder.289- **macOS Docker workaround is applied automatically.** On macOS with Docker, `macos_docker.config` is written to the reproducibility bundle and passed to Nextflow. It sets `stageInMode = "copy"` (avoids VirtioFS EDEADLK), `--platform linux/amd64` (Rosetta emulation), and routes STAR `_STARtmp` to the container's `/tmp` (avoids VirtioFS FIFO limitation). Output directories under `/tmp` emit a WARNING — use a path under HOME.290291## Safety292293- **Local-first**: User FASTQs and outputs remain on the local filesystem.294- **Strict preflight**: Nextflow is never invoked if validation fails.295- **No hallucinated outputs**: Only artifacts confirmed on disk are reported.296- **Disclaimer**: Every report includes the ClawBio medical disclaimer.297298## Agent Boundary299300The agent dispatches and explains; this skill executes.301302**Agent**: Interpret the user's preprocessing intent, choose the preset, and verify that `handoff_available` is `true` in `result.json` before routing to downstream skills.303304**Skill**: Validate environment and inputs, run the pipeline with controlled parameters, write all provenance and reproducibility artifacts, and report the detected `preferred_h5ad`.305306## Chaining Partners307308| Skill | When to chain |309|-------|--------------|310| `scrna-orchestrator` | After a successful run, pass `preferred_h5ad` for clustering, QC, and markers |311| `scrna-embedding` | Pass `preferred_h5ad` for scVI/scANVI batch integration and latent embeddings |312| `multiqc-reporter` | Re-aggregate QC across multiple wrapper runs |313314## Maintenance315316**Review cadence**: After each nf-core/scrnaseq major release. Check `NEXTFLOW_MIN_VERSION` (`schemas.py`), `SUPPORTED_PRESETS`, `SUPPORTED_PROFILES`, and this SKILL.md for accuracy.317318**Staleness signals**:319- Preflight rejects a Nextflow version that the current pipeline supports → update `NEXTFLOW_MIN_VERSION` in `schemas.py` and `reproducibility/pinned_versions.json`.320- New aligners appear upstream but are absent from `PRESET_ALIGNERS` → add to `schemas.py` and update tests.321- The VirtioFS macOS workaround (`stageInMode = "copy"`) is only necessary while Apple Silicon runs Docker via QEMU. Remove `_write_macos_docker_config` when a native arm64 Docker runtime eliminates VirtioFS deadlocks.322323**Deprecation criteria**: Deprecate if nf-core/scrnaseq releases a Python SDK with equivalent preflight, params, and provenance APIs.324325## Citations326327- [nf-core/scrnaseq 4.1.0](https://nf-co.re/scrnaseq/4.1.0)328- [nf-core/scrnaseq usage](https://nf-co.re/scrnaseq/4.1.0/docs/usage/)329- [nf-core/scrnaseq parameters](https://nf-co.re/scrnaseq/4.1.0/parameters/)330- [nf-core/scrnaseq output](https://nf-co.re/scrnaseq/4.1.0/docs/output/)331- [Nextflow](https://www.nextflow.io/)332- [Alevin-fry / Simpleaf](https://simpleaf.readthedocs.io/)333- [STARsolo](https://github.com/alexdobin/STAR/blob/master/docs/STARsolo.md)334- [kb-python / BUStools](https://www.kallistobus.tools/)