Motif Scanning
Overview
This skill enables comprehensive motif scanning using HOMER tools for genomic peak files. It scans genomic regions for specific transcription factor binding motifs using position-specific scoring matrices and identifies exact motif locations. To perform motif scanning:
- Always refer to the Inputs & Outputs section to check inputs and build the output architecture.
- Genome assembly: Always returned from user feedback (hg38, mm10, hg19, mm9, etc), never determined by yourself.
- Check chromosome names: Standardize chromosome names to format with "chr" (1 -> chr1, MT -> chrM).
- Prepare motif files: Position-specific scoring matrices (PSSM) in HOMER format, saved in ${HOMER_data}/knownTFs/motifs/${tf}.motif, and "tf" should be in lower case.
- Set scanning parameters: Region size, score thresholds, output format
- Run HOMER motif scanning command
When to use this skill
- Scan for potential binding sites for a certain TF in the whole genome or in specific genomic regions, like promoters of a gene list or peaks from ChIP-seq or ATAC-seq.
- Scanning ChIP-seq or ATAC-seq peaks for known motifs to validate TF binding specificity.
- Testing whether co-factor motifs (e.g., TAL1, KLF1, SPI1) co-occur within TF-bound or accessible regions to infer cooperative binding.
- Evaluating motif distribution patterns relative to genomic landmarks such as transcription start sites (TSS) or enhancers.
- Generating motif-annotated BED files for visualization in genome browsers or subsequent feature analysis.
Inputs & Outputs
Inputs
(1) Peak formats supported
- BED files: Standard genomic interval format
- narrowPeak: ENCODE narrow peak format
- broadPeak: ENCODE broad peak format
- HOMER peak files: Output from HOMER peak calling
(2) Motif formats supported
- HOMER motif format: Position-specific scoring matrices
- MEME motif format: MEME suite motif format
- TRANSFAC format: TRANSFAC database format
Outputs
${sample}_known_motif_scan/
results/
combined_motifs.txt # combined motif hits from all TFs
### Option 1: Scan motif in the specific genomic regions
${sample}_motif_find.txt
${sample}_motif_find.bed
### Option 2: Scan motif in the genome
${sample}.genomewide.txt
${sample}.genomewide.bed
### Option 3: Annotate peaks with motif hits
${sample}.anno_motif.txt
${sample}.motif_pos.bed (if `mbed` is True)
logs/ # analysis logs
motif_scan.log
Decision Tree
Step 0 — Gather Required Information from the User
Before calling any tool, ask the user:
- Sample name (
sample): used as prefix and for the output directory ${sample}_known_motif_scan.
- Genome assembly (
genome): e.g. hg38, mm10, danRer11.
- Never guess or auto-detect.
Step 1: Initialize Project
- Make director for this project:
Call:
mcp__project-init-tools__project_init
with:
sample: the user-provided sample name
task: known_motif_scan
The tool will:
- Create
${sample}_known_motif_scan directory.
- Get the full path of the
${sample}_known_motif_scan directory, which will be used as ${proj_dir}.
Step 2: Prepare genome file for homer
Call:
mcp__homer-tools__check_genome_installation
With:
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
The tool will:
- Check if the genome is installed in HOMER.
- If not, install the genome.
Step 3 (Optional): Standardize chromosome names for BED files
This step is optional. Only perform this step if the input file is a BED file. If the input file is a gene list, skip this step.
From 1 format to chr1 format
From MT format to chrM format
Call:
mcp__file-format-tools__standardize_bed_chrom_names
with:
input_bed: the user-provided BED file
output_bed: the path to save the standardized BED file
The tool will:
- Standardize the chromosome names in the BED file.
- Return the path of the standardized BED file.
Step 4: Prepare motif file for a certain TF
Here are two options depending on the user's request. Pick one of them based on the user's request.
- Locate motif file for a certain TF
- Use a custom motif file
Option 1: Locate motif file for a certain TF or a set of TFs
If the user provides a TF name or a set of TF names instead of a motif file, locate the motif file for the TF.
Call:
mcp__homer-tools__locate_motif_file
With:
proj_dir: directory to save the known motif scan results. In this skill, it is the full path of the ${sample}_known_motif_scan directory returned by mcp__project-init-tools__project_init
TF_name: the user-provided TF name or a set of TF names separated by comma, e.g. TF1,TF2,TF3
motif_type: Typically do not need to specify for model organisms. If the user provides data in "insects", "plants", "rna", "worms", "yeast", choose one as the appropriate motif type.
The tool will:
- Locate the motif file for the TF.
- Return the path of the motif file.
Option 2: Use a custom motif file
If the user provides a custom motif file, use the custom motif file. If the custom motif file is in MEME format, convert it to HOMER format:
Call:
mcp__file-format-tools__meme_to_homer
With:
proj_dir: directory to save the known motif scan results. In this skill, it is the full path of the ${sample}_known_motif_scan directory returned by mcp__project-init-tools__project_init
meme_file: the user-provided MEME motif file
The tool will:
- Convert the MEME motif file to HOMER motif file.
- Return the path of the HOMER motif file.
Step 5: Scan motif
Here are 3 options depending on the user's request. Pick one of them based on the user's request.
- Scan motif in the specific genomic regions
- Scan motif in the genome
- Annotate peaks with motif hits
Option 1: Scan motif in the specific genomic regions
- If the user provides a specific genomic regions file, scan the motif in the specific genomic regions:
Call:
mcp__homer-tools__find_motifs
With:
sample: the user-provided sample name
proj_dir: directory to save the known motif scan results. In this skill, it is the full path of the ${sample}_known_motif_scan directory returned by mcp__project-init-tools__project_init
input_file: the user-provided file containing genome regions. May end with .bed, .narrowPeak, .broadPeak, etc.
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
size: region size for motif finding for genome regions, typically 200-500bp for transcription factors (default: 200). If the input file is a gene list, set to None.
mask: mask repeat regions for cleaner motif analysis (default: True)
threads: number of processors to use (default: 4)
num_motifs: number of motifs to find (default: 25)
lengths: motif lengths to search (default: 8,10,12)
find: the path to the motif file. May be the motif file returned by mcp__homer-tools__locate_motif_file. This parameter must be set for this step.
nomotif: True to not use de novo motif finding
The tool will:
- Scan for potential binding sites for a certain TF in the genome regions in the bed file or the promoters of the genes in the gene list.
- Return the path of the known motif scan results under
${proj_dir}/results/ directory:
"{sample}_motif_find.txt" (To get this, find parameter must be set)
- Convert the results to BED format:
Call:
mcp__homer-tools__homer_pos2bed
With:
pos_file: the path to the known motif scan results. It will be under ${proj_dir}/results/ directory, and ends with .motif.txt.
The tool will:
- Convert the known motif scan results to BED format.
- Return the path of the converted BED file under
${proj_dir}/results/ directory:
"{sample}_motif_find.bed"
Option 2: Scan motif in the genome
Call:
mcp__homer-tools__scan_motif_genome_wide
With:
sample: the user-provided sample name
proj_dir: directory to save the known motif scan results. In this skill, it is the full path of the ${sample}_known_motif_scan directory returned by mcp__project-init-tools__project_init
motif_file: the path to the motif file. May be the motif file returned by mcp__homer-tools__locate_motif_file.
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
mask: mask repeat regions for cleaner motif analysis (default: True)
threads: number of processors to use (default: 4)
The tool will:
- Scan for potential binding sites for a certain TF in the genome.
- Return the path of the known motif scan results under
${proj_dir}/results/ directory:
- Convert the results to BED format:
Call:
mcp__homer-tools__homer_pos2bed
With:
pos_file: the path to the known motif scan results. It will be under ${proj_dir}/results/ directory, and ends with .genomewide.txt.
The tool will:
- Convert the known motif scan results to BED format.
- Return the path of the converted BED file under
${proj_dir}/results/ directory:
Option 3: Annotate peaks with motif hits
- Annotate peaks with motif hits:
Call:
mcp__homer-tools__annotate_peaks_motif_scan
With:
sample: the user-provided sample name
proj_dir: directory to save the known motif scan results. In this skill, it is the full path of the ${sample}_known_motif_scan directory returned by mcp__project-init-tools__project_init
peakfile: the user-provided peak file in BED format. May end with .bed, .narrowPeak, .broadPeak, etc.
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
motif_file: the path to the motif file. May be the motif file returned by mcp__homer-tools__locate_motif_file.
size: region size around peak centers (default: 200)
nmotifs: number of motifs to report per peak (default: None)
mbed: output motif hits in BED format (default: True). If True, a .motif_pos.bed file will be created under ${proj_dir}/results/ directory.
mscore: include motif scores in the output (default: False)
cpu: number of processors for parallel processing (default: 1)
bedgraph: output in bedGraph format (default: False)
hist: include histogram output with given number of bins (default: None)
The tool will:
- Annotate peaks with motif hits.
- Return the path of the known motif scan results under
${proj_dir}/results/ directory:
${sample}.anno_motif.txt
${sample}.motif_pos.bed (if mbed is True)
Quality Control and Best Practices
Pre-processing Steps
- Filter peaks: Remove low-quality or artifact peaks
- Size selection: Use appropriate region size (-size parameter)
- Motif quality: Use high-quality position-specific scoring matrices
- Score thresholds: Set appropriate motif score cutoffs
Parameter Optimization
- Region size: Typically 200-500bp for transcription factors
- Number of motifs: Report top 1-5 motifs per peak
- Score thresholds: Use default or optimize based on motif quality
- Threads: Use available CPU cores for faster processing
Important Metrics
- Motif score: Position-specific scoring matrix match score
- Position: Exact genomic location of motif match
- Strand: DNA strand where motif was found
- Sequence: Actual DNA sequence at motif location
Troubleshooting
Common Issues
- No motif hits found: Check motif file format and region size
- Memory errors: Reduce region size or use fewer threads
- Slow performance: Use
-cpu option for parallel processing
- Genome not found: Verify genome assembly name and installation
Error Handling
- Ensure HOMER is properly installed and configured
- Check that genome data is downloaded and accessible
- Verify input file formats and chromosome naming
- Ensure motif files are in correct format
- Check sufficient disk space for output files
1---2name: motif-scanning3description: This skill identifies the locations of known transcription factor (TF) binding motifs within genomic regions such as ChIP-seq or ATAC-seq peaks. It utilizes HOMER to search for specific sequence motifs defined by position-specific scoring matrices (PSSMs) from known motif databases. Use this skill when you need to detect the presence and precise genomic coordinates of known TF binding motifs within experimentally defined regions such as ChIP-seq or ATAC-seq peaks.4---56# Motif Scanning78## Overview910This skill enables comprehensive motif scanning using HOMER tools for genomic peak files. It scans genomic regions for specific transcription factor binding motifs using position-specific scoring matrices and identifies exact motif locations. To perform motif scanning:1112- Always refer to the **Inputs & Outputs** section to check inputs and build the output architecture.13- Genome assembly: Always returned from user feedback (hg38, mm10, hg19, mm9, etc), never determined by yourself.14- Check chromosome names: Standardize chromosome names to format with "chr" (1 -> chr1, MT -> chrM).15- Prepare motif files: Position-specific scoring matrices (PSSM) in HOMER format, saved in ${HOMER_data}/knownTFs/motifs/${tf}.motif, and "tf" should be in lower case.16- Set scanning parameters: Region size, score thresholds, output format17- Run HOMER motif scanning command1819---2021## When to use this skill2223- Scan for potential binding sites for a certain TF in the whole genome or in specific genomic regions, like promoters of a gene list or peaks from ChIP-seq or ATAC-seq.24- Scanning ChIP-seq or ATAC-seq peaks for known motifs to validate TF binding specificity.25- Testing whether co-factor motifs (e.g., TAL1, KLF1, SPI1) co-occur within TF-bound or accessible regions to infer cooperative binding.26- Evaluating motif distribution patterns relative to genomic landmarks such as transcription start sites (TSS) or enhancers.27- Generating motif-annotated BED files for visualization in genome browsers or subsequent feature analysis.2829---3031## Inputs & Outputs3233### Inputs34(1) Peak formats supported35- **BED files**: Standard genomic interval format36- **narrowPeak**: ENCODE narrow peak format37- **broadPeak**: ENCODE broad peak format38- **HOMER peak files**: Output from HOMER peak calling39(2) Motif formats supported40- **HOMER motif format**: Position-specific scoring matrices41- **MEME motif format**: MEME suite motif format42- **TRANSFAC format**: TRANSFAC database format434445### Outputs46```bash47${sample}_known_motif_scan/48 results/49 combined_motifs.txt # combined motif hits from all TFs5051 ### Option 1: Scan motif in the specific genomic regions52 ${sample}_motif_find.txt53 ${sample}_motif_find.bed5455 ### Option 2: Scan motif in the genome56 ${sample}.genomewide.txt57 ${sample}.genomewide.bed5859 ### Option 3: Annotate peaks with motif hits60 ${sample}.anno_motif.txt61 ${sample}.motif_pos.bed (if `mbed` is True)6263 logs/ # analysis logs 64 motif_scan.log65```66---6768## Decision Tree6970### Step 0 — Gather Required Information from the User7172Before calling any tool, **ask the user**:73741. Sample name (`sample`): used as prefix and for the output directory `${sample}_known_motif_scan`.752. Genome assembly (`genome`): e.g. `hg38`, `mm10`, `danRer11`. 76 - **Never** guess or auto-detect.7778---7980### Step 1: Initialize Project81821. Make director for this project:8384Call:85- `mcp__project-init-tools__project_init`8687with:88- `sample`: the user-provided sample name89- `task`: known_motif_scan9091The tool will:92- Create `${sample}_known_motif_scan` directory.93- Get the full path of the `${sample}_known_motif_scan` directory, which will be used as `${proj_dir}`.9495---9697### Step 2: Prepare genome file for homer9899Call:100- `mcp__homer-tools__check_genome_installation`101102With:103- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`104105The tool will:106- Check if the genome is installed in HOMER.107- If not, install the genome.108109---110111### Step 3 (Optional): Standardize chromosome names for BED files112113This step is optional. Only perform this step if the input file is a BED file. If the input file is a gene list, skip this step.114115From `1` format to `chr1` format116From `MT` format to `chrM` format117118Call:119- `mcp__file-format-tools__standardize_bed_chrom_names`120121with:122- `input_bed`: the user-provided BED file123- `output_bed`: the path to save the standardized BED file124125The tool will:126- Standardize the chromosome names in the BED file.127- Return the path of the standardized BED file.128129---130131### Step 4: Prepare motif file for a certain TF132133Here are two options depending on the user's request. Pick one of them based on the user's request.1341. Locate motif file for a certain TF1352. Use a custom motif file136137#### Option 1: Locate motif file for a certain TF or a set of TFs138139If the user provides a TF name or a set of TF names instead of a motif file, locate the motif file for the TF.140141Call:142- `mcp__homer-tools__locate_motif_file`143144With:145- `proj_dir`: directory to save the known motif scan results. In this skill, it is the full path of the `${sample}_known_motif_scan` directory returned by `mcp__project-init-tools__project_init`146- `TF_name`: the user-provided TF name or a set of TF names separated by comma, e.g. `TF1,TF2,TF3`147- `motif_type`: Typically do not need to specify for model organisms. If the user provides data in "insects", "plants", "rna", "worms", "yeast", choose one as the appropriate motif type.148149The tool will:150- Locate the motif file for the TF.151- Return the path of the motif file.152153#### Option 2: Use a custom motif file154155If the user provides a custom motif file, use the custom motif file. If the custom motif file is in MEME format, convert it to HOMER format:156157Call:158- `mcp__file-format-tools__meme_to_homer`159160With:161- `proj_dir`: directory to save the known motif scan results. In this skill, it is the full path of the `${sample}_known_motif_scan` directory returned by `mcp__project-init-tools__project_init`162- `meme_file`: the user-provided MEME motif file163164The tool will:165- Convert the MEME motif file to HOMER motif file.166- Return the path of the HOMER motif file.167168---169170### Step 5: Scan motif171172Here are 3 options depending on the user's request. Pick one of them based on the user's request.1731. Scan motif in the specific genomic regions1742. Scan motif in the genome1753. Annotate peaks with motif hits176177#### Option 1: Scan motif in the specific genomic regions1781791. If the user provides a specific genomic regions file, scan the motif in the specific genomic regions:180181Call:182- `mcp__homer-tools__find_motifs`183184With:185- `sample`: the user-provided sample name186- `proj_dir`: directory to save the known motif scan results. In this skill, it is the full path of the `${sample}_known_motif_scan` directory returned by `mcp__project-init-tools__project_init`187- `input_file`: the user-provided file containing genome regions. May end with `.bed`, `.narrowPeak`, `.broadPeak`, etc.188- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`189- `size`: region size for motif finding for genome regions, typically 200-500bp for transcription factors (default: 200). If the input file is a gene list, set to None.190- `mask`: mask repeat regions for cleaner motif analysis (default: True)191- `threads`: number of processors to use (default: 4)192- `num_motifs`: number of motifs to find (default: 25)193- `lengths`: motif lengths to search (default: 8,10,12)194- `find`: the path to the motif file. May be the motif file returned by `mcp__homer-tools__locate_motif_file`. This parameter must be set for this step.195- `nomotif`: `True` to not use de novo motif finding196197The tool will:198- Scan for potential binding sites for a certain TF in the genome regions in the bed file or the promoters of the genes in the gene list.199- Return the path of the known motif scan results under `${proj_dir}/results/` directory:200 - `"{sample}_motif_find.txt"` (To get this, `find` parameter must be set)2012022. Convert the results to BED format:203204Call:205- `mcp__homer-tools__homer_pos2bed`206207With:208- `pos_file`: the path to the known motif scan results. It will be under `${proj_dir}/results/` directory, and ends with `.motif.txt`.209210The tool will:211- Convert the known motif scan results to BED format.212- Return the path of the converted BED file under `${proj_dir}/results/` directory:213 - `"{sample}_motif_find.bed"`214215216#### Option 2: Scan motif in the genome217218Call:219`mcp__homer-tools__scan_motif_genome_wide`220221With:222- `sample`: the user-provided sample name223- `proj_dir`: directory to save the known motif scan results. In this skill, it is the full path of the `${sample}_known_motif_scan` directory returned by `mcp__project-init-tools__project_init`224- `motif_file`: the path to the motif file. May be the motif file returned by `mcp__homer-tools__locate_motif_file`.225- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`226- `mask`: mask repeat regions for cleaner motif analysis (default: True)227- `threads`: number of processors to use (default: 4)228229The tool will:230- Scan for potential binding sites for a certain TF in the genome.231- Return the path of the known motif scan results under `${proj_dir}/results/` directory:232 - `${sample}.genomewide.txt`2332342352. Convert the results to BED format:236237Call:238- `mcp__homer-tools__homer_pos2bed`239240With:241- `pos_file`: the path to the known motif scan results. It will be under `${proj_dir}/results/` directory, and ends with `.genomewide.txt`.242243The tool will:244- Convert the known motif scan results to BED format.245- Return the path of the converted BED file under `${proj_dir}/results/` directory:246 - `${sample}.genomewide.bed`247248249#### Option 3: Annotate peaks with motif hits2502511. Annotate peaks with motif hits:252253Call:254`mcp__homer-tools__annotate_peaks_motif_scan`255256With:257- `sample`: the user-provided sample name258- `proj_dir`: directory to save the known motif scan results. In this skill, it is the full path of the `${sample}_known_motif_scan` directory returned by `mcp__project-init-tools__project_init`259- `peakfile`: the user-provided peak file in BED format. May end with `.bed`, `.narrowPeak`, `.broadPeak`, etc.260- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`261- `motif_file`: the path to the motif file. May be the motif file returned by `mcp__homer-tools__locate_motif_file`.262- `size`: region size around peak centers (default: 200)263- `nmotifs`: number of motifs to report per peak (default: None)264- `mbed`: output motif hits in BED format (default: True). If True, a `.motif_pos.bed` file will be created under `${proj_dir}/results/` directory.265- `mscore`: include motif scores in the output (default: False)266- `cpu`: number of processors for parallel processing (default: 1)267- `bedgraph`: output in bedGraph format (default: False)268- `hist`: include histogram output with given number of bins (default: None)269270The tool will:271- Annotate peaks with motif hits.272- Return the path of the known motif scan results under `${proj_dir}/results/` directory:273 - `${sample}.anno_motif.txt`274 - `${sample}.motif_pos.bed` (if `mbed` is True)275276277278279## Quality Control and Best Practices280281### Pre-processing Steps2821. **Filter peaks**: Remove low-quality or artifact peaks2832. **Size selection**: Use appropriate region size (-size parameter)2843. **Motif quality**: Use high-quality position-specific scoring matrices2854. **Score thresholds**: Set appropriate motif score cutoffs286287### Parameter Optimization288- **Region size**: Typically 200-500bp for transcription factors289- **Number of motifs**: Report top 1-5 motifs per peak290- **Score thresholds**: Use default or optimize based on motif quality291- **Threads**: Use available CPU cores for faster processing292293### Important Metrics294- **Motif score**: Position-specific scoring matrix match score295- **Position**: Exact genomic location of motif match296- **Strand**: DNA strand where motif was found297- **Sequence**: Actual DNA sequence at motif location298299300## Troubleshooting301302### Common Issues3031. **No motif hits found**: Check motif file format and region size3042. **Memory errors**: Reduce region size or use fewer threads3053. **Slow performance**: Use `-cpu` option for parallel processing3064. **Genome not found**: Verify genome assembly name and installation307308### Error Handling309- Ensure HOMER is properly installed and configured310- Check that genome data is downloaded and accessible311- Verify input file formats and chromosome naming312- Ensure motif files are in correct format313- Check sufficient disk space for output files314