Functional Enrichment (HOMER + R)
Overview
- Validate input: Accept BED/peak files with genomic coordinates or gene lists; check format and genome assembly.
- Map regions to genes: Convert regions to a unique gene set using HOMER
annotatePeaks.pl.
- Run GO enrichment: Use HOMER
findGO.pl (or annotatePeaks.pl -go) for BP/MF/CC.
- Run KEGG enrichment: Use HOMER
findGO.pl -kegg (or annotatePeaks.pl -kegg).
- Collect outputs: Save tidy tables for downstream plotting and a compact summary of top terms.
- Visualize in R: Create barplots and dotplots (GO/KEGG) with
ggplot2 from standardized outputs.
- QC & troubleshooting: Provide checks for genome mismatch, chromosome naming, and low-signal inputs.
Inputs & Outputs
Inputs (choose one):
Option 1: Input is a genomic region file (BED/narrowPeak/broadPeak)
Genomic region formats supported:
- BED files: Standard genomic interval format
- narrowPeak: narrow peak format
- broadPeak: broad peak format
Option 2: Input is a gene list (txt)
gene_list.txt with one official gene symbol per line (no header). And an optional gene_list_background.txt with one official gene symbol per line (no header).
Outputs (directory layout):
${sample}_functional_enrichment/
results/
${sample}.anno_genomic_features.txt
${sample}.anno_genomic_features_stats.txt
biological_process.txt
cellular_component.txt
molecular_function.txt
kegg.txt
biocyc.txt
chromosome.txt
cosmic.txt
interactions.txt
interpro.txt
gene3d.txt
pathwayInteractionDB.txt
pfam.txt
prints.txt
prosite.txt
reactome.txt
smpdb.txt
wikipathways.txt
gwas.txt
lipidmaps.txt
msigdb.txt
smart.txt
tables/
${sample}.gene_list.txt
go_bp.tsv
go_mf.tsv
go_cc.tsv
kegg.tsv
logs/
${sample}.anno_genomic_features.log # if genome region file is provided
findGO.log
Decision Tree
Step 0 — Gather Required Information from the User
Before calling any tool, ask the user:
- Sample name (
sample): used as prefix and for the output directory ${sample}_functional_enrichment.
- Genome assembly (
genome): e.g. hg38, mm10, danRer11.
- Never guess or auto-detect.
Step 1: Initialize Project
- Make director for this project:
Call:
mcp__project-init-tools__project_init
with:
sample: the user-provided sample name
task: de_novo_motif_discovery
The tool will:
- Create
${sample}_functional_enrichment directory.
- Get the full path of the
${sample}_functional_enrichment directory, which will be used as ${proj_dir}.
Step 2: Prepare genome file for homer
Call:
mcp__homer-tools__check_genome_installation
With:
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
The tool will:
- Check if the genome is installed in HOMER.
- If not, install the genome.
Step 3 (Optional): Standardize chromosome names for BED files
This step is optional. Only perform this step if the input file is a BED file. If the input file is a gene list, skip this step.
From 1 format to chr1 format
From MT format to chrM format
Call:
mcp__file-format-tools__standardize_bed_chrom_names
with:
input_bed: the user-provided BED file
output_bed: the path to save the standardized BED file
The tool will:
- Standardize the chromosome names in the BED file.
- Return the path of the standardized BED file.
Step 4 (Optional): Convert gene ID to gene symbol
This step is optional. Only perform this step if the input file is a gene list file. If the input file is a BED file, skip this step.
Call:
mcp__mygene-tools__convert_gene_ids_mygene
With:
input_ids_file: the user-provided gene list file. May end with .txt.
scopes: the source ID type for mygene (e.g., 'ensembl.gene', 'symbol', 'entrezgene', 'uniprot', or a comma-separated list).
fields: the comma-separated target fields to retrieve from mygene (e.g., 'symbol,ensembl.gene,uniprot,entrezgene').
species: the species for mygene (e.g., 'human', 'mouse', 'zebrafish', or NCBI taxon ID like '9606').
out_file: the path to save the converted gene list file. In this skill, it is the full path of the ${sample}_functional_enrichment directory returned by mcp__project-init-tools__project_init
batch_size: the batch size for mygene.querymany (default 1000).
The tool will:
- Convert the gene ID to gene symbol.
- Return the path of the converted gene list file.
Step 5: GO enrichment analysis
Option 1: from genomic regions file
Only if the input file is a BED file. If the input file is a gene list, call tools in Option 2.
- annotate the genomic regions using Homer's
annotatePeaks.pl with -go option. If user also provides a background genome region file, like a control peak file, also call this tool for the background genome region file. Use a different ${sample} as the sample name for the background sample.
Call:
mcp__homer-tools__annotate_genomic_features
With:
sample: the user-provided sample name
proj_dir: directory to save the genomic feature annotation results. In this skill, it is the full path of the ${sample}_functional_enrichment directory returned by mcp__project-init-tools__project_init
regions_bed: the user-provided regions file in BED format. May end with .bed, .narrowPeak, .broadPeak, etc.
genome: the user-provided genome assembly, e.g. hg38, mm10, danRer11
ann: "custom homer annotation file (created by assignGenomeAnnotation.pl), (default: None).
size_given: keep original region sizes (default: True)
cpg: include CpG information (default: False)
go: True to perform GO enrichment analysis.
The tool will:
- Annotate the genomic regions using Homer's
annotatePeaks.pl.
- Return the path of the annotated regions file under
${proj_dir}/results/ directory, and the path to the log file under ${proj_dir}/logs/ directory.
${proj_dir}/results/${sample}.anno_genomic_features.txt
${proj_dir}/results/${sample}.anno_genomic_features_stats.txt
${proj_dir}/logs/${sample}.anno_genomic_features.log
- (optional) extract the genes from the annotated regions file if neccessary for future analysis or the target gene list is requested by user. If not requested, skip this step.
Call:
mcp__file-format-tools__extract_gene_list
With:
sample: the user-provided sample name
proj_dir: directory to save the genomic feature annotation results. In this skill, it is the full path of the ${sample}_functional_enrichment directory returned by mcp__project-init-tools__project_init
The tool will:
- Extract the genes from the annotated regions file.
- Return the path of the gene list file under
${proj_dir}/tables/ directory.
${proj_dir}/tables/${sample}.gene_list.txt
Option 2: from gene list file
Only if the input file is a gene list file. If the input file is a BED file, call tools in Option 1.
Call:
mcp__homer-tools__gene_function_enrichment
With:
sample: the user-provided sample name
proj_dir: directory to save the GO & KEGG enrichment results. In this skill, it is the full path of the ${sample}_functional_enrichment directory returned by mcp__project-init-tools__project_init
gene_list_file: the user-provided gene list file. May end with .txt.
organism: the user-provided organism name, e.g. human, mouse, zebrafish, etc.
background_gene_list_file: the user-provided background gene list file. May end with .txt. If not provided, set this parameter to None.
The tool will:
- Find the GO enrichment for the gene list.
- Return the path of the GO & KEGG enrichment results under
${proj_dir}/results/ directory.
${proj_dir}/results/biological_process.txt
${proj_dir}/results/kegg.txt
- ... other GO and KEGG enrichment results files.
- Return the path of the log file under
${proj_dir}/logs/ directory.
${proj_dir}/logs/${sample}.find_go_and_kegg_enrichment.log
Alternative direct from BED
annotatePeaks.pl peaks.bed hg38 -go results/{run}/tables/go_dir -genomeOntology
annotatePeaks.pl peaks.bed hg38 -kegg results/{run}/tables/kegg_dir
Notes & Best Practices
- Genome & naming: Ensure the HOMER genome key matches the species; chromosome naming must be consistent (
chr1 vs 1).
- BED format: Tab-delimited, ≥3 columns, 0-based coordinates, no header.
- Multiple testing: Prefer FDR (BH) if provided; otherwise fallback to P-value.
- Background set:
-bg helps reduce bias; choose a reasonable universe (e.g., all expressed or all accessible regions → genes).
- Direct-from-BED:
annotatePeaks.pl -go/-kegg is convenient; the gene-list route yields uniform TSVs for plotting.
Troubleshooting
- Many NAs after annotation: Check genome version, chromosome naming, BED formatting, and headers.
- Empty/weak enrichment: Ensure sufficient genes (suggest ≥50), verify species of symbols, tune thresholds or background.
- Column name drift: HOMER versions may differ; adjust R column mappings if needed.
1---2name: functional-enrichment3description: Perform GO and KEGG functional enrichment using HOMER from genomic regions (BED/narrowPeak/broadPeak) or gene lists, and produce R-based barplot/dotplot visualizations. Use this skill when you want to perform GO and KEGG functional enrichment using HOMER from genomic regions or just want to link genomic region to genes.4---56# Functional Enrichment (HOMER + R)78## Overview910- **Validate input**: Accept BED/peak files with genomic coordinates or gene lists; check format and genome assembly.11- **Map regions to genes**: Convert regions to a unique gene set using HOMER `annotatePeaks.pl`.12- **Run GO enrichment**: Use HOMER `findGO.pl` (or `annotatePeaks.pl -go`) for BP/MF/CC.13- **Run KEGG enrichment**: Use HOMER `findGO.pl -kegg` (or `annotatePeaks.pl -kegg`).14- **Collect outputs**: Save tidy tables for downstream plotting and a compact summary of top terms.15- **Visualize in R**: Create barplots and dotplots (GO/KEGG) with `ggplot2` from standardized outputs.16- **QC & troubleshooting**: Provide checks for genome mismatch, chromosome naming, and low-signal inputs.1718## Inputs & Outputs1920### Inputs (choose one):21#### Option 1: Input is a genomic region file (BED/narrowPeak/broadPeak)22Genomic region formats supported:23- **BED files**: Standard genomic interval format24- **narrowPeak**: narrow peak format25- **broadPeak**: broad peak format2627#### Option 2: Input is a gene list (txt)28- `gene_list.txt` with one official gene symbol per line (no header). And an optional `gene_list_background.txt` with one official gene symbol per line (no header).2930### Outputs (directory layout):31```bash32${sample}_functional_enrichment/33 results/34 ${sample}.anno_genomic_features.txt35 ${sample}.anno_genomic_features_stats.txt36 biological_process.txt37 cellular_component.txt 38 molecular_function.txt 3940 kegg.txt 41 biocyc.txt 42 chromosome.txt 43 cosmic.txt44 interactions.txt 45 interpro.txt46 gene3d.txt47 pathwayInteractionDB.txt48 pfam.txt49 prints.txt 50 prosite.txt 51 reactome.txt52 smpdb.txt53 wikipathways.txt5455 gwas.txt 56 lipidmaps.txt 57 msigdb.txt 58 smart.txt5960 tables/61 ${sample}.gene_list.txt62 go_bp.tsv63 go_mf.tsv64 go_cc.tsv65 kegg.tsv66 logs/67 ${sample}.anno_genomic_features.log # if genome region file is provided68 findGO.log69```707172## Decision Tree737475### Step 0 — Gather Required Information from the User7677Before calling any tool, **ask the user**:78791. Sample name (`sample`): used as prefix and for the output directory `${sample}_functional_enrichment`.802. Genome assembly (`genome`): e.g. `hg38`, `mm10`, `danRer11`. 81 - **Never** guess or auto-detect.8283---8485### Step 1: Initialize Project86871. Make director for this project:8889Call:90- `mcp__project-init-tools__project_init`9192with:93- `sample`: the user-provided sample name94- `task`: de_novo_motif_discovery9596The tool will:97- Create `${sample}_functional_enrichment` directory.98- Get the full path of the `${sample}_functional_enrichment` directory, which will be used as `${proj_dir}`.99100---101102103### Step 2: Prepare genome file for homer104105Call:106- `mcp__homer-tools__check_genome_installation`107108With:109- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`110111The tool will:112- Check if the genome is installed in HOMER.113- If not, install the genome.114115116---117118119### Step 3 (Optional): Standardize chromosome names for BED files120121This step is optional. Only perform this step if the input file is a BED file. If the input file is a gene list, skip this step.122123From `1` format to `chr1` format124From `MT` format to `chrM` format125126Call:127- `mcp__file-format-tools__standardize_bed_chrom_names`128129with:130- `input_bed`: the user-provided BED file131- `output_bed`: the path to save the standardized BED file132133The tool will:134- Standardize the chromosome names in the BED file.135- Return the path of the standardized BED file.136137138---139140### Step 4 (Optional): Convert gene ID to gene symbol141142This step is optional. Only perform this step if the input file is a gene list file. If the input file is a BED file, skip this step.143144Call:145- `mcp__mygene-tools__convert_gene_ids_mygene`146147With:148- `input_ids_file`: the user-provided gene list file. May end with `.txt`.149- `scopes`: the source ID type for mygene (e.g., 'ensembl.gene', 'symbol', 'entrezgene', 'uniprot', or a comma-separated list).150- `fields`: the comma-separated target fields to retrieve from mygene (e.g., 'symbol,ensembl.gene,uniprot,entrezgene').151- `species`: the species for mygene (e.g., 'human', 'mouse', 'zebrafish', or NCBI taxon ID like '9606').152- `out_file`: the path to save the converted gene list file. In this skill, it is the full path of the `${sample}_functional_enrichment` directory returned by `mcp__project-init-tools__project_init`153- `batch_size`: the batch size for mygene.querymany (default 1000).154155The tool will:156- Convert the gene ID to gene symbol.157- Return the path of the converted gene list file.158159---160161162### Step 5: GO enrichment analysis163164#### Option 1: from genomic regions file165166Only if the input file is a BED file. If the input file is a gene list, call tools in Option 2.1671681. annotate the genomic regions using Homer's `annotatePeaks.pl` with `-go` option. If user also provides a background genome region file, like a control peak file, also call this tool for the background genome region file. Use a different `${sample}` as the sample name for the background sample.169170Call:171`mcp__homer-tools__annotate_genomic_features`172173With:174- `sample`: the user-provided sample name175- `proj_dir`: directory to save the genomic feature annotation results. In this skill, it is the full path of the `${sample}_functional_enrichment` directory returned by `mcp__project-init-tools__project_init`176- `regions_bed`: the user-provided regions file in BED format. May end with `.bed`, `.narrowPeak`, `.broadPeak`, etc.177- `genome`: the user-provided genome assembly, e.g. `hg38`, `mm10`, `danRer11`178- `ann`: "custom homer annotation file (created by assignGenomeAnnotation.pl), (default: None).179- `size_given`: keep original region sizes (default: True)180- `cpg`: include CpG information (default: False)181- `go`: `True` to perform GO enrichment analysis.182183The tool will:184- Annotate the genomic regions using Homer's `annotatePeaks.pl`.185- Return the path of the annotated regions file under `${proj_dir}/results/` directory, and the path to the log file under `${proj_dir}/logs/` directory.186 - `${proj_dir}/results/${sample}.anno_genomic_features.txt`187 - `${proj_dir}/results/${sample}.anno_genomic_features_stats.txt`188 - `${proj_dir}/logs/${sample}.anno_genomic_features.log`189190---1911922. (optional) extract the genes from the annotated regions file if neccessary for future analysis or the target gene list is requested by user. If not requested, skip this step.193194Call:195`mcp__file-format-tools__extract_gene_list`196197With:198- `sample`: the user-provided sample name199- `proj_dir`: directory to save the genomic feature annotation results. In this skill, it is the full path of the `${sample}_functional_enrichment` directory returned by `mcp__project-init-tools__project_init`200201The tool will:202- Extract the genes from the annotated regions file.203- Return the path of the gene list file under `${proj_dir}/tables/` directory.204 - `${proj_dir}/tables/${sample}.gene_list.txt`205206207---208209#### Option 2: from gene list file210211Only if the input file is a gene list file. If the input file is a BED file, call tools in Option 1.212213Call:214`mcp__homer-tools__gene_function_enrichment`215216With:217- `sample`: the user-provided sample name218- `proj_dir`: directory to save the GO & KEGG enrichment results. In this skill, it is the full path of the `${sample}_functional_enrichment` directory returned by `mcp__project-init-tools__project_init`219- `gene_list_file`: the user-provided gene list file. May end with `.txt`.220- `organism`: the user-provided organism name, e.g. `human`, `mouse`, `zebrafish`, etc.221- `background_gene_list_file`: the user-provided background gene list file. May end with `.txt`. If not provided, set this parameter to `None`.222223The tool will:224- Find the GO enrichment for the gene list.225- Return the path of the GO & KEGG enrichment results under `${proj_dir}/results/` directory.226 - `${proj_dir}/results/biological_process.txt`227 - `${proj_dir}/results/kegg.txt`228 - ... other GO and KEGG enrichment results files.229- Return the path of the log file under `${proj_dir}/logs/` directory.230 - `${proj_dir}/logs/${sample}.find_go_and_kegg_enrichment.log`231232233---234235236237238> **Alternative direct from BED** 239> `annotatePeaks.pl peaks.bed hg38 -go results/{run}/tables/go_dir -genomeOntology` 240> `annotatePeaks.pl peaks.bed hg38 -kegg results/{run}/tables/kegg_dir`241242243## Notes & Best Practices244245- **Genome & naming**: Ensure the HOMER genome key matches the species; chromosome naming must be consistent (`chr1` vs `1`).246- **BED format**: Tab-delimited, ≥3 columns, 0-based coordinates, no header.247- **Multiple testing**: Prefer FDR (BH) if provided; otherwise fallback to P-value.248- **Background set**: `-bg` helps reduce bias; choose a reasonable universe (e.g., all expressed or all accessible regions → genes).249- **Direct-from-BED**: `annotatePeaks.pl -go/-kegg` is convenient; the gene-list route yields uniform TSVs for plotting.250251## Troubleshooting252253- **Many NAs after annotation**: Check genome version, chromosome naming, BED formatting, and headers.254- **Empty/weak enrichment**: Ensure sufficient genes (suggest ≥50), verify species of symbols, tune thresholds or background.255- **Column name drift**: HOMER versions may differ; adjust R column mappings if needed.