Source: https://github.com/aipoch/medical-research-skills
XGBoost Modeling And Feature Importance Ranking
Use this skill to train an XGBoost model from a tabular dataset and export both feature importance ranking tables and feature importance plots.
Use This Skill When
- You need a command-line XGBoost workflow in R for tabular data.
- You need a reproducible train-test split, model training, and evaluation.
- You need feature importance ranking outputs as both a table and a figure.
- You need automatic one-hot encoding for categorical predictors.
- Your data may contain a first unnamed sample ID column such as
V1 that should not enter the model.
Do Not Use This Skill When
- Your classification target has more than 2 classes.
- Your input is not tabular CSV, TXT, or TSV data.
- You need causal interpretation, mechanism claims, or policy, business, or clinical conclusions.
- You only need narrative interpretation or triage of an existing result rather than model training.
Primary Command
Rscript scripts/main.R \
--data_file <input_file> \
--target_var <target_column> \
--task_type <auto|classification|regression> \
--output_dir <output_dir>
Prerequisites
Rscript is available in the shell.
- Required R packages:
optparse, data.table, Matrix, xgboost.
- Install missing packages with
Rscript -e 'install.packages(c("optparse", "data.table", "Matrix", "xgboost"), repos="https://cloud.r-project.org")'.
Core Arguments
| Argument |
Required |
Description |
--data_file |
Yes |
Input CSV, TXT, or TSV file |
--target_var |
Yes |
Target column used for modeling |
--task_type |
No |
auto, classification, or regression. Default auto |
--output_dir |
No |
Output directory, default ./XGBoost_Results |
--ignore_vars |
No |
Comma-separated columns to exclude from predictors |
--positive_class |
No |
Positive class label for binary classification |
--test_size |
No |
Test set proportion between 0 and 1, default 0.2 |
--seed |
No |
Random seed, default 123 |
--nrounds |
No |
Maximum boosting rounds, default 300 |
--max_depth |
No |
Tree depth, default 6 |
--eta |
No |
Learning rate, default 0.1 |
--subsample |
No |
Row sampling ratio, default 0.8 |
--colsample_bytree |
No |
Column sampling ratio, default 0.8 |
--min_child_weight |
No |
Minimum child weight, default 1 |
--gamma |
No |
Minimum split loss reduction, default 0 |
--lambda |
No |
L2 regularization, default 1 |
--alpha |
No |
L1 regularization, default 0 |
--early_stopping_rounds |
No |
Early stopping rounds, default 20 |
--importance_metric |
No |
gain, cover, or frequency. Default gain |
--top_n |
No |
Number of features to plot, default 20 |
--output_format |
No |
Table format: csv or txt, default csv |
--output_prefix |
No |
Output filename prefix, default xgboost |
Input Requirements
- The input file must contain the target column.
- Predictor columns can be numeric, integer, logical, character, or factor-like text.
- Character and factor predictors are one-hot encoded automatically.
- A first unnamed identifier column such as
V1 is automatically excluded when it contains unique sample IDs.
- Rows with missing target values are removed before training.
- For classification, exactly 2 classes are required.
- For regression, the target column must be numeric.
- Each class should have at least 2 rows so both training and test sets can be created.
Example input:
,fustat,CAMK2N2,GGT6,GPR161,RAB26,RIBC2
TCGA-C5-A1M5,1,2.248291938,5.274690305,2.825215762,3.121114894,5.35318565
TCGA-EA-A5O9,0,3.346176843,5.404368414,2.604616977,0.629473197,4.429314674
TCGA-C5-A3HL,0,3.363100974,5.363314779,4.124799581,4.127228806,4.916596068
Minimal Workflow
- Confirm the input file exists and the target column name is correct.
- Run
scripts/main.R with --data_file and --target_var.
- Check the output directory for
table/feature_importance_* and figure/feature_importance_*.
If you omit --data_file or --target_var, the script exits with SKILL_MISSING_INPUT.
Outputs
Expected output structure:
<output_dir>/
├── table/
├── figure/
└── data/
Primary outputs:
table/<output_prefix>_feature_importance.csv
table/<output_prefix>_model_performance.csv
figure/<output_prefix>_feature_importance_<importance_metric>.png
Additional outputs:
Feature importance table fields include:
Rank
Feature
Gain
Cover
Frequency
SelectedMetric
SelectedValue
Feature Importance Metrics
gain: Average contribution to loss reduction. Recommended for most ranking use cases.
cover: Relative sample coverage contributed by a feature.
frequency: How often the feature is used in splits.
Read These Files When Needed
| Need |
File |
| XGBoost method details and importance interpretation |
references/algorithm.md |
| More CLI examples |
references/cli-guide.md |
| Error diagnosis |
references/troubleshooting.md |
| Main execution entry point |
scripts/main.R |
| Bundled test data |
tests/data/ |
Quick Examples
Auto-detected binary classification on dt_sample1.csv:
Rscript scripts/main.R \
--data_file tests/data/dt_sample1.csv \
--target_var fustat \
--task_type auto \
--output_dir tests/output_binary
Binary classification on dt_sample2.csv:
Rscript scripts/main.R \
--data_file tests/data/dt_sample2.csv \
--target_var fustat \
--task_type classification \
--importance_metric gain \
--output_dir tests/output_gain
Character-label classification on dt_sample3.txt:
Rscript scripts/main.R \
--data_file tests/data/dt_sample3.txt \
--target_var Group \
--task_type classification \
--positive_class high \
--top_n 15 \
--output_dir tests/output_group
Validation
Rscript scripts/main.R --help
Rscript scripts/main.R \
--data_file tests/data/dt_sample1.csv \
--target_var fustat \
--task_type classification \
--output_dir tests/validation_output
After running analysis, verify that these files exist:
tests/validation_output/table/xgboost_feature_importance.csv
tests/validation_output/table/xgboost_model_performance.csv
tests/validation_output/figure/xgboost_feature_importance_gain.png
Common Errors
SKILL_FILE_NOT_FOUND: Input file path is wrong or inaccessible.
SKILL_MISSING_COLUMNS: The target column is missing.
SKILL_INVALID_DATA: Data is malformed, the target type is unsuitable, classification has more or fewer than 2 classes, or too few usable rows remain.
SKILL_INVALID_PARAMETER: An argument value is invalid.
SKILL_DEPENDENCY_MISSING: A required R package such as xgboost is unavailable.
If the issue is not obvious, read references/troubleshooting.md.
1---2name: xgboost-analysis3description: Use when building XGBoost models on tabular data and returning feature importance ranking outputs. Supports binary classification and regression with automatic task detection, train-test split, performance tables, feature importance ranking tables, and PNG importance plots.4license: MIT5---6> **Source**: [https://github.com/aipoch/medical-research-skills](https://github.com/aipoch/medical-research-skills)
7
8# XGBoost Modeling And Feature Importance Ranking
9
10Use this skill to train an XGBoost model from a tabular dataset and export both feature importance ranking tables and feature importance plots.
11
12## Use This Skill When
13
14- You need a command-line XGBoost workflow in R for tabular data.
15- You need a reproducible train-test split, model training, and evaluation.
16- You need feature importance ranking outputs as both a table and a figure.
17- You need automatic one-hot encoding for categorical predictors.
18- Your data may contain a first unnamed sample ID column such as `V1` that should not enter the model.
19
20## Do Not Use This Skill When
21
22- Your classification target has more than 2 classes.
23- Your input is not tabular CSV, TXT, or TSV data.
24- You need causal interpretation, mechanism claims, or policy, business, or clinical conclusions.
25- You only need narrative interpretation or triage of an existing result rather than model training.
26
27## Primary Command
28
29```bash
30Rscript scripts/main.R \
31 --data_file <input_file> \
32 --target_var <target_column> \
33 --task_type <auto|classification|regression> \
34 --output_dir <output_dir>
35```
36
37## Prerequisites
38
39- `Rscript` is available in the shell.
40- Required R packages: `optparse`, `data.table`, `Matrix`, `xgboost`.
41- Install missing packages with `Rscript -e 'install.packages(c("optparse", "data.table", "Matrix", "xgboost"), repos="https://cloud.r-project.org")'`.
42
43## Core Arguments
44
45| Argument | Required | Description |
46|----------|----------|-------------|
47| `--data_file` | Yes | Input CSV, TXT, or TSV file |
48| `--target_var` | Yes | Target column used for modeling |
49| `--task_type` | No | `auto`, `classification`, or `regression`. Default `auto` |
50| `--output_dir` | No | Output directory, default `./XGBoost_Results` |
51| `--ignore_vars` | No | Comma-separated columns to exclude from predictors |
52| `--positive_class` | No | Positive class label for binary classification |
53| `--test_size` | No | Test set proportion between 0 and 1, default `0.2` |
54| `--seed` | No | Random seed, default `123` |
55| `--nrounds` | No | Maximum boosting rounds, default `300` |
56| `--max_depth` | No | Tree depth, default `6` |
57| `--eta` | No | Learning rate, default `0.1` |
58| `--subsample` | No | Row sampling ratio, default `0.8` |
59| `--colsample_bytree` | No | Column sampling ratio, default `0.8` |
60| `--min_child_weight` | No | Minimum child weight, default `1` |
61| `--gamma` | No | Minimum split loss reduction, default `0` |
62| `--lambda` | No | L2 regularization, default `1` |
63| `--alpha` | No | L1 regularization, default `0` |
64| `--early_stopping_rounds` | No | Early stopping rounds, default `20` |
65| `--importance_metric` | No | `gain`, `cover`, or `frequency`. Default `gain` |
66| `--top_n` | No | Number of features to plot, default `20` |
67| `--output_format` | No | Table format: `csv` or `txt`, default `csv` |
68| `--output_prefix` | No | Output filename prefix, default `xgboost` |
69
70## Input Requirements
71
72- The input file must contain the target column.
73- Predictor columns can be numeric, integer, logical, character, or factor-like text.
74- Character and factor predictors are one-hot encoded automatically.
75- A first unnamed identifier column such as `V1` is automatically excluded when it contains unique sample IDs.
76- Rows with missing target values are removed before training.
77- For classification, exactly 2 classes are required.
78- For regression, the target column must be numeric.
79- Each class should have at least 2 rows so both training and test sets can be created.
80
81Example input:
82
83```csv
84,fustat,CAMK2N2,GGT6,GPR161,RAB26,RIBC2
85TCGA-C5-A1M5,1,2.248291938,5.274690305,2.825215762,3.121114894,5.35318565
86TCGA-EA-A5O9,0,3.346176843,5.404368414,2.604616977,0.629473197,4.429314674
87TCGA-C5-A3HL,0,3.363100974,5.363314779,4.124799581,4.127228806,4.916596068
88```
89
90## Minimal Workflow
91
921. Confirm the input file exists and the target column name is correct.
932. Run `scripts/main.R` with `--data_file` and `--target_var`.
943. Check the output directory for `table/feature_importance_*` and `figure/feature_importance_*`.
95
96If you omit `--data_file` or `--target_var`, the script exits with `SKILL_MISSING_INPUT`.
97
98## Outputs
99
100Expected output structure:
101
102```text
103<output_dir>/
104├── table/
105├── figure/
106└── data/
107```
108
109Primary outputs:
110
111- `table/<output_prefix>_feature_importance.csv`
112- `table/<output_prefix>_model_performance.csv`
113- `figure/<output_prefix>_feature_importance_<importance_metric>.png`
114
115Additional outputs:
116
117- `session_info.txt`
118
119Feature importance table fields include:
120
121- `Rank`
122- `Feature`
123- `Gain`
124- `Cover`
125- `Frequency`
126- `SelectedMetric`
127- `SelectedValue`
128
129## Feature Importance Metrics
130
131- `gain`: Average contribution to loss reduction. Recommended for most ranking use cases.
132- `cover`: Relative sample coverage contributed by a feature.
133- `frequency`: How often the feature is used in splits.
134
135## Read These Files When Needed
136
137| Need | File |
138|------|------|
139| XGBoost method details and importance interpretation | `references/algorithm.md` |
140| More CLI examples | `references/cli-guide.md` |
141| Error diagnosis | `references/troubleshooting.md` |
142| Main execution entry point | `scripts/main.R` |
143| Bundled test data | `tests/data/` |
144
145## Quick Examples
146
147Auto-detected binary classification on `dt_sample1.csv`:
148
149```bash
150Rscript scripts/main.R \
151 --data_file tests/data/dt_sample1.csv \
152 --target_var fustat \
153 --task_type auto \
154 --output_dir tests/output_binary
155```
156
157Binary classification on `dt_sample2.csv`:
158
159```bash
160Rscript scripts/main.R \
161 --data_file tests/data/dt_sample2.csv \
162 --target_var fustat \
163 --task_type classification \
164 --importance_metric gain \
165 --output_dir tests/output_gain
166```
167
168Character-label classification on `dt_sample3.txt`:
169
170```bash
171Rscript scripts/main.R \
172 --data_file tests/data/dt_sample3.txt \
173 --target_var Group \
174 --task_type classification \
175 --positive_class high \
176 --top_n 15 \
177 --output_dir tests/output_group
178```
179
180## Validation
181
182```bash
183Rscript scripts/main.R --help
184```
185
186```bash
187Rscript scripts/main.R \
188 --data_file tests/data/dt_sample1.csv \
189 --target_var fustat \
190 --task_type classification \
191 --output_dir tests/validation_output
192```
193
194After running analysis, verify that these files exist:
195
196- `tests/validation_output/table/xgboost_feature_importance.csv`
197- `tests/validation_output/table/xgboost_model_performance.csv`
198- `tests/validation_output/figure/xgboost_feature_importance_gain.png`
199
200## Common Errors
201
202- `SKILL_FILE_NOT_FOUND`: Input file path is wrong or inaccessible.
203- `SKILL_MISSING_COLUMNS`: The target column is missing.
204- `SKILL_INVALID_DATA`: Data is malformed, the target type is unsuitable, classification has more or fewer than 2 classes, or too few usable rows remain.
205- `SKILL_INVALID_PARAMETER`: An argument value is invalid.
206- `SKILL_DEPENDENCY_MISSING`: A required R package such as `xgboost` is unavailable.
207
208If the issue is not obvious, read `references/troubleshooting.md`.