Train a sentence-transformers Model
Overview
This skill trains or fine-tunes sentence-transformers models across three model classes:
SentenceTransformer (bi-encoder; dense or static embedding model) — for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal.
CrossEncoder (reranker; pair scoring) — for two-stage retrieval / pair classification.
SparseEncoder (SPLADE; sparse vectors over vocabulary) — for learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene).
This SKILL.md is a router, not a manual. It tells you which references and example scripts to load for your task. The actual content — recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting — lives in references/ and scripts/.
Do not synthesize a training script from this file alone. Open the matching train_<type>_example.py in this skill's scripts folder and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, force=True, seed, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.
When to Use
Use this skill when the user needs to:
- Train or fine-tune a dense embedding model (retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal).
- Train or fine-tune a reranker / cross-encoder for two-stage retrieval or pair classification.
- Train or fine-tune a SPLADE / sparse encoder for learned-sparse retrieval or inverted-index backends.
- Fine-tune an existing sentence-transformers checkpoint on a custom dataset.
- Add LoRA, distillation, Matryoshka, multi-dataset, multilingual, or static-embedding variants.
Trigger keywords: embedding model training, fine-tune sentence-transformers, train reranker, train cross-encoder, train SPLADE, sparse encoder training, retrieval model fine-tuning, sentence similarity model training, bi-encoder training, dense retrieval training.
Prerequisites
pip install "sentence-transformers[train]>=5.0"
# For multimodal [SentenceTransformer], add the relevant extra:
# pip install "sentence-transformers[train,image]>=5.0"
# pip install "sentence-transformers[train,audio]>=5.0"
# pip install "sentence-transformers[train,video]>=5.0"
pip install trackio # optional tracker; or wandb / tensorboard / mlflow
hf auth login # or set HF_TOKEN with write scope (for Hub push)
GPU strongly recommended. CPU works only for demos and [SentenceTransformer] StaticEmbedding.
Procedure
Step 1 — Identify the model type
| Tag |
Class |
What it does |
When to pick |
| [SentenceTransformer] |
SentenceTransformer (bi-encoder) |
Maps each input to a fixed-dim dense vector |
Retrieval, similarity, clustering, classification, paraphrase mining, dedup |
| [CrossEncoder] |
CrossEncoder (reranker) |
Scores (query, passage) pairs jointly |
Two-stage retrieval (rerank top-100 from bi-encoder), pair classification |
| [SparseEncoder] |
SparseEncoder (SPLADE) |
Sparse vectors over the vocabulary |
Learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene) |
Tiebreakers when the request is ambiguous:
- "embedding model" / "vector search" / "similarity" → [SentenceTransformer]
- "rerank" / "ranker" / "two-stage" → [CrossEncoder]
- "SPLADE" / "sparse" / "inverted index" → [SparseEncoder]
- If still unclear, ask the user.
Step 2 — Load required reading (in full, before writing any code)
Do not triage by perceived relevance. Read every file listed for your type.
Per-type — always required
[SentenceTransformer]
references/losses_sentence_transformer.md — loss-to-data-shape mapping; BatchSamplers.NO_DUPLICATES requirement for MNRL-family; Cached* ↔ gradient_checkpointing incompatibility.
references/evaluators_sentence_transformer.md — evaluator-to-task mapping; metric_for_best_model key construction (named vs unnamed); per-evaluator primary_metric values.
references/model_architectures.md — encoder vs decoder vs static vs Router pipelines; pooling rules (mean / cls / lasttoken); auto-mean-pooling behavior for fresh-start MLM bases.
scripts/train_sentence_transformer_example.py — production template; copy this as your starting point.
[CrossEncoder]
references/losses_cross_encoder.md — pointwise / pairwise / listwise / distillation; pos_weight derivation; activation_fn=Identity() mandatory for non-BCE losses (silent eval-rank collapse otherwise).
references/evaluators_cross_encoder.md — CrossEncoderRerankingEvaluator recipe; named-evaluator key format eval_{name}_{primary_metric}.
scripts/train_cross_encoder_example.py — production template; copy this as your starting point.
[SparseEncoder]
references/losses_sparse_encoder.md — SpladeLoss wrapper requirement; FLOPS regularizer weights; smoke-test active-dim ramp behavior.
references/evaluators_sparse_encoder.md — SparseNanoBEIREvaluator (English-only) and the in-domain alternative; eval_{name}_{primary_metric} key format.
scripts/train_sparse_encoder_example.py — production template; copy this as your starting point.
Cross-cutting — always required (regardless of task)
references/training_args.md — TrainingArguments knobs, precision rules (load fp32 + autocast bf16/fp16; never torch_dtype=bfloat16), warmup_steps (float) vs deprecated warmup_ratio, save_steps must be a multiple of eval_steps for load_best_model_at_end, schedulers, HPO, tracker, resume, hub-push variants.
references/dataset_formats.md — column-matching rules (label name auto-detection; column-order-not-name); reshaping recipes; hard-negative mining options.
references/base_model_selection.md — discovery commands; per-type model namespaces; ModernBERT-family max_seq_length=8192 trap; datasets >= 4 script-loader rejection; non-English starting-point shortcuts.
references/troubleshooting.md — symptom-indexed failure recipes. Skim the section headings on every run, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.
Cross-cutting — load when applicable
references/hardware_guide.md — Required for >24GB models, multi-GPU, or HF Jobs runs. VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors.
references/hf_jobs_execution.md — Required when running on HF Jobs.
references/prompts_and_instructions.md — Required when using prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding query: / passage: style prefixes.
Variant scripts (open when the task matches)
- [SentenceTransformer]:
scripts/train_sentence_transformer_matryoshka_example.py
scripts/train_sentence_transformer_multi_dataset_example.py
scripts/train_sentence_transformer_with_lora_example.py
scripts/train_sentence_transformer_distillation_example.py
scripts/train_sentence_transformer_make_multilingual_example.py
scripts/train_sentence_transformer_static_embedding_example.py
- [CrossEncoder]:
scripts/train_cross_encoder_distillation_example.py
scripts/train_cross_encoder_listwise_example.py
- [SparseEncoder]:
scripts/train_sparse_encoder_distillation_example.py
- Hard-negative mining CLI:
scripts/mine_hard_negatives.py
Step 3 — Copy the production template
Open the matching train_<type>_example.py in this skill's scripts folder (scripts/train_sentence_transformer_example.py, scripts/train_cross_encoder_example.py, or scripts/train_sparse_encoder_example.py) and copy it as your starting point. Do not write a training script from scratch or from a synthesized snippet.
Step 4 — Replace placeholders with the user's task
Replace MODEL_NAME, DATASET_NAME, RUN_NAME, the loss, and the evaluator with the user's task.
- Cross-check loss/data-shape match against the matching losses file (
references/losses_sentence_transformer.md, references/losses_cross_encoder.md, or references/losses_sparse_encoder.md).
- Cross-check the
metric_for_best_model key against the matching evaluators file (references/evaluators_sentence_transformer.md, references/evaluators_cross_encoder.md, or references/evaluators_sparse_encoder.md) (named evaluators format the key as eval_{name}_{primary_metric}).
Step 5 — Smoke-test before any long run
Set max_steps=1 with a tiny dataset slice. The production templates show one common pattern using a SMOKE_TEST environment variable.
# Example smoke-test invocation (pattern from the production template)
$env:SMOKE_TEST = "1"
python scripts/train_sentence_transformer_example.py
Step 6 — Run the full training
Remove-Item Env:SMOKE_TEST # clear smoke-test flag
python scripts/train_sentence_transformer_example.py
Step 7 — After the run
- Append results to
logs/experiments.md.
- Propose iteration if the verdict is weak or marginal.
Constraints the produced script must satisfy
These are non-negotiable contracts. Implementation lives in the production templates and references — do not reinvent.
- Capture the pre-training evaluator score as
baseline_eval before trainer.train().
- Emit a single end-of-run verdict line:
VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=...
A monitor scrapes for this line.
- Silence noisy loggers — set
httpx, httpcore, huggingface_hub, urllib3, filelock, fsspec to WARNING. Otherwise HF download URLs flood the agent's context.
- Tee logs to
logs/{RUN_NAME}.log.
- End with
model.push_to_hub(...) wrapped in try/except.
- Smoke-test before any long run (
max_steps=1 + tiny dataset slice). The production templates show one common pattern (SMOKE_TEST env var).
- [CrossEncoder] Include
EarlyStoppingCallback(patience>=3) — CE rerankers often peak mid-training and regress.
- [SparseEncoder] Log
query_active_dims / corpus_active_dims on the verdict line; high nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. ..._query_active_dims); use suffix matching to pluck them — see the SPARSE production template for the exact pattern.
Defaults
Override only if the user specifies otherwise:
- Local execution. Pitch HF Jobs only if local hardware can't fit the job.
- Single run. After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in
references/training_args.md (Experimentation section).
- Public Hub push at end-of-run, wrapped in try-except. On HF Jobs (ephemeral env) ALSO enable in-trainer push (
push_to_hub=True + hub_strategy="every_save"); details in references/hf_jobs_execution.md.
Pitfalls
- Do not synthesize a training script from this file alone. Always start from the matching
train_<type>_example.py in this skill's scripts folder. Prior agent runs have repeatedly missed load-bearing scaffolding (autocast helper, model-card class, logger silencing, force=True, seed, TF32, version-compatible imports, named-evaluator metric handling) when rolling their own.
- [CrossEncoder]
activation_fn=Identity() is mandatory for non-BCE losses. Using the default sigmoid activation with a non-BCE loss causes silent eval-rank collapse.
- [SentenceTransformer]
Cached* losses are incompatible with gradient_checkpointing. Do not combine them.
- [SentenceTransformer] MNRL-family losses require
BatchSamplers.NO_DUPLICATES. Using the default batch sampler will silently degrade or break training.
torch_dtype=bfloat16 must never be set at load time. Load in fp32 and use autocast (bf16/fp16) during training. See references/training_args.md.
save_steps must be a multiple of eval_steps when using load_best_model_at_end. Otherwise the trainer raises or silently never loads the best checkpoint.
- ModernBERT-family models default to
max_seq_length=8192. This is a VRAM trap — explicitly set max_seq_length to your actual needs. See references/base_model_selection.md.
datasets >= 4 rejects some script-based dataset loaders. Pin or work around per references/base_model_selection.md.
- Named evaluators format
metric_for_best_model as eval_{name}_{primary_metric}, not eval_{primary_metric}. Mismatching this key means the trainer never selects the best model.
- [SparseEncoder] high nDCG with collapsed sparsity is not a win. Always check
query_active_dims / corpus_active_dims on the verdict line.
- HF download URLs flood the agent's context unless you silence
httpx, httpcore, huggingface_hub, urllib3, filelock, fsspec to WARNING.
- Skim
references/troubleshooting.md section headings on every run, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that are cheaper to recognize before they fire.
Verification
After the training run completes, verify the following:
Verdict line emitted. Check the log tail for the single-line verdict:
Select-String -Path "logs\*.log" -Pattern "^VERDICT:"
Expected output:
logs/{RUN_NAME}.log:123:VERDICT: WIN | score=0.842 | baseline=0.710 | delta=+0.132
Baseline was captured before training. Confirm baseline_eval appears in the log before trainer.train() output:
Select-String -Path "logs\*.log" -Pattern "baseline_eval"
Best model loaded. If load_best_model_at_end=True, confirm the trainer logged loading the best checkpoint:
Select-String -Path "logs\*.log" -Pattern "best model"
Hub push attempted. Confirm push_to_hub was called (success or caught exception):
Select-String -Path "logs\*.log" -Pattern "push_to_hub|Upload|model pushed"
[SparseEncoder] sparsity logged. Confirm active dims appear on the verdict line:
Select-String -Path "logs\*.log" -Pattern "active_dims"
[CrossEncoder] early stopping fired (if peak was mid-training). Confirm callback activity:
Select-String -Path "logs\*.log" -Pattern "EarlyStopping|patience"
Experiments log updated:
Test-Path "logs\experiments.md"
Select-String -Path "logs\experiments.md" -Pattern $RUN_NAME
Related skills
references/training_args.md — deep dive on TrainingArguments knobs, precision, schedulers, HPO, resume, hub-push variants.
references/hardware_guide.md — VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors.
references/hf_jobs_execution.md — running on HF Jobs (ephemeral environment, in-trainer push).
references/prompts_and_instructions.md — prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic) and query: / passage: prefixes.
1---2name: train-sentence-transformers3description: Trains or fine-tunes sentence-transformers bi-encoders, CrossEncoder rerankers, and SparseEncoder/SPLADE using the bundled references and example scripts. Use when the user asks to train embeddings, rerankers, or SPLADE on custom data. Not for inference-only embedding calls, Hugging Face paper lookup, or model-card eval tables.4license: Apache-2.05---6
7# Train a sentence-transformers Model
8
9## Overview
10
11This skill trains or fine-tunes sentence-transformers models across three model classes:
12
13- **`SentenceTransformer`** (bi-encoder; dense or static embedding model) — for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal.
14- **`CrossEncoder`** (reranker; pair scoring) — for two-stage retrieval / pair classification.
15- **`SparseEncoder`** (SPLADE; sparse vectors over vocabulary) — for learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene).
16
17**This SKILL.md is a router, not a manual.** It tells you which references and example scripts to load for your task. The actual content — recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting — lives in `references/` and `scripts/`.
18
19**Do not synthesize a training script from this file alone.** Open the matching `train_<type>_example.py` in this skill's scripts folder and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, `force=True`, `seed`, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.
20
21## When to Use
22
23Use this skill when the user needs to:
24
25- Train or fine-tune a **dense embedding model** (retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal).
26- Train or fine-tune a **reranker / cross-encoder** for two-stage retrieval or pair classification.
27- Train or fine-tune a **SPLADE / sparse encoder** for learned-sparse retrieval or inverted-index backends.
28- Fine-tune an existing sentence-transformers checkpoint on a custom dataset.
29- Add LoRA, distillation, Matryoshka, multi-dataset, multilingual, or static-embedding variants.
30
31**Trigger keywords:** embedding model training, fine-tune sentence-transformers, train reranker, train cross-encoder, train SPLADE, sparse encoder training, retrieval model fine-tuning, sentence similarity model training, bi-encoder training, dense retrieval training.
32
33## Prerequisites
34
35```powershell
36pip install "sentence-transformers[train]>=5.0"
37# For multimodal [SentenceTransformer], add the relevant extra:
38# pip install "sentence-transformers[train,image]>=5.0"
39# pip install "sentence-transformers[train,audio]>=5.0"
40# pip install "sentence-transformers[train,video]>=5.0"
41pip install trackio # optional tracker; or wandb / tensorboard / mlflow
42hf auth login # or set HF_TOKEN with write scope (for Hub push)
43```
44
45GPU strongly recommended. CPU works only for demos and `[SentenceTransformer]` `StaticEmbedding`.
46
47## Procedure
48
49### Step 1 — Identify the model type
50
51| Tag | Class | What it does | When to pick |
52|---|---|---|---|
53| **[SentenceTransformer]** | `SentenceTransformer` (bi-encoder) | Maps each input to a fixed-dim dense vector | Retrieval, similarity, clustering, classification, paraphrase mining, dedup |
54| **[CrossEncoder]** | `CrossEncoder` (reranker) | Scores `(query, passage)` pairs jointly | Two-stage retrieval (rerank top-100 from bi-encoder), pair classification |
55| **[SparseEncoder]** | `SparseEncoder` (SPLADE) | Sparse vectors over the vocabulary | Learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene) |
56
57**Tiebreakers when the request is ambiguous:**
58
59- "embedding model" / "vector search" / "similarity" → **[SentenceTransformer]**
60- "rerank" / "ranker" / "two-stage" → **[CrossEncoder]**
61- "SPLADE" / "sparse" / "inverted index" → **[SparseEncoder]**
62- If still unclear, ask the user.
63
64### Step 2 — Load required reading (in full, before writing any code)
65
66**Do not triage by perceived relevance. Read every file listed for your type.**
67
68#### Per-type — always required
69
70**[SentenceTransformer]**
71
72- `references/losses_sentence_transformer.md` — loss-to-data-shape mapping; `BatchSamplers.NO_DUPLICATES` requirement for MNRL-family; `Cached*` ↔ `gradient_checkpointing` incompatibility.
73- `references/evaluators_sentence_transformer.md` — evaluator-to-task mapping; `metric_for_best_model` key construction (named vs unnamed); per-evaluator `primary_metric` values.
74- `references/model_architectures.md` — encoder vs decoder vs static vs Router pipelines; pooling rules (mean / cls / lasttoken); auto-mean-pooling behavior for fresh-start MLM bases.
75- `scripts/train_sentence_transformer_example.py` — production template; **copy this as your starting point**.
76
77**[CrossEncoder]**
78
79- `references/losses_cross_encoder.md` — pointwise / pairwise / listwise / distillation; `pos_weight` derivation; `activation_fn=Identity()` mandatory for non-BCE losses (silent eval-rank collapse otherwise).
80- `references/evaluators_cross_encoder.md` — `CrossEncoderRerankingEvaluator` recipe; named-evaluator key format `eval_{name}_{primary_metric}`.
81- `scripts/train_cross_encoder_example.py` — production template; **copy this as your starting point**.
82
83**[SparseEncoder]**
84
85- `references/losses_sparse_encoder.md` — `SpladeLoss` wrapper requirement; FLOPS regularizer weights; smoke-test active-dim ramp behavior.
86- `references/evaluators_sparse_encoder.md` — `SparseNanoBEIREvaluator` (English-only) and the in-domain alternative; `eval_{name}_{primary_metric}` key format.
87- `scripts/train_sparse_encoder_example.py` — production template; **copy this as your starting point**.
88
89#### Cross-cutting — always required (regardless of task)
90
91- `references/training_args.md` — `TrainingArguments` knobs, precision rules (load fp32 + autocast bf16/fp16; never `torch_dtype=bfloat16`), `warmup_steps` (float) vs deprecated `warmup_ratio`, `save_steps` must be a multiple of `eval_steps` for `load_best_model_at_end`, schedulers, HPO, tracker, resume, hub-push variants.
92- `references/dataset_formats.md` — column-matching rules (label name auto-detection; column-order-not-name); reshaping recipes; hard-negative mining options.
93- `references/base_model_selection.md` — discovery commands; per-type model namespaces; ModernBERT-family `max_seq_length=8192` trap; `datasets >= 4` script-loader rejection; non-English starting-point shortcuts.
94- `references/troubleshooting.md` — symptom-indexed failure recipes. **Skim the section headings on every run, even a healthy one.** The "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.
95
96#### Cross-cutting — load when applicable
97
98- `references/hardware_guide.md` — **Required for >24GB models, multi-GPU, or HF Jobs runs.** VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors.
99- `references/hf_jobs_execution.md` — **Required when running on HF Jobs.**
100- `references/prompts_and_instructions.md` — **Required when using prompt-tuned bases** (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding `query: ` / `passage: ` style prefixes.
101
102#### Variant scripts (open when the task matches)
103
104- **[SentenceTransformer]:**
105 - `scripts/train_sentence_transformer_matryoshka_example.py`
106 - `scripts/train_sentence_transformer_multi_dataset_example.py`
107 - `scripts/train_sentence_transformer_with_lora_example.py`
108 - `scripts/train_sentence_transformer_distillation_example.py`
109 - `scripts/train_sentence_transformer_make_multilingual_example.py`
110 - `scripts/train_sentence_transformer_static_embedding_example.py`
111- **[CrossEncoder]:**
112 - `scripts/train_cross_encoder_distillation_example.py`
113 - `scripts/train_cross_encoder_listwise_example.py`
114- **[SparseEncoder]:**
115 - `scripts/train_sparse_encoder_distillation_example.py`
116- **Hard-negative mining CLI:** `scripts/mine_hard_negatives.py`
117
118### Step 3 — Copy the production template
119
120Open the matching `train_<type>_example.py` in this skill's scripts folder (`scripts/train_sentence_transformer_example.py`, `scripts/train_cross_encoder_example.py`, or `scripts/train_sparse_encoder_example.py`) and copy it as your starting point. Do not write a training script from scratch or from a synthesized snippet.
121
122### Step 4 — Replace placeholders with the user's task
123
124Replace `MODEL_NAME`, `DATASET_NAME`, `RUN_NAME`, the loss, and the evaluator with the user's task.
125
126- Cross-check loss/data-shape match against the matching losses file (`references/losses_sentence_transformer.md`, `references/losses_cross_encoder.md`, or `references/losses_sparse_encoder.md`).
127- Cross-check the `metric_for_best_model` key against the matching evaluators file (`references/evaluators_sentence_transformer.md`, `references/evaluators_cross_encoder.md`, or `references/evaluators_sparse_encoder.md`) (named evaluators format the key as `eval_{name}_{primary_metric}`).
128
129### Step 5 — Smoke-test before any long run
130
131Set `max_steps=1` with a tiny dataset slice. The production templates show one common pattern using a `SMOKE_TEST` environment variable.
132
133```powershell
134# Example smoke-test invocation (pattern from the production template)
135$env:SMOKE_TEST = "1"
136python scripts/train_sentence_transformer_example.py
137```
138
139### Step 6 — Run the full training
140
141```powershell
142Remove-Item Env:SMOKE_TEST # clear smoke-test flag
143python scripts/train_sentence_transformer_example.py
144```
145
146### Step 7 — After the run
147
1481. Append results to `logs/experiments.md`.
1492. Propose iteration if the verdict is weak or marginal.
150
151## Constraints the produced script must satisfy
152
153These are non-negotiable contracts. Implementation lives in the production templates and references — do not reinvent.
154
1551. **Capture the pre-training evaluator score** as `baseline_eval` **before** `trainer.train()`.
1562. **Emit a single end-of-run verdict line:**
157 ```
158 VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=...
159 ```
160 A monitor scrapes for this line.
1613. **Silence noisy loggers** — set `httpx`, `httpcore`, `huggingface_hub`, `urllib3`, `filelock`, `fsspec` to `WARNING`. Otherwise HF download URLs flood the agent's context.
1624. **Tee logs** to `logs/{RUN_NAME}.log`.
1635. **End with `model.push_to_hub(...)` wrapped in `try/except`.**
1646. **Smoke-test before any long run** (`max_steps=1` + tiny dataset slice). The production templates show one common pattern (`SMOKE_TEST` env var).
1657. **[CrossEncoder]** Include `EarlyStoppingCallback(patience>=3)` — CE rerankers often peak mid-training and regress.
1668. **[SparseEncoder]** Log `query_active_dims` / `corpus_active_dims` on the verdict line; high nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. `..._query_active_dims`); use suffix matching to pluck them — see the SPARSE production template for the exact pattern.
167
168## Defaults
169
170Override only if the user specifies otherwise:
171
172- **Local execution.** Pitch HF Jobs only if local hardware can't fit the job.
173- **Single run.** After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in `references/training_args.md` (Experimentation section).
174- **Public Hub push at end-of-run, wrapped in try-except.** On HF Jobs (ephemeral env) ALSO enable in-trainer push (`push_to_hub=True` + `hub_strategy="every_save"`); details in `references/hf_jobs_execution.md`.
175
176## Pitfalls
177
178- **Do not synthesize a training script from this file alone.** Always start from the matching `train_<type>_example.py` in this skill's scripts folder. Prior agent runs have repeatedly missed load-bearing scaffolding (autocast helper, model-card class, logger silencing, `force=True`, `seed`, TF32, version-compatible imports, named-evaluator metric handling) when rolling their own.
179- **[CrossEncoder] `activation_fn=Identity()` is mandatory for non-BCE losses.** Using the default sigmoid activation with a non-BCE loss causes silent eval-rank collapse.
180- **[SentenceTransformer] `Cached*` losses are incompatible with `gradient_checkpointing`.** Do not combine them.
181- **[SentenceTransformer] MNRL-family losses require `BatchSamplers.NO_DUPLICATES`.** Using the default batch sampler will silently degrade or break training.
182- **`torch_dtype=bfloat16` must never be set at load time.** Load in fp32 and use autocast (bf16/fp16) during training. See `references/training_args.md`.
183- **`save_steps` must be a multiple of `eval_steps`** when using `load_best_model_at_end`. Otherwise the trainer raises or silently never loads the best checkpoint.
184- **ModernBERT-family models default to `max_seq_length=8192`.** This is a VRAM trap — explicitly set `max_seq_length` to your actual needs. See `references/base_model_selection.md`.
185- **`datasets >= 4` rejects some script-based dataset loaders.** Pin or work around per `references/base_model_selection.md`.
186- **Named evaluators format `metric_for_best_model` as `eval_{name}_{primary_metric}`**, not `eval_{primary_metric}`. Mismatching this key means the trainer never selects the best model.
187- **[SparseEncoder] high nDCG with collapsed sparsity is not a win.** Always check `query_active_dims` / `corpus_active_dims` on the verdict line.
188- **HF download URLs flood the agent's context** unless you silence `httpx`, `httpcore`, `huggingface_hub`, `urllib3`, `filelock`, `fsspec` to WARNING.
189- **Skim `references/troubleshooting.md` section headings on every run**, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that are cheaper to recognize before they fire.
190
191## Verification
192
193After the training run completes, verify the following:
194
1951. **Verdict line emitted.** Check the log tail for the single-line verdict:
196 ```powershell
197 Select-String -Path "logs\*.log" -Pattern "^VERDICT:"
198 ```
199 Expected output:
200 ```
201 logs/{RUN_NAME}.log:123:VERDICT: WIN | score=0.842 | baseline=0.710 | delta=+0.132
202 ```
203
2042. **Baseline was captured before training.** Confirm `baseline_eval` appears in the log before `trainer.train()` output:
205 ```powershell
206 Select-String -Path "logs\*.log" -Pattern "baseline_eval"
207 ```
208
2093. **Best model loaded.** If `load_best_model_at_end=True`, confirm the trainer logged loading the best checkpoint:
210 ```powershell
211 Select-String -Path "logs\*.log" -Pattern "best model"
212 ```
213
2144. **Hub push attempted.** Confirm `push_to_hub` was called (success or caught exception):
215 ```powershell
216 Select-String -Path "logs\*.log" -Pattern "push_to_hub|Upload|model pushed"
217 ```
218
2195. **[SparseEncoder] sparsity logged.** Confirm active dims appear on the verdict line:
220 ```powershell
221 Select-String -Path "logs\*.log" -Pattern "active_dims"
222 ```
223
2246. **[CrossEncoder] early stopping fired (if peak was mid-training).** Confirm callback activity:
225 ```powershell
226 Select-String -Path "logs\*.log" -Pattern "EarlyStopping|patience"
227 ```
228
2297. **Experiments log updated:**
230 ```powershell
231 Test-Path "logs\experiments.md"
232 Select-String -Path "logs\experiments.md" -Pattern $RUN_NAME
233 ```
234
235## Related skills
236
237- `references/training_args.md` — deep dive on `TrainingArguments` knobs, precision, schedulers, HPO, resume, hub-push variants.
238- `references/hardware_guide.md` — VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors.
239- `references/hf_jobs_execution.md` — running on HF Jobs (ephemeral environment, in-trainer push).
240- `references/prompts_and_instructions.md` — prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic) and `query: ` / `passage: ` prefixes.