DeePMD-kit Training
Use this skill to guide DeePMD-kit model training without loading every model-specific recipe up front.
The workflow is intentionally progressive:
- Understand the user's data, target accuracy, compute budget, and deployment backend.
- Choose an appropriate model family.
- Read only the reference file for the selected model under
models/.
- Generate or edit
input.json, run training, monitor, freeze, and test.
Progressive disclosure protocol
Do not start by reading every model document. First classify the request:
- If the user already named a model, read only that model reference.
- If the user asks for a recommendation, collect the decision inputs below, choose a model, then read only the selected reference.
- If model-specific parameters are not needed yet, stay in this top-level workflow.
Available model references:
| Model reference |
Read when |
models/se-e2-a.md |
The user wants a classical DeepPot-SE baseline, broad compatibility, or a smaller/established production model. |
models/dpa3.md |
The user wants a high-accuracy DPA3/LAM workflow, large/diverse datasets, dynamic neighbor selection, or pretrained DPA3-style training. |
Model selection
Ask only for missing information that changes the choice. Prefer reasonable defaults when the answer is obvious from context.
Key inputs:
- Data format and size: deepmd/npy, deepmd/hdf5, mixed type, number of systems/frames/elements.
- Target: quick baseline, production accuracy, large atomic model, transfer/fine-tuning, or deployment in MD.
- Compute: CPU/GPU, available memory, single-node vs. distributed training.
- Backend/deployment: PyTorch/TensorFlow/JAX/Paddle training; LAMMPS, Python inference, or other downstream use.
- Labels: energy/force only or also virial/stress.
- System diversity: single chemistry/phase vs. diverse multi-domain datasets.
Recommended defaults:
- Choose se_e2_a for a robust baseline, small to medium systems, compatibility-focused workflows, or when compute is limited.
- Choose DPA3 for high accuracy on diverse datasets, LAM-style training, or when the user explicitly asks for DPA3, DPA-3, LiGS, dynamic neighbor selection, or pretrained DPA3 variants.
Common workflow
1. Confirm environment
dp --version
For PyTorch training, use dp --pt ...; for TensorFlow, use dp ...; for other backends, confirm the installed backend first.
2. Confirm training data
Training data should be in DeePMD format, typically deepmd/npy or deepmd/hdf5. If the user has raw electronic-structure outputs, convert them first with dpdata before writing the training input.
Minimum information needed to build input.json:
type_map
- training system paths
- validation system paths
- whether virial labels are present and should be trained
- target number of steps or accuracy/time budget
- model choice
3. Read the selected model reference
After selecting a model, read the corresponding file under models/ and apply its model-specific configuration, hyperparameters, and caveats.
4. Train
dp --pt train input.json
Use the backend-specific command if not using PyTorch.
Restart from a checkpoint when needed:
dp --pt train input.json --restart model.ckpt.pt
5. Monitor
Training progress is usually written to lcurve.out. Check for:
- decreasing validation RMSE
- NaN or exploding losses
- train/validation divergence
- learning-rate schedule behaving as expected
6. Freeze and test
dp --pt freeze -o model.pth
dp --pt test -m model.pth -s /path/to/test_system -n 30
Adjust the backend flags and output extension for non-PyTorch models.
Agent checklist
References
1---2name: deepmd-train3description: Train DeePMD-kit models with progressive disclosure. Use when the user wants to train a DeePMD-kit potential, prepare an input.json, choose between model families such as se_e2_a/DeepPot-SE and DPA3, run `dp train`, monitor learning curves, freeze checkpoints, or test trained models. Start with model selection and read only the selected model reference under `models/` when model-specific configuration is needed.4license: LGPL-3.0-or-later5---6
7# DeePMD-kit Training
8
9Use this skill to guide DeePMD-kit model training without loading every model-specific recipe up front.
10The workflow is intentionally progressive:
11
121. Understand the user's data, target accuracy, compute budget, and deployment backend.
131. Choose an appropriate model family.
141. Read only the reference file for the selected model under [`models/`](models/).
151. Generate or edit `input.json`, run training, monitor, freeze, and test.
16
17## Progressive disclosure protocol
18
19Do not start by reading every model document. First classify the request:
20
21- If the user already named a model, read only that model reference.
22- If the user asks for a recommendation, collect the decision inputs below, choose a model, then read only the selected reference.
23- If model-specific parameters are not needed yet, stay in this top-level workflow.
24
25Available model references:
26
27| Model reference | Read when |
28| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
29| [`models/se-e2-a.md`](models/se-e2-a.md) | The user wants a classical DeepPot-SE baseline, broad compatibility, or a smaller/established production model. |
30| [`models/dpa3.md`](models/dpa3.md) | The user wants a high-accuracy DPA3/LAM workflow, large/diverse datasets, dynamic neighbor selection, or pretrained DPA3-style training. |
31
32## Model selection
33
34Ask only for missing information that changes the choice. Prefer reasonable defaults when the answer is obvious from context.
35
36Key inputs:
37
38- Data format and size: deepmd/npy, deepmd/hdf5, mixed type, number of systems/frames/elements.
39- Target: quick baseline, production accuracy, large atomic model, transfer/fine-tuning, or deployment in MD.
40- Compute: CPU/GPU, available memory, single-node vs. distributed training.
41- Backend/deployment: PyTorch/TensorFlow/JAX/Paddle training; LAMMPS, Python inference, or other downstream use.
42- Labels: energy/force only or also virial/stress.
43- System diversity: single chemistry/phase vs. diverse multi-domain datasets.
44
45Recommended defaults:
46
47- Choose **se_e2_a** for a robust baseline, small to medium systems, compatibility-focused workflows, or when compute is limited.
48- Choose **DPA3** for high accuracy on diverse datasets, LAM-style training, or when the user explicitly asks for DPA3, DPA-3, LiGS, dynamic neighbor selection, or pretrained DPA3 variants.
49
50## Common workflow
51
52### 1. Confirm environment
53
54```bash
55dp --version
56```
57
58For PyTorch training, use `dp --pt ...`; for TensorFlow, use `dp ...`; for other backends, confirm the installed backend first.
59
60### 2. Confirm training data
61
62Training data should be in DeePMD format, typically deepmd/npy or deepmd/hdf5. If the user has raw electronic-structure outputs, convert them first with dpdata before writing the training input.
63
64Minimum information needed to build `input.json`:
65
66- `type_map`
67- training system paths
68- validation system paths
69- whether virial labels are present and should be trained
70- target number of steps or accuracy/time budget
71- model choice
72
73### 3. Read the selected model reference
74
75After selecting a model, read the corresponding file under [`models/`](models/) and apply its model-specific configuration, hyperparameters, and caveats.
76
77### 4. Train
78
79```bash
80dp --pt train input.json
81```
82
83Use the backend-specific command if not using PyTorch.
84
85Restart from a checkpoint when needed:
86
87```bash
88dp --pt train input.json --restart model.ckpt.pt
89```
90
91### 5. Monitor
92
93Training progress is usually written to `lcurve.out`. Check for:
94
95- decreasing validation RMSE
96- NaN or exploding losses
97- train/validation divergence
98- learning-rate schedule behaving as expected
99
100### 6. Freeze and test
101
102```bash
103dp --pt freeze -o model.pth
104dp --pt test -m model.pth -s /path/to/test_system -n 30
105```
106
107Adjust the backend flags and output extension for non-PyTorch models.
108
109## Agent checklist
110
111- [ ] Model was selected before reading model-specific details.
112- [ ] Only the selected model reference was loaded.
113- [ ] Training/validation data paths exist or are clearly marked as placeholders.
114- [ ] `type_map` matches the data and model/pretrained checkpoint.
115- [ ] Virial loss is enabled only when virial labels are available and desired.
116- [ ] Backend command matches the selected model and installed DeePMD-kit environment.
117- [ ] The generated `input.json` is valid JSON.
118- [ ] Training was monitored via `lcurve.out` or equivalent logs.
119- [ ] Final model was frozen and tested when requested.
120
121## References
122
123- [Training documentation](https://docs.deepmodeling.com/projects/deepmd/en/latest/train/training.html)
124- [Training input documentation](https://docs.deepmodeling.com/projects/deepmd/en/latest/train/train-input.html)
125- [Model documentation](https://docs.deepmodeling.com/projects/deepmd/en/latest/model/index.html)