# PyTorch Training Configuration and Evaluation

> Configure PyTorch training scripts with specific evaluation metrics (Precision, Recall, F1), tunable hyperparameters (batch size, warmup, optimizer type, weight decay, attention dropout), and a custom GELU activation function.

- Skill: `ecnu-icalk/pytorch-training-configuration-and-evaluation` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ecnu-icalk/pytorch-training-configuration-and-evaluation`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ecnu-icalk/pytorch-training-configuration-and-evaluation/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: ECNU-ICALK (https://skillmd.com/u/ecnu-icalk)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/ecnu-icalk/pytorch-training-configuration-and-evaluation

---


# PyTorch Training Configuration and Evaluation

Configure PyTorch training scripts with specific evaluation metrics (Precision, Recall, F1), tunable hyperparameters (batch size, warmup, optimizer type, weight decay, attention dropout), and a custom GELU activation function.

## Prompt

# Role & Objective
Configure PyTorch training scripts to include specific evaluation metrics, tunable hyperparameters, and a custom GELU activation function.

# Operational Rules & Constraints
1. **Evaluation Metrics**: Modify the evaluation function to compute Precision, Recall, and F1 score using `sklearn.metrics` with `average='macro'`.
2. **Hyperparameters**: Define and utilize the following variables for tuning:
   - `batch_size`
   - `warmup_steps`
   - `optimizer_type` (e.g., "AdamW", "SGD")
   - `weight_decay`
   - `attention_dropout_rate`
3. **Activation Function**: Implement the `gelu_new` activation function using the formula: `0.5 * x * (1 + torch.tanh(torch.sqrt(2 / torch.pi) * (x + 0.044715 * torch.pow(x, 3))))`.
4. **Model Configuration**: Apply `attention_dropout_rate` to the `nn.TransformerEncoderLayer` and use `optimizer_type` to configure the optimizer (AdamW or SGD).

# Anti-Patterns
- Do not use the default accuracy metric alone; always include Precision, Recall, and F1.
- Do not hardcode hyperparameters; use the specified variables.

## Triggers

- modify evaluation function
- add hyperparameters
- compute F1 score
- add gelu_new
- tune batch size

