# Ml Experiment Reproducibility

> Pin seeds/config/dataset versions and provide a deterministic rerun path.

- Skill: `authenticfake/ml-experiment-reproducibility` (Agent Skill)
- Install (CLI): `npx skillmds@latest add authenticfake/ml-experiment-reproducibility`
- Raw SKILL.md: https://api.skillmd.com/api/skills/authenticfake/ml-experiment-reproducibility/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- Author: authenticfake (https://skillmd.com/u/authenticfake)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/authenticfake/ml-experiment-reproducibility

---


# Skill: ML Experiment Reproducibility

## Intent

Ensure ML, data science, model evaluation, and experiment requirements are reproducible, measurable, and traceable.

This skill separates product code from experiments and prevents unverifiable model-quality claims.

## Use when

Use this skill when a REQ touches ML models, datasets, training, fine-tuning, feature engineering, evaluation metrics, model comparison, notebooks, pipelines, data quality, experiment tracking, or batch inference.

## Do not use when

Do not use this skill for generic LLM prompt/RAG work unless the REQ includes ML datasets, model metrics, training, offline evaluation, or experiment comparison.

## Signals

- The REQ mentions ML, model training, fine-tuning, dataset, feature, label, metric, accuracy, precision, recall, F1, ROC, drift, experiment, notebook, inference, pipeline, validation split, baseline, or model registry.
- Acceptance criteria include measurable model quality.
- Generated files include notebooks, data loaders, evaluation scripts, model wrappers, or dataset fixtures.

## Required behavior

- Define datasets, fixtures, or sample data boundaries explicitly.
- Define metrics and thresholds before implementation.
- Keep training/experiment code separate from production inference code when practical.
- Make evaluation commands reproducible.
- Record assumptions about data availability, privacy, sampling, and labels.
- Include deterministic smoke tests for data loading and metric computation.
- Avoid requiring large datasets or GPUs for local blocking checks unless explicitly required.

## Forbidden behavior

- Do not claim model improvement without baseline and metric evidence.
- Do not hardcode local absolute dataset paths.
- Do not put private or production data into generated fixtures.
- Do not make local tests depend on GPU availability unless the REQ explicitly requires it.
- Do not mix notebook-only exploration with production code without extraction into testable modules.
- Do not hide data quality gaps behind successful code execution.

## Evidence required

- Evaluation command or script with documented inputs.
- Metric definitions and thresholds.
- Small deterministic fixture or synthetic dataset when real data is unavailable.
- Tests for metric computation, data validation, or inference wrapper behavior.
- HOWTO explaining local smoke evaluation and optional full evaluation.
- Notes describing baseline, limitations, and data assumptions.

## Repair guidance

- If metrics are missing, add a minimal evaluation script and documented thresholds.
- If data paths are hardcoded, move them to configuration.
- If only a notebook exists, extract reusable logic into a module and test it.
- If full data is unavailable, create a small representative fixture and mark full evaluation as external/non-blocking.
- If model quality is unknown, downgrade claims and document the required next evaluation.

## Gate implications

Gate should block promotion when:
- Model quality is part of acceptance criteria but no metric evidence exists.
- Evaluation cannot be reproduced.
- Tests require unavailable private data or hardware.
- Generated code uses hardcoded local data paths.
- Production inference behavior is untested.

Gate may allow non-blocking warnings when:
- Full-scale evaluation requires external infrastructure but local smoke evaluation passes.
- Drift monitoring is documented as future work outside the current REQ scope.

## Examples

- A classifier REQ includes fixture data, metric computation tests, baseline threshold, and HOWTO eval commands.
- A batch inference REQ separates loader, predictor, and writer with deterministic unit tests.
- A data quality REQ defines validation rules and failure examples.

## Non-examples

- A notebook that prints a high accuracy score without data or command reproducibility.
- A model wrapper that requires a production dataset to import.
- A fine-tuning REQ with no baseline, no metrics, and no eval script.
---

# CLike Promotable KIT Enforcement Layer

## Purpose

This layer makes the skill operational for CLike `/kit` generation.

The goal is not to produce plausible code. The goal is to produce candidate artifacts that can be evaluated, repaired, and promoted through EVAL and GATE with minimal human rework.

## Promotable Code Obligations

When this skill is selected for a REQ, the KIT must:

- respect `main_module_boundary`;
- respect `functional_scope` and `technical_scope`;
- generate the smallest complete implementation slice;
- prefer repository-native conventions over invented abstractions;
- produce source files only under the target KIT source root;
- produce tests only under the target KIT test root;
- keep canonical `src/`, `test/`, and `tests/` read-only during candidate generation;
- document any intentional limitation instead of pretending completeness;
- avoid broad rewrites unless explicitly required by the REQ.

## Required Candidate Artifacts

The KIT should produce or update:

```text
runs/kit/<REQ-ID>/src/
runs/kit/<REQ-ID>/test/
runs/kit/<REQ-ID>/ci/LTC.json
runs/kit/<REQ-ID>/ci/HOWTO.md
runs/kit/<REQ-ID>/docs/KIT_<REQ-ID>.md
```

If the REQ is documentation-only or policy-only, the KIT must explicitly state why source/test artifacts are not required.

## Code Shape Expectations

Generated code should favor:

- explicit boundaries;
- dependency injection or constructor/function injection where practical;
- small cohesive modules;
- deterministic local behavior;
- typed schemas/contracts when the stack supports them;
- error paths that are visible and testable;
- safe defaults;
- clear adapter seams for external systems.

Generated code must avoid:

- hidden global state;
- hardcoded environment assumptions;
- silent fallbacks;
- fake success;
- speculative framework layers;
- broad unrelated refactors;
- acceptance criteria implemented only in prose.

## Test Expectations

Tests must map to acceptance criteria.

Prefer:

- deterministic unit tests;
- contract tests around adapters and payloads;
- failure-path tests;
- local fake/simulator tests for external dependencies;
- smoke checks only when deeper tests are not possible.

Avoid:

- placeholder tests;
- tests that only import modules when behavior is required;
- tests that require production credentials;
- network-dependent blocking tests unless explicitly scoped.

## LTC Expectations

`ci/LTC.json` must be valid JSON and include enough information for EvalRunner or a local agent to execute checks.

It should include:

- target `req_id`;
- lane/runtime profile when known;
- blocking local commands;
- optional external commands;
- report paths when available;
- environment-blocked status for unavailable infrastructure;
- gate-relevant policy hints.

## HOWTO Expectations

`ci/HOWTO.md` must be clear enough for a developer to run without guessing.

It should include:

- where to run commands from;
- prerequisites;
- local commands;
- expected result;
- troubleshooting;
- required environment variables;
- optional external validation steps;
- limitations and non-goals.

## Gate Impact

GATE should BLOCK promotion when this selected skill is materially violated.

Blocking examples:

- source is not mapped to the target REQ;
- acceptance-critical behavior has no test or executable evidence;
- LTC/HOWTO are missing for runnable code;
- production services or credentials are required for local blocking checks;
- selected capability obligations are ignored;
- generated files modify forbidden canonical roots;
- code claims completeness without evidence.

GATE may WARN when:

- optional external validation is not available but a deterministic local contract check exists;
- documentation is thin but executable evidence is complete;
- future hardening is correctly documented as out of scope.

