Tiny Model Creator
Create a tiny random model for the requested architecture without loading the original model weights. The artifact must preserve the real architecture and execute the requested task path.
If repairing a tiny model after a validation or benchmark failure, start from the exact failing local artifact, creation script, traceback, and task reproducer. Repair that artifact and rerun the same command; do not generate a different model and use its success as evidence for the failed one.
Step 1 — Inspect the original model
Download or load configuration and lightweight code/processor assets only. Record:
model_type,architectures,auto_map, and Transformers metadata;- nested text, vision, audio, and projector configurations;
- hidden-size, head-count, grouping, rotary, cache, vocabulary, and special token invariants;
- tokenizer, processor, chat-template, and remote-code files required to load;
- the correct model or pipeline class and execution interface for the task.
Determine and record the Transformers version or version range supported by the original model, especially for trust-remote-code architectures. Create and validate the tiny model using a compatible version; do not silently substitute another model implementation because the active Transformers version lacks the architecture.
Inspect precision fields under every relevant configuration level. Remote
models may use dtype, torch_dtype, or both, including separate values in
vision, text, audio, or projector sub-configs.
Step 2 — Write a reusable constructor
Create create_tiny_model.py in the designated working directory. It must:
- Load the original configuration without loading original weights.
- Reduce layers, hidden dimensions, intermediate sizes, vocabulary, image resolution or patch counts, experts, and similar scale parameters.
- Preserve divisibility and coupling invariants such as head dimensions, grouped-query attention, projector sizes, vision/text bridges, cache dimensions, and special-token IDs.
- Instantiate random weights through the real architecture class.
- Save every required config, tokenizer, processor, chat template, generation config, and remote-code asset.
- Reuse a completed cached output directory on repeated calls.
Define maximum parameter-count and model-memory budgets, then verify both after construction so reducing layer count cannot be offset by widening other dimensions. Estimate weight memory from each parameter's element count and element size, and reject a candidate that exceeds either budget.
Do not reuse a cache merely because config.json and a weight file exist.
Before returning it, validate a cache-format/version marker and all critical
configuration invariants, including architecture identity, dimensions,
special tokens, processor assets, and nested precision fields. Rebuild the
cache when the generator logic or required invariants change.
Keep construction logic easy to adapt into
_create_tiny_<model_type>_model() in tests/openvino/utils_tests.py.
Step 3 — Verify architecture identity
Compare the original and tiny configurations. Preserve:
model_typeandarchitectures;- task-relevant sub-config types and component roles;
- cache/stateful and position-ID contracts;
- VLM processor classes, placeholder/image tokens, and merge contracts;
- MoE/expert topology, even when expert counts are reduced.
When deliberately forcing a test model to float32, update and verify every
effective precision field used by the remote configuration (dtype and/or
torch_dtype, including nested sub-configs). Reload the saved model and check
its actual parameter dtypes; editing an ignored config key is not sufficient.
Never rename the model type, substitute a nearby architecture, or remove a component merely to make export pass.
If the tiny model's model_type, architectures, or task-relevant component
identity differs from the original model, stop and report the mismatch. A tiny
model that executes successfully through another architecture is not a valid
fixture.
Step 4 — Validate the real task path
Reload the saved directory through its documented Transformers or pipeline API and execute the requested task.
Follow the task-specific tiny-model validation and output-validity instructions
supplied for <task>.
If task execution fails, repair the violated configuration invariant, recreate the model, and rerun it. Loading, saving, or a forward pass alone is not success.
The final validation evidence must load the exact output directory returned by the creator, execute the requested task, and include the command and output. Do not validate one directory and return a different cached or previously generated artifact.
Rules
- Do not upload the tiny model to Hugging Face.
- Do not edit installed packages or the virtual environment.
- Do not modify system files or system-wide package installations.
- Use a deterministic seed where supported.
- Avoid original weight downloads and large generated artifacts.
- Never commit machine-specific absolute paths.
- Cache repository-test fixtures so test collection does not rebuild them unnecessarily.
Report
Report the output directory, script path, configuration comparison, parameter count, exact task execution command, output, and any dependency or remote-code constraints.