Generate C/C++/CUDA Code from an AI Model
Generate deployable C/C++ or CUDA code from an AI model using MATLAB Coder or
GPU Coder. The workflow follows a common pattern regardless of model framework:
load, inspect, write entry-point, generate MEX, verify, then generate production code.
When to Use
- User wants to generate C/C++/CUDA code from an AI model (PyTorch, LiteRT)
- User has a model file (.pt2, .tflite) and wants to load it into MATLAB
- User wants MEX acceleration for an AI model
- User wants to generate CUDA code or GPU-accelerated MEX from an AI model
- User wants to deploy an AI model to hardware
- User wants to use a PyTorch or LiteRT model in Simulink (simulation or code generation)
- User wants to verify AI model numerics between the source framework and MATLAB
When NOT to Use
- General MATLAB Coder usage (codegen syntax, config tuning, writing codegen-ready code)
- Editable dlnetwork for Deep Learning Toolbox workflows (quantization, compression, transfer learning) — use
importNetworkFromPyTorch (PyTorch), importNetworkFromTensorFlow (SavedModel), or importNetworkFromKeras (.keras/.h5) which return a dlnetwork. For deployment of an editable dlnetwork with model compression (INT8 quantization via dlquantizer, pruning, projection) or exportNetworkToSimulink workflows — use matlab-deploy-embedded-ai (Pattern 1).
- Training or fine-tuning — this skill is for inference code generation only
Supported Frameworks
| Framework |
Model format |
Load function |
Status |
| PyTorch |
.pt2 |
loadPyTorchExportedProgram |
Supported (R2026a+) |
| LiteRT / TFLite |
.tflite |
loadLiteRTModel |
Supported (R2026a+) |
For PyTorch-specific details (API routing, entry-point pattern, export workflow,
data layout, common mistakes): see references/pytorch-workflow.md.
For LiteRT-specific details (API routing, entry-point pattern, variable-size inputs,
Simulink integration, conventions): see references/litert-workflow.md.
For converting TensorFlow/Keras/.h5 to .tflite: see
references/tensorflow-to-litert-conversion.md.
Generic Workflow
The code generation workflow follows the same steps for any framework:
1. Load and Inspect
Load the model and check its input/output specifications to determine expected
shapes and types.
2. Write Entry-Point Function
Create a codegen-compatible entry-point function that:
- Loads the model from a file path
- Runs inference on an input
- Returns the output
The model file path must be wrapped with coder.Constant so it's known at
compile time.
3. Verify Numerics
Compare MATLAB inference output against the source framework to confirm correct
loading. Use the same input data in both environments and compare with tolerance.
4. Generate MEX (First!)
Always generate MEX before lib/exe to verify on the host machine:
CPU MEX:
cfg = coder.config("mex");
codegen -config cfg -args {coder.Constant("model_file"), input} entryPoint
CUDA MEX (GPU acceleration):
cfg = coder.gpuConfig("mex");
codegen -config cfg -args {coder.Constant("model_file"), input} entryPoint
For CPU MEX SIMD acceleration (SIMDAcceleration = 'Full' for AVX2 on
Intel/AMD), see the matlab-generate-code skill. For the DNN-
inference-specific MEX AVX2 ceiling, see references/dnn-codegen-options.md.
5. Verify MEX Output
Compare MEX output against MATLAB reference using matlab.unittest with
tolerance:
refOut = entryPoint("model_file", input);
mexOut = entryPoint_mex("model_file", input);
testCase = matlab.unittest.TestCase.forInteractiveUse;
testCase.verifyThat(mexOut, matlab.unittest.constraints.IsEqualTo(refOut, ...
'Within', matlab.unittest.constraints.AbsoluteTolerance(single(1e-5))));
6. Generate Library/Executable
Once MEX is verified, generate production code:
cfgLib = coder.config("lib");
cfgLib.TargetLang = "C++"; % set to "C++" for C++ output; default is "C"
codegen -config cfgLib -args {coder.Constant("model_file"), input} entryPoint
For DLL: coder.config("dll"). For executable: coder.config("exe").
CUDA variants: Replace coder.config with coder.gpuConfig.
Performance tuning:
- Generic knobs (SIMD instruction sets, reduction-loop vectorization,
multithreaded loops): see the
matlab-generate-code skill.
- MATLAB Coder ↔ Simulink Coder property naming duality and
slbuild
set_param patterns: see the matlab-deploy-embedded-code skill.
- DNN-inference-specific knobs (
DLTargetLibrary / DeepLearningConfig to
disable third-party DL libraries, LargeConstantGeneration to serialize
weights to data files): see references/dnn-codegen-options.md.
7. Use in Simulink
For Simulink integration, use the dedicated PyTorch ExportedProgram block from
dlosslib — set ModelFilePath to the .pt2 file and it auto-detects
input/output shapes. No entry-point function or coder.Constant needed.
Pre/post-processing can be done with Simulink blocks around the dedicated block.
If you need everything in a single block, use a MATLAB Function block with
loadPyTorchExportedProgram + invoke (same pattern as the entry-point, but
the model path is a string literal — no coder.Constant).
Both paths support slbuild code generation (requires fixed-step solver + ERT
or GRT target). See references/simulink-workflow.md for full details.
8. Deploy to Hardware (Optional — requires Embedded Coder)
For embedded deployment, use the same entry-point function with an Embedded Coder
configuration. See the matlab-deploy-embedded-code skill for ERT config,
hardware settings, PIL/SIL verification, and target-specific options.
Ask the user to install the skill if it is not installed
Key Functions
| Function |
Purpose |
Package |
Since |
coder.Constant |
Make argument a compile-time constant |
MATLAB Coder |
R2011a |
coder.gpuConfig |
Create GPU (CUDA) code generation config |
GPU Coder |
R2017b |
codegen |
Generate code |
MATLAB Coder |
R2011a |
loadPyTorchExportedProgram |
Load .pt2 into MATLAB |
MATLAB Coder Support Package for PyTorch and LiteRT Models |
R2026a |
loadLiteRTModel |
Load .tflite into MATLAB |
MATLAB Coder Support Package for PyTorch and LiteRT Models |
R2026a |
Conventions
- Always check the model's input specifications for correct input shape and type
- Always generate MEX first, verify, then proceed to lib/exe
- Always use
coder.Constant for the model file path argument
- Input data is typically single-precision (check model input specs to confirm)
- Do NOT use
importNetworkFromPyTorch, importNetworkFromTensorFlow, or importNetworkFromKeras in this skill — they return a dlnetwork on a different path. If the user needs a dlnetwork for quantization, projection, pruning, or exportNetworkToSimulink before code generation, route to matlab-deploy-embedded-ai instead
References
references/pytorch-workflow.md — Full PyTorch-specific workflow: API routing,
entry-point pattern, export guidance, common mistakes, and conventions. Consult
for any PyTorch/.pt2 model code generation task. Links to deeper PyTorch
references (API signatures, data layout, numeric verification, supported models).
references/export-pytorch-models.md — Exporting an eager-mode PyTorch model to
.pt2 with torch.export (upstream of loading). Consult when the user has a
PyTorch model but no .pt2 file yet, or hits torch.export SerializeError /
kwarg-mismatch errors. Links to pytorch-export-patterns.md (per-source
templates) and pytorch-export-gotchas.md (torch 2.11 serialization fixes).
references/simulink-workflow.md — Simulink integration: dedicated PyTorch
ExportedProgram block (Path A) vs MATLAB Function block (Path B), block mask
parameters, code generation config, and key differences from command-line codegen.
references/litert-workflow.md — Full LiteRT-specific workflow: API routing
(loadLiteRTModel → inputSpecifications → invoke), entry-point pattern,
variable-size input handling, Simulink integration, and conventions. Consult
for any LiteRT/.tflite model code generation task.
references/tensorflow-to-litert-conversion.md — Converting TensorFlow
SavedModel/Keras/.h5 to .tflite via the Python tf.lite.TFLiteConverter API.
Consult when the user has a TensorFlow model but no .tflite file yet.
references/litert-numeric-verification.md — Verifying MEX numerics against
MATLAB reference for LiteRT models (tolerance guidance, common mismatches).
references/codegen-workflow.md — Shared code generation steps for both
PyTorch and LiteRT: MEX generation, MEX verification, library/executable
targets, coder.Constant usage, and coder.DeepLearningConfig notes.
references/dnn-codegen-options.md — DNN-INFERENCE-SPECIFIC codegen
options: DLTargetLibrary / DeepLearningConfig('none') for the plain-C
DL path, LargeConstantGeneration for serializing large DNN weights to
data files, and MEX SIMD ceiling in a DNN-inference context. Read this
when the generic file's knobs need DNN-specific framing (e.g., "the MEX
SIMD cap matters because inference is the target").
See Also
matlab-generate-code — Generic MATLAB Coder tuning (SIMD instruction sets, OpenMP multi-threading, reduction-loop vectorization). Consult for performance options not specific to DNN inference.
matlab-deploy-embedded-ai — dlnetwork-based codegen with model compression (quantization, pruning, projection) and exportNetworkToSimulink workflows (Pattern 1). Use it when the source is an editable dlnetwork in MATLAB rather than a .pt2 / .tflite file.
matlab-optimize-gpu-codegen — CUDA-target codegen tuning (coder.gpuConfig, kernel fusion, memory-hierarchy options) when the deployment target is an NVIDIA GPU rather than CPU or embedded hardware.
Copyright 2026 The MathWorks, Inc.
1---2name: matlab-deploy-ai-model3description: Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. Use when the user wants to integrate an AI model into an application with code generation as the end goal — generating MEX, CUDA MEX, static library, dynamic library, or executable — or using the model in Simulink for simulation and code generation. Covers PyTorch ExportedProgram (.pt2) via loadPyTorchExportedProgram and LiteRT (.tflite) via loadLiteRTModel (R2026a+). Keywords: PyTorch, torch, .pt2, ExportedProgram, loadPyTorchExportedProgram, invoke, codegen, MEX, CUDA, GPU, C, C++, deploy, AI model, deep learning model, LiteRT, TFLite, TensorFlow Lite, Simulink, slbuild, PyTorch ExportedProgram block, MATLAB Function block, dlosslib, loadLiteRTModel.4license: https://www.mathworks.com/content/dam/mathworks/license/pmrl/lic5---67# Generate C/C++/CUDA Code from an AI Model89Generate deployable C/C++ or CUDA code from an AI model using MATLAB Coder or10GPU Coder. The workflow follows a common pattern regardless of model framework:11load, inspect, write entry-point, generate MEX, verify, then generate production code.1213## When to Use1415- User wants to generate C/C++/CUDA code from an AI model (PyTorch, LiteRT)16- User has a model file (.pt2, .tflite) and wants to load it into MATLAB17- User wants MEX acceleration for an AI model18- User wants to generate CUDA code or GPU-accelerated MEX from an AI model19- User wants to deploy an AI model to hardware20- User wants to use a PyTorch or LiteRT model in Simulink (simulation or code generation)21- User wants to verify AI model numerics between the source framework and MATLAB2223## When NOT to Use2425- **General MATLAB Coder usage** (codegen syntax, config tuning, writing codegen-ready code)26- **Editable dlnetwork for Deep Learning Toolbox workflows** (quantization, compression, transfer learning) — use `importNetworkFromPyTorch` (PyTorch), `importNetworkFromTensorFlow` (SavedModel), or `importNetworkFromKeras` (`.keras`/`.h5`) which return a `dlnetwork`. For deployment of an editable `dlnetwork` with model compression (INT8 quantization via `dlquantizer`, pruning, projection) or `exportNetworkToSimulink` workflows — use `matlab-deploy-embedded-ai` (Pattern 1).27- **Training or fine-tuning** — this skill is for inference code generation only2829## Supported Frameworks3031| Framework | Model format | Load function | Status |32|-----------|-------------|---------------|--------|33| PyTorch | `.pt2` | `loadPyTorchExportedProgram` | Supported (R2026a+) |34| LiteRT / TFLite | `.tflite` | `loadLiteRTModel` | Supported (R2026a+) |3536For PyTorch-specific details (API routing, entry-point pattern, export workflow,37data layout, common mistakes): see `references/pytorch-workflow.md`.3839For LiteRT-specific details (API routing, entry-point pattern, variable-size inputs,40Simulink integration, conventions): see `references/litert-workflow.md`.41For converting TensorFlow/Keras/.h5 to `.tflite`: see42`references/tensorflow-to-litert-conversion.md`.4344## Generic Workflow4546The code generation workflow follows the same steps for any framework:4748### 1. Load and Inspect4950Load the model and check its input/output specifications to determine expected51shapes and types.5253### 2. Write Entry-Point Function5455Create a codegen-compatible entry-point function that:56- Loads the model from a file path57- Runs inference on an input58- Returns the output5960The model file path must be wrapped with `coder.Constant` so it's known at61compile time.6263### 3. Verify Numerics6465Compare MATLAB inference output against the source framework to confirm correct66loading. Use the same input data in both environments and compare with tolerance.6768### 4. Generate MEX (First!)6970Always generate MEX before lib/exe to verify on the host machine:7172**CPU MEX:**73```matlab74cfg = coder.config("mex");75codegen -config cfg -args {coder.Constant("model_file"), input} entryPoint76```7778**CUDA MEX (GPU acceleration):**79```matlab80cfg = coder.gpuConfig("mex");81codegen -config cfg -args {coder.Constant("model_file"), input} entryPoint82```8384For CPU MEX SIMD acceleration (`SIMDAcceleration = 'Full'` for AVX2 on85Intel/AMD), see the `matlab-generate-code` skill. For the DNN-86inference-specific MEX AVX2 ceiling, see `references/dnn-codegen-options.md`.8788### 5. Verify MEX Output8990Compare MEX output against MATLAB reference using `matlab.unittest` with91tolerance:9293```matlab94refOut = entryPoint("model_file", input);95mexOut = entryPoint_mex("model_file", input);96testCase = matlab.unittest.TestCase.forInteractiveUse;97testCase.verifyThat(mexOut, matlab.unittest.constraints.IsEqualTo(refOut, ...98 'Within', matlab.unittest.constraints.AbsoluteTolerance(single(1e-5))));99```100101### 6. Generate Library/Executable102103Once MEX is verified, generate production code:104105```matlab106cfgLib = coder.config("lib");107cfgLib.TargetLang = "C++"; % set to "C++" for C++ output; default is "C"108codegen -config cfgLib -args {coder.Constant("model_file"), input} entryPoint109```110111For DLL: `coder.config("dll")`. For executable: `coder.config("exe")`.112113**CUDA variants:** Replace `coder.config` with `coder.gpuConfig`.114115**Performance tuning:**116- Generic knobs (SIMD instruction sets, reduction-loop vectorization,117 multithreaded loops): see the `matlab-generate-code` skill.118- MATLAB Coder ↔ Simulink Coder property naming duality and `slbuild`119 `set_param` patterns: see the `matlab-deploy-embedded-code` skill.120- DNN-inference-specific knobs (`DLTargetLibrary` / `DeepLearningConfig` to121 disable third-party DL libraries, `LargeConstantGeneration` to serialize122 weights to data files): see `references/dnn-codegen-options.md`.123124### 7. Use in Simulink125126For Simulink integration, use the dedicated `PyTorch ExportedProgram` block from127`dlosslib` — set `ModelFilePath` to the `.pt2` file and it auto-detects128input/output shapes. No entry-point function or `coder.Constant` needed.129130Pre/post-processing can be done with Simulink blocks around the dedicated block.131If you need everything in a single block, use a MATLAB Function block with132`loadPyTorchExportedProgram` + `invoke` (same pattern as the entry-point, but133the model path is a string literal — no `coder.Constant`).134135Both paths support `slbuild` code generation (requires fixed-step solver + ERT136or GRT target). See `references/simulink-workflow.md` for full details.137138### 8. Deploy to Hardware (Optional — requires Embedded Coder)139140For embedded deployment, use the same entry-point function with an Embedded Coder141configuration. See the `matlab-deploy-embedded-code` skill for ERT config,142hardware settings, PIL/SIL verification, and target-specific options.143Ask the user to install the skill if it is not installed144145## Key Functions146147| Function | Purpose | Package | Since |148|----------|---------|---------|-------|149| `coder.Constant` | Make argument a compile-time constant | MATLAB Coder | R2011a |150| `coder.gpuConfig` | Create GPU (CUDA) code generation config | GPU Coder | R2017b |151| `codegen` | Generate code | MATLAB Coder | R2011a |152| `loadPyTorchExportedProgram` | Load .pt2 into MATLAB | MATLAB Coder Support Package for PyTorch and LiteRT Models | R2026a |153| `loadLiteRTModel` | Load .tflite into MATLAB | MATLAB Coder Support Package for PyTorch and LiteRT Models | R2026a |154155## Conventions156157- Always check the model's input specifications for correct input shape and type158- Always generate MEX first, verify, then proceed to lib/exe159- Always use `coder.Constant` for the model file path argument160- Input data is typically single-precision (check model input specs to confirm)161- Do NOT use `importNetworkFromPyTorch`, `importNetworkFromTensorFlow`, or `importNetworkFromKeras` in this skill — they return a `dlnetwork` on a different path. If the user needs a `dlnetwork` for quantization, projection, pruning, or `exportNetworkToSimulink` before code generation, route to `matlab-deploy-embedded-ai` instead162163## References164165- `references/pytorch-workflow.md` — Full PyTorch-specific workflow: API routing,166 entry-point pattern, export guidance, common mistakes, and conventions. Consult167 for any PyTorch/.pt2 model code generation task. Links to deeper PyTorch168 references (API signatures, data layout, numeric verification, supported models).169- `references/export-pytorch-models.md` — Exporting an eager-mode PyTorch model to170 `.pt2` with `torch.export` (upstream of loading). Consult when the user has a171 PyTorch model but no `.pt2` file yet, or hits `torch.export` `SerializeError` /172 kwarg-mismatch errors. Links to `pytorch-export-patterns.md` (per-source173 templates) and `pytorch-export-gotchas.md` (torch 2.11 serialization fixes).174- `references/simulink-workflow.md` — Simulink integration: dedicated PyTorch175 ExportedProgram block (Path A) vs MATLAB Function block (Path B), block mask176 parameters, code generation config, and key differences from command-line codegen.177- `references/litert-workflow.md` — Full LiteRT-specific workflow: API routing178 (`loadLiteRTModel` → `inputSpecifications` → `invoke`), entry-point pattern,179 variable-size input handling, Simulink integration, and conventions. Consult180 for any LiteRT/.tflite model code generation task.181- `references/tensorflow-to-litert-conversion.md` — Converting TensorFlow182 SavedModel/Keras/.h5 to `.tflite` via the Python `tf.lite.TFLiteConverter` API.183 Consult when the user has a TensorFlow model but no `.tflite` file yet.184- `references/litert-numeric-verification.md` — Verifying MEX numerics against185 MATLAB reference for LiteRT models (tolerance guidance, common mismatches).186- `references/codegen-workflow.md` — Shared code generation steps for both187 PyTorch and LiteRT: MEX generation, MEX verification, library/executable188 targets, `coder.Constant` usage, and `coder.DeepLearningConfig` notes.189- `references/dnn-codegen-options.md` — DNN-INFERENCE-SPECIFIC codegen190 options: `DLTargetLibrary` / `DeepLearningConfig('none')` for the plain-C191 DL path, `LargeConstantGeneration` for serializing large DNN weights to192 data files, and MEX SIMD ceiling in a DNN-inference context. Read this193 when the generic file's knobs need DNN-specific framing (e.g., "the MEX194 SIMD cap matters because inference is the target").195196## See Also197198- `matlab-generate-code` — Generic MATLAB Coder tuning (SIMD instruction sets, OpenMP multi-threading, reduction-loop vectorization). Consult for performance options not specific to DNN inference.199- `matlab-deploy-embedded-ai` — `dlnetwork`-based codegen with model compression (quantization, pruning, projection) and `exportNetworkToSimulink` workflows (Pattern 1). Use it when the source is an editable `dlnetwork` in MATLAB rather than a `.pt2` / `.tflite` file.200- `matlab-optimize-gpu-codegen` — CUDA-target codegen tuning (`coder.gpuConfig`, kernel fusion, memory-hierarchy options) when the deployment target is an NVIDIA GPU rather than CPU or embedded hardware.201202----203204Copyright 2026 The MathWorks, Inc.205206----