# Ort Build

> Build ONNX Runtime from source. Use this skill when asked to build, compile, or generate CMake files for ONNX Runtime.

- Skill: `microsoft/ort-build` (Agent Skill)
- Install (CLI): `npx skillmds@latest add microsoft/ort-build`
- Raw SKILL.md: https://api.skillmd.com/api/skills/microsoft/ort-build/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: Microsoft (https://skillmd.com/u/microsoft)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/microsoft/ort-build

---


# Building ONNX Runtime

The build scripts `build.sh` (Linux/macOS) and `build.bat` (Windows) delegate to `tools/ci_build/build.py`.

## Build phases

Three phases, controlled by flags:

- `--update` — generate CMake build files
- `--build` — compile (add `--parallel` to speed this up)
- `--test` — run tests

For native builds, if none are specified (and `--skip_tests` is not passed), **all three run by default**. For cross-compiled builds, the default is `--update` + `--build` only.

### When to use `--update`

You need `--update` when:
- First build in a new build directory
- New source files are added (some CMake targets use glob patterns, others use explicit file lists — re-run to pick up new files either way)
- CMake configuration changes (new flags, updated CMakeLists.txt)

You do **not** need `--update` when only modifying existing `.cc`/`.h` files — just use `--build`. Skipping it saves time.

## Examples

```bash
# Full build (update + build + test)
./build.sh --config Release --parallel
.\build.bat --config Release --parallel     # Windows

# Just regenerate CMake files
./build.sh --config Release --update

# Just compile (skip CMake regeneration and tests)
./build.sh --config Release --build --parallel

# Just run tests (after a prior build)
./build.sh --config Release --test

# Build with CUDA execution provider
./build.sh --config Release --parallel --use_cuda --cuda_home /usr/local/cuda --cudnn_home /usr/local/cuda

# Configure and build the WebGPU execution provider as a shared library (Windows)
.\build.bat --config RelWithDebInfo --build_dir .\build\WGPU --use_webgpu --build_shared_lib --update --build --parallel

# Incrementally rebuild the same WebGPU configuration after changing existing source files
.\build.bat --config RelWithDebInfo --build_dir .\build\WGPU --use_webgpu --build_shared_lib --build --parallel

# Build Python wheel
./build.sh --config Release --parallel --build_wheel

# Build a specific CMake target (much faster than a full build)
./build.sh --config Release --build --parallel --target onnxruntime_common

# Load flags from an option file (one flag per line)
./build.sh "@./custom_options.opt" --build --parallel
```

## Key flags

| Flag | Description |
|------|-------------|
| `--config` | `Debug`, `MinSizeRel`, `Release`, or `RelWithDebInfo` |
| `--parallel` | Enable parallel compilation (recommended) |
| `--skip_tests` | Skip running tests after build |
| `--build_wheel` | Build the Python wheel package |
| `--use_cuda` | Enable CUDA EP. Requires `--cuda_home`/`--cudnn_home` or `CUDA_HOME`/`CUDNN_HOME` env vars. On Windows, only `cuda_home`/`CUDA_HOME` is validated. |
| `--target T` | Build a specific CMake target (requires `--build`; e.g., `onnxruntime_common`, `onnxruntime_test_all`) |
| `--use_webgpu` | Enable WebGPU EP. To run its tests locally on Linux without a GPU, see the `webgpu-local-testing` skill. |
| `--cmake_extra_defines onnxruntime_QUICK_BUILD=ON` | Faster CUDA build: instantiates a reduced kernel set. **Side effect:** Flash is compiled for head_dim 128 only, so most attention shapes fall back to **MEA** (changes which attention kernel is compiled/dispatched). Don't use it to characterize Flash-vs-arch behavior. |
| `--build_dir` | Build output directory |

## Build output path

Default: `build/<Platform>/<Config>/` where Platform is `Linux`, `MacOS`, or `Windows`.

With Visual Studio multi-config generators, the config name appears twice (e.g., `build/Windows/Release/Release/`).

It may be customized with `--build_dir`.
For example, `--build_dir .\build\WGPU --config RelWithDebInfo` creates the CMake build tree at
`build/WGPU/RelWithDebInfo/`; Visual Studio places final binaries in its `RelWithDebInfo/` subdirectory.
The `--build_shared_lib` flag in the WebGPU example is optional and is only needed when building the ONNX Runtime DLL.

## Agent tips

- **Activate a Python virtual environment** before building. See "Python > Virtual environment" in `AGENTS.md`.
- **Build flags can silently reroute which kernel/code path executes.** A build option can
  change *which* kernel is compiled, and therefore which code path actually runs — so a CI
  failure can live in a different code path than your local build exercises. Before
  hypothesizing a hardware- or algorithm-specific cause (e.g. "this GPU arch miscomputes"),
  first identify **which kernel actually ran** for the failing configuration (see the
  `ort-test` skill → "Verify which path/kernel actually executed"). Concrete instance:
  `onnxruntime_QUICK_BUILD=ON` compiles FlashAttention for head_dim 128 only, so most
  attention shapes silently dispatch to Memory-Efficient Attention instead of Flash —
  details in the `cuda-attention-kernel-patterns` skill.
- **Prefer `python tools/ci_build/build.py` directly** over `build.bat`/`build.sh` when redirecting output. The `.bat` wrapper runs in `cmd.exe`, which breaks PowerShell redirection.
- **Redirect output to a file** (e.g., `> build_log.txt 2>&1`). Build output is large and will overflow terminal buffers.
- **Run builds in the background** — a full build can take tens of minutes to over an hour. Poll the log for `"Build complete"` or errors.
- **Use `--parallel`** by default unless the user says otherwise.
- Ask the user what they want to build (config, execution providers, wheel, etc.) if not clear from their prompt.

