Venice API Overview
Venice.ai is an OpenAI-compatible inference platform for text, image, audio, video, and embeddings. One API — two ways to pay: a traditional API key (Pro account), or a wallet (x402, USDC on Base, no account required).
Use when
- You're writing code against
api.venice.ai for the first time.
- You need to decide between API-key and x402/wallet authentication.
- You want a quick map of which endpoint to call for which task.
- You need to understand the common response headers (
X-Balance-Remaining, PAYMENT-REQUIRED, etc.).
Base URL
All endpoints live under:
https://api.venice.ai/api/v1
The OpenAPI spec is distributed at outerface/swagger.yaml (current version 20260420.235001).
Authentication (pick one per request)
| Scheme |
Header |
Best for |
BearerAuth |
Authorization: Bearer <VENICE_API_KEY> |
Server-side apps, dashboards, usage analytics, bundled credits |
siwx (x402) |
X-Sign-In-With-X: <base64 SIWE JSON> |
No account, pay-as-you-go with USDC on Base, serverless / agents |
Every inference endpoint accepts either — see venice-auth.
# Bearer
curl https://api.venice.ai/api/v1/models \
-H "Authorization: Bearer $VENICE_API_KEY"
# x402 / SIWE (one-liner via the SDK)
import { VeniceClient } from 'venice-x402-client'
const v = new VeniceClient(process.env.WALLET_KEY)
await v.models.list()
Endpoint map
Inference
| Category |
Endpoints |
Skill |
| Chat |
POST /chat/completions |
venice-chat |
| Responses (Alpha) |
POST /responses |
venice-responses |
| Embeddings |
POST /embeddings |
venice-embeddings |
| Image gen |
POST /image/generate, POST /images/generations, GET /image/styles |
venice-image-generate |
| Image edit |
POST /image/edit, POST /image/multi-edit, POST /image/upscale, POST /image/background-remove |
venice-image-edit |
| TTS |
POST /audio/speech |
venice-audio-speech |
| STT |
POST /audio/transcriptions |
venice-audio-transcription |
| Music (async) |
POST /audio/quote, /audio/queue, /audio/retrieve, /audio/complete |
venice-audio-music |
| Video (async) |
POST /video/quote, /video/queue, /video/retrieve, /video/complete, /video/transcriptions |
venice-video |
Catalog
| Category |
Endpoints |
Skill |
| Models |
GET /models, /models/traits, /models/compatibility_mapping |
venice-models |
| Characters |
GET /characters, /characters/{slug}, /characters/{slug}/reviews |
venice-characters |
Account, billing, wallet
| Category |
Endpoints |
Skill |
| API keys |
`GET |
POST |
| Billing |
GET /billing/balance, /billing/usage, /billing/usage-analytics |
venice-billing |
| x402 wallet |
GET /x402/balance/{wallet}, POST /x402/top-up, GET /x402/transactions/{wallet} |
venice-x402 |
Utility
| Category |
Endpoints |
Skill |
| Crypto RPC proxy |
GET /crypto/rpc/networks, POST /crypto/rpc/{network} |
venice-crypto-rpc |
| Augment |
POST /augment/text-parser, /augment/scrape, /augment/search |
venice-augment |
Response headers to watch
| Header |
When |
Meaning |
X-Balance-Remaining |
x402 inference success |
USDC credits left, e.g. "4.230000" |
X-RateLimit-Limit-* / X-RateLimit-Remaining-* |
all inference |
your current per-minute/day caps |
PAYMENT-REQUIRED |
402 on x402 inference |
base64 JSON with top-up + SIWX challenge (x402 v2) |
Content-Encoding |
200 when client sent Accept-Encoding: gzip, br |
compression (embeddings, chat) |
Pricing model at a glance
- Pricing is dynamic per request, metered in USD.
- Paid inference endpoints in the spec carry an
x-payment-info block with min and max bounds in USD (typically min: 0.001, max: 10.00; higher for bulk video/audio). Read-only discovery routes like GET /models, /models/traits, and /models/compatibility_mapping do not.
- Pro (Bearer) accounts draw from DIEM (staked credits), USD balance, and bundled credits in priority order.
- x402 (wallet) users draw from a prepaid USDC credit balance on Base.
- The authoritative per-model price is on
GET /models → model_spec.pricing (when present — video models omit it; use /video/quote for video pricing) (see venice-models).
Standard error shape
Every error body follows one of:
{ "error": "Human-readable message" }
or, for 400 validation errors:
{ "error": "...", "details": { "fieldName": { "_errors": ["Field is required"] } } }
402 on x402 adds structured topUpInstructions and siwxChallenge. See venice-errors for the full table and retry strategy.
OpenAI compatibility — what works and what doesn't
- Drop-in:
/chat/completions, /embeddings, /images/generations, /audio/speech, /audio/transcriptions, /models.
- Ignored but accepted for compat:
user, store.
- Venice-only extensions live under:
venice_parameters (chat completions)
venice_parameters is rejected on /responses — use headers / native fields instead
- Model feature suffixes (e.g.
zai-org-glm-5-1:enable_web_search=on, kimi-k2-6:strip_thinking_response=true&disable_thinking=true) flip venice_parameters via the model ID — see venice-chat.
Versioning
info.version in swagger.yaml is a timestamp (YYYYMMDD.HHMMSS). There is no /v2; features roll forward on the single /api/v1 surface and are guarded by:
- Alpha/Beta tags in endpoint descriptions (e.g.
/responses, Billing).
x-guidance / model capability flags on /models.
- Always check the model's
model_spec.capabilities from GET /models for feature flags (supportsWebSearch, supportsReasoning, supportsE2EE, supportsXSearch, supportsMultipleImages, supportsFunctionCalling, supportsAudioInput, supportsVideoInput, …) before relying on a feature.
Fast start checklist
- Read
venice-auth and choose Bearer vs x402.
GET /models — pick a model and note its model_spec.constraints and model_spec.pricing.
- Wire up one happy-path call from the matching skill.
- Add error handling using
venice-errors (402, 422, 429).
- Hook up observability via
X-Balance-Remaining / /billing/usage / /x402/transactions.
1---2name: venice-venice-api-overview3description: High-level map of the Venice.ai API - base URL, authentication modes, endpoint categories, response headers, pricing model, error shape, and versioning. Load this first when starting any Venice integration.4license: MIT5---6
7# Venice API Overview
8
9Venice.ai is an OpenAI-compatible inference platform for text, image, audio, video, and embeddings. One API — two ways to pay: a traditional **API key** (Pro account), or a **wallet** (x402, USDC on Base, no account required).
10
11## Use when
12
13- You're writing code against `api.venice.ai` for the first time.
14- You need to decide between API-key and x402/wallet authentication.
15- You want a quick map of which endpoint to call for which task.
16- You need to understand the common response headers (`X-Balance-Remaining`, `PAYMENT-REQUIRED`, etc.).
17
18## Base URL
19
20All endpoints live under:
21
22```
23https://api.venice.ai/api/v1
24```
25
26The OpenAPI spec is distributed at `outerface/swagger.yaml` (current version `20260420.235001`).
27
28## Authentication (pick one per request)
29
30| Scheme | Header | Best for |
31|---|---|---|
32| `BearerAuth` | `Authorization: Bearer <VENICE_API_KEY>` | Server-side apps, dashboards, usage analytics, bundled credits |
33| `siwx` (x402) | `X-Sign-In-With-X: <base64 SIWE JSON>` | No account, pay-as-you-go with USDC on Base, serverless / agents |
34
35Every inference endpoint accepts **either** — see [`venice-auth`](../venice-auth/SKILL.md).
36
37```bash
38# Bearer
39curl https://api.venice.ai/api/v1/models \
40 -H "Authorization: Bearer $VENICE_API_KEY"
41
42# x402 / SIWE (one-liner via the SDK)
43import { VeniceClient } from 'venice-x402-client'
44const v = new VeniceClient(process.env.WALLET_KEY)
45await v.models.list()
46```
47
48## Endpoint map
49
50### Inference
51
52| Category | Endpoints | Skill |
53|---|---|---|
54| Chat | `POST /chat/completions` | [`venice-chat`](../venice-chat/SKILL.md) |
55| Responses (Alpha) | `POST /responses` | [`venice-responses`](../venice-responses/SKILL.md) |
56| Embeddings | `POST /embeddings` | [`venice-embeddings`](../venice-embeddings/SKILL.md) |
57| Image gen | `POST /image/generate`, `POST /images/generations`, `GET /image/styles` | [`venice-image-generate`](../venice-image-generate/SKILL.md) |
58| Image edit | `POST /image/edit`, `POST /image/multi-edit`, `POST /image/upscale`, `POST /image/background-remove` | [`venice-image-edit`](../venice-image-edit/SKILL.md) |
59| TTS | `POST /audio/speech` | [`venice-audio-speech`](../venice-audio-speech/SKILL.md) |
60| STT | `POST /audio/transcriptions` | [`venice-audio-transcription`](../venice-audio-transcription/SKILL.md) |
61| Music (async) | `POST /audio/quote`, `/audio/queue`, `/audio/retrieve`, `/audio/complete` | [`venice-audio-music`](../venice-audio-music/SKILL.md) |
62| Video (async) | `POST /video/quote`, `/video/queue`, `/video/retrieve`, `/video/complete`, `/video/transcriptions` | [`venice-video`](../venice-video/SKILL.md) |
63
64### Catalog
65
66| Category | Endpoints | Skill |
67|---|---|---|
68| Models | `GET /models`, `/models/traits`, `/models/compatibility_mapping` | [`venice-models`](../venice-models/SKILL.md) |
69| Characters | `GET /characters`, `/characters/{slug}`, `/characters/{slug}/reviews` | [`venice-characters`](../venice-characters/SKILL.md) |
70
71### Account, billing, wallet
72
73| Category | Endpoints | Skill |
74|---|---|---|
75| API keys | `GET|POST|DELETE /api_keys`, `/api_keys/{id}`, `/api_keys/rate_limits`, `/api_keys/rate_limits/log`, `/api_keys/generate_web3_key` | [`venice-api-keys`](../venice-api-keys/SKILL.md) |
76| Billing | `GET /billing/balance`, `/billing/usage`, `/billing/usage-analytics` | [`venice-billing`](../venice-billing/SKILL.md) |
77| x402 wallet | `GET /x402/balance/{wallet}`, `POST /x402/top-up`, `GET /x402/transactions/{wallet}` | [`venice-x402`](../venice-x402/SKILL.md) |
78
79### Utility
80
81| Category | Endpoints | Skill |
82|---|---|---|
83| Crypto RPC proxy | `GET /crypto/rpc/networks`, `POST /crypto/rpc/{network}` | [`venice-crypto-rpc`](../venice-crypto-rpc/SKILL.md) |
84| Augment | `POST /augment/text-parser`, `/augment/scrape`, `/augment/search` | [`venice-augment`](../venice-augment/SKILL.md) |
85
86## Response headers to watch
87
88| Header | When | Meaning |
89|---|---|---|
90| `X-Balance-Remaining` | x402 inference success | USDC credits left, e.g. `"4.230000"` |
91| `X-RateLimit-Limit-*` / `X-RateLimit-Remaining-*` | all inference | your current per-minute/day caps |
92| `PAYMENT-REQUIRED` | `402` on x402 inference | base64 JSON with top-up + SIWX challenge (x402 v2) |
93| `Content-Encoding` | `200` when client sent `Accept-Encoding: gzip, br` | compression (embeddings, chat) |
94
95## Pricing model at a glance
96
97- Pricing is **dynamic per request**, metered in USD.
98- Paid inference endpoints in the spec carry an `x-payment-info` block with `min` and `max` bounds in USD (typically `min: 0.001`, `max: 10.00`; higher for bulk video/audio). Read-only discovery routes like `GET /models`, `/models/traits`, and `/models/compatibility_mapping` do not.
99- Pro (Bearer) accounts draw from **DIEM** (staked credits), **USD** balance, and **bundled credits** in priority order.
100- x402 (wallet) users draw from a prepaid **USDC credit balance** on Base.
101- The authoritative per-model price is on `GET /models` → `model_spec.pricing` (when present — video models omit it; use `/video/quote` for video pricing) (see [`venice-models`](../venice-models/SKILL.md)).
102
103## Standard error shape
104
105Every error body follows one of:
106
107```json
108{ "error": "Human-readable message" }
109```
110
111or, for 400 validation errors:
112
113```json
114{ "error": "...", "details": { "fieldName": { "_errors": ["Field is required"] } } }
115```
116
117`402` on x402 adds structured `topUpInstructions` and `siwxChallenge`. See [`venice-errors`](../venice-errors/SKILL.md) for the full table and retry strategy.
118
119## OpenAI compatibility — what works and what doesn't
120
121- Drop-in: `/chat/completions`, `/embeddings`, `/images/generations`, `/audio/speech`, `/audio/transcriptions`, `/models`.
122- Ignored but accepted for compat: `user`, `store`.
123- Venice-only extensions live under:
124 - `venice_parameters` (chat completions)
125 - `venice_parameters` is **rejected** on `/responses` — use headers / native fields instead
126- Model feature suffixes (e.g. `zai-org-glm-5-1:enable_web_search=on`, `kimi-k2-6:strip_thinking_response=true&disable_thinking=true`) flip `venice_parameters` via the model ID — see [`venice-chat`](../venice-chat/SKILL.md#model-feature-suffixes).
127
128## Versioning
129
130- `info.version` in `swagger.yaml` is a timestamp (`YYYYMMDD.HHMMSS`). There is **no** `/v2`; features roll forward on the single `/api/v1` surface and are guarded by:
131 - **Alpha/Beta** tags in endpoint descriptions (e.g. `/responses`, Billing).
132 - `x-guidance` / model capability flags on `/models`.
133- Always check the model's `model_spec.capabilities` from `GET /models` for feature flags (`supportsWebSearch`, `supportsReasoning`, `supportsE2EE`, `supportsXSearch`, `supportsMultipleImages`, `supportsFunctionCalling`, `supportsAudioInput`, `supportsVideoInput`, …) before relying on a feature.
134
135## Fast start checklist
136
1371. Read [`venice-auth`](../venice-auth/SKILL.md) and choose Bearer vs x402.
1382. `GET /models` — pick a model and note its `model_spec.constraints` and `model_spec.pricing`.
1393. Wire up one happy-path call from the matching skill.
1404. Add error handling using [`venice-errors`](../venice-errors/SKILL.md) (402, 422, 429).
1415. Hook up observability via `X-Balance-Remaining` / `/billing/usage` / `/x402/transactions`.