Opper Multimodal

Use the Opper multimodal and realtime surfaces — everything beyond text. Covers media generation (images, audio speech / TTS, transcription / STT, video, OCR), the /v3/files storage API and file_id reuse, vision / PDF input on chat models, and realtime two-way voice / audio over WebSocket (wss://api.opper.ai/v3/realtime, browser tickets via /v3/realtime-sessions, function-scoped /v3/realtime/{name}). Use this skill whenever the user wants to generate or edit an image, do text-to-speech or speech-to-text, transcribe audio, generate video, run OCR on a PDF / image, upload or reuse media files, send images / PDFs to a model, or build a voice / realtime app on Opper — even if they only say "Opper". For text / chat, structured output, migration, and server-side tools like web_search, use the `opper-api` skill instead. For any endpoint signature or payload question, fetch the live OpenAPI spec at https://api.opper.ai/v3/openapi.yaml first.

opper-ai bc00c05 16.1 KB Updated

File contents

opper-ai/opper-skills/tree/main/opper-multimodal commit bc00c054f4

Frequently asked questions

npx skillmds@latest add opper-ai/opper-multimodal