Groq Inference

Ultra-fast LLM inference on custom LPU hardware. OpenAI-compatible API at api.groq.com. Lowest latency in the industry (500-1000+ tok/s). Supports chat completions, vision, audio (Whisper STT + TTS), tool calling, JSON mode, and streaming. Free tier available. Inference only — no training.

synthetic-sciences 6eb53c1 12.1 KB Updated

File contents

synthetic-sciences/openscience/tree/main/backend/cli/skills/ml-inference/groq commit 6eb53c11a6

Frequently asked questions

npx skillmds@latest add synthetic-sciences/groq-inference