Mlx Local Inference

Full local AI inference stack on Apple Silicon Macs via MLX. Includes: LLM chat (Qwen3-14B, Gemma3-12B), speech-to-text ASR (Qwen3-ASR, Whisper), text embeddings (Qwen3-Embedding 0.6B/4B), OCR (PaddleOCR-VL), TTS (Qwen3-TTS), and an automatic transcription daemon with LLM correction. All models run locally via MLX with OpenAI-compatible APIs. Use when the user needs local AI capabilities: text generation, speech recognition, embeddings/vector search, OCR, text-to-speech, or batch audio transcription — without cloud API calls.

johnalbertini14-glitch Updated 1 repo stars

File contents

johnalbertini14-glitch/openclaw-skills/tree/main/skills/bendusy/mlx-local-inference commit 2bb2018dc0

Frequently asked questions

npx skillmds@latest add johnalbertini14-glitch/mlx-local-inference