AI Open Source Models
mlx-serve
Run any LLM locally on your Mac faster than LM Studio or Ollama, with chat, coding agents, images, video, and voice, all fully offline and open source.
mlx-serve
What is mlx-serve?
mlx-serve is a free, open-source macOS application that runs large language models (LLMs) entirely on your Mac using Apple's MLX framework, providing faster inference than alternatives like LM Studio and Ollama. It supports a wide range of models, including DeepSeek V4 Flash, Gemma 4, Qwen, and more, and offers features like speculative decoding, voice cloning, image generation, and agent sandboxing.
mlx-serve vs Similar AI Tools
| Pricing Model | Free | Free | Free | Free |
| Free Credits | ||||
| Key Features |
|
|
|
|
| Pros |
|
|
|
|
| Cons |
|
|
|
|
| Best For |
|
|
|
|
How to use mlx-serve?
- 1Download the MLX Core app from GitHub Releases or install via Homebrew.
- 2Launch the app and use the Model Browser to download any supported model.
- 3Start chatting in the built-in interface or connect external apps via the OpenAI/Anthropic compatible API.
- 4Use the ⌃Space launcher for quick access, or configure voice and Telegram bot for remote control.
mlx-serve Key Features
- Local inference on Apple Silicon with native Metal acceleration
- Speculative decoding (Prompt Lookup Decoding, drafter models, MTP) for up to 2x faster generation
- Image generation (Krea-2-Turbo, FLUX.2) and photo editing from text instructions
- Video generation (LTX-Video 2.3) with synced audio and talking characters
- Voice cloning and text-to-speech (Qwen3-TTS) from a few seconds of audio
- Agent mode with 10 built-in tools (shell, file, search, browse, etc.) and MCP server support
- Agent Sandbox using Apple Virtualization for isolated command execution
- Ollama-compatible API for drop-in replacement of Ollama apps
- Claude Code integration for local coding agents
- KV-cache quantization and continuous batching for multiple concurrent chats
mlx-serve Use Cases
- Running local LLMs for private AI assistance without internet
- Developing and testing AI agents with sandboxed tool execution
- Replacing cloud APIs for coding agents (Claude Code, Cursor, Continue)
- Generating images, video, and audio on-device for creative projects
- Building custom AI assistants with voice control and scheduling
mlx-serve Pricing & Free Credits
mlx-serve currently operates on a Free model.
This tool is completely free to use
mlx-serve Pros & Cons
Pros
- Faster than LM Studio and Ollama on identical models
- Completely local and private – no data leaves your Mac
- Supports a huge range of models (MLX, GGUF, DeepSeek V4 Flash)
- Includes image, video, voice generation and editing
- Lightweight (~4.5 MB) and easy to install as a native Mac app
- Seamless integration with existing tools via OpenAI/Anthropic/Ollama APIs
Cons
- Mac-only (requires Apple Silicon)
- Some advanced features (e.g., DeepSeek V4) need high memory (96 GB+)
- No built-in cloud sync or multi-device support
What is mlx-serve best for?
- Mac users who want private, high-performance local AI
- Developers building AI agents and coding assistants
- Content creators needing offline image/video/audio generation
- Privacy-conscious individuals and organizations