The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
macos metal mtp mlx inference-engine apple-silicon local-llm llm-inference local-ai qwen speculative-decoding openai-compatible claude-code qwen3-next anthropic-compatible native-mtp mtplx qwen3-8 qwen3-8-flash-next flash-next
-
Updated
Sep 24, 2026 - Python