Qwen
Qwen3.5 0.8B, 4B, 9B, 27B, 35B-A3B · Qwen3.6 27B, 35B-A3B
A broad multimodal range for chat, image understanding, and tool-enabled workflows. Start small for speed, then move up when you want more capable reasoning on a Mac with more memory.
Every LocalEngine model runs on your own Apple hardware. Download the one that fits your task, then chat, reason, inspect images, or work with tools without sending your data to a cloud service.
Qwen3.5 0.8B, 4B, 9B, 27B, 35B-A3B · Qwen3.6 27B, 35B-A3B
A broad multimodal range for chat, image understanding, and tool-enabled workflows. Start small for speed, then move up when you want more capable reasoning on a Mac with more memory.
Gemma 4 E2B, E4B, 12B, 26B-A4B, 31B · E4B MLX 8-bit
Instruction-tuned vision-language models for clear, useful image-and-text conversations. The MLX 8-bit E4B edition is tailored for efficient Apple Silicon use on macOS.
Qwythos 9B v2 · Q4_K_M and BF16
A 9B vision-language model with multi-token prediction. Choose the Q4_K_M download for a practical local setup, or BF16 when you have the memory for a higher-fidelity model.
Ornith 1.0 9B and 35B
A text-only agentic coding model, post-trained on Gemma 4 and Qwen 3.5, with tool calling. It is a focused choice when your work is code and structured tasks rather than images.
Laguna XS 2.1 · Q4_K_M and BF16
A 33B mixture-of-experts coding and agentic model with 3B active parameters per token. Q4_K_M needs at least 36 GB of unified memory; BF16 needs at least 72 GB.
Qwen3.5 0.8B, 2B, and 4B
Vision-language models that make iPhone and iPad chat feel capable without leaving the device. 0.8B is the recommended first download; 2B is the balanced step up; 4B is for newer devices.
Gemma 4 E2B and E4B
Instruction-tuned models for on-device image-and-text conversations. E2B is a great starting point; E4B is the larger option when your device can spare more storage and memory.
MiniCPM5 1B
A compact text-only model with long-context support and tool calling. It is small and fast for practical, private chat when you do not need image input.
Start with the smallest model that covers your task. You can always switch models in LocalEngine as your needs grow.
On iPhone or iPad: start with Qwen3.5 0.8B for a quick, image-aware first experience.
On a Mac: choose Qwen or Gemma for versatile image and text chat.
For coding: pick Ornith, or Laguna XS when you have a high-memory Mac.