EXO
exo labs · Run AI models locally across a device cluster, with automatic discovery, distributed inference and a dashboard for model and cluster management.
About
EXO connects devices into a local AI cluster so that models can use resources across multiple machines rather than being limited to one device. The project is maintained by exo labs and uses MLX for inference and distributed communication.
Cluster and model workflow
The official repository describes automatic device discovery and topology-aware model placement based on device resources and network links. Tensor parallelism and supported RDMA-over-Thunderbolt configurations are intended to accelerate distributed inference. A built-in dashboard lets users manage the cluster and chat with models; the macOS application runs in the background on the Mac.
Integration
EXO exposes interfaces compatible with OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama APIs. It also supports custom models from Hugging Face. Version 1.0.71 adds Kimi K2.6, additional model cards and provider-recommended sampling defaults, and addresses M5 vision-model and distributed-output issues.
macOS requirements
The recorded Homebrew 1.0.71 cask targets Apple-silicon Macs and lists macOS 15 or later. The current upstream README separately states macOS Tahoe 26.2 or later for the macOS app and its RDMA setup. These sources differ: do not assume the recorded cask's minimum OS guarantees current-app compatibility. RDMA requires compatible Thunderbolt 5 hardware and OS configuration. The app may request permission to modify system settings and install a network profile.
EXO is released under Apache-2.0. The screenshot attached to this listing is the official macOS application image, not a generated mockup.
Sources
New What’s new in 1.0.71 Apr 23 · OpenNavo editorial
- ModelsAdds Kimi K2.6 with multimodality and additional quantization cards.
- InferenceUses provider-recommended sampling defaults including min_p and top_k.
- IntegrationAdds a Pi integration tab.