Llama
GGML organization and Llama contributors · A macOS menu-bar controller for local GGUF models, llama.cpp inference, built-in web chat and compatible model APIs.
Описание
Llama is a macOS menu-bar application for running local language models with llama.cpp. It suits developers and local-AI users who want a small controller, model downloads, a built-in chat interface and an API that other applications can use.
Models and API
The app starts a local server on port 9931, with API endpoints under /v1. It uses an installed llama.cpp engine or installs a prebuilt engine for the Mac. It discovers compatible existing models, recommends models that fit the machine and installs GGUF models from Hugging Face. Models load when requested and unload when idle. Use the built-in Web UI, compatible chat/editor clients, coding agents or direct API requests. Version 0.44.0 also supports decision models: the /v1/systemone endpoint returns probabilities for typed-question options, and compatible models show a Decision chip. It can download available DFlash draft heads; the publisher reports up to 50% faster generation than MTP, which is a workload-dependent claim. A request-building page supports OpenAI chat, OpenAI Responses and Anthropic-style requests, with images, tools and structured-output settings where the selected model supports them.
Installation and resources
Download the official Llama DMG and put the app in Applications, or use brew install --cask llama-app. The current cask requires an Apple-silicon arm64 Mac and macOS 15 or later. The catalog now lists 0.44.0, while the current package download link still points to the historical 0.43.0 DMG (1,226,629 bytes). The official 0.44.0 Llama.dmg release asset is 1,257,897 bytes; these are distinct artifacts. Model weights and downloaded inference engines require much more disk space than this controller. Model size, quantization, context length and concurrency determine memory and performance requirements; the app's recommendations are useful but not a guarantee every model will run well.
Network, accounts and cost
Inference runs on the Mac. Downloading models or a prebuilt engine requires network access and is separate from local inference. Ordinary localhost use does not require a cloud inference subscription. Hugging Face repositories may be public or gated; their own access terms and model licenses apply. The app is MIT-licensed, while model weights and external tools have separate licenses. Tailscale access requires the separately installed and signed-in Tailscale service, with its own account and terms.
Network exposure and practical cautions
By default, the server is reachable only on the Mac. Settings can bind it to Tailscale or to the current network. The README warns that This network binds all interfaces and the server has no password by default; use that mode only on a trusted network, and avoid combining it with agent mode on a network you do not own. Local inference is not a promise that deliberately enabled network clients or downloads generate no traffic. Check who can reach the server and what connected agents can do.
Custom model overrides in ~/.config/llama/models.user.ini can override the app's memory-derived settings. Keep copies before changing them, and review ignored-option notices. External clients can supply private prompts or use tool-capable agents; check their permissions and generated actions. The inspected sources do not establish a blanket Accessibility or Full Disk Access requirement. Grant access only to the model/configuration locations and workflows you intend to use.
Sources: Official project, storage and network behavior, 0.44.0 release, llama.cpp server API, MIT license.
Новая версия Что нового в 0.44.0 1 янв. 1 г. · Редакция OpenNavo
- ModelsRun decision models through /v1/systemone with option probabilities.
- EngineUpdate to b11429; switch immediately when no model is loaded.
- SettingsMerge Downloads and Command into Advanced.