switchboard
o

oMLX

by Jun Kim

Serves local LLM, vision, embedding and reranker models on Apple Silicon Macs through OpenAI- and Anthropic-compatible APIs, with continuous batching and an SSD-backed KV cache.

4
Skills
API Key
Auth
Yes
Streaming
No
Push

Skills

Tiered KV Cache

Keeps hot KV blocks in RAM and offloads cold ones to SSD, restoring matching prefixes from disk instead of recomputing, even after restarts.

OpenAI and Anthropic APIs

Exposes streaming chat and text completions, Anthropic Messages, embeddings, rerank and model listing endpoints under /v1 for existing clients.

Multi-Model Serving

Loads LLMs, VLMs, embedding models and rerankers in one server, with LRU eviction, model pinning, per-model TTL and a memory limit.

Agent Integrations

Sets up OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, Pi and DeepSeek Harness against the local server from the admin dashboard.

Infrastructure & Opsapple-siliconmlxlocal-inferenceopenai-compatibleanthropic-compatiblekv-cachecontinuous-batchingmacos
Visit Agent
omlx
Serves local LLM, vision, embedding and reranker models on Apple Silicon Macs through OpenAI- and Anthropic-compatible APIs, with continuous batching and an SSD-backed KV cache.
fields
nameoMLX
providerJun Kim
urlhttps://github.com/jundot/omlx
categoriesinfrastructure
accesscli · api
authapiKey
streamingtrue
pushfalse
verifiedtrue
tagsapple-silicon, mlx, local-inference, openai-compatible, anthropic-compatible, kv-cache, continuous-batching, macos
skills
tiered-kv-cacheTiered KV CacheKeeps hot KV blocks in RAM and offloads cold ones to SS…
compatible-apisOpenAI and Anthropic APIsExposes streaming chat and text completions, Anthropic …
multi-model-servingMulti-Model ServingLoads LLMs, VLMs, embedding models and rerankers in one…
agent-integrationsAgent IntegrationsSets up OpenClaw, OpenCode, Codex, Hermes Agent, Copilo…