oMLX
by Jun Kim
Serves local LLM, vision, embedding and reranker models on Apple Silicon Macs through OpenAI- and Anthropic-compatible APIs, with continuous batching and an SSD-backed KV cache.
Skills
Tiered KV Cache
Keeps hot KV blocks in RAM and offloads cold ones to SSD, restoring matching prefixes from disk instead of recomputing, even after restarts.
OpenAI and Anthropic APIs
Exposes streaming chat and text completions, Anthropic Messages, embeddings, rerank and model listing endpoints under /v1 for existing clients.
Multi-Model Serving
Loads LLMs, VLMs, embedding models and rerankers in one server, with LRU eviction, model pinning, per-model TTL and a memory limit.
Agent Integrations
Sets up OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, Pi and DeepSeek Harness against the local server from the admin dashboard.
Related Agents
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
Magnitude
Runs open-weight models locally on kernels tuned to your hardware and connects them to coding agents through the magnit…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…
Osaurus
Runs AI agents entirely on your Mac — a native Apple Silicon runtime for local MLX models with shared memory, file acce…