Colibri
by JustVugg
Runs large Mixture-of-Experts models such as GLM, DeepSeek V4, and Kimi on consumer hardware by streaming routed experts from SSD, with an OpenAI-compatible local API.
Skills
Expert Streaming
Keeps dense weights resident in RAM and streams routed experts from disk on demand, with LRU and pinned hot-store caches.
OpenAI-Compatible API
Serves chat and completion endpoints with SSE streaming and tool calls through coli serve or coli web.
Local Cluster Mode
Spreads routed expert execution across worker machines via a coordinator, keeping routing and KV state local.
Multi-SSD Mirrors
Streams model copies from more than one drive, validating mirrors at startup and falling back to the primary on errors.
Related Agents
ds4
Runs DeepSeek V4 Flash, GLM 5.x, and Qwen models locally on Metal, CUDA, or ROCm with a native engine, a CLI, a coding…
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
Magnitude
Runs open-weight models locally on kernels tuned to your hardware and connects them to coding agents through the magnit…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…