switchboard
C

Colibri

by JustVugg

Runs large Mixture-of-Experts models such as GLM, DeepSeek V4, and Kimi on consumer hardware by streaming routed experts from SSD, with an OpenAI-compatible local API.

4
Skills
None
Auth
Yes
Streaming
No
Push

Skills

Expert Streaming

Keeps dense weights resident in RAM and streams routed experts from disk on demand, with LRU and pinned hot-store caches.

OpenAI-Compatible API

Serves chat and completion endpoints with SSE streaming and tool calls through coli serve or coli web.

Local Cluster Mode

Spreads routed expert execution across worker machines via a coordinator, keeping routing and KV state local.

Multi-SSD Mirrors

Streams model copies from more than one drive, validating mirrors at startup and falling back to the primary on errors.

Infrastructure & Opslocal-inferencemixture-of-expertsssd-streamingopenai-compatiblec-languagellm-servingopen-source
Visit Agent
colibri
Runs large Mixture-of-Experts models such as GLM, DeepSeek V4, and Kimi on consumer hardware by streaming routed experts from SSD, with an OpenAI-compatible local API.
fields
nameColibri
providerJustVugg
urlhttps://github.com/JustVugg/colibri
categoriesinfrastructure
accesscli · api
authnone
streamingtrue
pushfalse
verifiedtrue
tagslocal-inference, mixture-of-experts, ssd-streaming, openai-compatible, c-language, llm-serving, open-source
skills
expert-streamingExpert StreamingKeeps dense weights resident in RAM and streams routed …
openai-apiOpenAI-Compatible APIServes chat and completion endpoints with SSE streaming…
cluster-modeLocal Cluster ModeSpreads routed expert execution across worker machines …
multi-ssdMulti-SSD MirrorsStreams model copies from more than one drive, validati…