switchboard
F

FreeToken

by FlashML

Serves large Mixture-of-Experts models such as DeepSeek-V4-Flash and GLM-5.2 on consumer NVIDIA GPUs, sharing work across GPU and CPU behind OpenAI- and Anthropic-compatible APIs.

3
Skills
None
Auth
No
Streaming
No
Push

Skills

Compatible API Server

Starts a local server with ft serve that exposes OpenAI /v1 endpoints, the Anthropic Messages API and the Responses API for a chosen model.

Coding Agent Launcher

Discovers the served model with ft launch, writes the coding agent's provider config, installs its CLI if missing and launches it.

MoE Expert Offload

Runs MoE layers with bandwidth-adaptive CPU-GPU co-execution and LRU expert caching, reallocating VRAM between experts and KV at runtime.

Infrastructure & Opsmoe-inferencelocal-inferenceconsumer-gpunvidia-rtxexpert-offloadopenai-compatibleanthropic-compatiblemodel-serving
Visit Agent
freetoken
Serves large Mixture-of-Experts models such as DeepSeek-V4-Flash and GLM-5.2 on consumer NVIDIA GPUs, sharing work across GPU and CPU behind OpenAI- and Anthropic-compatible APIs.
fields
nameFreeToken
providerFlashML
urlhttps://github.com/FlashML-org/FreeToken
categoriesinfrastructure
accesscli · api
authnone
streamingfalse
pushfalse
verifiedtrue
tagsmoe-inference, local-inference, consumer-gpu, nvidia-rtx, expert-offload, openai-compatible, anthropic-compatible, model-serving
skills
api-serverCompatible API ServerStarts a local server with ft serve that exposes OpenAI…
agent-launcherCoding Agent LauncherDiscovers the served model with ft launch, writes the c…
moe-offloadMoE Expert OffloadRuns MoE layers with bandwidth-adaptive CPU-GPU co-exec…