FreeToken
by FlashML
Serves large Mixture-of-Experts models such as DeepSeek-V4-Flash and GLM-5.2 on consumer NVIDIA GPUs, sharing work across GPU and CPU behind OpenAI- and Anthropic-compatible APIs.
Skills
Compatible API Server
Starts a local server with ft serve that exposes OpenAI /v1 endpoints, the Anthropic Messages API and the Responses API for a chosen model.
Coding Agent Launcher
Discovers the served model with ft launch, writes the coding agent's provider config, installs its CLI if missing and launches it.
MoE Expert Offload
Runs MoE layers with bandwidth-adaptive CPU-GPU co-execution and LRU expert caching, reallocating VRAM between experts and KV at runtime.
Related Agents
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…
Magnitude
Runs open-weight models locally on kernels tuned to your hardware and connects them to coding agents through the magnit…
oMLX
Serves local LLM, vision, embedding and reranker models on Apple Silicon Macs through OpenAI- and Anthropic-compatible…