Runs DeepSeek V4 Flash, GLM 5.x, and Qwen models locally on Metal, CUDA, or ROCm with a native engine, a CLI, a coding agent, and an OpenAI-compatible server.
Skills
Local API Server
Serves OpenAI-style chat, Responses, completions, and Anthropic-style messages endpoints with tools and SSE streaming.
Native Coding Agent
Runs ds4-agent directly on the model without an HTTP server, using the model's native tool format and saved sessions.
SSD Streaming
Streams model weights from SSD so smaller Macs can run large models, such as full GLM 5.x on 128 GB systems.
Multi-GPU Serving
Uses multiple CUDA cards, including Ada Lovelace L40S, as a multi-user LLM server for the supported models.
Related Agents
Colibri
Runs large Mixture-of-Experts models such as GLM, DeepSeek V4, and Kimi on consumer hardware by streaming routed expert…
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
Magnitude
Runs open-weight models locally on kernels tuned to your hardware and connects them to coding agents through the magnit…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…