switchboard

Runs DeepSeek V4 Flash, GLM 5.x, and Qwen models locally on Metal, CUDA, or ROCm with a native engine, a CLI, a coding agent, and an OpenAI-compatible server.

4
Skills
None
Auth
Yes
Streaming
No
Push

Skills

Local API Server

Serves OpenAI-style chat, Responses, completions, and Anthropic-style messages endpoints with tools and SSE streaming.

Native Coding Agent

Runs ds4-agent directly on the model without an HTTP server, using the model's native tool format and saved sessions.

SSD Streaming

Streams model weights from SSD so smaller Macs can run large models, such as full GLM 5.x on 128 GB systems.

Multi-GPU Serving

Uses multiple CUDA cards, including Ada Lovelace L40S, as a multi-user LLM server for the supported models.

Infrastructure & Opslocal-inferencedeepseekmetalcudassd-streamingopenai-compatiblellm-servingopen-source
Visit Agent
ds4
Runs DeepSeek V4 Flash, GLM 5.x, and Qwen models locally on Metal, CUDA, or ROCm with a native engine, a CLI, a coding agent, and an OpenAI-compatible server.
fields
nameds4
providerSalvatore Sanfilippo (antirez)
urlhttps://github.com/antirez/ds4
categoriesinfrastructure
accesscli · api
authnone
streamingtrue
pushfalse
verifiedtrue
tagslocal-inference, deepseek, metal, cuda, ssd-streaming, openai-compatible, llm-serving, open-source
skills
local-serverLocal API ServerServes OpenAI-style chat, Responses, completions, and A…
native-agentNative Coding AgentRuns ds4-agent directly on the model without an HTTP se…
ssd-streamingSSD StreamingStreams model weights from SSD so smaller Macs can run …
multi-gpuMulti-GPU ServingUses multiple CUDA cards, including Ada Lovelace L40S, …