Strata
by Niko1221
Runs the Qwen3.8-Flash-Next model on a 12 GB+ consumer GPU and serves it locally over OpenAI- and Anthropic-compatible APIs, for people connecting chat apps and coding agents.
Skills
OpenAI-Compatible API
Serves chat completions with streaming and tool calls at http://127.0.0.1:8080/v1.
Anthropic Messages API
Accepts Anthropic Messages requests so Claude Code can point ANTHROPIC_BASE_URL at the local server.
Responses API
Serves the OpenAI Responses API statelessly for Codex CLI and similar clients.
Image Input
Reads attached pictures in chat or through the API when image support is enabled at setup.
Management MCP Server
Lets an AI assistant install, start, stop, and check Strata through a bundled MCP server.
Related Agents
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
FreeToken
Serves large Mixture-of-Experts models such as DeepSeek-V4-Flash and GLM-5.2 on consumer NVIDIA GPUs, sharing work acro…
Magnitude
Runs open-weight models locally on kernels tuned to your hardware and connects them to coding agents through the magnit…
OpenSandbox
Runs AI-agent workloads in isolated Docker or Kubernetes sandboxes, exposing sandbox lifecycle, command, filesystem, an…