Redis creator ships ds4 to run 284B parameter AI models locally

Salvatore Sanfilippo just dropped DwarfStar 4, a C inference engine that runs massive frontier models on high-memory local machines.

Tools · Source: Hacker News

What happened

Salvatore Sanfilippo, widely known as the creator of Redis, just released DwarfStar 4. It is a narrow C inference engine built specifically for high-memory machines. The tool targets Apple Silicon, NVIDIA CUDA, and AMD ROCm platforms. It allows developers to run massive frontier open-weight models entirely locally. Supported models currently include DeepSeek V4 and V4.1 Flash, GLM 5.x, and Qwen3.8 Flash Next. The engine handles both text and vision inputs natively.

The architecture relies on asymmetric 2-bit quantization. This specific compression targets the routed experts in mixture-of-experts models while keeping the critical shared paths precise. This approach allows a massive 284-billion-parameter model like DeepSeek V4 Flash to actually fit on consumer hardware without becoming lobotomized. The creator explicitly states this is not a generic GGUF runner. The engine follows a small set of model families and validates each supported layout end to end.

The software stack bundles a command line interface, local HTTP APIs, and a native coding agent. A standout feature is how it handles the KV cache. The cache is keyed by the SHA1 hash of the rendered prompt prefix and saved directly to the solid state drive. If you restart the server, a matching prefix is simply reloaded from disk instead of being recomputed. The local server also speaks standard OpenAI and Anthropic API formats out of the box.

Key facts

Why it matters

This release fundamentally changes the economics of AI development for software builders. You no longer need to pay massive cloud providers for inference on frontier models. If you have a 64GB Mac or a dedicated Linux box, you can run persistent coding agents locally with zero recurring latency costs. The disk-backed KV cache makes long-context agent workloads highly efficient. Developers can load massive codebases into the context window once and never pay the prefill penalty again after a restart.

The second-order effect here is the true decentralization of high-end AI capabilities. High-performance local hardware is shifting from a luxury to a one-time capital expense that replaces a recurring cloud tax. Tools like Codex, Claude Code, and OpenCode can now point directly to your localhost server. This gives founders and enterprise engineers total privacy over their proprietary codebases while still utilizing state-of-the-art mixture-of-experts models. We are moving away from remote serving as the only viable path for massive models.

For builders

Zero-cost local coding agents

You can point OpenCode or Claude Code directly to your local ds4 server. Builders get frontier coding assistance without paying per-token API fees or leaking proprietary source code to external cloud providers.

Instant prompt recovery via SSD

The engine saves long prompt prefixes to your solid state drive. Founders running massive system prompts save huge amounts of compute time because server restarts completely skip the expensive prefill phase.

Strict hardware requirements limit access

You need serious hardware to actually play this game. If you do not have at least 64GB of unified memory on an Apple Silicon Mac or a DGX Spark, you are stuck paying cloud providers.

My take

Antirez understands developer experience better than almost anyone in the industry. Building a narrow engine tailored for specific models instead of a bloated generic runner is exactly the right move. Local AI is finally moving from a slow toy for hobbyists to a serious, privacy-first tool for founders.

Original reporting: Hacker News. This is my rewrite and opinion.

More AI news for builders