release
Apr 24, 2026
By Teun
OpenClaw v2026.4.24 Ships Voice Calls, DeepSeek V4, and Azure Speech TTS
The biggest release of the month adds full-agent voice calls, DeepSeek V4 Flash and Pro models, Azure Speech as a bundled TTS provider, and the Crestodian first-run TUI helper for smoother CLI onboarding.
OpenClaw v2026.4.24 adds one of its most visible features yet: full-agent voice calls. The release, published on April 24, also brings support for DeepSeek V4 Flash and Pro models, bundles Azure Speech as a text-to-speech provider, and introduces Crestodian, a first-run terminal helper designed to make command-line onboarding less clumsy.
The voice-calling feature is the headline addition. In practical terms, it means an OpenClaw agent is no longer limited to text chat or typed input, but can participate in a phone-style call flow as part of a broader automation setup. That matters for teams building customer-facing assistants, internal helpdesks, or workflows where spoken interaction is more natural than filling out a form or opening a chat window.
OpenClaw is positioning this as a full-agent capability rather than a narrow add-on. That suggests the system is meant to handle the call as part of the agent’s normal execution path, instead of routing audio through a separate tool that sits outside the main orchestration layer. For automation builders, that kind of integration is usually the difference between a demo and something that can be wired into a real workflow.
The release also adds DeepSeek V4 Flash and Pro model support. DeepSeek has become a prominent name in the open model ecosystem, and the availability of multiple variants gives users a choice between faster, lighter responses and a more capable model tier, depending on the task. For agent builders, that can matter because not every step in a workflow needs the same level of model quality or latency.
Model support is only one part of the picture, though. OpenClaw also ships Azure Speech as a bundled text-to-speech, or TTS, provider. TTS converts generated text into spoken audio, which is the piece that makes voice output usable in a call or assistant flow. Bundling Azure Speech means users do not have to assemble that provider separately before testing or deploying a voice-enabled agent.
The Azure addition also fits how many teams already operate. Microsoft’s cloud services are common in enterprise environments, and speech APIs are often adopted where reliability, infrastructure integration, or procurement rules matter as much as raw model performance. By including Azure Speech in the release, OpenClaw gives those users a familiar path for spoken output without forcing them to stitch together their own provider stack.
Crestodian is aimed at a different pain point: first-run friction. Anyone who has set up a CLI-based AI tool knows that the initial experience can be the part most likely to fail, especially when credentials, config files, and environment variables are involved. A terminal UI helper can guide users through that setup interactively, which is useful for developers who want to get to a working installation quickly without reading a long setup guide first.
The combination of voice, model support, speech output, and onboarding support suggests this release is about making OpenClaw feel less like a collection of advanced parts and more like a ready-to-use agent platform. The release notes on GitHub describe it as the month’s biggest update, and the new features span both the runtime behavior of agents and the setup experience around them.
That mix is notable because voice agents are often held back less by model quality than by the plumbing around them. Audio input, audio output, model routing, and first-time configuration all have to work together before a voice system feels usable, and v2026.4.24 pushes on each of those pieces at once.