Inworld Launches Realtime TTS-2 Voice Model Family for Controllable, Realtime Speech
Inworld, a research lab and inference provider focused on realtime AI for consumer-facing applications, today launched
Press Release Disclaimer: This is a press release distributed through the XPR Media network. It has not been independently verified by our newsroom.

![]()
Inworld, a research lab and inference provider focused on realtime AI for consumer-facing applications, today launched the TTS-2 voice model family, comprising Realtime TTS-2 and Realtime TTS-2 Flash. The models are designed to let developers control how generated speech is delivered while choosing the quality, latency, volume, and cost profile that fits their application.
Realtime TTS-2 is the family’s primary quality model. Realtime TTS-2 Flash is built for latency, high-volume use, and cost-sensitive production workloads. Both support natural-language delivery instructions, voice cloning, and the same language coverage, allowing developers to change the serving profile without giving up control over the performance.
Consumer AI applications face a different infrastructure equation from conventional software. Every conversation, spoken lesson, generated story turn, or character exchange creates another speech, model, and compute cost. Users judge the complete response rather than its individual components, which makes quality, latency, reliability, and economics part of the same product decision.
“We are obsessed with how voice AI feels, not just how it sounds. Realtime voice is the most natural way for people to communicate with AI, because it is the most natural way people communicate with each other. Voice is how we actually connect. We built Realtime TTS-2 to make that connection feel real,” said Kylan Gibbs, CEO and co-founder of Inworld.
Developers can direct tone, pacing, and expression in the prompt
Developers can steer delivery by placing natural-language instructions directly inside the text, using directions such as [calm, reassuring] or [excited]. Inline cues including [laugh], [breathe], [sigh], [cough], and [yawn] are rendered as sounds rather than spoken as words. Adjustable pauses give developers another way to shape timing and intent.
Realtime TTS-2 also supports text-based voice design, allowing developers to describe an original voice without supplying reference audio. Instant, zero-shot voice cloning can create a voice from five to 15 seconds of authorized input audio. Developers are responsible for obtaining the consent and rights required to use any source recording.
Delivery steering is available across more than 100 languages. Cross-language voice support extends across more than 200 languages and more than 500 dialects, with the models designed to preserve voice identity as the language changes.
TTS-2 Flash cuts time to first byte to 25ms
Realtime TTS-2 delivers latency under 100ms at the 99th percentile, as measured under Inworld’s launch test conditions. Realtime TTS-2 Flash recorded 25ms time to first byte in Coval benchmarking and is intended for applications where response speed, sustained volume, or unit cost shapes the product architecture. Both models support WebSocket streaming.
The TTS-2 family can be used as a standalone speech layer or with Inworld’s broader set of modular APIs, including Realtime STT, Realtime API, Realtime Router, Realtime Inference, and Compute. The Realtime API is compatible with the OpenAI Realtime protocol, giving developers a familiar integration path while preserving control of their application and user relationship.
“Inworld’s TTS-2 marks a real step forward in emotionally expressive voice synthesis. When combined with the conversational intelligence of LiveKit agents, it enables interactions that feel genuinely human, responsive, nuanced, and alive in ways that feel natural,” said David Zhao, CTO of LiveKit.
Inworld’s Realtime TTS, Realtime STT, and Realtime API are available natively inside LiveKit Agents. LiveKit handles the live media and agent orchestration around an interaction, including interruptions and turn-taking, while developers can select Inworld components for speech input, model access, inference, routing, managed sessions, and speech output.
Built for consumer products where every interaction adds cost
Open-ended consumer applications create voice requirements that prerecorded dialogue cannot meet. Latitude’s Voyage, for example, lets players build AI-driven role-playing worlds and interact with characters outside fixed dialogue trees. Each new choice can create a scene that needs an original line, delivered in the right voice and with the right dramatic intent.
“The AI Native games of the future need characters that you can deeply connect with. In building our AI native game platform Voyage, voice models that offer the full control and emotional complexity to make characters actually feel real is one of the biggest pieces missing. TTS 2 is a significant advance in helping make that future a reality,” said Nick Walton, CEO of Latitude, the company behind AI Dungeon and Voyage.
Inworld’s published customer results show how speech infrastructure can affect product economics and behavior. Talkpal, a conversational language-learning product serving more than five million learners across more than 80 languages, reported 40 percent lower TTS cost, 7 percent higher feature usage, and 4 percent higher retention in a four-week A/B test after integrating Inworld Realtime TTS. Status, an AI-powered social simulation from Wishroll, reported an approximately 95 percent reduction in AI cost while serving more than 500,000 daily active users with consistent sub-second performance at peak load.
“Cost is the first thing standing in these developers’ way, so it is the first thing we are taking down,” Gibbs said. “But a consumer app does not succeed on cheap inference alone. These teams need help with growth and distribution. We are building toward a model where a small team can come to Inworld and get the technology, the economics, and the support to build a real consumer business.”
Inworld power applications with 100+ million daily users, support more than one billion users a day, and serve more than 10 trillion LLMs tokens per month.
Available now through the Inworld API
Realtime TTS-2 and Realtime TTS-2 Flash are available through Inworld’s API. On-demand TTS pricing starts at $25 per million characters. Monthly plans reduce the listed rate to $20 per million characters on the Creator plan, $17.50 on Builder, $15 on Developer, and $12.50 on Growth.
Developers can review TTS capabilities and documentation at docs.inworld.ai/tts/capabilities and current pricing at inworld.ai/pricing.
About Inworld
Inworld is a research lab and inference provider focused on realtime AI for consumer-facing applications. Inworld builds first-party speech models, serves LLMs, and operates the inference behind modular APIs, enabling developers of high-volume consumer applications to meet the quality, latency, reliability, and cost bar their users require. Inworld serves more than 10 trillion LLM tokens per month.
View source version on businesswire.com: https://www.businesswire.com/news/home/20260902226008/en/
Media gallery

