Fundraising Fox

Inworld AI

Mountain View, US · Founded 2021 · 85 employees on LinkedIn · 24 known investors

Inworld AI provides a full-duplex audio streaming platform that enables real-time conversational AI interactions with intelligent turn-taking, function calling, and multi-model routing capabilities for developers building conversational applications.

Also known as Inworld · Theai, Inc.

Founders & leadership

Inworld AI was founded in 2021 by Ilya Gelfenbeyn, Kylan Gibbs, and Michael Ermolenko.

IGIlya Gelfenbeyn
Ilya GelfenbeyninCo-founder · Executive Chairman
KGKylan Gibbs
Kylan GibbsinCEO & Co-Founder
MEMichael Ermolenko
Michael ErmolenkoinCo-founder · Chief Technology Officer

Investors · 24

Also in the syndicate · 6

Accelerator Investments LLCBITKRAFT VenturesCRVFirst Spark VenturesM12SK Telecom Venture Capital

Valuation · disclosed

Disclosed events
$500Mvaluation at Series A extensionAug 2023
filing ↗

Source: SEC prospectus filings, and round valuations the company or its investors disclosed — follow each entry's link for the claim.

Company profile

researched Aug 2026

Inworld AI (legal entity Theai, Inc.) is a research lab and inference provider for realtime AI aimed at consumer-facing applications. Its platform combines first-party speech models with model routing and inference capacity behind a single API and billing relationship. The product set comprises Realtime TTS (the TTS-2 research preview and TTS-2 Flash, plus TTS 1.5 Max and 1.5 Mini), Realtime STT, an end-to-end speech-to-speech Realtime API, Realtime Inference (Inworld-optimized open-source models such as Gemma 4, DeepSeek V3.2/V4 and GLM-5.1/5.2), a Realtime Router that reaches 220+ LLMs across first-party and third-party tracks (OpenAI, Anthropic, Google, xAI, Meta, Mistral, DeepSeek, Qwen, Groq and DeepInfra), and Compute, which provides dedicated capacity for traffic-heavy customers.

The Realtime API streams speech in and speech out over a single WebSocket or WebRTC connection, with full-duplex audio, context-aware turn detection with adjustable eagerness, mid-session function calling, dynamic context management, and compatibility with the OpenAI Realtime API. Speech features include instant and professional voice cloning from short reference audio, text-based voice design, natural-language voice steering across dimensions such as emotion, articulation, intonation, volume, pitch, range, speed and vocal style, non-verbal tags, word/character/phoneme alignment for lipsync, and cross-lingual voice identity across 200+ languages. Realtime STT adds diarization, custom vocabularies, voice profiling, semantic and acoustic VAD, word-level timestamps and multilingual support across 30+ languages. Deployment options span hosted cloud, self-managed VPC and on-premise, with zero-data-retention configurations for regulated workloads.

The company began with a focus on AI-driven virtual characters for games, metaverse, VR/AR and brand experiences, offering a no-code character studio and integrations with Unreal and Unity, and a Character Engine orchestrating multiple models for personality, emotion, long-term memory and dialogue animation. It later positioned around the Inworld Runtime for scaling consumer AI applications before framing itself as a realtime voice AI infrastructure provider.

Founding story

The company was founded in July 2021 by Ilya Gelfenbeyn, Michael Ermolenko and Kylan Gibbs, with a third-party profile also listing Yuvakiran Arthala as a founder. The founders previously built the conversational AI platform API.AI, acquired by Google and renamed Dialogflow, and worked on generative models at Google and DeepMind, where the team led product for LLMs. According to CEO Kylan Gibbs, the company was started because AI was accruing to business automation while consumer experiences lagged, beginning with AI agents for gaming and media.

Business model

Inworld sells API access to its speech models, Realtime API, LLM Router and dedicated compute to developers, priced per usage tier (for example a Growth plan and enterprise-scale rates) with a single billing relationship across the stack. A third-party profile describes a revenue model based on subscription services and partnerships with developers integrating AI capabilities, with scalable subscription plans; dedicated GPUs are offered on a per-GPU-hour basis.

Usage- and plan-based pricing across products: per-million-characters for realtime TTS, per-hour for realtime STT, no markup added on routed LLM spend, a stated percentage of the public rate for realtime inference, and per-GPU-hour pricing for dedicated GPUs starting from $5. Rates are described as falling further at enterprise scale.

Traction

Customers cited by the company include Status by Wishroll (reported to reach 1M users in 19 days and a 95% AI cost reduction), OtherHalf, Bible Chat, Particle, Luvu, Talkpal, Playroom, AstroBeam, Isekai Zero and Latitude. A third-party profile reports over 500,000 daily active users and clients primarily in North America and Europe across gaming, media and training. Partners and integrations referenced include LiveKit, Vapi, Stream and k-ID.

Latest developments

A company resource page published 2026-02-19 describes six products — Realtime TTS, Realtime STT, the Realtime API, Realtime Inference, the Realtime Router and Compute — including the Realtime TTS-2 research preview and TTS-2 Flash, routing to 220+ LLMs, and cumulative funding of more than $125 million. The website cites provider rates as of June 2026 in publishing price comparisons and links to an explanation of its price cuts.

Full profile — market position, technology, go-to-market, geography, history

Market position

Inworld positions itself as a lower-cost alternative to point vendors across the voice stack, publishing comparisons against ElevenLabs on TTS pricing and Deepgram on STT pricing and against typical LLM gateway markups. Company and third-party materials state its voice models rank first on the Artificial Analysis Speech Arena. At the time of its 2023 raise it was described as the best-funded startup at the intersection of AI and gaming.

A vertically integrated stack that combines first-party speech models, provider-agnostic LLM routing and underlying inference capacity under one API and one bill, with emphasis on realtime latency, natural-language steerability of speech output, voice cloning and design, broad language coverage, and pricing set below comparable single-purpose providers.

Technology

Co-designed models and serving infrastructure for low-latency inference: first-party TTS models reporting sub-100ms time-to-first-byte for Realtime TTS-2 and 25ms for TTS-2 Flash, streaming STT, and full-duplex speech-to-speech over WebSocket/WebRTC. A router layer provides automatic failover, A/B testing and routing strategies based on cost, latency, user tier, region or custom metadata, with observability across the attempt chain. Earlier technology centered on the Inworld Character Engine, which orchestrated multiple machine learning models for cognition, memory, emotion and multimodal character expression, with hybrid inference architecture and game-engine integrations.

Go-to-market

Self-serve developer onboarding and documentation alongside an enterprise sales motion, supported by published pricing plans, integrations and partnerships with voice-agent and infrastructure vendors (LiveKit, Vapi, Stream, k-ID), availability of the Runtime through the Azure marketplace, and published customer references from AI-native startups.

Developers and companies building consumer-facing AI applications, including AI companions and social apps, learning and education, health and wellness, agentic workforce tools, games and interactive media, avatar experiences, language-learning apps and phone agents; historically also game studios and media companies building NPCs and virtual characters.

Geography

Headquartered in Mountain View, California, United States, with the 2022 Series A announcement issued from San Francisco. A third-party profile states clients are served primarily in North America and Europe, and the platform supports global deployment across 200+ languages.

History

Founded in July 2021 in Mountain View, California, Inworld raised pre-seed and seed rounds totaling nearly $20 million, closing the seed in March 2022. It released a beta product, hired Academy Award winner John Gaeta as Chief Creative Officer, and was selected as one of six companies for the 2022 Disney Accelerator. In August 2022 it closed a $50 million Series A led by Section 32 and Intel Capital, bringing total funding to roughly $70 million. In August 2023 it raised a further $50 million led by Lightspeed Venture Partners at a valuation above $500 million, passing $100 million in total funding, with proceeds earmarked for R&D, hiring, infrastructure and an open-source version of its character engine. By 2025 the company was described as having raised over $120 million and having built the Inworld Runtime for scaling consumer applications; by 2026 its public positioning had shifted to realtime voice AI models, routing and inference, with more than $125 million raised.

Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Customer milestone: Status by Wishroll usersFeb 20261,000,000 users
Daily active usersJan 2026500,000 users
Dedicated GPU priceJun 2026$5
Languages supported (STT)Feb 202630 languages
Languages supported (TTS)Feb 2026200 languages
LLM routing markupJun 20260%
LLMs available via Realtime RouterFeb 2026220 models
Post-money valuationSep 2025$500M
Realtime STT price (Growth plan)Jun 2026$0.1
Realtime TTS price (Growth plan)Jun 2026$12.5
Realtime TTS-2 Flash time to first byteFeb 202625 milliseconds
Realtime TTS-2 time to first byteFeb 2026100 milliseconds
Total funding raisedFeb 2026$125M

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Founder mafia

2 people who came through Inworld AI went on to found or lead other companies.

Competitors · 9

by search overlap
OpenRouter533 shared keywordsOpenRouter operates an AI gateway that allows developers to access and compare hundreds of language models from multiple providers in a single platform. The service eliminates vendor lock-in while offering improved pricing, uptime, and reliability for companies and developers.
ElevenLabs523 shared keywordsElevenLabs is an AI research and product company that builds voice generation and audio processing platforms for creating human-like voice interactions with technology. The company serves individuals with accessibility needs, nonprofits, and general users across healthcare, education, and culture.
OpenAI490 shared keywordsOpenAI is an AI research and deployment company focused on developing artificial general intelligence (AGI) with emphasis on safety and beneficial outcomes for humanity.
Speechify448 shared keywordsSpeechify is a text-to-speech application that converts written content—including textbooks, PDFs, emails, and web pages—into audio for users to listen to. The platform serves students and professionals with reading difficulties, particularly those with dyslexia and ADHD, enabling them to access educational and professional materials through audio.
Murf325 shared keywordsMurf offers an AI voice generator and text-to-speech studio that turns scripts or pre-recorded audio into human-sounding AI voiceovers, with a library of 200+ voices across accents and 20+ languages. It targets businesses creating product demos, marketing videos, and e-learning content.
MiniMax275 shared keywordsMiniMax is a general AI company that develops proprietary multimodal foundation models handling text, audio, image, video, and music, with capabilities in code generation, agents, and long-context processing. It offers AI-native products such as MiniMax Code, MiniMax Hub, MiniMax Audio, and Talkie, along with an open platform for enterprises and developers.
Canva269 shared keywordsDesign platform; Blackbird invested when it was just an idea and in every round since.
Hugging Face262 shared keywordsHugging Face is a collaboration platform that hosts and provides access to machine learning models, datasets, and applications. It offers both open-source tools for the ML community and paid compute and enterprise solutions for teams building AI applications.
ReadSpeaker244 shared keywordsReadSpeaker provides AI-powered text-to-speech technology offering 300+ voices in 90+ languages, delivered through plugins, APIs, SDKs, and cloud or on-premise servers. It serves organizations across sectors such as education, government, gaming, publishing, and transportation for accessibility, content reading, and custom branded voices.

Companies competing with Inworld AI for the same Google search keywords, organic and paid, via search-intersection analysis.

Timeline · 7

launches, deals, and filings
Feb 2026
Realtime TTS-2 and TTS-2 Flash

Inworld's Realtime TTS family includes TTS-2 (research preview) and TTS-2 Flash, offering steerable, expressive speech with sub-100ms and 25ms time-to-first-byte respectively, voice cloning, text-based voice design and 200+ language support.

source ↗

Sep 2025
Runtime made available through the Azure marketplace with M12 support

CEO Kylan Gibbs said M12, Microsoft's venture fund, helped accelerate availability of Inworld Runtime through the Azure marketplace.

source ↗

Aug 2023
Raises $50M at over $500M valuation led by Lightspeed

Lightspeed Venture Partners led a $50 million round valuing Inworld at more than $500 million, with participation from Stanford University, Samsung Next, Microsoft's M12, First Spark Ventures and LG Technology Ventures. Proceeds were earmarked for R&D, hiring, infrastructure and launching an open-source version of the character engine.

$50M source ↗

Aug 2022
Closes $50M Series A led by Section 32 and Intel Capital

Inworld AI announced the close of a $50 million Series A to expand its developer platform for AI-driven virtual characters in gaming, metaverse, entertainment and brand experiences, bringing total funding to approximately $70 million.

$50M source ↗

Jan 2022
Hires John Gaeta as Chief Creative Officer

After closing its seed round in March 2022, Inworld hired Academy Award winner John Gaeta as Chief Creative Officer.

source ↗

Jan 2022
Selected for the 2022 Disney Accelerator

Inworld was one of six companies selected for the 2022 Disney Accelerator class, focused on immersive experiences including AR, NFTs and AI characters.

source ↗

Jan 2022
Beta product release

Inworld released the beta version of its developer platform for creating AI-driven virtual characters, featuring a no-code studio and Unreal and Unity integrations.

source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

In the news

Research sources · 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Inworld AI do?
Inworld AI is a research lab and inference provider offering realtime text-to-speech, speech-to-text and LLM routing APIs.
Who founded Inworld AI?
Inworld AI was founded by Ilya Gelfenbeyn, Kylan Gibbs, Michael Ermolenko in 2021.
Who are Inworld AI's investors?
Inworld AI's investors include Bitkraft Esports Ventures, Brave Capital, CRV (Charles River Ventures), Dentsu Ventures, Disney Accelerator, First Spark Ventures, Llc, Intel Capital, Kleiner Perkins and 10 more.
Where is Inworld AI headquartered?
Inworld AI is headquartered in Mountain View, US.