Together AI
Founded 2022 · 427 employees on LinkedIn · 15 known investors
Find your way into Together AI
782 people in our graph share verified history with the Together AI team — schools, employers, funds. One of them is your warm intro.
Together AI provides infrastructure for building real-time voice agents, combining speech-to-text, LLM, and text-to-speech models on co-located GPU infrastructure for low-latency conversations. It offers access to multiple voice models (such as MiniMax, Rime, Deepgram, OpenAI, and Cartesia) through a single API, aimed at developers deploying production-scale voice applications.
Also known as Together AI · Together Computer
Founders & leadership
Together AI was founded in 2022 by Vipul Ved Prakash, Ce Zhang, Chris Ré, and Percy Liang.


Investors · 15
Also in the syndicate · 3
Funding
SEC filings, press & company announcements$1.6B disclosed across 2 of 3 rounds · 2023–2026
- $800MSeries CJul 2026Source ↗
- $800MraisedJul 2026 · 4 sourcesSource ↗
- Undisclosed amountSeries ANov 2023
Kleiner Perkins (lead), 137 Ventures, Definition Capital, Emergence Capital, Factory, Greycroft, Long Journey Ventures, Lux Capital, NEA, NVIDIA, Prosperity 7, SCB10x, SV Angel
Source ↗
Source: company announcements and press reports — follow each round's link for the claim.
Company profile
researched Aug 2026Together AI positions itself as an "AI native cloud": a full-stack platform covering inference, accelerated compute and model shaping for teams building AI applications. Its inference products include serverless inference for running open-source models on demand, batch inference for asynchronous processing of large workloads (stated to scale to 30 billion tokens per model with any serverless model or private deployment), provisioned throughput with reserved capacity, token-based pricing and a stated 99% uptime SLA plus drop-in API compatibility, dedicated model inference on isolated infrastructure, and dedicated container inference aimed at generative media workloads such as video, audio and image models.
On the compute side, the company offers accelerated compute ranging from self-serve instant clusters to thousands of GPUs, optimized with the Together Kernel Collection, including on-demand B200 availability on Together GPU Clusters. Adjacent infrastructure services include code sandboxes for building development environments for AI apps and agents, and managed storage combining object storage and parallel filesystems for AI workloads with no egress fees. Model shaping is represented by fine-tuning of open-source models for production use, positioned as improving accuracy, reducing hallucinations and controlling model behavior without customers managing training infrastructure. The platform claims 2x faster inference, 60% lower cost through workload-specific optimization, and 90% faster pre-training via the Together Kernel Collection.
The company also runs a research organization, Together Research, that publishes work spanning inference optimization, GPU kernels, model architecture and agents — for example FlashAttention-4, Mamba-3, cache-aware prefill–decode disaggregation, consistency diffusion language models, speculative decoding for reinforcement-learning rollouts, and agent benchmarks and datasets such as ParallelKernelBench, DSGym, EinsteinArena and CoderForge-Preview. Papers are co-authored with researchers affiliated with universities including Princeton, CMU, Stanford-affiliated authors, UC Berkeley, Seoul National University and Georgia Tech, and Together AI hosts an AI Native Conf and presents work at ICML.
Business model
Together AI sells cloud infrastructure and platform services for AI workloads, spanning on-demand and dedicated inference, GPU clusters, sandboxes, storage and fine-tuning, with self-serve sign-up alongside a sales-led route for larger deployments.
Sources describe usage- and capacity-based offerings: serverless, on-demand inference with no long-term commitments; batch inference for asynchronous workloads; provisioned throughput sold as committed capacity with token-based pricing; and dedicated model or container deployments.
Latest developments
Recent announcements on the company site include a partnership with Y Combinator to deliver a dedicated YC GPU cluster, on-demand B200 GPUs on Together GPU Clusters, serving of MiniMax-M3 for efficient inference, and product and research announcements made at the AI Native Conf.
▸Full profile — market position, technology, go-to-market
Market position
Together AI presents itself as a full-stack alternative to general-purpose clouds for AI workloads, competing on inference speed, cost and access to open-source models.
The company ties its product claims directly to in-house research output, citing 2x faster inference, 60% lower cost via workload-specific optimization and 90% faster pre-training with the Together Kernel Collection, and offers a single platform spanning serverless and dedicated inference, GPU clusters, storage, sandboxes and fine-tuning.
Technology
The platform is built around systems research on inference and training efficiency: the Together Kernel Collection for GPU kernel optimization, techniques such as FlashAttention-4, cache-aware prefill–decode disaggregation for long-context serving, distribution-aware speculative decoding, consistency diffusion language models, and looped/state-space model architectures (Mamba-3, Parcae). Serving includes recent open models such as MiniMax-M3, and compute includes NVIDIA B200 GPUs.", "go_to_market">
Go-to-market
Developers and AI teams running open-source models in production, including those with large-scale batch workloads, generative media (video, audio, image) workloads, agentic applications, and organizations needing multi-thousand-GPU training or pre-training capacity.
Compiled by commissioned research from 1 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Founder mafia
The Together AI mafia →6 people who came through Together AI went on to found or lead other companies.
Timeline · 3
launches, deals, and filingsTogether AI and Y Combinator announced a partnership to deliver the first dedicated Y Combinator GPU cluster.
The company announced on-demand availability of B200 GPUs within Together GPU Clusters.
Together AI announced it is now serving the MiniMax-M3 model for efficient inference.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
AI Startups That Raised Funding in August 2026: Big Roundsthebusinessperspective.com · Aug 2026▸Research sources · 1
primary sources listed
- Together AItogether.xyz · web
1 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Together AI do?
- Together AI operates a full-stack "AI native cloud" for inference, GPU compute and model fine-tuning, grounded in its own systems research.
- Who founded Together AI?
- Together AI was founded by Vipul Ved Prakash, Ce Zhang, Chris Ré, Percy Liang in 2022.
- Who are Together AI's investors?
- Together AI's investors include Lux Capital, SCB 10X, 137 Ventures, Aramco Ventures, Definition, Emergence Capital, Greycroft, Kleiner Perkins and 4 more.
- How much funding has Together AI raised?
- Together AI has disclosed $1.6B raised across 2 of its 3 known rounds.







