Fundraising Fox

Inferact

5 known investors

inferact.ai β†—

Inferact develops inference infrastructure for large language models, building on the open-source vLLM engine to make model serving cheaper and faster. The company serves AI teams deploying models at scale, from research labs to hyperscalers.

Also known as vLLM (open-source project commercialized by Inferact)

Founders & leadership

RW
Roger WangFounder
WK
Woosuk KwonFounder

Investors Β· 5

Funding

SEC filings, press & company announcements

Source: company announcements and press reports β€” follow each round's link for the claim.

Valuation Β· disclosed

Disclosed events
$800Mvaluation at SeedJan 2026
filing β†—

Source: SEC prospectus filings, and round valuations the company or its investors disclosed β€” follow each entry's link for the claim.

Company profile

researched Aug 2026

Inferact is an AI infrastructure company founded by the creators and core maintainers of vLLM, an open-source large language model inference engine. Its stated mission is to grow vLLM into the world's AI inference engine and to make inference cheaper and faster. The company frames inference as an unsolved and worsening problem: models are growing larger, architectures are proliferating (mixture-of-experts, multimodal, agentic), and hardware is fragmenting across more accelerators and programming models, widening the gap between model capability and the systems that serve them. Inferact argues inference is shifting from a fraction of compute to the majority through test-time compute, reinforcement-learning training loops and synthetic data generation.

The company positions itself at the intersection of model innovation and hardware diversity, working with model vendors for day-zero support of new architectures and with hardware vendors integrating new silicon. Its long-term goal is to make deploying a frontier model at scale as simple as spinning up a serverless database, absorbing infrastructure complexity into its platform rather than requiring dedicated infrastructure teams. Sequoia Capital describes Inferact as supporting 200+ model architectures across major accelerator platforms with day-zero compatibility, while the company's own site cites vLLM support for 500+ model architectures and 200+ accelerator types.

Inferact was founded in 2025 and is based in San Francisco, California. Named founding members include Simon Mo (CEO), Woosuk Kwon, Kaichao You, Roger Wang, Joseph Gonzalez and Ion Stoica. In January 2026 it announced a $150 million seed round at an $800 million valuation.

Founding story

The founding team consists of the creators and core maintainers of vLLM, the open-source LLM inference engine incubated in 2023 at Ion Stoica's UC Berkeley lab. Founding members listed on the company site are Simon Mo, Woosuk Kwon, Kaichao You, Roger Wang, Joseph Gonzalez and Ion Stoica. The team says it has maintained the engine since its first commit and deployed it at frontier scale in both research and production, and started Inferact to build a universal inference layer that can run any model on any chip.

Business model

Inferact commercializes the open-source vLLM inference engine, building enterprise inference systems and a platform for serving models across hardware while contributing optimizations back to the open-source project. Sources state funding will be used to commercialize vLLM, scale the platform for enterprise AI inference, and expand adoption among existing and new customers; specific pricing or contract structures are not disclosed in the sources.

Not disclosed in the sources beyond an intent to commercialize vLLM and scale a platform for enterprise AI inference.

Traction

Traction is primarily ecosystem-based: vLLM is described as the most popular open-source LLM inference engine, with 2,000+ contributors, 500+ supported model architectures and 200+ supported accelerator types, and is said to power inference at global scale for frontier labs, hyperscalers and startups. Reported users of vLLM include Amazon's cloud service and shopping app. Investor validation includes a $150 million seed at an $800 million valuation.

Latest developments

In January 2026 Inferact announced its $150 million seed round at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed Venture Partners with Sequoia Capital, Altimeter Capital, Redpoint Ventures and ZhenFund participating, and appeared on the a16z podcast to describe its universal inference layer vision. A company news entry dated August 12, 2026 announces full certification of Kimi K3 vendor verification and the availability of its production inference stack.

β–ΈFull profile β€” market position, technology, go-to-market, geography, history, risks & controversies

Market position

Inferact is presented as the commercial vehicle behind vLLM, described in sources as the most popular and de facto standard open-source LLM inference engine. TechCrunch frames its debut as mirroring the commercialization of the SGLang project as RadixArk, which reportedly raised at a $400 million valuation led by Accel, and notes investor interest in inference technologies as AI attention shifts from training to deployment.

Inferact's differentiation rests on the founding team's stewardship of vLLM since its first commit, the project's position at the intersection of model vendors and hardware vendors (day-zero model support and silicon integrations), and a commitment to keeping inference infrastructure open source rather than behind proprietary walls.

Technology

The core technology is vLLM, an open-source LLM inference engine that the company says supports 500+ model architectures, runs on 200+ accelerator types, and is built by 2,000+ contributors. Inferact works with model vendors for day-zero support of new architectures such as mixture-of-experts, multimodal and agentic models, and with hardware vendors to cover new accelerators. The stated aim is a universal, open-source inference layer that runs any model on any chip across deployment environments, with performance optimizations contributed upstream.

Go-to-market

Inferact's go-to-market is anchored in the existing vLLM open-source community and its adoption by frontier labs, hyperscalers and startups; the company says it exists to supercharge vLLM adoption and will flow its optimizations back to the community. It also engages via investor-hosted media such as the a16z podcast and is recruiting engineers and researchers.

AI teams deploying models in production, including frontier research labs, hyperscalers and cloud providers, and startups serving large user bases. TechCrunch reports CEO Simon Mo told Bloomberg that existing vLLM users include Amazon's cloud service and the shopping app.

Geography

Inferact is described as San Francisco, California-based. The sources do not detail additional offices or international operations.

History

vLLM was incubated in 2023 at the UC Berkeley lab of Databricks co-founder Ion Stoica, alongside the SGLang project. The project is described as currently managed by the PyTorch Foundation. Inferact was founded in 2025 by vLLM's creators and core maintainers, who state they have been stewards of the engine since its first commit. Sequoia lists both founding and partnering in 2025. The company publicly announced itself and a $150 million seed round on January 22, 2026, the same day a16z published a podcast episode with cofounders Woosuk Kwon and Simon Mo. A company news item dated August 12, 2026 announces full certification of Kimi K3 vendor verification.

Risks & controversies

The sources do not report controversies. Contextual risks noted include competition from similarly positioned commercializations of open-source inference projects such as RadixArk (SGLang), and the tension inherent in monetizing an open-source project that the company pledges to keep open and community-driven.

Compiled by commissioned research from 8 cited public sources β€” announcements, filings, and press listed under research sources below.

Key figures

latest reported
Accelerator types supported by vLLMJan 2026200 accelerator types (200+)
Model architectures supported (per Sequoia)Jan 2026200 architectures (200+)
Model architectures supported by vLLMJan 2026500 architectures (500+)
Valuation (post-money, seed round)Jan 2026$800M
VLLM open-source contributorsJan 20262,000 contributors (2,000+)

Company-reported or press-reported figures, each dated to when it was claimed β€” not independently audited.

Competitors Β· 5

by search overlap
NVIDIA5 shared keywordsNVIDIA designs GPUs, AI computing hardware, and software platforms for AI development, data centers, autonomous vehicles, and robotics, serving developers, researchers, and enterprises. Its offerings include AI models, power architecture for AI factories, and compute infrastructure used across industries such as manufacturing, healthcare, and automotive.
Infer4 shared keywordsInfer provides software infrastructure for insurance agencies to manage their operations and client data.
Cloudfactory3 shared keywordsCloudFactory provides a platform and consulting services to develop, deploy, and operate AI systems in production, combining quality data preparation, model fine-tuning, human validation, and inference oversight. It serves enterprises in high-stakes sectors such as healthcare, finance, robotics/embodied AI, retail, logistics, manufacturing, and agriculture.
Inception3 shared keywordsInception Labs develops diffusion-based large language models (dLLMs) designed for production applications. Their Mercury model family offers faster inference speeds than traditional autoregressive LLMs while maintaining quality, targeting developers and enterprises building AI systems.
Hugging Face3 shared keywordsHugging Face is a collaboration platform that hosts and provides access to machine learning models, datasets, and applications. It offers both open-source tools for the ML community and paid compute and enterprise solutions for teams building AI applications.

Companies competing with Inferact for the same Google search keywords, organic and paid, via search-intersection analysis.

Timeline Β· 4

launches, deals, and filings
Aug 2026
Inferact announces full certification of Kimi K3 vendor verification

Company news item stating Inferact announced full certification of Kimi K3 vendor verification, bringing its production inference stack to everyone.

source β†—

Jan 2026
Inferact raises $150M seed at $800M valuation to commercialize vLLM

Inferact announced its transition of the open-source vLLM project into a VC-backed company, raising $150 million in seed funding at an $800 million valuation, co-led by Andreessen Horowitz and Lightspeed Venture Partners with participation from Sequoia Capital, Altimeter Capital, Redpoint Ventures and ZhenFund. Funds are earmarked for commercializing vLLM, scaling the enterprise inference platform and expanding customer adoption.

$150M source β†—

Jan 2026
Founders discuss Inferact and a universal inference layer on a16z podcast

a16z published a podcast episode with general partner Matt Bornstein and Inferact cofounders Woosuk Kwon and Simon Mo covering how models are run in production, the origins of vLLM, and Inferact's vision for a universal open-source inference layer across hardware and model architectures.

source β†—

Jan 2025
Sequoia Capital partners with Inferact

Sequoia Capital's company page lists Inferact as founded in 2025 and 'Partnered 2025', with Lauren Reeder as the responsible partner.

source β†—

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

In the news

β–ΈResearch sources Β· 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Inferact do?
Inferact, founded by the creators of open-source vLLM, builds inference infrastructure to serve large AI models faster and cheaper.
Who founded Inferact?
Inferact was founded by Roger Wang, Woosuk Kwon.
Who are Inferact's investors?
Inferact's investors include Andreessen Horowitz, Redpoint Ventures, Sequoia Capital, Striker Venture Partners, Laude Ventures.