OctoAI
acquired by NVIDIASanta Clara, US · Founded 2019 · 5 known investors
Find your way into OctoAI
Sign up to find every way you can get introduced to OctoAI
- Your email: who you already know at OctoAI, and who your contacts know
- 2nd- and 3rd-degree connections: people you know who worked, studied or invested with the OctoAI team
- Your LinkedIn: the connections who can make the intro for you
We rank every route by how warm it is, so you raise on warm intros instead of cold emails.
Acquired by NVIDIA September 2024 · terms undisclosed — Reported $165M-$250M (unconfirmed by company) · source ↗
The page is NVIDIA's site (octo.ai now redirects to NVIDIA content), presenting AI computing platforms, GPU hardware, and related tools for developers, researchers, and industries such as media, manufacturing, healthcare, and supply chains. It references acquired company Deci and products spanning accelerated computing and quantum development.
Also known as Octo AI · OctoAI · OctoML
Founders & leadership
OctoAI was founded in 2019 by Jason Knight.
Board


Investors · 5
How we know: Amplify Partners's portfolio page · funding news · the VCSheet dataset · research · chatforest.com · Not right? Tell us
How we know: Madrona Venture Group's portfolio page · the VCSheet dataset · research · chatforest.com · Not right? Tell us
How we know: Madrona Venture Labs's portfolio page · Not right? Tell us
How we know: funding news · research · chatforest.com · Not right? Tell us
Reported raises · per SEC filings
Form D private placements$19.6M disclosed across 2 of 6 rounds · 2019–2021
▶$15.7MraisedApr 2020 · 3 investors · Other TechnologyRule 506(b)
- Luis CezeExecutive Officer, Director
- Jason KnightExecutive Officer, Director
- Matt McIlwainDirector
- Offering amount
- $15.7M
- Amount sold
- $15.7M
- Minimum investment
- $1
- First sale
- Mar 2020
- Incorporated
- Corporation, Delaware, 2019
- Federal exemptions
- 06b
▶$3.9MraisedOct 2019 · 14 investors · Other TechnologyRule 506(b)
- Luis CezeExecutive Officer, Director
- Matt McIlwainDirector
- Jason KnightDirector
- Offering amount
- $3.9M
- Amount sold
- $3.9M
- Minimum investment
- $1
- First sale
- Oct 2019
- Incorporated
- Corporation, Delaware, 2019
- Federal exemptions
- 06b
Source: SEC EDGAR Form D. Amounts as filed; amended filings shown once at their latest values.
Valuation · disclosed
Disclosed eventsSource: SEC prospectus filings, and round valuations the company or its investors disclosed — follow each entry's link for the claim.
Company profile
researched Aug 2026OctoAI, originally named OctoML, was founded in 2019 as a spinout of the Paul G. Allen School of Computer Science at the University of Washington in Seattle. The founding team comprised the creators of Apache TVM, an open-source deep learning compiler stack used to deploy neural networks across heterogeneous hardware including NVIDIA GPUs, AWS Inferentia, Intel CPUs, ARM processors and edge devices. The company's initial thesis was that machine learning models perform very differently depending on target hardware and that hand-tuning for each target is labor-intensive; OctoML aimed to commercialize TVM-based automated compilation and optimization so enterprises could deploy faster and cheaper models without in-house compiler engineering teams.
The first product, Octomizer, accepted ONNX, PyTorch or TensorFlow models, let users select a target (cloud GPU, edge device, specific CPU), and returned an optimized deployable artifact. It combined Apache TVM with ONNX Runtime, NVIDIA TensorRT and Intel OpenVINO, with claimed improvements of up to 10x over unoptimized deployment. In June 2023 the company launched OctoAI, described as a self-optimizing compute service for AI, and reoriented around hosting foundation models — Llama 2 (7B/13B/70B and chat variants), Code Llama Instruct, Mistral 7B Instruct, and Stable Diffusion XL with custom LoRA asset support — behind an OpenAI-compatible API, along with a bring-your-own-model option for fine-tuned Llama 2 variants, a Python SDK and REST API. A later product, OctoStack, was described as a toolkit for deploying, running and scaling generative AI models on a hardware-agnostic basis.
NVIDIA acquired OctoAI in late September 2024. On 26 September 2024 customers were emailed notice that commercial availability of the services would wind down and that access to all OctoAI services and accounts would terminate effective 31 October 2024, giving developers roughly five weeks to migrate. No successor product was offered; a migration guide pointed users to competing inference providers.
Founding story
OctoML was founded in 2019 out of the Paul G. Allen School of Computer Science at the University of Washington in Seattle by the creators of Apache TVM. Named founders include Luis Ceze (CEO), a UW computer science professor and ACM Fellow who retained his faculty position throughout the company's existence; Tianqi Chen (CTO), a UW PhD and creator of both Apache TVM and XGBoost; Jason Knight (CPO), covering product strategy and go-to-market; Jared Roesch (Chief Architect), a core TVM contributor working on compiler architecture; and Thierry Moreau (VP of Technology Partnerships), responsible for hardware integration and ecosystem work.
Business model
OctoAI sold AI infrastructure software and services to developers and enterprises: initially a model compilation and optimization service (Octomizer), later an API-based inference platform hosting open-source and customer-supplied foundation models, and the OctoStack toolkit for deploying and scaling generative AI models across hardware.
Sources describe a commercial developer platform accessed via API and SDK with a free sign-up tier; specific pricing terms are not stated in the sources.
Traction
Within 10 months of the OctoAI product launch, the platform reported more than 25,000 developers and over 100 high-growth business customers.
Latest developments
NVIDIA acquired OctoAI in late September 2024 in a deal one source describes as reportedly worth hundreds of millions of dollars. Customers were notified on 26 September 2024 that commercial services would wind down, with all service access and accounts terminated on 31 October 2024 and no successor product; a migration guide directed users to Fireworks AI, Together AI, Anyscale and Groq. The octo.ai domain now serves NVIDIA's corporate website.
▸Full profile — market position, technology, go-to-market, geography, history, risks & controversies
Market position
Positioned as a hardware-agnostic inference and model-optimization provider competing with other generative AI inference platforms. One source notes that the market for ML compilation tooling proved narrower than its funding implied, since enterprises with genuinely diverse hardware were mostly large technology companies with internal teams, while most other companies standardized on NVIDIA GPUs via cloud providers. Its hardware-agnostic technology was characterized as the strategic rationale for NVIDIA's acquisition.
Founder-level authorship of Apache TVM, hardware-agnostic optimization spanning NVIDIA, AMD, AWS Inferentia, Intel, Qualcomm and ARM targets, an OpenAI-compatible API with bring-your-own-model and custom LoRA support, and a developer-first tooling experience.
Technology
The core technology was Apache TVM, an open-source deep learning compiler stack created by the founding team that automatically optimizes neural network computation for diverse hardware targets. Octomizer layered TVM together with ONNX Runtime, NVIDIA TensorRT and Intel OpenVINO to benchmark and tune models, with claimed gains up to 10x versus unoptimized deployment. After the 2023 pivot, TVM-based optimization was applied server-side to hosted models exposed through an OpenAI-compatible API. Public repositories under the octoml GitHub organization include forks and implementations of mlc-llm, FlashInfer, EAGLE-1/EAGLE-2 and llama-recipes, plus SDK/documentation configuration, cookbooks and reference solutions.
Go-to-market
Developer-led self-service sign-up with free trial access, a Python SDK, REST API and public cookbooks and reference solutions on GitHub; enterprise sales to ML engineering teams; and partner-led distribution, including early AWS Partner status with joint Amazon EKS deployment guides and hardware partner announcements around the Series C.
Initially ML engineers and enterprises spending engineering cycles on deployment optimization; after the 2023 pivot, application and product developers integrating open-source foundation models without managing GPU infrastructure.
Geography
Headquartered in Seattle, Washington, where it spun out of the University of Washington; services were delivered as a cloud platform.
History
Founded in 2019 as OctoML, the company raised a seed round in October 2019 and Series A, B and C rounds through November 2021, reaching roughly $132 million raised at a post-Series C valuation of about $900 million. Its first phase (2019–2022) centered on the Octomizer model-optimization service and AWS partnership work. In June 2023 it pivoted to hosted generative AI inference under the OctoAI brand. NVIDIA acquired the company in late September 2024 and commercial services were terminated on 31 October 2024. Public GitHub activity under the octoml organization largely stops in September–October 2024.
Risks & controversies
The commercial platform was shut down roughly five weeks after the acquisition announcement, leaving customers with production applications a short migration window and no successor offering. Sources also note the underlying ML compilation market proved smaller than the company's funding and roughly $900 million post-Series C valuation implied. Note that several similarly named entities are unrelated to this company, including Octo, the federal IT modernization services firm founded by Mehul Sanghani, and consumer apps at octoai.space and octo-ai.app.
Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Timeline · 5
launches, deals, and filingsOctoAI's commercial platform access and customer accounts were deactivated on 31 October 2024, with no successor product offered.
On 26 September 2024 OctoAI emailed customers announcing the wind down of commercial availability of its services and termination of account access effective 31 October 2024, with a migration guide pointing to Fireworks AI, Together AI, Anyscale and Groq.
NVIDIA acquired OctoAI in late September 2024, in a deal reported to be worth hundreds of millions of dollars; the acquisition was attributed to OctoAI's hardware-agnostic inference technology.
OctoML launched OctoAI, a self-optimizing compute service for AI hosting foundation models such as Llama 2, Mistral, Code Llama and Stable Diffusion XL behind an OpenAI-compatible API.
OctoML was founded in 2019, spun out of the Paul G. Allen School of Computer Science at the University of Washington by the creators of Apache TVM.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
▸Research sources · 8
primary sources listed
- Octoocto.ai · web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does OctoAI do?
- OctoAI (formerly OctoML) was a University of Washington spinout selling generative AI inference infrastructure; NVIDIA acquired it in 2024.
- Who founded OctoAI?
- OctoAI was founded by Jason Knight in 2019.
- Who are OctoAI's investors?
- OctoAI's investors include Amplify Partners, Madrona Venture Group, Madrona Venture Labs, Addition, Tiger Global Management.
- How much funding has OctoAI raised?
- OctoAI has disclosed $19.6M raised across 2 of its 6 known rounds.
- Who acquired OctoAI?
- OctoAI was acquired by NVIDIA.
- Where is OctoAI headquartered?
- OctoAI is headquartered in Santa Clara, US.


