Luminal
YC S25San Francisco, US Β· Founded 2025 Β· 7 employees Β· Hiring Β· 11 known investors
Luminal provides AI model compilation and inference optimization software that compiles models into optimized native code for GPUs and ASICs, delivering 2-3x faster inference performance compared to traditional runtime inference engines.
Also known as Luminal AI Β· Luminal AI, Inc.
Founders & leadershipΒ· Y Combinator alumni (S25)
Luminal was founded in 2025 by Joe Fioti, Matthew Gunton, and Jake Stevens.
Investors Β· 11
Also in the syndicate Β· 8
Funding
SEC filings, press & company announcements- Undisclosed amountSeedNov 2025 Β· 3 sources
Felicis Ventures (lead), A Capital, Album VC, Ben Porterfield, Craft Ventures, Crosslink Capital, Dmitry Dakhnovsky, Gradient Ventures, Guillermo Rauch, Hack VC, Innovation Endeavors, Liquid 2 Ventures, Magnetic Ventures, Matt Garratt, Pareto Holdings, Paul Graham, Pelion Venture Partners, Ronny Conway, TSVC Capital, UpHonest Capital, Y Combinator
Source β
Source: company announcements and press reports β follow each round's link for the claim.
The YC applicationΒ· Summer 2025 batch
Alumni Q&A
- What is the core problem you are solving? Why is this a big problem? What made you decide to work on it?
- The founders observe that most companies running their own AI models spend more than necessary by not optimizing them, often paying thousands of dollars extra per month while prioritizing speed to market. Luminal aims to handle hardware and cloud infrastructure optimization so teams can focus on building models.
- What is your long-term vision? If you truly succeed, what will be different about the world?
- The founders aim to change how machine learning engineers work with the underlying compute layer for AI, arguing that current hardware outpaces the low-level software. They believe improving this software layer would reduce NVIDIA's advantage and broaden access to AI compute.
Summarized from the founders' answers on their Y Combinator profile.
Company profile
researched Aug 2026Luminal is a San Francisco-based AI infrastructure company building a compiler and runtime stack that makes machine learning models run faster on GPUs, CPUs and ASICs. Rather than interpreting models dynamically at runtime as conventional inference engines do, Luminal lowers a model (from PyTorch or Hugging Face) into a minimal graph intermediate representation β a pure dataflow graph β then applies hardware-aware fusion, tiling, memory-planning and scheduling passes before emitting GPU kernels or ASIC instructions ahead of time. The company describes this compiler-first design as eliminating runtime overhead and claims 2-3x higher throughput than existing inference engines on standard benchmarks.
The open-source core, published as the Rust crate luminal-ai/luminal, is described as a high-performance general-purpose inference compiler. Its architecture reduces all computation to 15 primitive operations (unary, binary, reduction and data-movement ops), models dynamic shapes as symbolic dimensions, and integrates with PyTorch as a torch.compile backend. Instead of hand-written heuristics or destructive rewrite rules of the kind used by XLA, torch.compile and TVM, Luminal treats optimization as a search problem: it generates large numbers of candidate graphs and kernels via rewrite rules and selects the fastest by measured runtime, an approach the company says can automatically rediscover complex optimizations such as Flash Attention. The repository documents native PyTorch support, integration of kernel libraries such as FlashInfer and cuBLASLt, low-precision dtypes (mxfp4, nvfp4, fp8), and multi-device parallelism topologies searched ahead of time.
On top of the compiler, Luminal markets a 'Hyperscale Inference OS' that schedules and load-balances inference workloads across heterogeneous clusters of CPUs, GPUs and ASICs, monitoring utilization, redistributing work in real time and booting or shutting down nodes as demand fluctuates.
Founding story
Co-founder Joe Fioti was working on chip design at Intel roughly three years before the November 2025 funding announcement when he concluded that the binding constraint on hardware adoption was software, not silicon: 'You can make the best hardware on earth, but if it's hard for developers to use, they're just not going to use it.' He founded Luminal in 2025 with Jake Stevens, previously at Apple, where he worked on iPhone imaging and had earlier founded a company with an exit and served as head of growth at a startup he grew to roughly $5M ARR, and Matthew Gunton, previously an engineer at Amazon working on systems that automatically detected and remediated issues in the global fulfillment and inventory network. Fioti's Intel work included AI accelerators and CPU microcode.
Business model
Luminal sells compute and inference capacity, positioning itself similarly to neo-cloud providers such as CoreWeave and Lambda Labs but differentiating through compiler-level optimization that extracts more performance from the infrastructure it operates. Its site offers two commercial paths: Luminal Cloud, a managed serverless inference service billed on usage with scale-to-zero and automatic batching; and on-prem or licensed cloud deployment with dedicated engineering support, custom kernel optimization and negotiated SLAs. The core compiler is distributed as an open-source project on GitHub under MIT/Apache licensing.
Usage-based pricing for managed serverless inference on Luminal Cloud ('pay only for what you use', scale to zero), and licensed on-prem or dedicated cloud deployments bundled with engineering support, custom kernel optimization and SLAs. The underlying compiler is open source.
Traction
The GitHub repository shows roughly 2.9k stars, 222 forks and 2,878 commits. The YC launch post states Luminal already powers research at Yale, production workloads at VC-backed startups and several research labs, and includes a runnable Llama 3 8B CUDA example. Vendor benchmarks claim Q8 Llama 3 8B at about 80% of theoretical maximum performance on an H100 and 36k tokens/sec on GPT-OSS 120B across 8xH100 SXM versus 28k for TensorRT-LLM and 26k for vLLM. Team size is listed at seven with three open engineering roles.
Latest developments
On November 17, 2025, Luminal announced a $5.3 million seed round led by Felicis Ventures with angel investment from Paul Graham, Guillermo Rauch and Ben Porterfield. A third-party aggregator lists total funding of $5.5M across two rounds, adding a $500K seed dated September 2025 with participation from A Capital, Album VC, Craft Ventures, Crosslink Capital, Gradient Ventures, Hack VC, Innovation Endeavors, Liquid 2 Ventures, Magnetic Ventures, Pareto Holdings, Pelion Venture Partners, TSVC Capital, UpHonest Capital, Y Combinator and several angels. The company is actively hiring compiler and cloud inference engineers in San Francisco.
βΈFull profile β market position, technology, go-to-market, geography, history, risks & controversies
Market position
TechCrunch places Luminal within a growing cohort of inference-optimization startups, alongside established inference providers such as Baseten and Together AI that have long specialized in optimization, and newer entrants including Tensormesh and Clarifai focused on narrower technical approaches. Its compute-selling business model is compared to neo-clouds such as CoreWeave and Lambda Labs, while its compiler competes conceptually with Nvidia's CUDA stack as well as XLA, torch.compile and TVM. Reported risks include competition from in-house optimization teams at large AI labs, which can tune for a single model family; Fioti argues the general-purpose case remains economically valuable even if six months of hand-tuning can beat a compiler.
Luminal's stated differentiators are ahead-of-time compilation instead of runtime interpretation; a search-based optimizer that explores millions of candidate graphs and kernels rather than relying on hand-tuned heuristics or pretrained LLMs, enabling automatic discovery of complex optimizations; a deliberately minimal RISC-style set of 15 primitive ops that keeps the core library small and portable; native Rust implementation that talks directly to accelerator APIs (CUDA, Metal) without containers, virtual environments or abstraction layers; hardware portability, since recompiling on new chips yields fresh optimal kernels; and drop-in PyTorch compatibility. TechCrunch frames the strategic bet as building out the software stack around largely open-source parts of CUDA at a time when GPUs remain scarce.
Technology
A Rust-based ahead-of-time inference compiler that lowers models to a minimal dataflow graph IR built from 15 primitive operations (Log2, Exp2, Sin, Sqrt, Recip; Add, Mul, Mod, LessThan; SumReduce, MaxReduce, Iota, Gather, Scatter, Cast), models dynamic shapes as symbolic dimensions, and applies fusion, tiling, memory planning and scheduling before emitting GPU kernels or ASIC instructions. Optimization is framed as a search over rewrite-rule-generated candidate graphs, benchmarked by measured runtime. It integrates with PyTorch as a compile backend (torch.compile with a luminal_cuda backend), interacts directly with CUDA and Metal APIs without containers or virtual environments, supports kernel libraries such as FlashInfer and cuBLASLt, low-precision dtypes (mxfp4, nvfp4, fp8) and multi-device parallelism topologies, and is validated against PyTorch for correctness. A separate scheduling layer load-balances inference across heterogeneous CPU/GPU/ASIC clusters with dynamic node provisioning.
Go-to-market
Go-to-market combines an open-source project (GitHub repository and Discord community) as a developer entry point with direct sales motions: a self-serve 'Get Started' path to Luminal Cloud and a 'Talk to an Engineer' path for licensed cloud or on-prem deployments. The YC launch post solicited AI researchers, ML engineers and companies running custom models directly via email and booked meetings, and highlighted the ability to reduce monthly compute bills. The website also offers early access to the platform.
AI researchers, ML engineers and companies that run their own custom models in production, including VC-backed startups and research labs; the YC launch post cites research at Yale and production workloads at VC-backed startups. Enterprise buyers wanting on-prem or licensed deployments with SLAs are a second segment.
Geography
Headquartered in San Francisco, California, with engineering roles advertised as San Francisco-based. Sources do not describe offices elsewhere.
History
The company was founded in 2025 by Joe Fioti (CEO), Jake Stevens and Matthew Gunton, and took part in Y Combinator's Summer 2025 batch, with Jared Friedman as primary partner. It published a YC launch post, 'Luminal - PyTorch for Production,' in July 2025. A third-party profile records a $500K seed round dated September 2025 with a broad syndicate of funds and angels. On November 17, 2025, Luminal announced a $5.3 million seed round led by Felicis Ventures with angel participation from Paul Graham, Guillermo Rauch and Ben Porterfield. As of the YC profile, the team numbered seven and the company was hiring compiler and cloud inference engineers in San Francisco.
Risks & controversies
Reported risk factors are competitive rather than legal or ethical: major AI labs run internal optimization teams that can tune for a single model family, while Luminal must adapt to arbitrary customer models; hand-tuning a model architecture over months can outperform any compiler. Performance figures such as the 2-3x and 3.2x-versus-vLLM claims and the GPT-OSS 120B benchmark are vendor-published and not independently verified in the sources. Funding totals also differ between sources ($5.3M announced seed per press coverage versus $5.5M across two rounds per an aggregator).
Compiled by commissioned research from 8 cited public sources β announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed β not independently audited.
Competitors Β· 3
by search overlapCompanies competing with Luminal for the same Google search keywords, organic and paid, via search-intersection analysis.
Timeline Β· 4
launches, deals, and filingsLuminal announced $5.3 million in seed funding led by Felicis Ventures, with angel investment from Paul Graham, Guillermo Rauch (Vercel CEO) and Ben Porterfield.
$5.3M source β
A $500K seed round listed with investors including A Capital, Album VC, Craft Ventures, Crosslink Capital, Gradient Ventures, Hack VC, Innovation Endeavors, Liquid 2 Ventures, Magnetic Ventures, Pareto Holdings, Pelion Venture Partners, TSVC Capital, UpHonest Capital, Y Combinator, and angels Matt Garratt, Dmitry Dakhnovsky and Ronny Conway.
$500K source β
Luminal published its YC launch post presenting an open-source ML compiler that generates CUDA kernels and offers one-line deployment, alongside serverless deployment with no idle costs or cold starts.
Luminal was part of Y Combinator's Summer 2025 batch; primary YC partner listed as Jared Friedman.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
βΈResearch sources Β· 8
primary sources listed
- Luminalluminal.com Β· web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Luminal do?
- YC-backed San Francisco startup building an open-source AI inference compiler and serverless inference cloud for GPUs and ASICs.
- Who founded Luminal?
- Luminal was founded by Joe Fioti, Matthew Gunton, Jake Stevens in 2025.
- Who are Luminal's investors?
- Luminal's investors include Felicis Ventures, Y Combinator, New Enterprise Associates (NEA).
- Where is Luminal headquartered?
- Luminal is headquartered in San Francisco, US.





