TrainLoop
YC W25San Francisco, US · Founded 2025 · 6 employees · 3 known investors
Find your way into TrainLoop
83 people in our graph share verified history with the TrainLoop team — schools, employers, funds. One of them is your warm intro.
TrainLoop is a post-training research lab that trains specialized AI models for enterprise customers in pharma, biotech, logistics, and banking. The company develops custom foundation models and agents optimized for long-horizon tasks and mission-critical workflows in high-risk industries.
Also known as Phoenix · Trainloop AI · TrainLoop AI
Founders & leadership· Y Combinator alumni (W25)
TrainLoop was founded in 2025 by Jackson Stokes and Mason Pierce.


Investors · 3
Company profile
researched Aug 2026TrainLoop is a San Francisco-based post-training research and product lab that builds specialized language and multimodal models for long-horizon tasks. It works with pharmaceutical, biotech, logistics and banking enterprises to convert domain expertise and proprietary customer data into custom models, pairing a research team studying long-horizon post-training methods with a deployment team that works alongside customer subject-matter experts.
The company's stated model areas are biological reasoning models for pharmaceutical and biotech customers (aimed at diagnosis and treatment-response work in drug development), reliable agents for high-risk industries such as financial services and logistics where decision accuracy in mission-critical workflows matters, and purpose-built multimodal reasoning systems for complex image understanding and document abstraction. Its four stated research directions are life sciences, continual training (methods that avoid catastrophic forgetting), information theory (capacity-aware objectives for stable, interpretable reasoning), and evaluation and interpretability.
At its Y Combinator launch the company positioned itself as "Reasoning Fine-Tuning": a platform making reinforcement learning-based fine-tuning accessible to developers, structured as data curation via a lightweight SDK that gathers training signals from production usage, training of a reward model that teaches an LLM the preferred outputs, and deployment of the resulting model through standard APIs. Third-party directory coverage describes the same workflow, including data collection, custom model deployment and data security.
Founding story
TrainLoop was founded in 2025 by Jackson Stokes (Founder/CEO) and Mason Pierce (Founder/CTO). Stokes previously worked on performance for Google's Gemini models and AlphaFold, with work described as supporting AI search summaries and systems including Waymo; Pierce led engineering at Second (YC W23), where he worked on large-scale enterprise codebase migrations and retrieval-augmented generation systems and encountered the difficulty of fine-tuning off-the-shelf models. The founders framed the company around the gap between internal model-training tooling at large labs such as Google and OpenAI and what is available to developers deploying models in production.
Business model
B2B. TrainLoop engages enterprises in a structured research-to-production collaboration: jointly defining research objectives from the customer's proprietary data and technical strengths, advancing models, and sustaining their performance in production. Earlier positioning centered on a self-serve developer platform with an SDK, managed training and API-based inference.
Traction
The company reports partnerships with NollaMD (differential diagnosis in visual medicine), Mercor and Pathos, and states that its models frequently achieve state-of-the-art or Pareto-optimal performance on their target tasks. Team size is listed as 6 in the Y Combinator profile.
Latest developments
TrainLoop's public positioning has shifted from a developer self-serve reasoning fine-tuning platform toward an enterprise post-training research and product lab, with recent partnership write-ups covering NollaMD, Mercor and Pathos and research notes on one-step model training with GRPO and the low-rank nature of learning GSM8K.
▸Full profile — market position, technology, go-to-market, geography, history, risks & controversies
Market position
Positioned in the LLM fine-tuning, post-training and model-customization segment alongside AI infrastructure, LLMOps and RLHF platforms, differentiating on domain-specific expert models, lightweight SDK integration and rapid deployment. It is categorized by third-party directories under AI observability, AI testing, MLOps and enterprise software.
The company emphasizes reinforcement-learning-driven fine-tuning that uses real production usage data and reward modeling rather than prompt engineering or basic supervised fine-tuning, delivered as a managed end-to-end workflow spanning data collection, training, deployment and data security, with models specialized to individual customer tasks.
Technology
Post-training methods including reinforcement learning (with Group Relative Policy Optimization cited as a frequent starting point), reward modeling built from real production usage signals collected via a lightweight SDK, and LoRA fine-tuning. Research directions cover continual training methods that avoid catastrophic forgetting, capacity-aware information-theoretic objectives for stable and interpretable reasoning, and tools for evaluating both external model behavior and internal representations. Delivered systems include multimodal reasoning models for image understanding and document abstraction, and agents for mission-critical workflows.
Go-to-market
Direct enterprise engagement via an "schedule an intro" consultation funnel on the company site, supported by published research notes and case-study-style partnership write-ups. The company was introduced to the developer market through a Y Combinator launch post and an alpha sign-up program.
Large enterprises in pharmaceuticals, biotechnology, logistics, banking and financial services; earlier positioning targeted developers and engineering teams deploying LLMs in production and needing domain-specific reliability for tasks such as code generation, compliance, legal and healthcare.
Geography
Headquartered in San Francisco, California, United States.
History
The company was founded in 2025 and participated in Y Combinator's Winter 2025 batch, launching publicly as "TrainLoop: Unlock Next-Level Reasoning through Fine-Tuning" with an alpha program and a developer-focused reinforcement learning fine-tuning platform. Its current public positioning is broader: a post-training research and product lab serving pharma, biotech, logistics and banking enterprises with custom expert models, supported by published research notes on topics such as Group Relative Policy Optimization and LoRA training dynamics, and named partnerships with NollaMD, Mercor and Pathos.
Risks & controversies
Specific customers beyond the named partnerships have not been publicly disclosed, and performance claims such as state-of-the-art or Pareto-optimal results are company-stated. Third-party directory listings publish revenue, valuation and funding figures that are explicitly labeled as estimates derived from industry averages.
Compiled by commissioned research from 6 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Competitors · 10
by search overlapCompanies competing with TrainLoop for the same Google search keywords, organic and paid, via search-intersection analysis.
Timeline · 5
launches, deals, and filingsTrainLoop is listed as an active Y Combinator company in the Winter 2025 batch, based in San Francisco, with David Lieb as primary partner.
Listed as a recent partnership, accompanying a write-up titled "Your knowledge work agent should be a coding agent".
Listed as a recent partnership, described as a new state of the art for differential diagnosis in visual medicine.
Listed as a recent partnership, accompanying a research note titled "Can We Train a Model in One Step?" on Group Relative Policy Optimization.
TrainLoop published a Y Combinator launch post introducing a reinforcement-learning fine-tuning platform with a three-line SDK for data curation, reward-model training and API-based deployment, and opened an alpha program.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
▸Research sources · 6
primary sources listed
- TrainLooptrainloop.ai · web
6 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does TrainLoop do?
- TrainLoop is a San Francisco post-training research and product lab that trains specialized AI models for long-horizon enterprise tasks.
- Who founded TrainLoop?
- TrainLoop was founded by Jackson Stokes, Mason Pierce in 2025.
- Who are TrainLoop's investors?
- TrainLoop's investors include Moonfire Ventures, Olive Tree Capital, Y Combinator.
- Where is TrainLoop headquartered?
- TrainLoop is headquartered in San Francisco, US.












