Fundraising Fox

TLDC

1 known investors

Find your way into TLDC

10 people in our graph share verified history with the TLDC team — schools, employers, funds. One of them is your warm intro.

Samir Kaulunlockedknows Daanish Khazi · together at Traba (overlapped)
×2knows the team · via Paper Instruments
×7knows the team · via Traba

Paper Instruments develops frontier AI models and tooling for knowledge work, including agent-first office document libraries and models tailored for white-collar professional services.

Also known as Paper Instruments · The LLM Data Company · TLDC

Founders & leadership

DKDaanish Khazi
Daanish Khaziin𝕏FounderCo-Founder at Paper Instruments

Investors · 1

Company profile

researched Aug 2026

TLDC (The LLM Data Company, also listed as Paper Instruments) is a San Francisco company founded in 2025 that builds evaluation tooling for language models and agents and post-trains models for knowledge work. Its initial product, doteval, is a workspace for authoring, versioning and executing evaluations as code against a YAML schema. It provides an editor-style interface comparable to a code IDE, AI-generated grading diffs in place of manual scoring, side-by-side comparison of eval runs across model checkpoints and prompts, and fine-grained rubrics with aligned graders. Eval specifications can be exported in one click for use as reward datasets in GRPO-style reinforcement learning or reinforcement fine-tuning, linking measurement directly to post-training.

The company frames the problem as the unreliability of generic LLM judges and manual, spreadsheet- or JSON-based evaluation workflows, which leave teams unable to tell whether a model upgrade or prompt change is a net improvement, and which fail to supply the well-specified reward datasets that modern RL techniques require. It has worked with frontier AI teams to benchmark complex model tasks, and secondary coverage reports use in legal, technical and safety-critical workflows.

Under the Paper Instruments name, the company also trains and releases domain models. In March 2026 it announced Kos-1 Lite, a medical reasoning model reported at 46.6% on HealthBench Hard and 66.6% on HealthBench, described as prioritizing clinical accuracy, reduced sycophancy, triage and deferral, and bedside manner, and as serving at a fraction of the cost of models above one trillion parameters.

Founding story

Founded in 2025 by Daanish Khazi (Founder/CEO), Gavin Bains and Joseph Besgen. Secondary coverage describes all three as former Traba engineers or employees, with prior experience at Tesla and Meta (Khazi), and Honey and Roland Berger (Besgen).

Business model

Developer-first software sold on a subscription basis, targeted at applied AI teams, alongside paid engagement work benchmarking performance for frontier AI teams and building evaluation datasets. Secondary coverage describes direct sales into mid-market enterprises with evaluation needs.

Subscription-based developer platform, per secondary coverage.

Traction

Works with frontier AI teams on benchmarking complex model tasks; secondary coverage from July 2025 reports usage by Perplexity, Diode and Cubic in legal, technical and safety-critical workflows. Team size listed at three. Its Kos-1 Lite model reported the top HealthBench Hard score among the compared models at 46.6%.

Latest developments

On 3 March 2026 the company announced Kos-1 Lite, its first medical reasoning model, reporting a state-of-the-art 46.6% on HealthBench Hard and 66.6% on HealthBench, trailing only GPT-5 High on the latter, and made the model available to try interactively.

Full profile — market position, technology, go-to-market, geography, history, risks & controversies

Market position

Positioned in AI infrastructure and evaluation tooling. Secondary coverage names Pi Labs and Haize Labs as competitors, along with internally built lab tooling, and characterizes the company as early-stage with early production usage among named AI companies.

The company positions its tooling as evaluation-as-infrastructure rather than a results viewer: structured, versioned, reusable eval specifications that can be authored collaboratively across engineering, product and legal, and then reused directly as reinforcement learning reward signals. For its medical model, it argues that general-purpose LLMs trade off groundedness and non-sycophancy for instruction-following on tasks such as coding, and that a domain-specific reasoning model can reach higher clinical benchmark scores at lower serving cost than much larger frontier models.

Technology

Evals-as-code on a YAML schema with versioning across model checkpoints, AI-generated diffs, aligned graders and fine-grained rubrics, plus export of eval specs as reward datasets for GRPO-style reinforcement learning and reinforcement fine-tuning. The company also post-trains large language models, releasing Kos-1 Lite, a reasoning model for clinical use evaluated on OpenAI's HealthBench, a benchmark built with 262 physicians across 60 countries using open-ended conversations.

Go-to-market

Direct engagement with applied AI and frontier model teams, including early access to doteval, help building evaluation datasets, and support for teams adopting GRPO or reinforcement fine-tuning; inbound via a founders contact address. Model releases are announced through the company's research blog with an interactive demo.

Teams building, fine-tuning or operating LLMs at scale, including frontier AI labs, product organizations choosing between models, AI startups building models for regulated industries, and research and open-source teams; secondary coverage cites mid-market enterprises and applied AI teams, with reported users including Perplexity, Diode and Cubic.

Geography

Headquartered in San Francisco, California.

History

The company was founded in 2025 in San Francisco by Daanish Khazi, Gavin Bains and Joseph Besgen and joined a 2025 Y Combinator batch (listed as Spring 2025 by YC, referenced as S25 elsewhere), with Diana Hu as primary partner. It launched publicly as The LLM Data Company with the doteval eval workspace and worked with frontier AI teams on benchmarking complex model tasks. The entity is also listed under the name Paper Instruments, described as post-training LLMs to accelerate knowledge work with open-source frontier models. In March 2026 it announced Kos-1 Lite, a medical reasoning model.

Risks & controversies

Secondary coverage flags early-stage go-to-market execution, heterogeneous customer requirements across domains, competition from evaluation tooling built internally at large labs, and the tension between flexibility and simplicity in product design.

Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
HealthBench Hard score (Kos-1 Lite)Mar 202646.6%
HealthBench score (Kos-1 Lite)Mar 202666.6%
Team sizeJan 20253 people

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Competitors · 9

by search overlap
Hugging Face9 shared keywordsHugging Face is a collaboration platform that hosts and provides access to machine learning models, datasets, and applications. It offers both open-source tools for the ML community and paid compute and enterprise solutions for teams building AI applications.
Atlan6 shared keywordsAtlan builds a "context layer" and metadata management platform, including a data catalog, data governance tools, and a Context Lakehouse, that supplies enterprise data context to AI agents and data teams. It serves large enterprises looking to make AI systems operate with knowledge of their data environment.
Netezza6 shared keywordsIBM is a global technology company whose business spans enterprise software (including Red Hat, HashiCorp, and Confluent), IT infrastructure such as mainframes, servers, and storage, and IT consulting services. The company is also investing heavily in quantum computing and AI-based enterprise offerings, including its Lightwell open-source software security clearinghouse and the Anderon quantum wafer foundry.
Databricks5 shared keywordsDatabricks provides a unified platform for data engineering, analytics, and artificial intelligence, enabling organizations to build and deploy data and AI applications on cloud infrastructure.
DataCamp5 shared keywordsDataCamp is a learning platform for teams that teaches data and AI skills through hands-on coursework accessible via web browser and mobile app. It serves enterprise customers and development teams seeking to build technical capabilities.
Herald4 shared keywordsHerald is an AI-powered incident detection and root cause analysis platform that automatically identifies anomalies in software systems and investigates issues before they impact customers. It integrates with existing observability and infrastructure tools to provide rapid diagnosis without requiring manual runbook knowledge.
NVIDIA4 shared keywordsNVIDIA designs GPUs, AI computing hardware, and software platforms for AI development, data centers, autonomous vehicles, and robotics, serving developers, researchers, and enterprises. Its offerings include AI models, power architecture for AI factories, and compute infrastructure used across industries such as manufacturing, healthcare, and automotive.
Uniphore4 shared keywordsUniphore provides a Business AI Cloud platform that enables enterprises to build, deploy, and scale autonomous AI agents across operations with built-in governance, security, and sovereignty controls. The platform serves organizations across financial services, insurance, healthcare, and telecommunications, helping them automate back-office and front-office workflows while maintaining compliance and avoiding vendor lock-in.
LiteLLM4 shared keywordsLiteLLM provides an AI gateway that simplifies access to over 100 large language models through a unified OpenAI-compatible interface, with built-in features for fallbacks, spend tracking, and load balancing. The platform serves developers and engineering teams integrating multiple LLM providers into their applications.

Companies competing with TLDC for the same Google search keywords, organic and paid, via search-intersection analysis.

Timeline · 3

launches, deals, and filings
Mar 2026
Launch of Kos-1 Lite medical reasoning model

Paper Instruments announced Kos-1 Lite, its first medical reasoning model, reporting 46.6% on HealthBench Hard and 66.6% on HealthBench, with an interactive demo. The company positions the model around clinical accuracy, reduced sycophancy, triage and deferral behavior, and lower serving cost than trillion-parameter frontier models.

source ↗

Jan 2025
Participation in Y Combinator batch

The company took part in a 2025 Y Combinator batch (listed as Spring 2025 on the YC company page and as S25 in secondary coverage), with Diana Hu as primary partner.

source ↗

Jan 2025
Launch of doteval eval workspace

The company launched doteval, a workspace for writing, versioning and executing evals-as-code against a YAML schema, with AI-generated diffs, run comparison across checkpoints, and one-click export of eval specs as reinforcement learning training sets.

source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

Research sources · 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does TLDC do?
San Francisco AI company building eval tooling for LLMs and agents and post-training domain models for knowledge work.
Who founded TLDC?
TLDC was founded by Daanish Khazi.
Who are TLDC's investors?
TLDC's investors include NextGen Venture Partners.