Fundraising Fox

Besimple AI

YC X25

San Francisco, US · Founded 2025 · 6 employees · Hiring · 7 known investors

BeSimple collects, licenses, and annotates high-quality conversational audio datasets across 15+ languages for training and evaluating audio and multi-modal AI models. The company serves AI teams building speech recognition, voice, and conversational AI applications.

Also known as Simple Annotation · Besimple · BeSimple AI

Founders & leadership· Y Combinator alumni (X25)

Besimple AI was founded in 2025 by Yi Zhong and Bill Wang.

YZYi Zhong
Yi ZhonginCo-FounderYi Zhong held product leadership roles at Meta, Microsoft, and Dropbox, working on large-scale AI systems deployment.
BWBill Wang
Bill WanginCo-FounderBill Wang worked at Meta, where he launched multiple products and led the GenAI Annotation team developing infrastructure for LLaMa training. Earlier at Meta, he managed engineering teams focused on connectivity improvements and SMS optimization across the platform.

Investors · 7

Also in the syndicate · 4

Multimodal VenturesPorterfield VenturesSurgepoint CapitalYCombinatorlead

Funding

SEC filings, press & company announcements
  • Undisclosed amountseed roundDec 2025

    YCombinator (lead)

    Source ↗

Source: company announcements and press reports — follow each round's link for the claim.

Company profile

researched Aug 2026

Besimple AI is a Y Combinator-backed company (Spring 2025 / X25 batch) building what it describes as the data layer for AI, starting with audio. It curates a proprietary set of conversational audio spanning a wide range of languages, dialects and accents, then uses human expert audio annotators together with its own annotation platform to process that audio for automatic speech recognition, including transcription and speaker diarization. The company states it holds over millions of hours of conversational data, sourced through a global network of independent contributors covering 15+ languages and diverse accents, and can run custom collections to a customer's specification, including role-plays and domain-specific conversations.

Alongside the dataset business, the company has offered a self-serve annotation product that generates a tailored annotation platform from raw data pasted or streamed in by the customer. That product supports text, chat, audio, video and LLM traces, drafts or imports annotation guidelines, provides an automated human-in-the-loop workflow, and uses LLM-based "AI judges" that learn from incoming annotations to evaluate live traffic and flag borderline cases for human review. Deployment options include on-premise installation and user management for internal subject-matter experts, external vendors or Besimple's own vetted annotators.

Delivery follows a defined onboarding sequence: a scoping conversation about hours, languages and scenarios; sample delivery within 48 hours for quality and metadata review; validation of samples in the customer's own training pipeline; production access to full datasets via API or S3; and scaling of annotation capacity from roughly 10 to 100+ annotators with monthly dataset expansions. The company also publishes audio-model benchmarks, including Vocal Affect Bench and Voice Code Bench.

Founding story

Founded in 2025 by Yi Zhong and Bill Wang, who met the problem while at Meta, where they spent years building data infrastructure and the annotation platform for the Llama team. Zhong is described as an AI product leader with prior roles at Meta, Microsoft and Dropbox; Wang previously led Meta's GenAI Annotation team, developing an in-house annotation platform for Llama training, and earlier managed an engineering organization working on connectivity for over 300 million users and Meta's SMS spend optimization.

Business model

B2B. Besimple AI sells access to conversational audio datasets and annotation services to AI teams, offering flexible licensing arrangements it describes as suited to both startups and enterprises, with delivery via API or S3 and optional custom collection and annotation work. It has also offered a self-serve annotation platform product with an enterprise and on-premise deployment option.

Data licensing and paid annotation/data-collection services, with licensing deals structured for startups and enterprises; an annotation platform product is also offered, including a waitlist-based self-serve path.

Traction

The company reports over millions of hours of conversational audio data and a global contributor network covering 15+ languages. Edexia, an AI grading company, is cited as a user of the annotation product. Team size is reported at six.

Latest developments

Announced a $3M seed round on 2025-11-26 funded by Y Combinator and several venture firms and angels, stated it was hiring, and published audio-model benchmarks including Voice Code Bench (2026-05-01) and Vocal Affect Bench (2026-07-07). Open roles include an audio QA lead contractor, a strategic projects lead for audio data, and a mobile engineer.

Full profile — market position, technology, go-to-market, geography, history

Market position

Positioned as a licensed, ethically sourced alternative to scraping unlicensed audio and to multi-month internal build-outs of acquisition and annotation infrastructure, emphasizing 48-hour sample turnaround and continuous dataset expansion. The founders frame the product as letting teams "spin up a Scale AI in 60 seconds."

Licensed and ethically sourced audio rather than scraped material, sample delivery within 48 hours instead of multi-month legal negotiation, a vetted expert annotator network, proprietary collection and annotation tooling, and founder experience building the annotation platform used for Meta's Llama models.

Technology

Proprietary audio data collection and annotation tooling: a contributor collection platform, an in-house annotation platform producing human-level transcription and speaker diarization for ASR, automatically generated task-specific annotation interfaces, guideline generation, human-in-the-loop workflow automation, and LLM-based AI judges for real-time evaluation of live traffic.

Go-to-market

Direct sales through demo bookings and founder-led outreach, a self-serve waitlist on the company website, distribution via the Y Combinator launch platform, and content in the form of published audio-model benchmarks.

AI labs, startups and enterprises training or evaluating speech recognition, voice and multimodal models and voice agents; also teams needing annotation infrastructure for text, chat, audio, video and LLM traces. Named user Edexia, an AI grading company, used Besimple to annotate decisions and improve its evaluations.

Geography

Headquartered in San Francisco per its Y Combinator profile, with job listings based in San Mateo, California (some remote in the US) and secondary reporting placing the team between Redwood City and San Francisco. Data collection relies on a global network of independent contributors covering 15+ languages, dialects and accents.

History

The company was founded in 2025 and participated in Y Combinator's Spring 2025 (X25) batch, with Nicolas Dessaigne as primary partner. Its YC launch introduced a product for generating a custom data annotation platform in about 60 seconds. On 2025-11-26 it announced a $3M seed round funded by Y Combinator, Surgepoint Capital, Porterfield Ventures, Amino Capital, WELIGHT Capital, Multimodal Ventures, Script Capital and angel investors. Its public positioning subsequently centered on conversational audio datasets for voice AI, and it published the Voice Code Bench and Vocal Affect Bench benchmarks.

Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Annotation capacity scalingJan 202610 to 100+ annotators
Conversational audio data volumeNov 2025over millions of hours
HeadcountAug 202615
Languages covered by contributor networkJan 202615+
Sample delivery timeJan 202648 hours
Team sizeJan 20266 employees

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Images

Besimple AI photoBesimple AI photo

Timeline · 5

launches, deals, and filings
Jul 2026
Vocal Affect Bench benchmark published

source ↗

May 2026
Voice Code Bench benchmark published

source ↗

Nov 2025
Besimple AI announces $3M seed round

The company announced it had raised $3M to build the data layer for AI, starting with audio, with funding from Y Combinator, Surgepoint Capital, Porterfield Ventures, Amino Capital, WELIGHT Capital, Multimodal Ventures, Script Capital and angel investors.

$3M source ↗

Jan 2025
Participation in Y Combinator Spring 2025 (X25) batch

Besimple AI joined Y Combinator's Spring 2025 batch, with Nicolas Dessaigne as primary partner.

source ↗

Jan 2025
YC launch of 60-second annotation platform builder

Launch post introducing a product that generates a custom annotation platform from pasted or streamed data, with auto-generated UI, guidelines, human-in-the-loop workflow and LLM-based AI judges; supports text, chat, audio, video and LLM traces with optional on-premise deployment.

source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

In the news

Research sources · 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Besimple AI do?
Besimple AI collects, licenses and annotates conversational audio datasets used to train and evaluate voice and multimodal AI models.
Who founded Besimple AI?
Besimple AI was founded by Yi Zhong, Bill Wang in 2025.
Who are Besimple AI's investors?
Besimple AI's investors include Y Combinator, TRAC (Third Round Analytics Capital), Script Capital.
Where is Besimple AI headquartered?
Besimple AI is headquartered in San Francisco, US.