Fundraising Fox

Reducto

YC W24

San Francisco, US · Founded 2023 · 30 employees · Hiring · 15 known investors

Reducto provides document processing APIs that parse, classify, extract, and edit data from documents at enterprise scale. The platform combines OCR and vision language models to transform unstructured documents into structured data for AI teams across finance, healthcare, legal, and other industries.

Also known as Remembrall

Founders & leadership· Y Combinator alumni (W24)

Reducto was founded in 2023 by Adit Abraham and Raunak Chowdhuri.

AAAdit Abraham
Adit Abrahamin𝕏Co-Founder & CEOAdit Abraham studied computer science at MIT, worked as a product manager on ads and search at Google, and conducted machine learning research at MIT's Media Lab before co-founding Reducto.30 Under 30 2026
RCRaunak Chowdhuri
Raunak Chowdhuriin𝕏FounderHe studied computer science at MIT and previously founded and operated a computational chemistry consulting business. He has published computer vision research during his high school years.30 Under 30 2026

Investors · 15

Also in the syndicate · 8

Andrew OfstadArash FerdowsiJJ FliegelmanKulveer TaggarLiquid 2 VenturesRalph GooteeRichard AbermanTracy Young

Funding

SEC filings, press & company announcements

$75M disclosed across 1 of 5 rounds · 2024–2025

Source: company announcements and press reports — follow each round's link for the claim.

Valuation · disclosed

Disclosed events
$600Mvaluation at $75m Series BOct 2025
filing ↗

Source: SEC prospectus filings, and round valuations the company or its investors disclosed — follow each entry's link for the claim.

Company profile

researched Aug 2026

MIT graduates Adit Abraham and Raunak Chowdhuri founded Reducto in 2023 after a weekend project exposed how poorly conventional OCR and general-purpose language models handled real enterprise documents. Abraham had worked in Google Ads/Search product management and MIT Media Lab machine-learning research; Chowdhuri had published computer-vision work from an early age and worked in MIT research and applied ML. They entered YC W24 and built an API-first ingestion layer that combines traditional OCR/computer vision, proprietary models and frontier vision-language models rather than relying on a single model. Reducto's Parse API preserves text, layout, tables, equations, handwriting, images and bounding boxes across 30+ file types and 100+ languages; Split and Classify segment mixed packets; Extract maps content to requested schemas with citations; Edit modifies source documents; Studio supports testing/evaluation; pipelines, webhooks, CLI and MCP connect the functions. Agentic OCR and Deep Extract add iterative review/correction loops intended to catch last-mile mistakes on long, dense documents. The business targets AI-native application companies and regulated enterprises that would otherwise maintain brittle OCR, chunking and extraction pipelines. Deployments range from multi-tenant cloud to customer VPC, on-premise and fully air-gapped environments, with SOC 2 Type II, HIPAA/BAA, data residency and zero-data-retention controls on higher tiers. Reducto raised an $8.4m seed led by First Round in October 2024, a $24.5m Series A led by Benchmark in April 2025 and a $75m Series B led by Andreessen Horowitz in October 2025, totaling $107.9m mathematically and $108m as announced. Other backers include Y Combinator, BoxGroup, SV Angel, Liquid 2, WndrCo and operator angels. The Series B valued it at roughly $600m according to Forbes. Processing crossed 250m pages by April 2025 and 1b pages by October 2025; 2026 company materials claim more than 4b cumulative pages and capacity of hundreds of millions per day. Named customers include Harvey, Scale AI, Vanta, Mercor, Rogo, Airtable, Anterior, August, LEA and Elysian, plus unnamed Fortune 10 and top-five hedge-fund deployments. Public Standard pricing includes the first 15,000 credits free and then $0.015/credit; a standard parse is one credit/page and batch parsing is 20% cheaper. Growth and Enterprise are negotiated. Competitive and diligence risks include rapid commoditization by frontier multimodal models, document heterogeneity, benchmark sponsorship bias, high GPU/inference costs, sensitive-data exposure, customer concentration in AI startups, and strong rivals including Google/AWS/Azure document AI, LlamaParse, Unstructured, Extend, LandingAI, Instabase, Rossum and Hyperscience.

Founding story

A weekend document-processing hack by two MIT friends exposed the accuracy gap between classic OCR and what AI applications needed. They named Reducto after the Harry Potter shattering spell, joined YC W24 and focused on converting messy human documents into faithful machine-usable structures.

Business model

Usage-based document-intelligence API and developer platform, with enterprise subscriptions and private/on-premise deployment.

Free initial credits followed by per-credit usage; negotiated Growth volume commitments; Enterprise contracts for dedicated throughput, support, security, SLA and private deployment.

Traction

Processed 250m+ cumulative pages by April 2025, more than 1b by October 2025 and claims 4b+ by July 2026; monthly processing grew 6x in the six months before Series B. Current materials claim hundreds of millions of pages/day capacity and 99.9%+ uptime. YC lists 30 employees. Named users include Harvey, Scale AI, Vanta, Mercor, Rogo and Airtable.

Latest developments

Deep Extract launched April 2026 and ranked first in the Reducto-commissioned, micro1-published LongExtractBench in June 2026. Current 2026 platform adds MCP/agent interfaces, Edit, Studio evaluations, regional/private deployment and claims more than 4b pages processed.

Full profile — market position, technology, go-to-market, geography, history, ownership, risks & controversies

Market position

Well-capitalized document-AI infrastructure vendor used by leading AI applications and regulated enterprises; positioned between hyperscaler OCR services, RAG ingestion libraries and full intelligent-document-processing suites.

Optimizes the entire document-ingestion stack for hard layouts and production reliability, exposes granular provenance and multiple workflow primitives, and can run in private or air-gapped environments rather than offering only commodity OCR.

Technology

Hybrid OCR/computer vision and VLM pipeline; layout reconstruction; table/equation/chart/handwriting interpretation; schema extraction with bounding-box citations; Agentic OCR and Deep Extract verification loops; packet splitting/classification; document editing; pipelines, webhooks, CLI, MCP and Studio; cloud/VPC/on-prem/air-gapped inference.

Go-to-market

Developer self-service and API trial, technical benchmarks/cookbooks, YC ecosystem adoption, founder-led sales, customer case studies and expansion into regulated enterprise/VPC/on-premise contracts.

AI product teams and enterprises converting PDFs, spreadsheets, presentations, scans and other complex files into structured, cited data or agent-ready context.

AI-native startups; legal and financial-services platforms; healthcare and insurance automation; real estate; government and high-security enterprises; Fortune-scale internal AI teams.

Geography

San Francisco headquarters; cloud endpoints and customers globally, EU/Australia data-residency options, and customer-controlled VPC/on-premise/air-gapped deployments.

History

Founded 2023; YC W24; Parse API combined computer vision with VLMs; $8.4m First Round-led seed in October 2024; Agentic OCR, Split/Classify/Extract expansion and $24.5m Benchmark-led Series A in April 2025; crossed 1b pages and raised $75m a16z-led Series B in October 2025; launched document Edit and Deep Extract in 2025-26; expanded enterprise deployment, MCP and benchmarks in 2026.

Ownership

Privately held by founders, employees and investors led across stages by First Round Capital, Benchmark and Andreessen Horowitz, with Y Combinator, BoxGroup and other seed backers; cap-table percentages are undisclosed.

Risks & controversies

No material public litigation or regulatory enforcement found. Key risks: multimodal frontier models may commoditize parsing; long-tail layouts and handwritten/low-quality scans can still fail; iterative agents increase latency and cost; customer data is highly sensitive; upstream model/subprocessor changes affect privacy and accuracy; private deployments are operationally complex; and sponsor-designed benchmarks require caveated interpretation. Reducto commissioned LongExtractBench and contributed methodology/ground-truth tooling, although micro1 sourced documents and performed reconciliation/diligence.

Compiled by commissioned research from 14 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Cumulative pages processedJul 20264,000,000,000 pages minimum company claim
Deep extract beta fieldsApr 202628,000,000 fields minimum
HeadcountAug 202686
Longextractbench recallJun 202699.6%
Monthly processing growth 6 monthsOct 20256 times
Team sizeAug 202630 people

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Founder mafia

3 people who came through Reducto went on to found or lead other companies.

Competitors · 5

by search overlap
Docsumo41 shared keywordsDocsumo provides APIs and document processing platform to extract data from unstructured documents across industries including insurance, banking, and logistics. The platform enables businesses to automate document data capture with 90%+ touchless processing capabilities.
LlamaIndex39 shared keywordsLlamaIndex provides document parsing and AI agent workflows that automate document-based administrative operations such as invoice matching, contract data extraction, form processing, and compliance reporting. It serves enterprise operations and developer teams building retrieval-augmented generation (RAG) and agent-based applications.
NanoNets30 shared keywordsNanonets provides AI agents that automate document-heavy, complex business processes by extracting structured data from documents like invoices and purchase orders, then executing workflows across finance, operations, and healthcare systems. The platform integrates with existing enterprise tools and uses a proprietary OCR model to handle edge cases that undermine agent reliability.
Netezza22 shared keywordsIBM is a global technology company whose business spans enterprise software (including Red Hat, HashiCorp, and Confluent), IT infrastructure such as mainframes, servers, and storage, and IT consulting services. The company is also investing heavily in quantum computing and AI-based enterprise offerings, including its Lightwell open-source software security clearinghouse and the Anderon quantum wafer foundry.
Lido20 shared keywordsLido is a document data extraction tool that pulls information from documents, including handwritten and scanned files in any layout or language, validates it against connected systems, and automates downstream workflows.

Companies competing with Reducto for the same Google search keywords, organic and paid, via search-intersection analysis.

Customers & partners

Named customers · 8

AirtableAnteriorHarveyLEAMercorRogoScale AIVanta

Relationships the company or its partners disclosed publicly — case studies, joint announcements, press.

Pricing

as listed Aug 2026
EnterpriseReductoLarge/high-security organizations · custom enterprise
custom enterprise
GrowthReductoScaling and regulated teams · custom usage commitment
custom usage commitment
StandardReductoDevelopers and early-stage teams · usage based
$0.015/credit

Public list pricing as researched from the company's own pricing pages; negotiated and enterprise terms vary.

Acquisitions · 1

Early investors' stakes continue via these deals
Opennote

Opennote is an AI-powered learning platform that consolidates lecture materials, notes, and readings in one place and provides contextual explanations, practice tools, and video tutorials to help students understand course content.

Timeline · 5

launches, deals, and filings
Jun 2026
LongExtractBench results published

Reducto ranked first on a commissioned benchmark published by micro1; sponsorship and methodology provenance are explicitly caveated.

source ↗

Apr 2026
Deep Extract launched

Agent-in-the-loop structured extraction iteratively verified and corrected requested fields.

source ↗

Oct 2025
One billion pages processed

Company reported cumulative page volume above one billion and 6x monthly growth in six months.

source ↗

Apr 2025
Agentic OCR framework

Introduced multi-pass VLM review and correction plus cheaper processing for simple pages.

source ↗

Jan 2024
Y Combinator W24

Reducto entered YC's Winter 2024 batch and launched developer-facing document parsing.

source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

Legal entities · 1

corporate structure
Reducto, Inc.United States (state not independently verified) · active

In the news

Research sources · 14

primary sources listed

14 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Reducto do?
Agentic document-intelligence infrastructure for parsing, classifying, splitting, extracting and editing complex unstructured files for AI applications and enterprise workflows.
Who founded Reducto?
Reducto was founded by Adit Abraham, Raunak Chowdhuri in 2023.
Who are Reducto's investors?
Reducto's investors include Andreessen Horowitz, BoxGroup, First Round Capital, SV Angel, WndrCo, Y Combinator, Benchmark.
How much funding has Reducto raised?
Reducto has disclosed $75M raised across 1 of its 5 known rounds.
Where is Reducto headquartered?
Reducto is headquartered in San Francisco, US.