Datology AI
Redwood City, US · Founded 2023 · 78 employees on LinkedIn · 19 known investors
DatologyAI provides automated data curation for AI model training, helping companies identify and remove low-quality data to improve model performance and reduce computational costs. The platform serves AI development teams and companies training custom machine learning models.
Also known as Datology · Datology AI, Inc. · DatologyAI · DatologyAI, Inc.
Founders & leadership
Datology AI was founded in 2023 by Ari Morcos and Bogdan Gaza.


Investors · 19
Also in the syndicate · 12
Funding
SEC filings, press & company announcements- Undisclosed amountSeries AMay 2024
Felicis (lead), Amazon Alexa Fund, Amplify Partners, M12, Outset Capital, Radical Ventures
Source ↗
Source: company announcements and press reports — follow each round's link for the claim.
Company profile
researched Aug 2026DatologyAI, Inc. develops software that automates the curation of datasets used to train AI models, including large language models and other generative systems. The platform identifies which data in a corpus is most valuable for a given model application, removes redundant or harmful samples, suggests augmentation with additional or synthetic data, and determines how data should be batched during training. According to the company and press coverage, it can process datasets at petabyte scale and is modality-agnostic, handling text, images, video, audio, tabular data and more specialized formats such as genomic and geospatial data [2][3][4].
The company positions its product as a four-stage "data refinery": cleaning (removing poorly formatted, empty, short or evaluation-contaminated records), curating (filtering by quality, task and taxonomy), creating (generating synthetic data for task relevance and diversity), and composing (mixing sources into staged training datasets). Outputs are intended for mid-training on open models or pre-training of custom foundation models, and deployments run inside the customer's own private cloud or on-premises infrastructure. The company states it does not sell or source data itself, only processes customer-supplied proprietary, public web, and licensed data [0][2][3][4].
Published customer results include a legal-domain mid-training engagement reporting a 5% improvement on LegalBench and 2.5% on general evaluations using 100B mid-training tokens against a 15T-token pre-trained base, and a frontier-class open-weights mixture-of-experts model (398B total parameters, 13B active) trained on roughly 17-20T tokens curated by Datology, which the company says served 3.37T tokens on OpenRouter in its first two months [0].
Founding story
Ari Morcos, who holds a PhD in neuroscience from Harvard, spent two years at DeepMind applying neuroscience-inspired techniques to understand and improve AI models and more than five years at Meta's AI lab studying the mechanisms underlying model behavior, work that largely involved manipulating training data. He founded DatologyAI to abstract away data preparation work in model training, together with co-founders Matthew Leavitt and Bogdan Gaza, a former engineering lead at Amazon and later Twitter [2][3].
Business model
Enterprise software sold to organizations that train their own AI models; the product is deployed into the customer's infrastructure (on-premises or virtual private cloud) rather than requiring data to be transferred to the vendor. The company states it does not sell data or tokens, only tooling that processes customer data [0][2][3].
Not specified in the available sources.
Traction
Reported customer engagements include a legal-domain mid-training project with Thomson Reuters-referenced evaluation gains and a partnership with Arcee AI on model building; the site cites a frontier open-weights MoE model trained on Datology-curated tokens that served 3.37T tokens on OpenRouter in its first two months. Total disclosed funding reported at $57.5 million as of May 2024 [0][6].
Latest developments
The company website lists research updates dated 2026, including work on automating data research (DataSmith), inducing concision in vision-language models via data curation, and pretraining with finetuning data, alongside case studies on legal-domain mid-training and a frontier open-weights MoE pre-training engagement [0].
▸Full profile — market position, technology, go-to-market, geography, history, risks & controversies
Market position
Operates in the AI training-data curation category. Coverage frames the company against data preparation and curation tools such as CleanLab, Lilac, Labelbox, YData and Galileo, with the company claiming broader scope in dataset size and modality coverage than those alternatives. Reporting also notes skepticism about fully automated curation, citing prior failures of algorithmically curated public datasets [2][3].
Claimed differentiation rests on scale (petabyte-level datasets), modality breadth (text, image, video, audio, tabular, genomic, geospatial), deployment inside the customer's own infrastructure so proprietary data is not transferred, and the position that the company processes customer data rather than selling datasets or tokens [0][2][3][4].
Technology
A pipeline that automatically analyzes large datasets to select high-value samples, reduce redundancy, balance long-tailed distributions, identify complex "concepts" requiring higher-quality samples, flag data likely to cause unintended model behavior, apply targeted augmentation and synthetic data generation, and optimize data mixing and batching for staged training. It scales to petabytes, supports text, image, video, audio, tabular, genomic and geospatial modalities, and runs within the customer's VPC or on-premises environment so data does not leave customer control [0][2][3][4].
Go-to-market
Direct enterprise sales with a "book a demo" motion on the company website, supported by published customer case studies, testimonials and public research output. At launch the company worked with a limited set of design-partner customers ahead of a wider platform release [0][3].
Enterprises and AI teams training custom or foundation models on their own data, including those building models from scratch for governance or compliance reasons, and model-builder companies performing mid-training or domain adaptation. Named customer references include Arcee AI and Thomson Reuters [0][2][3].
Geography
United States; profiles list the company in San Francisco, California and, in another directory, Redwood City, California [4][6].
History
Founded in 2023 by Ari Morcos (CEO), Matthew Leavitt (CSO) and Bogdan Gaza (CTO) [2][4][6]. In February 2024 the company emerged publicly with an $11.65 million seed round led by Amplify Partners, at which point it said it was working with a limited number of customers ahead of a broader platform release later that year [2][3]. A $46 million Series A led by Felicis Ventures was reported in May 2024, bringing total reported funding to $57.5 million and funding hiring of researchers and engineers plus infrastructure expansion [6]. The company's site subsequently published research updates and customer case studies covering pre-training and mid-training engagements [0].
Risks & controversies
Press coverage raises skepticism about automated dataset curation generally, noting that algorithmically curated datasets have previously required withdrawal (LAION) and that automatically filtered models can still produce toxic output; the company says its tooling augments rather than replaces manual curation. Effectiveness claims at the time of the seed round were largely vendor-asserted and unverified externally [2][3].
Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Competitors · 9
by search overlapCompanies competing with Datology AI for the same Google search keywords, organic and paid, via search-intersection analysis.
Timeline · 2
launches, deals, and filingsSeries A round with participation from the Amazon Alexa Fund, Microsoft's M12, Elad Gil and follow-on investment from Amplify Partners and Radical Ventures; proceeds earmarked for hiring researchers and engineers and expanding platform infrastructure. Reported to bring total funding to $57.5 million.
$46M source ↗
Seed financing with participation from Radical Ventures, Conviction Capital, Outset Capital and Quiet Capital, plus angel investors including Jeff Dean, Yann LeCun, Adam D'Angelo, Cohere co-founders Aidan Gomez and Ivan Zhang, Contextual AI founder Douwe Kiela, Naveen Rao and Jascha Sohl-Dickstein.
$11.7M source ↗
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
▸Research sources · 8
primary sources listed
- Datology AIdatologyai.com · web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Datology AI do?
- DatologyAI builds automated data curation tooling that turns raw enterprise and web data into training datasets for AI models.
- Who founded Datology AI?
- Datology AI was founded by Ari Morcos, Bogdan Gaza in 2023.
- Who are Datology AI's investors?
- Datology AI's investors include Felicis Ventures, Outset Capital, Radical Ventures, Amplify Partners, M12 (Microsoft's Venture Fund), Amazon Alexa Fund, Amplify Partners CV.
- Where is Datology AI headquartered?
- Datology AI is headquartered in Redwood City, US.




