Fundraising Fox

Datology AI

Redwood City, US · Founded 2023 · 78 employees on LinkedIn · 19 known investors

DatologyAI provides automated data curation for AI model training, helping companies identify and remove low-quality data to improve model performance and reduce computational costs. The platform serves AI development teams and companies training custom machine learning models.

Also known as Datology · Datology AI, Inc. · DatologyAI · DatologyAI, Inc.

Founders & leadership

Datology AI was founded in 2023 by Ari Morcos and Bogdan Gaza.

AMAri Morcos
Ari MorcosinCEO and Co-FounderAri Morcos is the co-founder and CEO of DatologyAI, which offers automated data curation to improve AI model training. He was previously a senior staff research scientist at Meta AI Research (FAIR), where he studied neural network computation and the properties of data underlying useful representations, and earlier worked at DeepMind in London.
BGBogdan Gaza
Bogdan GazainCo-Founder & CTOCo-founder and CTO of DatologyAI, which provides automated data curation for AI model training to help teams improve model performance and cut computational costs.

Investors · 19

Also in the syndicate · 12

Adam D'AngeloAidan GomezConviction CapitalDouwe KielaGeoffrey HintonIvan ZhangJascha Sohl-DicksteinJeff DeanM12Naveen RaoQuiet Capital Management LLCYann LeCun

Funding

SEC filings, press & company announcements

Source: company announcements and press reports — follow each round's link for the claim.

Company profile

researched Aug 2026

DatologyAI, Inc. develops software that automates the curation of datasets used to train AI models, including large language models and other generative systems. The platform identifies which data in a corpus is most valuable for a given model application, removes redundant or harmful samples, suggests augmentation with additional or synthetic data, and determines how data should be batched during training. According to the company and press coverage, it can process datasets at petabyte scale and is modality-agnostic, handling text, images, video, audio, tabular data and more specialized formats such as genomic and geospatial data [2][3][4].

The company positions its product as a four-stage "data refinery": cleaning (removing poorly formatted, empty, short or evaluation-contaminated records), curating (filtering by quality, task and taxonomy), creating (generating synthetic data for task relevance and diversity), and composing (mixing sources into staged training datasets). Outputs are intended for mid-training on open models or pre-training of custom foundation models, and deployments run inside the customer's own private cloud or on-premises infrastructure. The company states it does not sell or source data itself, only processes customer-supplied proprietary, public web, and licensed data [0][2][3][4].

Published customer results include a legal-domain mid-training engagement reporting a 5% improvement on LegalBench and 2.5% on general evaluations using 100B mid-training tokens against a 15T-token pre-trained base, and a frontier-class open-weights mixture-of-experts model (398B total parameters, 13B active) trained on roughly 17-20T tokens curated by Datology, which the company says served 3.37T tokens on OpenRouter in its first two months [0].

Founding story

Ari Morcos, who holds a PhD in neuroscience from Harvard, spent two years at DeepMind applying neuroscience-inspired techniques to understand and improve AI models and more than five years at Meta's AI lab studying the mechanisms underlying model behavior, work that largely involved manipulating training data. He founded DatologyAI to abstract away data preparation work in model training, together with co-founders Matthew Leavitt and Bogdan Gaza, a former engineering lead at Amazon and later Twitter [2][3].

Business model

Enterprise software sold to organizations that train their own AI models; the product is deployed into the customer's infrastructure (on-premises or virtual private cloud) rather than requiring data to be transferred to the vendor. The company states it does not sell data or tokens, only tooling that processes customer data [0][2][3].

Not specified in the available sources.

Traction

Reported customer engagements include a legal-domain mid-training project with Thomson Reuters-referenced evaluation gains and a partnership with Arcee AI on model building; the site cites a frontier open-weights MoE model trained on Datology-curated tokens that served 3.37T tokens on OpenRouter in its first two months. Total disclosed funding reported at $57.5 million as of May 2024 [0][6].

Latest developments

The company website lists research updates dated 2026, including work on automating data research (DataSmith), inducing concision in vision-language models via data curation, and pretraining with finetuning data, alongside case studies on legal-domain mid-training and a frontier open-weights MoE pre-training engagement [0].

Full profile — market position, technology, go-to-market, geography, history, risks & controversies

Market position

Operates in the AI training-data curation category. Coverage frames the company against data preparation and curation tools such as CleanLab, Lilac, Labelbox, YData and Galileo, with the company claiming broader scope in dataset size and modality coverage than those alternatives. Reporting also notes skepticism about fully automated curation, citing prior failures of algorithmically curated public datasets [2][3].

Claimed differentiation rests on scale (petabyte-level datasets), modality breadth (text, image, video, audio, tabular, genomic, geospatial), deployment inside the customer's own infrastructure so proprietary data is not transferred, and the position that the company processes customer data rather than selling datasets or tokens [0][2][3][4].

Technology

A pipeline that automatically analyzes large datasets to select high-value samples, reduce redundancy, balance long-tailed distributions, identify complex "concepts" requiring higher-quality samples, flag data likely to cause unintended model behavior, apply targeted augmentation and synthetic data generation, and optimize data mixing and batching for staged training. It scales to petabytes, supports text, image, video, audio, tabular, genomic and geospatial modalities, and runs within the customer's VPC or on-premises environment so data does not leave customer control [0][2][3][4].

Go-to-market

Direct enterprise sales with a "book a demo" motion on the company website, supported by published customer case studies, testimonials and public research output. At launch the company worked with a limited set of design-partner customers ahead of a wider platform release [0][3].

Enterprises and AI teams training custom or foundation models on their own data, including those building models from scratch for governance or compliance reasons, and model-builder companies performing mid-training or domain adaptation. Named customer references include Arcee AI and Thomson Reuters [0][2][3].

Geography

United States; profiles list the company in San Francisco, California and, in another directory, Redwood City, California [4][6].

History

Founded in 2023 by Ari Morcos (CEO), Matthew Leavitt (CSO) and Bogdan Gaza (CTO) [2][4][6]. In February 2024 the company emerged publicly with an $11.65 million seed round led by Amplify Partners, at which point it said it was working with a limited number of customers ahead of a broader platform release later that year [2][3]. A $46 million Series A led by Felicis Ventures was reported in May 2024, bringing total reported funding to $57.5 million and funding hiring of researchers and engineers plus infrastructure expansion [6]. The company's site subsequently published research updates and customer case studies covering pre-training and mid-training engagements [0].

Risks & controversies

Press coverage raises skepticism about automated dataset curation generally, noting that algorithmically curated datasets have previously required withdrawal (LAION) and that automatically filtered models can still produce toxic output; the company says its tooling augments rather than replaces manual curation. Effectiveness claims at the time of the seed round were largely vendor-asserted and unverified externally [2][3].

Compiled by commissioned research from 8 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Company stageJan 2024Series A
Customer MoE model sizeJan 2026398B total parameters, 13B active
HeadcountAug 202678
Tokens curated for frontier open-weights MoE customer modelJan 202617,000,000,000,000 tokens
Tokens served by customer model on OpenRouter in first two monthsJan 20263,370,000,000,000 tokens
Total funding raisedMay 2024$57.5M

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Competitors · 9

by search overlap
Databricks12 shared keywordsDatabricks provides a unified platform for data engineering, analytics, and artificial intelligence, enabling organizations to build and deploy data and AI applications on cloud infrastructure.
Netezza12 shared keywordsIBM is a global technology company whose business spans enterprise software (including Red Hat, HashiCorp, and Confluent), IT infrastructure such as mainframes, servers, and storage, and IT consulting services. The company is also investing heavily in quantum computing and AI-based enterprise offerings, including its Lightwell open-source software security clearinghouse and the Anderon quantum wafer foundry.
Data11 shared keywordsdata.world offers a data marketplace that presents an organization's most-used data products in an e-commerce-style shopping experience, with domain-based organization and curated views. It targets enterprise business users and data teams, helping consumers discover and evaluate published data products while supporting data mesh and domain-oriented data management.
Datadog9 shared keywordsDatadog is a cloud-based monitoring and security platform that provides observability, analytics, and protection for infrastructure, applications, and data across enterprises. The platform serves development, security, and operations teams managing cloud-native and hybrid environments.
Hugging Face8 shared keywordsHugging Face is a collaboration platform that hosts and provides access to machine learning models, datasets, and applications. It offers both open-source tools for the ML community and paid compute and enterprise solutions for teams building AI applications.
Alation8 shared keywordsAlation builds a data catalog platform with machine learning and generative AI capabilities that helps enterprises find, understand, and trust their data. The platform serves large organizations and Fortune 1000 companies that depend on data for critical business decisions and AI applications.
DataCamp7 shared keywordsDataCamp is a learning platform for teams that teaches data and AI skills through hands-on coursework accessible via web browser and mobile app. It serves enterprise customers and development teams seeking to build technical capabilities.
Generative AI for Tabular Data7 shared keywordsMOSTLY AI offers a Data Intelligence Platform and an open-source Synthetic Data SDK that let enterprise data teams generate high-fidelity, privacy-safe synthetic data, mock data, and simulations while analyzing production data. Built on its TabularARGN model with differential privacy, it targets enterprise organizations for AI development, testing, and secure data sharing.
Datum6 shared keywordsDatum provides infrastructure and networking services for cloud platforms, offering edge computing capabilities, private networks, and direct interconnection technology typically available only to large cloud providers.

Companies competing with Datology AI for the same Google search keywords, organic and paid, via search-intersection analysis.

Timeline · 2

launches, deals, and filings
May 2024
$46M Series A led by Felicis Ventures

Series A round with participation from the Amazon Alexa Fund, Microsoft's M12, Elad Gil and follow-on investment from Amplify Partners and Radical Ventures; proceeds earmarked for hiring researchers and engineers and expanding platform infrastructure. Reported to bring total funding to $57.5 million.

$46M source ↗

Feb 2024
DatologyAI announces $11.65M seed round led by Amplify Partners

Seed financing with participation from Radical Ventures, Conviction Capital, Outset Capital and Quiet Capital, plus angel investors including Jeff Dean, Yann LeCun, Adam D'Angelo, Cohere co-founders Aidan Gomez and Ivan Zhang, Contextual AI founder Douwe Kiela, Naveen Rao and Jascha Sohl-Dickstein.

$11.7M source ↗

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

In the news

Research sources · 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Datology AI do?
DatologyAI builds automated data curation tooling that turns raw enterprise and web data into training datasets for AI models.
Who founded Datology AI?
Datology AI was founded by Ari Morcos, Bogdan Gaza in 2023.
Who are Datology AI's investors?
Datology AI's investors include Felicis Ventures, Outset Capital, Radical Ventures, Amplify Partners, M12 (Microsoft's Venture Fund), Amazon Alexa Fund, Amplify Partners CV.
Where is Datology AI headquartered?
Datology AI is headquartered in Redwood City, US.