Fundraising Fox

Diffbot

Menlo Park, US Β· 4 known investors

diffbot.com β†—

Diffbot operates a web-scale knowledge graph using AI to extract and structure facts from billions of web pages through techniques like natural language processing, entity linking, and image recognition. The company serves enterprises and data-dependent organizations that need reliable, structured information about entities and relationships across the public web.

Also known as Diffbot Technologies

AI & Machine LearningData & InfrastructureEnterprise SoftwareSaaS

Founders & leadership

MTMike Tung
Mike TungCEO

Investors Β· 4

Company profile

researched Aug 2026

Diffbot develops AI models that read unstructured web content and convert it into structured, linked knowledge. Rather than producing flat data dumps typical of rule-based scraping, Diffbot resolves extracted facts to entities that are connected as nodes in a graph, so that a mention of an organization in a news article is linked to the same entity found on other sites. The company describes its mission as "Knowledge as a Service" β€” building an autonomous system capable of synthesizing human knowledge β€” and states that it crawls the entire public web and operates an automated Knowledge Graph.

The product suite consists of several APIs: Extract, which uses computer-vision models to classify and render an arbitrary URL and return clean structured fields without rules or site-specific scrapers; Crawl, which walks an entire site from a seed URL to produce a single structured dataset; the Natural Language API, which identifies entities, relationships, facts and sentiment in raw text and resolves them against the Knowledge Graph; the Search/DQL API for querying the Knowledge Graph of people, organizations, products and articles as a database; and Enhance, which augments existing customer records with public web data. The Knowledge Graph exposes typed entities with detailed ontologies (for example, article entities carry fields such as author, publisher region, categories, sentiment, tags and crawl timestamp), and each entity has a persistent diffbotUri identifier.

More recently Diffbot has introduced a Web Search API that returns live structured search results and can be self-hosted on customer hardware, positioned as search without ads or telemetry. The company states the packaged local index is roughly 4TB, compressed from approximately 150TB of web data, and publishes comparative latency and accuracy figures against Tavily, Exa and Firecrawl. Diffbot also offers agent-oriented tooling, including "diffbot-skills" for agent harnesses and an MCP server for its DQL and Enhance APIs, and in January 2025 announced a language model grounded in its Knowledge Graph with citation grounding and a privacy-oriented design.

Business model

Diffbot sells programmatic access to web data extraction and its Knowledge Graph as APIs used by developers and enterprises. Access tiers include a free plan introduced in April 2024 that replaced a two-week trial and provides a monthly usage allowance. The Web Search product can additionally be deployed on customer-owned hardware for self-hosting.

Sources indicate paid API access with a free entry tier and an option to self-host the web search index; specific pricing is not stated in the material provided.

Traction

Diffbot states its APIs are used by companies globally at large scale, citing DuckDuckGo, ProQuo AI and Contingent as users. The Knowledge Graph is described as containing billions of people, organizations, products and articles. Other disclosed figures are technical benchmarks rather than commercial metrics; no revenue or customer-count data appears in the sources.

Latest developments

The most recent disclosed developments are the January 2025 launch of a Knowledge-Graph-backed language model emphasizing citation grounding and privacy, and the introduction of a Web Search API with an option to run a roughly 4TB index on customer hardware. A June 2025 blog post describes building an MCP server for Diffbot's DQL and Enhance APIs, and documentation promotes diffbot-skills for AI agent harnesses.

β–ΈFull profile β€” market position, technology, go-to-market, geography, history, risks & controversies

Market position

Diffbot positions itself as operating the world's largest automated Knowledge Graph and the largest structured database of the public web, and as one of few independent full-web crawlers outside Google and Bing. In web search APIs it benchmarks itself against Tavily, Exa and Firecrawl, competing on latency and self-hostability while its own published accuracy figure trails Exa.

Key differentiators claimed in the sources are rule-free extraction driven by computer vision rather than site-specific scrapers, resolution of extracted facts into a linked entity graph rather than flat dumps, an independent full-web crawl on self-owned hardware, and a web search index that customers can own and self-host without ads or telemetry.

Technology

Diffbot runs an independent crawl of the public web β€” separate from Google and Bing β€” on custom-assembled hardware in its own California data center. Its stack combines computer-vision-based page classification and rendering for rule-free extraction with natural language processing for entity and relation extraction, and applies entity linking, coreference resolution, knowledge fusion, knowledge inference and information retrieval to build the Knowledge Graph. The graph underpins GraphRAG-style applications and a Knowledge-Graph-grounded language model announced in January 2025. Company-published benchmarks report a 240 ms P90 latency for the Web Search API versus 998 ms (Tavily), 1,231 ms (Exa) and 1,335 ms (Firecrawl), with a stated accuracy of 0.65 against 0.8 for Exa; Extract returns results in roughly 300 ms.

Go-to-market

Developer-led distribution through public API documentation, an API reference site, a free plan, tutorials and a company blog, plus integrations aimed at AI agent frameworks (diffbot-skills, an MCP server for DQL and Enhance). Product pages include interactive demos that issue live API calls. The company publicizes launches through press releases and maintains presence on LinkedIn, Bluesky and Mastodon.

Companies worldwide that need large-scale structured data from the public web, including search and e-commerce products, business intelligence and predictive analytics vendors, supply chain risk tools, and developers building AI agents and RAG pipelines. Named users include DuckDuckGo (product data structuring for shopping search), ProQuo AI (organization data for predictive business development) and Contingent (news data for supply chain insight).

Geography

Headquartered in Menlo Park, United States, with its own data center in California; the company states it crawls the entire public web and serves customers internationally.

History

Public press releases trace the company from a 2011 production API applying visual learning to web content, a Page Classifier developer tool in 2012, and a computer-vision Product API in 2013, to the Discussions API in 2015 and a BBC Click television feature the same year. Funding milestones include $2 million from technology veterans announced in May 2012 and a $10M Series A in February 2016, when the team was described as 12 AI engineers. The Knowledge Graph launched in August 2018, a Knowledge-as-a-Service NLP system in September 2020, a free plan in April 2024, and a Knowledge-Graph-grounded language model in January 2025, followed by the self-hostable Web Search API. Documentation states the company has maintained its APIs for over 15 years.

Risks & controversies

Diffbot's own published benchmark shows its search accuracy (0.65) below Exa (0.8), indicating a latency-versus-accuracy tradeoff. The company frames current web search as ad-monetized and scraped for API access, a market dynamic it competes against. No litigation, regulatory action or other controversies appear in the provided sources.

Compiled by commissioned research from 8 cited public sources β€” announcements, filings, and press listed under research sources below.

Key figures

latest reported
API maintenance track recordJan 2025APIs maintained for over 15 years
Blog posts publishedJan 2025143 posts
Extract API typical response timeJan 2025300 ms
Self-hosted web index sizeJan 20254 TB
Team size at Series AFeb 201612 engineers
Web Search API accuracy (as published by company)Jan 20250.7
Web Search API P90 query latencyJan 2025240 ms

Company-reported or press-reported figures, each dated to when it was claimed β€” not independently audited.

Founder mafia

3 people who came through Diffbot went on to found or lead other companies.

Competitors Β· 8

by search overlap
Firecrawl237 shared keywordsFirecrawl provides an API to search, scrape, and interact with web content at scale, converting live internet data into clean markdown and structured data optimized for large language models, agents, and AI-native applications.
Elastic228 shared keywordsElastic develops search-powered software for enterprises, offering products for search, observability, and security built on Elasticsearch. The company positions its technology to help organizations use their data for AI applications.
Cloudflare Turnstile222 shared keywordsCloudflare provides a global cloud network platform delivering security, performance, and development services through sixty-plus integrated services including SASE, application security, and full-stack development infrastructure.
Apify188 shared keywordsApify is a marketplace and platform for web scraping and data extraction tools, offering pre-built automation actors and custom development services. It enables users to extract real-time data from websites, social media, and e-commerce platforms, and serves developers, businesses, and AI applications requiring structured web data.
SEMrush176 shared keywordsSemrush provides competitive intelligence, SEO, keyword research, and content generation tools for growth marketers and small teams to improve visibility and optimize their acquisition funnel across Google and AI search.
Browse AI143 shared keywordsBrowse AI is a no-code web scraping and monitoring platform that lets users extract, monitor, and integrate data from websites using prebuilt scrapers, workflows, scheduled runs, and human-behavior emulation. It also offers managed end-to-end data extraction services and serves users worldwide across use cases such as ecommerce pricing, real estate, job listings, and SEO research.
Netezza127 shared keywordsIBM is a global technology company whose business spans enterprise software (including Red Hat, HashiCorp, and Confluent), IT infrastructure such as mainframes, servers, and storage, and IT consulting services. The company is also investing heavily in quantum computing and AI-based enterprise offerings, including its Lightwell open-source software security clearinghouse and the Anderon quantum wafer foundry.
Ahrefs118 shared keywordsAhrefs is an SEO and marketing platform that provides tools and data for keyword research, competitor analysis, content creation, site auditing, rank tracking, and monitoring brand visibility across search engines and AI chatbots. It serves marketers and enterprises, offering an API, custom dashboards, and AI-driven insights for improving search rankings and traffic.

Companies competing with Diffbot for the same Google search keywords, organic and paid, via search-intersection analysis.

Timeline Β· 12

launches, deals, and filings
Jan 2025
Launch of factually grounded language model

Diffbot announced an LLM backed by its web Knowledge Graph, emphasizing citation grounding, accuracy and privacy-first design.

source β†—

Jan 2025
Web Search API, including self-hostable index

Diffbot introduced a Web Search API that can be self-hosted, packaging a web index of about 4TB from roughly 150TB of web data, with no ads or telemetry.

source β†—

Apr 2024
Free plan introduced

Diffbot introduced a free plan replacing its previous two-week trial, including a monthly allowance of usage.

source β†—

Sep 2020
Knowledge as a Service natural language processing system unveiled

Diffbot introduced a natural language processing product positioned as an industry-first 'Knowledge as a Service' NLP system.

source β†—

Aug 2018
Diffbot Knowledge Graph launched

Diffbot launched its Knowledge Graph, an AI-curated, structured and searchable database of web knowledge aimed at enterprises.

source β†—

Feb 2016
Diffbot raises $10M Series A

Diffbot announced a $10M Series A; the release described a team of 12 AI engineers building a database of structured information offered as 'knowledge-as-a-service'.

$10M source β†—

Jun 2015
Featured on BBC's Click

Diffbot was featured on the BBC television program Click following a visit to Diffbot's headquarters earlier that month.

source β†—

Mar 2015
Discussions API launched

Diffbot released the Discussions API for extracting and tracking user-generated content in forums, comments and reviews.

source β†—

Jul 2013
Product API launched

Diffbot launched a Product API that uses computer vision to automatically extract data from product web pages, turning e-commerce sites into product databases.

source β†—

Aug 2012
Page Classifier developer tool launched

Diffbot released a Page Classifier tool using a visual learning robot to identify the type of page content behind any URL.

source β†—

May 2012
Technology veterans invest $2 million

Diffbot announced $2 million from technology-industry investors, to be used to expand the team and infrastructure and add investors and advisors.

$2M source β†—

Aug 2011
Production API released for visual learning web extraction

Diffbot announced a production API letting developers apply visual learning technology to tie web content to context, structure and action.

source β†—

Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.

In the news

β–ΈResearch sources Β· 8

primary sources listed

8 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Diffbot do?
Diffbot builds AI models that crawl the public web and turn pages into structured, linked data via APIs and a Knowledge Graph.
Who founded Diffbot?
Diffbot was founded by Mike Tung.
Who are Diffbot's investors?
Diffbot's investors include Amplify Partners, Felicis Ventures, Webb Investment Network, Bloomberg Beta.
Where is Diffbot headquartered?
Diffbot is headquartered in Menlo Park, US.