Diffbot
Menlo Park, US Β· 4 known investors
Diffbot operates a web-scale knowledge graph using AI to extract and structure facts from billions of web pages through techniques like natural language processing, entity linking, and image recognition. The company serves enterprises and data-dependent organizations that need reliable, structured information about entities and relationships across the public web.
Also known as Diffbot Technologies
Founders & leadership

Investors Β· 4
Company profile
researched Aug 2026Diffbot develops AI models that read unstructured web content and convert it into structured, linked knowledge. Rather than producing flat data dumps typical of rule-based scraping, Diffbot resolves extracted facts to entities that are connected as nodes in a graph, so that a mention of an organization in a news article is linked to the same entity found on other sites. The company describes its mission as "Knowledge as a Service" β building an autonomous system capable of synthesizing human knowledge β and states that it crawls the entire public web and operates an automated Knowledge Graph.
The product suite consists of several APIs: Extract, which uses computer-vision models to classify and render an arbitrary URL and return clean structured fields without rules or site-specific scrapers; Crawl, which walks an entire site from a seed URL to produce a single structured dataset; the Natural Language API, which identifies entities, relationships, facts and sentiment in raw text and resolves them against the Knowledge Graph; the Search/DQL API for querying the Knowledge Graph of people, organizations, products and articles as a database; and Enhance, which augments existing customer records with public web data. The Knowledge Graph exposes typed entities with detailed ontologies (for example, article entities carry fields such as author, publisher region, categories, sentiment, tags and crawl timestamp), and each entity has a persistent diffbotUri identifier.
More recently Diffbot has introduced a Web Search API that returns live structured search results and can be self-hosted on customer hardware, positioned as search without ads or telemetry. The company states the packaged local index is roughly 4TB, compressed from approximately 150TB of web data, and publishes comparative latency and accuracy figures against Tavily, Exa and Firecrawl. Diffbot also offers agent-oriented tooling, including "diffbot-skills" for agent harnesses and an MCP server for its DQL and Enhance APIs, and in January 2025 announced a language model grounded in its Knowledge Graph with citation grounding and a privacy-oriented design.
Business model
Diffbot sells programmatic access to web data extraction and its Knowledge Graph as APIs used by developers and enterprises. Access tiers include a free plan introduced in April 2024 that replaced a two-week trial and provides a monthly usage allowance. The Web Search product can additionally be deployed on customer-owned hardware for self-hosting.
Sources indicate paid API access with a free entry tier and an option to self-host the web search index; specific pricing is not stated in the material provided.
Traction
Diffbot states its APIs are used by companies globally at large scale, citing DuckDuckGo, ProQuo AI and Contingent as users. The Knowledge Graph is described as containing billions of people, organizations, products and articles. Other disclosed figures are technical benchmarks rather than commercial metrics; no revenue or customer-count data appears in the sources.
Latest developments
The most recent disclosed developments are the January 2025 launch of a Knowledge-Graph-backed language model emphasizing citation grounding and privacy, and the introduction of a Web Search API with an option to run a roughly 4TB index on customer hardware. A June 2025 blog post describes building an MCP server for Diffbot's DQL and Enhance APIs, and documentation promotes diffbot-skills for AI agent harnesses.
βΈFull profile β market position, technology, go-to-market, geography, history, risks & controversies
Market position
Diffbot positions itself as operating the world's largest automated Knowledge Graph and the largest structured database of the public web, and as one of few independent full-web crawlers outside Google and Bing. In web search APIs it benchmarks itself against Tavily, Exa and Firecrawl, competing on latency and self-hostability while its own published accuracy figure trails Exa.
Key differentiators claimed in the sources are rule-free extraction driven by computer vision rather than site-specific scrapers, resolution of extracted facts into a linked entity graph rather than flat dumps, an independent full-web crawl on self-owned hardware, and a web search index that customers can own and self-host without ads or telemetry.
Technology
Diffbot runs an independent crawl of the public web β separate from Google and Bing β on custom-assembled hardware in its own California data center. Its stack combines computer-vision-based page classification and rendering for rule-free extraction with natural language processing for entity and relation extraction, and applies entity linking, coreference resolution, knowledge fusion, knowledge inference and information retrieval to build the Knowledge Graph. The graph underpins GraphRAG-style applications and a Knowledge-Graph-grounded language model announced in January 2025. Company-published benchmarks report a 240 ms P90 latency for the Web Search API versus 998 ms (Tavily), 1,231 ms (Exa) and 1,335 ms (Firecrawl), with a stated accuracy of 0.65 against 0.8 for Exa; Extract returns results in roughly 300 ms.
Go-to-market
Developer-led distribution through public API documentation, an API reference site, a free plan, tutorials and a company blog, plus integrations aimed at AI agent frameworks (diffbot-skills, an MCP server for DQL and Enhance). Product pages include interactive demos that issue live API calls. The company publicizes launches through press releases and maintains presence on LinkedIn, Bluesky and Mastodon.
Companies worldwide that need large-scale structured data from the public web, including search and e-commerce products, business intelligence and predictive analytics vendors, supply chain risk tools, and developers building AI agents and RAG pipelines. Named users include DuckDuckGo (product data structuring for shopping search), ProQuo AI (organization data for predictive business development) and Contingent (news data for supply chain insight).
Geography
Headquartered in Menlo Park, United States, with its own data center in California; the company states it crawls the entire public web and serves customers internationally.
History
Public press releases trace the company from a 2011 production API applying visual learning to web content, a Page Classifier developer tool in 2012, and a computer-vision Product API in 2013, to the Discussions API in 2015 and a BBC Click television feature the same year. Funding milestones include $2 million from technology veterans announced in May 2012 and a $10M Series A in February 2016, when the team was described as 12 AI engineers. The Knowledge Graph launched in August 2018, a Knowledge-as-a-Service NLP system in September 2020, a free plan in April 2024, and a Knowledge-Graph-grounded language model in January 2025, followed by the self-hostable Web Search API. Documentation states the company has maintained its APIs for over 15 years.
Risks & controversies
Diffbot's own published benchmark shows its search accuracy (0.65) below Exa (0.8), indicating a latency-versus-accuracy tradeoff. The company frames current web search as ad-monetized and scraped for API access, a market dynamic it competes against. No litigation, regulatory action or other controversies appear in the provided sources.
Compiled by commissioned research from 8 cited public sources β announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed β not independently audited.
Founder mafia
3 people who came through Diffbot went on to found or lead other companies.
Competitors Β· 8
by search overlapCompanies competing with Diffbot for the same Google search keywords, organic and paid, via search-intersection analysis.
Timeline Β· 12
launches, deals, and filingsDiffbot announced an LLM backed by its web Knowledge Graph, emphasizing citation grounding, accuracy and privacy-first design.
Diffbot introduced a Web Search API that can be self-hosted, packaging a web index of about 4TB from roughly 150TB of web data, with no ads or telemetry.
Diffbot introduced a free plan replacing its previous two-week trial, including a monthly allowance of usage.
Diffbot introduced a natural language processing product positioned as an industry-first 'Knowledge as a Service' NLP system.
Diffbot launched its Knowledge Graph, an AI-curated, structured and searchable database of web knowledge aimed at enterprises.
Diffbot announced a $10M Series A; the release described a team of 12 AI engineers building a database of structured information offered as 'knowledge-as-a-service'.
$10M source β
Diffbot was featured on the BBC television program Click following a visit to Diffbot's headquarters earlier that month.
Diffbot released the Discussions API for extracting and tracking user-generated content in forums, comments and reviews.
Diffbot launched a Product API that uses computer vision to automatically extract data from product web pages, turning e-commerce sites into product databases.
Diffbot released a Page Classifier tool using a visual learning robot to identify the type of page content behind any URL.
Diffbot announced $2 million from technology-industry investors, to be used to expand the team and infrastructure and add investors and advisors.
$2M source β
Diffbot announced a production API letting developers apply visual learning technology to tie web content to context, structure and action.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
βΈResearch sources Β· 8
primary sources listed
- Diffbotdiffbot.com Β· web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Diffbot do?
- Diffbot builds AI models that crawl the public web and turn pages into structured, linked data via APIs and a Knowledge Graph.
- Who founded Diffbot?
- Diffbot was founded by Mike Tung.
- Who are Diffbot's investors?
- Diffbot's investors include Amplify Partners, Felicis Ventures, Webb Investment Network, Bloomberg Beta.
- Where is Diffbot headquartered?
- Diffbot is headquartered in Menlo Park, US.




