Definity
Chicago, US Β· 22 employees on LinkedIn Β· 4 known investors
Definity provides an AI agent platform for data engineering on Lakehouse and Spark environments, offering cost optimization, in-motion observability, incident troubleshooting, and code-change validation for data pipelines. It targets enterprise data engineering teams and deploys within the customer's own cloud or on-prem environment.
Also known as definity
Founders & leadership


Investors Β· 4
Also in the syndicate Β· 1
Funding
SEC filings, press & company announcements$12M disclosed across 1 of 3 rounds Β· 2024β2026
- $12MSeries AApr 2026 Β· 3 sources
GreatPoint Ventures (lead), Dynatrace, Hyde Park Venture Partners, StageOne Ventures
Source β - Undisclosed amountSeries AMar 2026
Accrete Health Partners (lead)
Source β
Source: company announcements and press reports β follow each round's link for the claim.
Company profile
researched Aug 2026definity develops an agentic data engineering platform for lakehouse and Apache Spark environments. The product provides runtime intelligence across four areas: cost optimization (auto-tuning of jobs, clusters and pipeline code, plus detection of compute and execution waste), in-motion observability (monitoring of data quality, pipeline health and infrastructure performance with AI-based anomaly detection and automated run preemption), agentic troubleshooting (root-cause analysis using data and job lineage and a context-aware AI assistant), and code-change validation (runtime-aware simulation of runs in staging within CI to validate code changes, platform upgrades and migrations).
The platform is installed centrally, one time, with no code changes to existing pipelines β a Spark agent JAR is attached to the Spark session and points at the definity server. It documents capabilities including resource waste and over-provisioning analysis, data skew and parallelization detection, pipeline downtime and missed-SLA identification, auto-generated data validation tests and contracts covering freshness, volume, completeness, schema and distribution, full input/output and internal job lineage, and tracking of environment, platform and schema changes. Supported deployment targets include Databricks (including Databricks Serverless), AWS EMR, GCP Dataproc and Spark on Kubernetes, in cloud or on-premises environments, and support has been extended to Spark Streaming.
Founding story
The company states it originated from its founders' own experience leading enterprise data engineering, platform and product teams.
Business model
Software platform sold to enterprise data organizations, deployed inside the customer's own environment. A free tier and self-service account creation are offered alongside sales-led engagements (demo booking), and definity Cloud is distributed through the AWS Marketplace. A one-week "Spark Cost & Health Assessment" serves as an entry offering.
Traction
The company reports growing adoption among large enterprises including Fortune 500 companies, and stated that revenue tripled over the six months preceding April 2026. Published customer outcomes include a 44% reduction in Spark platform cost with 9x ROI in five months, a 58% cut in EMR platform cost with 74% shorter pipeline run-times, and a platform upgrade accelerated by six months with 50% faster workload validations. Aggregate figures cited across enterprise deployments are 40%+ infrastructure cost reduction, 90% prevented data incidents, 25% increased developer velocity and 50% faster deploys and upgrades.
Latest developments
In April 2026 the company announced a $12M Series A led by GreatPoint Ventures, with Dynatrace, StageOne Ventures and Hyde Park Venture Partners participating, bringing total funding to $16.5M, and launched its agentic data engineering platform. Earlier in April 2026 it extended support to Spark Streaming, and in August 2026 it made agentic optimization and observability for Databricks Serverless generally available. In January 2026 it introduced a Product Advisory Board.
βΈFull profile β market position, technology, go-to-market, geography, history
Market position
Positions itself as runtime infrastructure for operating modern lakehouse and Spark platforms, contrasting its approach with existing tools that monitor isolated aspects such as performance or data quality without a unified view and that react after incidents occur.
Emphasizes in-motion, execution-level intelligence that acts during pipeline runtime rather than issuing reactive alerts, unified coverage across infrastructure, pipeline and data layers, zero-code-change central instrumentation, and full in-environment deployment for security.
Technology
An in-motion architecture that runs directly inside production pipelines. A Spark agent is configured centrally at the SparkSession level, requiring no pipeline code changes, and observes jobs during execution to capture full-stack signals across infrastructure behavior, pipeline execution and data characteristics. These runtime signals feed AI agents and an AI assistant that perform anomaly detection, auto-tuning, root-cause analysis and CI validation, exposed in part through a runtime MCP interface. The system runs entirely within the customer environment so that data does not leave it, and supports both cloud and on-premises Spark.
Go-to-market
Direct enterprise sales supported by demo requests and a free tier, distribution via AWS Marketplace, a partnership with Databricks, conference talks (Data + AI Summit 2025 and 2026), webinars, interviews and a technical blog, plus a Product Advisory Board of enterprise data and platform leaders.
Enterprise data engineering and data platform teams running large-scale lakehouse and Spark deployments, including Fortune 500 organizations.
Geography
Headquartered in Chicago, Illinois; the Series A was covered as an Israeli high-tech funding round. The product deploys in customer cloud or on-premises environments.
History
Public milestones include the February 2025 launch of a Spark Cost & Health Assessment, an April 2025 partnership with Databricks for observability and optimization across the Databricks and Spark ecosystem, the August 2025 availability of definity Cloud on AWS Marketplace, the January 2026 formation of a Product Advisory Board, April 2026 Spark Streaming support and the Series A and platform launch, and August 2026 general availability of Databricks Serverless support.
Compiled by commissioned research from 8 cited public sources β announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed β not independently audited.
Timeline Β· 8
launches, deals, and filingsDatabricks Serverless support became generally available, providing a unified way to observe and optimize workloads across the Databricks platform.
definity announced a $12M Series A led by GreatPoint Ventures with participation from Dynatrace and existing investors StageOne Ventures and Hyde Park Venture Partners, bringing total funding to $16.5M. Proceeds are intended for expanding operations and development efforts.
$12M source β
Alongside the Series A, definity launched its agentic data engineering platform, positioned as runtime infrastructure for operating lakehouse and Spark data platforms.
definity extended execution-level observability and optimization to Spark Streaming for continuous data pipelines.
definity introduced a Product Advisory Board of enterprise data and platform leaders to help shape its roadmap and voice of the customer.
definity Cloud was made available on the AWS Marketplace for lakehouse and Spark observability and optimization.
definity announced a partnership with Databricks to bring full-stack data observability, performance tuning and proactive validation across the Databricks and Spark ecosystem.
definity introduced a one-week assessment to identify Spark platform waste, inefficiencies and potential savings.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
βΈResearch sources Β· 8
primary sources listed
- Definitydefinity.ai Β· web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Definity do?
- definity builds an agentic data engineering platform that optimizes cost and reliability for Lakehouse and Spark pipelines.
- Who founded Definity?
- Definity was founded by Ohad Raviv, Tom Bar-Yacov, Roy Daniel.
- Who are Definity's investors?
- Definity's investors include GreatPoint Ventures, Hyde Park Venture Partners, StageOne Ventures.
- How much funding has Definity raised?
- Definity has disclosed $12M raised across 1 of its 3 known rounds.
- Where is Definity headquartered?
- Definity is headquartered in Chicago, US.


