Moss
YC F25San Francisco, US Β· Founded 2024 Β· 5 employees Β· Hiring Β· 2 known investors
Moss provides a retrieval system optimized for production AI applications, delivering sub-10ms search latency without vector databases. The platform targets developers building voice AI, copilots, and real-time systems where milliseconds impact user experience.
Also known as InferEdge Inc. Β· Moss (moss.dev) Β· usemoss
Founders & leadershipΒ· Y Combinator alumni (F25)
Moss was founded in 2024 by Sri Raghu Malireddi and Harsha Nalluru.
Investors Β· 2
Company profile
researched Aug 2026Moss develops a real-time semantic search runtime aimed at conversational and multimodal AI applications. Rather than operating as a hosted vector database, the product embeds retrieval and embedding generation inside the customer's own process, removing the network round trip that the company says typically adds 200-500 ms to a query. Moss states end-to-end retrieval latency of under 10 ms and positions the runtime for voice agents, copilots, chat interfaces, documentation and knowledge search, and offline or on-device applications where retrieval sits on the critical path.
The system supports hybrid retrieval combining semantic and keyword search, built-in embedding models (with the option to supply an external model), metadata filtering with operators such as $eq, $and, $in and $near, and a managed cloud layer (Moss Cloud) for storing, syncing and distributing indexes. SDKs are published for Python (3.10+), TypeScript/Node.js (20+), Elixir and C (libmoss), alongside a separate WebAssembly package (@moss-dev/moss-web) for client-side search in the browser, a CLI for index management, and data connectors for SQLite, MongoDB, MySQL and Supabase. Y Combinator's profile describes the underlying vector index as built in Rust and WebAssembly and running natively across browsers, mobile devices and servers.
Moss publishes framework integrations for LangChain, DSPy, LlamaIndex, Pipecat, LiveKit, Vapi, ElevenLabs, Strands Agents, CrewAI, Haystack, AutoGen, Mastra, Langflow, Pydantic AI and the Vercel AI SDK, and its open-source repository includes worked examples for voice agents (an airline PNR ambient-retrieval demo and a multi-agent mortgage lending flow), classification, custom embeddings and browser/WASM use.
Founding story
Founded in 2024 by Sri Raghu Malireddi (Founder and CEO) and Harsha Nalluru (co-founder and CTO), who had known each other for more than eight years. Sri previously led ML work at Grammarly and Microsoft, shipping LLM and personalization systems across Office, Bing and Grammarly, including personalization work cited as driving 300% retention growth for Grammarly Keyboard and scaling models to 40M+ DAUs; he has published at conferences including ACL and holds patents in real-time ML. Harsha was a tech lead at Microsoft where he architected the core Azure SDK stack supporting 400+ cloud services and 100M+ weekly npm downloads. The founders say the idea came from repeatedly encountering retrieval lag while building large-scale agentic systems at Microsoft and Grammarly [7].
Business model
Moss distributes SDKs and an open-source repository while operating a managed cloud service; developers sign up at moss.dev for a project_id and project_key, with a free tier available and self-serve onboarding described as requiring no credit card. Indexes can be created and managed through a web portal or programmatically via the SDKs and CLI [0][2][7].
Sources describe a free tier plus paid usage tied to project credentials for the managed Moss Cloud index layer; specific pricing is not disclosed in the available material. The company reported three paying customers at the time of its Y Combinator launch [2][7].
Traction
The company website reports 250,000+ installs and use in production by teams building real-time AI systems. The public GitHub repository shows 666 stars, 91 forks and 257 commits. At its Y Combinator launch, Moss reported six enterprise design partners, three paying customers and seven further prospects evaluating, with usage and revenue described as growing roughly 100% week over week, plus production pilots achieving sub-10 ms retrieval and 70-90% token savings versus traditional pipelines [0][2][7].
Latest developments
Moss publishes a reference architecture guide, "The Production AI Stack," and a live latency demo environment on its site. The GitHub repository documents newer capabilities including Elixir and C SDKs, a browser/WASM SDK, database connectors for SQLite, MongoDB, MySQL and Supabase, a CLI, and voice-agent example applications. The company is hiring a Founding Customer Success Manager and an SDK Software Engineer [0][2][7].
βΈFull profile β market position, technology, go-to-market, geography, history, risks & controversies
Market position
Moss positions itself against hosted vector databases, publishing benchmarks that compare its latency with Pinecone, Qdrant and ChromaDB and claiming up to 100x faster performance. It frames its category as a retrieval runtime embedded in the agent process rather than a managed vector store, and states that agents can still use existing LLM frameworks and voice stacks alongside it [0][2].
Retrieval and embedding run in-process rather than over a network call to an external vector database, which the company presents as the source of its single-digit-millisecond latency; deployment targets include browser, edge, device and cloud with fully local, offline-capable execution; and operations are simplified by removing cluster management, HNSW tuning and sharding [0][2].
Technology
A search runtime rather than a database: documents are indexed and loaded into the application process, so queries avoid network hops and cluster or HNSW tuning. It combines semantic and keyword retrieval, built-in embedding inference, metadata filtering, and a WebAssembly build for browser execution; Y Combinator describes an optimized vector index implemented in Rust and WebAssembly that runs across browsers, mobile devices and servers, with 100% local execution supporting offline indexing and querying [0][2][7].
Go-to-market
Developer-led distribution through open-source SDKs, a public GitHub repository, documentation, Discord and a blog, combined with direct engineering outreach ("Talk to an Engineer", latency test demos and booked demos). Moss also partners with voice AI orchestration providers, embedding its retrieval layer in their context pipelines, and recruits enterprise design partners [0][2][7].
Teams building conversational and voice AI, copilots, chat interfaces and real-time or multimodal agents, including platform and founding engineers, infrastructure leads, agent/voice product managers, security and compliance stakeholders needing local context, and data/ML engineers evaluating embeddings and index configurations [0][7].
Geography
Headquartered in San Francisco per the Y Combinator profile, with job listings for San Francisco and remote US roles [7].
History
Moss was founded in 2024 and participated in Y Combinator's Fall 2025 batch, with Pete Koomen as primary partner. Its launch materials describe a managed portal at usemoss.dev alongside JavaScript and Python SDKs; the current site and repository show an expanded SDK set, database connectors, a CLI and a WebAssembly browser build. Team size is listed as 5 [2][7].
Risks & controversies
Most quantitative claims β latency, install counts, benchmark comparisons against Pinecone, Qdrant and ChromaDB, and customer and growth figures β originate from Moss's own website, repository and Y Combinator launch post rather than independent verification. The published benchmark was run on a single MacBook Pro (M4 Pro, 24GB) over 100,000 documents and measures Moss with built-in embedding while competitors used an external embedding service and cloud-hosted search, which limits comparability. No disclosed funding amount appears in the sources. The name is shared with an unrelated Berlin-based spend management fintech, creating identification risk in databases and press coverage [0][2][7].
Compiled by commissioned research from 8 cited public sources β announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed β not independently audited.
Competitors Β· 2
by search overlapCompanies competing with Moss for the same Google search keywords, organic and paid, via search-intersection analysis.
Timeline Β· 5
launches, deals, and filingsThe public repository documents a WebAssembly browser SDK (@moss-dev/moss-web), Elixir and C (libmoss) SDKs in addition to Python and TypeScript, database connectors for SQLite, MongoDB, MySQL and Supabase, a CLI for index management, and hybrid semantic plus keyword search with metadata filtering.
Moss published a guide describing a reference architecture for building real-time AI systems, promoted on its homepage.
Moss appears in Y Combinator's company directory as part of the Fall 2025 batch, categorized under Developer Tools, SaaS and AI, with Pete Koomen as primary partner.
Moss states it is working closely with voice AI orchestration companies Pipecat (Daily.co) and LiveKit, embedding its retrieval layer in their real-time retrieval and context pipelines.
Public launch post introducing Moss as a runtime delivering sub-10 ms lookups, instant index updates and no infrastructure overhead, with a managed portal at usemoss.dev and JavaScript and Python SDKs.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
In the news
I'm Glad That Moss Is On Console Now, But It's Not The Samekotaku.com Β· Jul 2026
Jaime Moss: WV can't keep funding schools like it's 1980 (Opinion) | Op-Ed Commentaries | wvgazettemail.comwvgazettemail.com Β· Jul 2026βΈResearch sources Β· 8
primary sources listed
- Mossmoss.dev Β· web
8 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Moss do?
- Moss is a sub-10ms semantic search runtime that embeds retrieval inside AI agents instead of using an external vector database.
- Who founded Moss?
- Moss was founded by Sri Raghu Malireddi, Harsha Nalluru in 2024.
- Who are Moss's investors?
- Moss's investors include Y Combinator, Tiger Global Management.
- Where is Moss headquartered?
- Moss is headquartered in San Francisco, US.



