/Companies

Edugorilla

Lucknow, IN · Founded 2019 · 430 employees on LinkedIn · 3 known investors

Find your way into Edugorilla

Sign up to see every warm intro you have to Edugorilla

  • Paths you didn't know you had: your email and LinkedIn already hold routes to the Edugorilla team. We find them for you
  • 2nd- and 3rd-degree connections: the friend-of-a-friend routes that take hours to manually find through your inbox or LinkedIn
  • Answers now, not in days: momentum is everything in a raise. Skip asking around whether someone knows someone

Days of research, done the moment you sign in, with every route ranked by how warm it is.

InfoBay.AI provides speech and voice datasets sourced from real dual-channel call center conversations, spanning 15+ industry verticals and low-resource global languages (including African, South Asian, and Southeast Asian languages). The data is used to train and improve voice AI, contact center AI, speaker diarization, and speech recognition models, with documented consent-chain provenance for enterprise buyers.

Also known as InfoBay · InfoBay.AI

Founders & leadership

Edugorilla was founded in 2019 by Rohit Manglik.

RMRohit Manglik
Rohit ManglikFounder and CEORohit Manglik is founder of Edugorilla (rebranded as InfoBay.AI), which provides speech and voice datasets from call center conversations across 15+ industry verticals and low-resource languages to train voice AI and speech recognition models. He has established global operations across multiple regions focused on data quality and AI infrastructure for large language models and intelligent automation.

Investors · 3

Company profile

researched Aug 2026

InfoBay.AI (operating at edugorilla.com) positions itself as a "training intelligence" layer for enterprise AI teams, supplying expert-verified training data, annotation infrastructure and model evaluation systems intended to improve factuality, reasoning and production reliability of large language models and other AI systems. Its stated service lines span pre-training data curation, supervised fine-tuning (SFT) dataset design with explicit reasoning traces created by domain subject-matter experts, RLHF and reward-modeling data, medical AI training data, code and reasoning datasets, and factuality auditing that documents hallucination-prone examples before and after remediation.

The company markets a proprietary corpus organized into fourteen collections covering audio, video, healthcare, textbooks, Q&A, coding, visual rendering, egocentric footage, theses, legal contracts, longitudinal time series, operational data, AI trust and safety, and real-world meeting audio. Headline corpus figures include 3.63M hours of multilingual audio (3.5M+ call center hours across roughly 50 languages, plus 57.5K+ podcast hours in 12-13 languages), ISBN-attributed academic textbooks in 15 languages covering 5,000+ subjects, 112M medical images (MRI, CT, X-ray) and 2.46M patient records across 20 specialties, and 12K coding codebases/DSA problems with 25M+ tokens across seven to nine programming languages. Audio assets carry metadata for gender, age, industry, channel, dialect and language, and are tracked for word error rate.

Business model

Business-to-business supply of licensed training datasets and data-engineering services to enterprise AI and model-development teams, delivered through corpus access and per-engagement annotation, dataset design and evaluation work. Engagements are described as beginning with a quality baseline, with deliverables measured against it via benchmark deltas, factuality improvement percentages and evaluation pass rates.

▸Full profile — market position, technology, go-to-market, geography

Market position

Presents itself as a specialist training-data and evaluation provider benchmarked against other data vendors, contrasting its expert-verified, domain-expert-created datasets with commodity crowdwork-based RLHF data and web-scraped pre-training corpora.

Claimed differentiators include real spontaneous call center speech (with disfluencies, interruptions and code-switching) rather than scripted or prompted audio; dual-channel separation of agent and customer tracks at a scale the company says no public dataset matches; native domain vocabulary from 15+ industry verticals such as banking, healthcare, insurance, agriculture, telecom, QSR, logistics, humanitarian aid, SaaS, real estate, solar energy and travel; coverage of low-resource languages including Swahili, Kinyarwanda, Luganda, Chichewa, Nepali, Bengali and Urdu; and curated rather than web-scraped textbook and coding data.

Technology

The offering centers on a proprietary multilingual data corpus plus structured annotation methodology and evaluation design rather than model building. Distinctive technical attributes cited include dual-channel call center recordings that place agent and customer on separate tracks for speaker-diarization ground truth, a four-step audio refining process, WER tracking, industry tagging, ISBN attribution and factuality scoring for textbook data, and age-stratified DICOM medical imaging with linked patient records.

Go-to-market

Direct enterprise sales motion driven from the website, with calls to action to request a model quality audit and to explore the corpus, supported by client testimonials referencing reduced hallucination rates and improved STEM benchmark reasoning accuracy.

Enterprise AI and machine-learning teams building or fine-tuning large language models, audio/voice AI, contact center AI and medical AI systems, including those operating in high-stakes regulated environments.

Geography

Data assets span global and regional languages including English (US, UK, India), Arabic, French, Portuguese (Brazil), Bahasa Indonesia, and African languages (Swahili, Kinyarwanda, Luganda, Ganda, Chichewa) alongside South Asian languages such as Hindi, Punjabi, Bengali, Nepali, Marathi, Tamil, Telugu, Kannada, Malayalam, Gujarati, Assamese, Odia, Mizo and Urdu.

Compiled by commissioned research from 1 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
Coding codebasesJan 202612,000 codebases
Corpus collectionsJan 202614 collections
Industry verticals covered by audio dataJan 202615+
Medical imagesJan 2026112,000,000 images
Multilingual audio corpusJan 20263,630,000 hours
Patient recordsJan 20262,460,000 records
Textbook languagesJan 202615 languages

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Related companies · 5

Companies working in the same space as Edugorilla.

▸Research sources · 1

primary sources listed

1 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Edugorilla do?
InfoBay.AI engineers expert-verified multilingual training data, annotation infrastructure and evaluation systems for enterprise AI models.
Who founded Edugorilla?
Edugorilla was founded by Rohit Manglik in 2019.
Who are Edugorilla's investors?
Edugorilla's investors include LVX Ventures, SucSEED Indovation, We Founder Circle Global Angels Fund.
Where is Edugorilla headquartered?
Edugorilla is headquartered in Lucknow, IN.