Fundraising Fox

Hyphenbox

Entrepreneur First '26

Bangalore, IN · Founded 2026 · 1 known investors

hyphenbox.com

Hyphenbox converts raw egocentric and multi-modal video into training-ready datasets for general-purpose robotics, providing frame-accurate action segmentation, 3D hand reconstruction, and full-body pose reconstruction. It serves frontier robotics labs and data-collection partners whose model development is bottlenecked on enriched, human-verified training data.

Also known as Hyphenbox

AI & Machine LearningData & InfrastructureDeveloper ToolsRobotics

Founders & leadership

Hyphenbox was founded in 2026 by Vishruth Nagendra and Shreyash Gupta.

VNVishruth Nagendra
Vishruth NagendrainFounder & CEOVishruth Nagendra studied Aerospace Engineering and Computer Science at IIT Bombay and has previously founded a Y Combinator-backed company. He started a medical data annotation business and headed data acquisition and controls for IIT Bombay Racing.
SGShreyash Gupta
Shreyash GuptainCTOShreyash Gupta has published robotics research covering SLAM, path planning, and perception systems for autonomous vehicles. At IIT Bombay he directed a team of over 40 engineers that produced India's first autonomous racecar.

Investors · 1

Company profile

researched Aug 2026

Hyphenbox positions itself as data infrastructure for physical AI and dexterous manipulation, converting raw egocentric and multi-modal capture into datasets that robotics models can train on directly. Its stated output combines three annotation primitives delivered through a single pipeline: dense action labels (frame-accurate action segmentation across long-horizon manipulation, including scene context, object state, and contact sequences), 2D-to-3D hand tracking (reconstruction described as millimeter-level, with 21 keypoints per hand, intended to remain stable under self-occlusion and close-range object interaction), and 2D-to-3D full-body pose (SMPL plus hands, 52 joints with root trajectory, described as physics-aware and temporally smooth, recovered from a single egocentric camera).

The company frames general-purpose robotics as a data problem rather than a model problem, arguing that capture volume is scaling faster than the supply of usable, training-ready data, and that generic labeling tools and off-the-shelf pose estimators were not designed for embodied, first-person footage. As supporting evidence, the site cites EgoScale / NVIDIA Research (arxiv.org) for the claim that the egocentric video used to pretrain NVIDIA Isaac GR00T N1.7 was explicitly action-labeled with dense 3D hand tracking and camera motion derived from raw footage.

The pipeline described runs from raw ego video (video, depth, and other signals captured in homes, factories, and retail environments), through pre-labeling by a proprietary vision-language model that proposes dense action segments and scene context, then 2D-to-3D hand and body reconstruction, then human-in-the-loop QA in which every label is reviewed by an expert reviewer rather than crowdsourced workers, producing a model-consumable dataset. Turnaround is described as days rather than months, with claims of millimeter-level 3D accuracy and 100% human verification. Several quantitative figures on the site (hours captured daily, hours used for GR00T pretraining, throughput multiplier per annotator) render as placeholder zeros in the fetched page and are therefore not recoverable."]

Business model

Hyphenbox operates as a service layer between data capture and model training, taking partners' raw egocentric footage and returning enriched, human-verified datasets. Its site frames the offering as sitting on the research side of the pipeline while collection partners continue to scale field capture.

The website presents an indicative per-hour-of-footage economics comparison, stating that training-ready datasets command roughly 6-7x the prevailing rate of raw egocentric capture, implying pricing tied to hours of footage processed. No explicit price list or contract structure is disclosed.

Traction

No verifiable customer names, contract values, or usage volumes are disclosed in the sources. Quantitative claims on the site regarding daily hours captured by partners and per-annotator throughput multiples appear as unpopulated placeholders in the fetched page.

Latest developments

The website references a 2026 horizon in its capture-versus-usable-data framing and cites NVIDIA Isaac GR00T N1.7 pretraining data as evidence for its thesis. No funding, customer, or partnership announcements appear in the available sources.

Full profile — market position, technology, go-to-market, geography, risks & controversies

Market position

The company describes itself as occupying the annotation layer between capture and training in the physical AI data supply chain, complementing rather than competing with data-collection operators.

Hyphenbox argues that generic labeling tools and off-the-shelf pose estimators were not built for embodied, first-person data, and differentiates on embodied-specific annotation primitives delivered as one pipeline, occlusion-stable millimeter-level 3D hand reconstruction, physics-aware full-body pose from a single egocentric camera, and 100% expert human verification rather than crowdsourced labeling.

Technology

The stack combines a proprietary vision-language model that pre-labels dense action segments and scene context with 2D-to-3D reconstruction models for hand and full-body pose, followed by expert human review. Hand output is 21 keypoints per hand at stated millimeter-level accuracy and designed to hold up under self-occlusion; body output is SMPL plus hands with 52 joints and root trajectory, described as physics-aware and temporally smooth from a single egocentric camera.

Go-to-market

Direct outreach and inbound consultation booking through the company website, with a call to action aimed specifically at organizations already collecting egocentric video, plus a direct email channel to the research team.

Frontier robotics labs, egocentric data-collection partners operating head-mounted rigs in homes, factories, and retail environments, and any organization whose roadmap is constrained by the availability of training-ready embodied data.

Geography

The public sources do not state office locations or served regions; the site refers generally to capture across homes, factories, and retail floors.

Risks & controversies

Publicly available information is limited to the company's own website, so all technical and commercial claims are self-reported and unverified by third parties. Note also that an unrelated consumer gift-box business has used the HyphenBox name and the hyphenbox.com domain in earlier coverage; that entity is a distinct company and its coverage was excluded from this profile.

Compiled by commissioned research from 4 cited public sources — announcements, filings, and press listed under research sources below.

Key figures

latest reported
3D hand keypoints tracked per handJan 202621 keypoints
Full-body pose joints reconstructed (SMPL + hands)Jan 202652 joints
Indicative price multiple of training-ready dataset vs raw egocentric capture, pJan 20266-7x
Share of labels passing expert human reviewJan 2026100%
Stated 3D reconstruction accuracyJan 2026millimeter-level

Company-reported or press-reported figures, each dated to when it was claimed — not independently audited.

Research sources · 4

primary sources listed

4 public sources were cited for this profile; the first-party ones are listed here.

Frequently asked questions

What does Hyphenbox do?
Hyphenbox turns raw egocentric and multi-modal video into training-ready annotated datasets for general-purpose robotics.
Who founded Hyphenbox?
Hyphenbox was founded by Vishruth Nagendra, Shreyash Gupta in 2026.
Who are Hyphenbox's investors?
Hyphenbox's investors include Entrepreneur First.
Where is Hyphenbox headquartered?
Hyphenbox is headquartered in Bangalore, IN.