Hyphenbox
Entrepreneur First '26Bangalore, IN · Founded 2026 · 1 known investors
Hyphenbox converts raw egocentric and multi-modal video into training-ready datasets for general-purpose robotics, providing frame-accurate action segmentation, 3D hand reconstruction, and full-body pose reconstruction. It serves frontier robotics labs and data-collection partners whose model development is bottlenecked on enriched, human-verified training data.
Also known as Hyphenbox
Founders & leadership
Hyphenbox was founded in 2026 by Vishruth Nagendra and Shreyash Gupta.


Investors · 1
Company profile
researched Aug 2026Hyphenbox positions itself as data infrastructure for physical AI and dexterous manipulation, converting raw egocentric and multi-modal capture into datasets that robotics models can train on directly. Its stated output combines three annotation primitives delivered through a single pipeline: dense action labels (frame-accurate action segmentation across long-horizon manipulation, including scene context, object state, and contact sequences), 2D-to-3D hand tracking (reconstruction described as millimeter-level, with 21 keypoints per hand, intended to remain stable under self-occlusion and close-range object interaction), and 2D-to-3D full-body pose (SMPL plus hands, 52 joints with root trajectory, described as physics-aware and temporally smooth, recovered from a single egocentric camera).
The company frames general-purpose robotics as a data problem rather than a model problem, arguing that capture volume is scaling faster than the supply of usable, training-ready data, and that generic labeling tools and off-the-shelf pose estimators were not designed for embodied, first-person footage. As supporting evidence, the site cites EgoScale / NVIDIA Research (arxiv.org) for the claim that the egocentric video used to pretrain NVIDIA Isaac GR00T N1.7 was explicitly action-labeled with dense 3D hand tracking and camera motion derived from raw footage.
The pipeline described runs from raw ego video (video, depth, and other signals captured in homes, factories, and retail environments), through pre-labeling by a proprietary vision-language model that proposes dense action segments and scene context, then 2D-to-3D hand and body reconstruction, then human-in-the-loop QA in which every label is reviewed by an expert reviewer rather than crowdsourced workers, producing a model-consumable dataset. Turnaround is described as days rather than months, with claims of millimeter-level 3D accuracy and 100% human verification. Several quantitative figures on the site (hours captured daily, hours used for GR00T pretraining, throughput multiplier per annotator) render as placeholder zeros in the fetched page and are therefore not recoverable."]
Business model
Hyphenbox operates as a service layer between data capture and model training, taking partners' raw egocentric footage and returning enriched, human-verified datasets. Its site frames the offering as sitting on the research side of the pipeline while collection partners continue to scale field capture.
The website presents an indicative per-hour-of-footage economics comparison, stating that training-ready datasets command roughly 6-7x the prevailing rate of raw egocentric capture, implying pricing tied to hours of footage processed. No explicit price list or contract structure is disclosed.
Traction
No verifiable customer names, contract values, or usage volumes are disclosed in the sources. Quantitative claims on the site regarding daily hours captured by partners and per-annotator throughput multiples appear as unpopulated placeholders in the fetched page.
Latest developments
The website references a 2026 horizon in its capture-versus-usable-data framing and cites NVIDIA Isaac GR00T N1.7 pretraining data as evidence for its thesis. No funding, customer, or partnership announcements appear in the available sources.
▸Full profile — market position, technology, go-to-market, geography, risks & controversies
Market position
The company describes itself as occupying the annotation layer between capture and training in the physical AI data supply chain, complementing rather than competing with data-collection operators.
Hyphenbox argues that generic labeling tools and off-the-shelf pose estimators were not built for embodied, first-person data, and differentiates on embodied-specific annotation primitives delivered as one pipeline, occlusion-stable millimeter-level 3D hand reconstruction, physics-aware full-body pose from a single egocentric camera, and 100% expert human verification rather than crowdsourced labeling.
Technology
The stack combines a proprietary vision-language model that pre-labels dense action segments and scene context with 2D-to-3D reconstruction models for hand and full-body pose, followed by expert human review. Hand output is 21 keypoints per hand at stated millimeter-level accuracy and designed to hold up under self-occlusion; body output is SMPL plus hands with 52 joints and root trajectory, described as physics-aware and temporally smooth from a single egocentric camera.
Go-to-market
Direct outreach and inbound consultation booking through the company website, with a call to action aimed specifically at organizations already collecting egocentric video, plus a direct email channel to the research team.
Frontier robotics labs, egocentric data-collection partners operating head-mounted rigs in homes, factories, and retail environments, and any organization whose roadmap is constrained by the availability of training-ready embodied data.
Geography
The public sources do not state office locations or served regions; the site refers generally to capture across homes, factories, and retail floors.
Risks & controversies
Publicly available information is limited to the company's own website, so all technical and commercial claims are self-reported and unverified by third parties. Note also that an unrelated consumer gift-box business has used the HyphenBox name and the hyphenbox.com domain in earlier coverage; that entity is a distinct company and its coverage was excluded from this profile.
Compiled by commissioned research from 4 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
▸Research sources · 4
primary sources listed
- Hyphenboxhyphenbox.com · web
4 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Hyphenbox do?
- Hyphenbox turns raw egocentric and multi-modal video into training-ready annotated datasets for general-purpose robotics.
- Who founded Hyphenbox?
- Hyphenbox was founded by Vishruth Nagendra, Shreyash Gupta in 2026.
- Who are Hyphenbox's investors?
- Hyphenbox's investors include Entrepreneur First.
- Where is Hyphenbox headquartered?
- Hyphenbox is headquartered in Bangalore, IN.
