Squeezebits
Acquired3 known investors
Seoul-based AI inference optimization startup whose compression and quantization software cuts the cost of running AI models.
Also known as SqueezeBits Β· SqueezeBits, Inc.
Investors Β· 3
Also in the syndicate Β· 1
Company profile
researched Aug 2026SqueezeBits is a Seoul, South Korea-based startup developing AI model lightweighting (compression) and inference optimization technology. Its core technique is quantization: reducing 32-bit model data to 4 bits or fewer while preserving model accuracy, which lowers the memory and compute required during inference and speeds up model execution. The company describes its work as building an ecosystem for designing, training and deploying highly compressed neural networks, and its technology is applied across mobile smartphones, laptops, edge devices and GPU cloud environments, supporting image, video, speech and natural-language models.
The company's products include the OwLite toolkit, which allows users who are not compression specialists to lighten, compare and analyze AI models, and Yetter, a generative AI API service built on its optimization inference engine for image and video generation with LLM services planned. SqueezeBits also publishes technical work on inference engines and hardware, including benchmarking of Intel Gaudi accelerators, guided decoding on vLLM and SGLang, disaggregated inference on Apple Silicon (NPU prefill with GPU decode), vocabulary trimming for small language models, and GraLoRA, a block-wise variant of LoRA fine-tuning. Its RoBoost Agent work integrates NVIDIA Cosmos world models to generate synthetic data for physical AI applications such as robotics and autonomous driving.
Founding story
The company was founded by researchers from the Neural Processing Unit (NPU) research team at POSTECH's graduate school. Its co-founders had published model-compression research at machine learning conferences including CVPR, NeurIPS and ICLR over the preceding seven years, with more than 70 international papers on deep learning acceleration, and had experience designing AI-specific hardware. Hyungjun Kim is a co-founder and chief executive officer. As of May 2024 Forbes described the company as founded two years earlier.
Business model
SqueezeBits sells AI optimization software, historically on a subscription basis, and also offers a generative AI API service (Yetter) built on its own inference engine. It has engaged in proof-of-concept and optimization projects with enterprise customers and in co-development partnerships with AI hardware vendors.
Subscription sales of optimization software, per Forbes, alongside a generative AI API service positioned as delivering image and video generation at reduced cost.
Traction
More than 20 companies, including Naver and SK Telecom, had completed technology verification (PoC) or projects with SqueezeBits as of January 2024. The company worked with global AI hardware firms including Intel and NVIDIA shortly after founding, and co-developed NPU model compression software with Rebellions in 2024.
Latest developments
In July 2026 Rebellions announced an agreement to acquire SqueezeBits as its first acquisition, following Rebellions' selection as the first direct investment target of Korea's National Growth Fund and its 2024 merger with Sapeon Korea; SqueezeBits is to continue operating independently. Recent product and research activity includes the Yetter generative AI API service, OwLite v2.5 with Qualcomm Neural Network support, Intel Gaudi optimization work, and RoBoost Agent synthetic data generation using NVIDIA Cosmos.
βΈFull profile β market position, technology, go-to-market, geography, history
Market position
Positioned as an AI inference optimization and model compression specialist working across multiple hardware platforms, with collaborations involving Intel, NVIDIA, Qualcomm and Rebellions; Rebellions selected it as its first acquisition target to extend from NPU hardware into optimization software and inference serving.
Founder team combining AI semiconductor design experience with published deep learning compression research, hardware-agnostic optimization spanning GPUs, NPUs, mobile and edge, and tooling that makes compression accessible to non-specialists.
Technology
Quantization of 32-bit model data to 4 bits or less while maintaining model performance, plus broader model compression and inference optimization. Product and research work spans the OwLite compression toolkit (with Qualcomm Neural Network support via Qualcomm AI Hub), the Yetter inference engine for diffusion models, vLLM and SGLang serving, guided decoding benchmarks, disaggregated inference on Apple Silicon NPUs and GPUs, vocabulary trimming for small language models, GraLoRA fine-tuning, and Intel Gaudi optimization for LLMs and diffusion models.
Go-to-market
The company combines direct enterprise engagements and PoCs with developer-community activity: technical blog publications and benchmarks, co-hosted workshops and meetups (Intel Gaudi with Lablup, vLLM with Rebellions, vLLM Korea meetups, a Modular Seoul developer meetup, and an Efficient AI model compression study group), and conference presence including a booth at GTC 2026 and events in Taipei and Singapore.
Companies deploying AI models in production that seek to reduce inference cost, including large Korean technology and telecom firms such as Naver and SK Telecom, AI semiconductor and hardware vendors, and developers running models on GPUs, NPUs and edge devices.
Geography
Headquartered in Seoul, South Korea, with stated plans to expand into overseas markets and conference and community activity in Korea, Taiwan, Singapore and the United States.
History
SqueezeBits raised about KRW 1 billion in seed funding before securing a KRW 2.5 billion pre-Series A round in January 2024 from Kakao Ventures, Samsung Next, POSCO Capital (POSCO Technology Investment) and POSTECH Holdings, with proceeds earmarked for strengthening its lightweighting technology and expanding internationally. By that point it had completed proofs of concept and projects with more than 20 companies, including Naver and SK Telecom, and had released the OwLite toolkit. From 2024 it worked with Rebellions on jointly developed model compression technology and dedicated software for Rebellions' NPU and on fostering an NPU-based open-source AI ecosystem in Korea. It launched the Yetter generative AI API service in 2025 and added Qualcomm Neural Network support to OwLite in v2.5. In July 2026 Rebellions announced it would acquire the company.
Compiled by commissioned research from 5 cited public sources β announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed β not independently audited.
Timeline Β· 7
launches, deals, and filingsAI semiconductor company Rebellions announced it will acquire SqueezeBits, its first acquisition, to extend from NPU hardware into optimization software and inference serving. SqueezeBits will continue to operate independently and will not be merged into Rebellions.
SqueezeBits introduced Yetter, a generative AI API service powered by its optimization inference engine, offering image and video services with LLM services planned.
OwLite v2.5 added official support for Qualcomm Neural Network (QNN) through integration with Qualcomm AI Hub.
SqueezeBits partnered with Intel to make Gaudi NPUs more usable in practice, optimizing LLMs and diffusion models, and co-hosted an Intel Gaudi hands-on workshop with Lablup.
SqueezeBits announced a KRW 2.5 billion pre-Series A round from Kakao Ventures, Samsung Next, POSCO Capital (POSCO Technology Investment) and POSTECH Holdings, to be used to strengthen its lightweighting technology and expand into overseas markets. Forbes valued the round at $1.8 million and noted KRW 1 billion raised earlier in seed funding.
$1.8M source β
SqueezeBits and Rebellions jointly developed model compression technology and dedicated software based on Rebellions' NPU and worked to foster an NPU-based open-source AI ecosystem for Korean developers.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
βΈResearch sources Β· 5
primary sources listed
- The official SqueezeBits Tech blogblog.squeezebits.com Β· web
5 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Squeezebits do?
- Seoul-based AI inference optimization startup whose compression and quantization software cuts the cost of running AI models.
- Who are Squeezebits's investors?
- Squeezebits's investors include Kakao Ventures, Samsung Next.
