Standard Intelligence PBC
San Francisco, US · 7 known investors
Standard Intelligence builds AI foundation models designed to explore and learn like humans, including a general computer action model that can navigate websites, complete CAD modeling sequences, and drive a car in real time, as well as an open base model for conversational speech.
Also known as SI · si.inc · Standard Intelligence
Investors · 7
Also in the syndicate · 3
Company profile
researched Aug 2026Standard Intelligence PBC (si.inc) is a San Francisco-based AI research company developing foundation models intended to explore and learn like humans, with a stated long-term mission of building aligned artificial general intelligence. Its principal published work is FDM-1, described as the first fully general computer action model: a foundation model for computer use trained directly on video rather than screenshots, which the company demonstrates navigating and fuzzing websites, performing multi-action CAD modeling sequences in Blender, and driving a car on public streets in San Francisco at 30 frames per second [0][1].
The company also released hertz-dev, an open-sourced 8.5-billion-parameter, full-duplex, audio-only autoregressive base model for interactive conversational speech. Hertz-dev comprises hertz-codec, a convolutional audio VAE encoding mono 16 kHz speech into an 8 Hz latent representation at a KL-regularized 1 kbps bitrate (5M encoder and 95M decoder parameters), and hertz-ar, a 40-layer, 8.4-billion-parameter decoder-only transformer with a 2,048-token context (roughly 4.5 minutes). The company describes it as the first publicly released base model for conversational audio, positioned as a starting point for downstream fine-tuning [0][3]. Separately, the company built "the heap," a 30-petabyte data storage cluster in downtown San Francisco, reportedly for under $500,000 [0].
Alongside model work, Standard Intelligence states it is investing in blue-sky research on alignment for "general learners," arguing that current alignment techniques are insufficient for models with human-level learning capabilities and that it is studying small versions of this problem in controlled environments [2]."]
Business model
No commercial product, pricing, or customer-facing offering is described in the available sources; activity to date consists of foundation-model research, open-sourced model weights (hertz-dev checkpoints released publicly), and internally built compute and storage infrastructure, funded by venture capital [0][2][3].
Not disclosed in the sources. A third-party directory publishes an estimated annual revenue figure of roughly $1.7M, but explicitly labels it a model-based estimate derived from comparable companies rather than reported data [4].
Traction
Reported traction is research output and infrastructure rather than commercial metrics: FDM-1 demonstrations across CAD, driving, and GUI fuzzing; an 11-million-hour labeled screen-recording corpus; a self-built 30-petabyte storage cluster; open-sourced hertz-dev checkpoints; and a $75M round with participation from named angels and advisors [0][1][2][3].
Latest developments
The most recent disclosed developments are the release of FDM-1 and a $75M funding round from Sequoia and Spark Capital, with partners Sonya Huang, Mikowai Ashwill, and Yasmin Razavi, plus angels and advisors Milan Kovac, Stanley Druckenmiller, and Andrej Karpathy. The company says the round unlocks several orders of magnitude of compute scaling for the FDM series and funds alignment research, and that it is hiring onto a six-person team [1][2].
▸Full profile — market position, technology, go-to-market, geography, history, risks & controversies
Market position
The company positions FDM-1 against the prevailing approach of fine-tuning vision-language models on contractor-annotated screenshots plus task-specific reinforcement learning environments, which it says limits agents to a few seconds of context; it cites OpenAI's Video PreTraining (VPT) work and VideoAgentTrek as prior IDM-based approaches with shorter context or no video context. For audio, it claims hertz-dev's real-world latency is roughly 2x lower than the previous state of the art and that hertz-codec outperforms SoundStream and EnCodec at 6 kbps while being on par with DAC at 8 kbps in subjective evaluations [1][3].
Claimed differentiators are training and inference directly on long-context video rather than screenshots, an unusually token-efficient video encoder, the ability to learn unsupervised from internet-scale video without contractor labeling (the largest cited open computer-action dataset is under 20 hours of 30 FPS video versus its 11-million-hour corpus), demonstrated transfer from screen tasks to real-world driving, and an openly released full-duplex conversational speech base model [1][3].
Technology
FDM-1 is trained with a three-stage recipe: an inverse dynamics model (IDM) is trained on 40,000 hours of contractor-labeled screen recordings; the IDM then labels an 11-million-hour screen-recording video corpus; and a forward dynamics model is autoregressively trained on next-action prediction, with an output token space of key presses and mouse-movement deltas. A central component is a video encoder that compresses nearly two hours of 30 FPS video into about 1M tokens, which the company claims is 50x more token-efficient than the prior state of the art and 100x more than OpenAI's encoder. The system uses OS checkpointing and forking VM infrastructure to enable test-time compute and state-space exploration. The self-driving demo forked openpilot's joystick mode and used under one hour of fine-tuning data. On the audio side, hertz-dev uses causal convolutions for streaming inference, 15-bit quantized phonetic latents via Finite Scalar Quantization, and duplex handling via two concatenated projection heads, achieving a stated 80ms theoretical average latency and 120ms real-world latency on a single RTX 4090 [1][3].
Go-to-market
Distribution to date is through public research publication and open-source release: hertz-dev weights and code are published on GitHub with downloadable checkpoints, and FDM-1 results are shared via demo videos and technical blog posts. The company solicits researcher and engineer applications and investor contact directly through its site and email [1][3].
The company frames FDM-1 as a model with long-context training needed to act as a coworker for CAD, finance, engineering, and eventually machine learning research, and highlights automated GUI testing/fuzzing as an application area; no named customers are disclosed [1].
Geography
Operations are based in San Francisco, United States, including the downtown San Francisco storage cluster and local self-driving demonstrations [0][1][2][4].
History
Public blog posts trace a progression from the hertz-dev conversational speech base model, released when the company described itself as a team of four in San Francisco, to FDM-1, a general computer action model trained on IDM-labeled video, followed by a $75M funding round from Sequoia and Spark Capital with the team at six people in San Francisco [2][3]. A third-party directory lists a founding year of 2023 [4].
Risks & controversies
No controversies are reported in the sources. Notable uncertainties: capability claims for FDM-1 and hertz-dev are self-published by the company and not independently verified in the material provided; the team is very small (four people at the hertz-dev release, six at the funding announcement); the company states current alignment techniques are insufficient for models with human-level learning capabilities; and third-party funding, revenue, and valuation figures conflict with or are explicitly estimated rather than reported (a directory cites $90M total raised and modeled revenue/valuation ranges versus the company's stated $75M round) [2][3][4].
Compiled by commissioned research from 7 cited public sources — announcements, filings, and press listed under research sources below.
Key figures
latest reportedCompany-reported or press-reported figures, each dated to when it was claimed — not independently audited.
Timeline · 4
launches, deals, and filingsStandard Intelligence announced $75M in new funding from Sequoia and Spark Capital, partnering with Sonya Huang, Mikowai Ashwill, and Yasmin Razavi, and adding angels and advisors including Milan Kovac, Stanley Druckenmiller, and Andrej Karpathy. The company said the round unlocks several orders of magnitude of compute scaling for the FDM model series and funds alignment research.
$75M source ↗
Constructed a 30-petabyte data storage cluster in downtown San Francisco for under $500,000.
Published FDM-1, a foundation model for computer use trained on IDM-labeled video from an 11-million-hour screen recording corpus, with demonstrations of CAD work in Blender, autonomous driving of a car in San Francisco via key presses, and GUI fuzzing that surfaced a bug in a mock banking app.
Released hertz-dev, an 8.5B-parameter full-duplex, audio-only autoregressive base model for interactive conversational speech, comprising hertz-codec and hertz-ar, with public weights/checkpoints on GitHub.
Dated company events from announcements, filings, and press; legal rows summarize public dockets and regulator releases.
▸Research sources · 7
primary sources listed
- Standard Intelligence PBCsi.inc · web
7 public sources were cited for this profile; the first-party ones are listed here.
Frequently asked questions
- What does Standard Intelligence PBC do?
- AI research company building video-pretrained computer action foundation models and open speech models, backed by $75M from Sequoia and Spark.
- Who are Standard Intelligence PBC's investors?
- Standard Intelligence PBC's investors include Sequoia Capital, Abstract Ventures, Spark Capital, Transpose Platform Management.
- Where is Standard Intelligence PBC headquartered?
- Standard Intelligence PBC is headquartered in San Francisco, US.
