Standard Intelligence Raises $75 Million for Computer-Use Models
The six-person startup is training AI to learn digital tasks from video, with efficiency and safety becoming the key tests for its approach.

SAN FRANCISCO, Calif. - Standard Intelligence has announced the closing of a $75 million financing round to accelerate the development of foundation models specifically designed to operate computer interfaces. The investment underscores a growing industry-wide push to transition artificial intelligence from a passive generator of text and images into an active agent capable of navigating digital environments. This round was led by Sequoia Capital and Spark Capital, two of the most prominent venture firms in Silicon Valley, with additional participation from a concentrated group of technology investors and high-profile researchers, including former Tesla AI director and OpenAI co-founder Andrej Karpathy.
The six-person San Francisco-based startup is entering a competitive field known as the Large Action Model (LAM) sector, where the goal is to bridge the gap between human intent and software execution. While traditional automation relies on rigid Application Programming Interfaces (APIs) or brittle scraping scripts, Standard Intelligence is developing FDM-1. This foundation digital model is trained to understand the causal relationship between a user's inputs and the resulting visual changes on a screen, essentially learning to 'see' and 'interact' with software in a manner similar to a human operator.
Central to the company's technical strategy is an approach that combines massive video datasets with an inverse-dynamics system. This system is designed to estimate the specific mouse movements, keyboard strokes, or interface actions that produced every visible result in a video sequence. By inferring the underlying actions behind the pixels, the company can bypass one of the most significant bottlenecks in AI development: the need for intensive, manual data labeling. Typically, training an agent to use software requires humans to painstakingly document every click, but Standard Intelligence’s method allows for more scalable self-supervised learning.
The scale of the data involved is substantial for a team of this size. Standard Intelligence stated that it has successfully assembled a library of 11 million hours of footage, providing the diverse visual training ground necessary for a model to generalize across different operating systems and applications. To parse this data, the company developed a proprietary video encoder capable of processing long sequences with high efficiency, addressing the computational costs that often plague high-resolution video analysis in machine learning workflows.
A key technical milestone highlighted by the company is its million-token context window. In practical terms, this allows the model to examine approximately two hours of video at a rate of 30 frames per second within a single processing pass. Maintaining such a large context window is critical for tasks that require long-term memory, such as navigating complex enterprise workflows where an action taken in one window may depend on information seen in a different application twenty minutes earlier. This capacity for long-form reasoning is often what separates simple macro-recorders from true digital agents.
Early demonstrations of the technology have showcased FDM-1 controlling sophisticated design software, an environment often considered difficult for AI due to its heavy reliance on precise spatial coordinates and idiosyncratic toolbars. Furthermore, the company has demonstrated the model's ability to quickly adapt to new web interfaces after only limited additional training. This suggests a level of flexibility that could solve the 'brittleness' problem, where a minor update to a website's layout typically breaks traditional automation tools.
The potential market for this technology is vast, particularly in the realm of repetitive work across enterprise applications that lack clean programming interfaces. Many legacy systems in finance, logistics, and healthcare do not offer modern APIs, forcing employees to spend hours manually transferring data between disparate screens. Standard Intelligence aims to automate these 'white-collar' manual tasks by creating a model that treats the user interface itself as the primary API, allowing the AI to work alongside or in place of human users in existing software ecosystems.
However, the move toward models that can use arbitrary software introduces significant technical and ethical risks. A model capable of clicking any button or entering any text may take damaging actions if it misunderstands a user’s ultimate goal or encounters a manipulated screen designed to deceive it. Unlike a chatbot that merely provides a wrong answer, a computer-use agent could potentially delete files, move funds, or expose sensitive data if it lacks robust guardrails. This unpredictability remains a primary concern for enterprise adoption of autonomous agents.
According to the company, the $75 million in new capital will primarily fund the acquisition of computing capacity, further research into model architectures, and the development of safety controls. The high cost of specialized hardware like H100 GPUs makes such a large round necessary for a startup attempting to train frontier-scale models from scratch. Even with a small headcount, the infrastructure costs associated with processing 11 million hours of video are significant, necessitating deep-pocketed backers who are willing to bet on high-risk, high-reward research.
Industry analysts note that Standard Intelligence enters a crowded market where tech giants like Google, Microsoft, and specialized startups like Adept are also racing to master computer-use models. The success of FDM-1 will likely depend on whether its efficiency claims hold up when subjected to independent comparisons. In the world of AI, speed and cost-per-task are often as important as raw capability, particularly for businesses looking to deploy these models at scale across thousands of workstations.
Beyond performance, the startup must prove that its models possess the common sense to ask for help before an uncertain action becomes an irreversible mistake. Developing a mechanism for the AI to 'pause and verify' is a key focus of current research, as dependable limits may be as vital as total task coverage for corporate clients. A digital agent that is 99% accurate but causes a catastrophic error once a month remains a liability for most large organizations.
As Standard Intelligence moves forward, the focus will likely shift from pure research and data gathering to real-world reliability tests. The participation of researchers like Karpathy suggests a strong technical foundation, but the transition from a laboratory setting to the messy, unpredictable environments of enterprise desktops will be the ultimate test. The company’s ability to maintain its lean structure while competing against the massive engineering teams of established AI labs will also be watched closely by the venture community.
For now, the $75 million round establishes Standard Intelligence as a significant player in the race to define the future of human-computer interaction. If the startup can deliver an agent that is both efficient enough to run economically and safe enough to operate autonomously, it could unlock a new wave of productivity in sectors that have remained largely untouched by the first generation of generative AI models. The next year will be critical as the company seeks to turn its FDM-1 prototype into a production-ready tool for its first wave of early adopters.
Sources
Written by
The Company Wire Staff
Reporting from The Company Wire newsroom. Staff bylines cover funding rounds, product launches and company news verified against primary sources.


