ai · May 1, 2026

Standard Intelligence raises $75M to develop efficient computer use models

SiliconANGLE News · View original source

Standard Intelligence raises $75M to develop efficient computer use models

Standard Intelligence Inc., a burgeoning artificial intelligence startup with a compact team of six, has successfully secured $75 million in funding to further its innovative endeavors. This funding round was spearheaded by notable venture capital firms Sequoia and Spark Capital, alongside contributions from several angel investors, including the distinguished AI researcher Andrej Karpathy.

The core of Standard Intelligence's offering is its foundation model, FDM-1, which is specifically tailored for computer use tasks. These tasks involve an AI's interaction with applications through their graphical interfaces. According to the company, FDM-1 is capable of executing a diverse array of activities, from scanning software for security vulnerabilities to operating computer-aided design (CAD) programs.

Traditionally, computer use models are trained using screenshots that depict human interactions with various applications. This process requires meticulous manual annotation, where each screenshot must be accompanied by a natural language description that explains the depicted actions. For instance, a sequence of screenshots illustrating an online shopping process would necessitate detailed descriptions of each step taken during the purchase.

In a departure from this conventional method, Standard Intelligence has opted to train FDM-1 using video footage rather than static screenshots. This innovative approach is complemented by the use of an inverse dynamics model (IMD), a type of neural network that automatically generates explanations for screenshots. This shift from manual to automated annotation significantly reduces costs, enabling researchers to compile much larger training datasets than previously feasible.

Standard Intelligence has amassed an impressive training dataset comprising 11 million hours of footage, vastly outpacing the next best open-source alternative. The expansion of an AI model's training dataset is crucial, as it typically enhances the quality of the model's outputs. Demonstrations of FDM-1 showcase its capabilities, including a video where it designs a metal component using a widely-used engineering application. In another instance, the model was fine-tuned for just one hour to control an autonomous vehicle via a website interface, successfully learning to drive the vehicle.

Another key feature of FDM-1 is its efficiency in utilizing hardware resources. Unlike many AI models that rely on complex chain-of-thought reasoning or external tools to perform tasks, FDM-1 operates with a streamlined approach. A significant aspect of its efficiency stems from its video encoder, which Standard Intelligence claims is 100 times more efficient than the alternative offered by OpenAI Group PBC.

A video encoder is a software component that converts video footage into mathematical representations that AI models can process. These representations can consume substantial memory storage. While it is possible to minimize the hardware requirements, this often comes at the expense of output quality. FDM-1’s encoder employs a masked compression objective, which selectively removes less important segments of the footage. This technique allows the model to maintain data quality while reducing memory usage. Standard Intelligence asserts that this encoder enables models with a context window of 1 million tokens to process two hours of 30 frames per second (FPS) video per prompt.

The newly acquired funding will be allocated towards enhancing the company’s computing capacity. Furthermore, Standard Intelligence aims to focus on developing AI safety guardrails that are specifically optimized for computer use models, ensuring that their technology remains secure and reliable as it evolves.

In summary, Standard Intelligence's innovative approach to training AI models, particularly through the use of video footage and automated annotation, positions it at the forefront of AI development for computer use tasks. With substantial funding and a clear vision for the future, the company is poised to make significant advancements in the field of artificial intelligence.

Why it matters

The advancements made by Standard Intelligence in developing FDM-1 could have far-reaching implications for creators and technologists. By leveraging video footage and automated annotation, the company is not only streamlining the training process for AI models but also expanding the potential applications of AI in various fields. This could lead to more efficient workflows in industries such as software development, engineering, and design, where AI can assist in complex tasks that require interaction with various applications.

Moreover, the efficiency of FDM-1 in utilizing hardware resources may democratize access to advanced AI capabilities. Smaller companies and individual creators, who may have previously been constrained by the need for high-end computing power, could now harness the potential of sophisticated AI tools without prohibitive costs. As AI becomes more accessible, it may foster innovation and creativity across diverse sectors, enabling a broader range of users to integrate AI into their workflows.

Additionally, the focus on developing AI safety guardrails is critical as the technology continues to evolve. Ensuring that AI systems operate safely and ethically is paramount, particularly as they become more integrated into everyday applications. Standard Intelligence's commitment to safety in its models could set a precedent for responsible AI development, encouraging other companies to prioritize safety measures in their own technologies.

Frequently asked questions

What is FDM-1?
FDM-1 is a foundation model developed by Standard Intelligence, optimized for computer use tasks, allowing AI to interact with applications through their graphical interfaces.
How does FDM-1 differ from traditional AI models?
Unlike traditional AI models that rely on screenshots and manual annotations, FDM-1 is trained on video footage and uses an inverse dynamics model for automatic annotation.
What are the implications of FDM-1's efficiency?
FDM-1's efficiency in using hardware resources may allow smaller companies and individual creators to access advanced AI capabilities without needing high-end computing power.

AI & art news in your inbox, daily

The day's top stories, summarized. Free, no spam, unsubscribe anytime.