Meta debuts internally developed AI chips for inference workloads
SiliconANGLE News · View original source

Meta Platforms Inc. has unveiled four custom-designed chips aimed at enhancing its internal artificial intelligence (AI) workloads. This announcement marks a significant step in the company’s ongoing efforts to develop specialized hardware tailored for AI applications, which are increasingly critical to its operations and user engagement strategies.
The latest update on Meta's processor development comes after a previous report in April 2024, when the company introduced the MTIA 200, a custom AI accelerator with an energy consumption of 90 watts. The new chips, however, demonstrate a substantial leap in capability, particularly with the most advanced of the four, the MTIA 500, which boasts a thermal design point of 1,700 watts.
Among the newly introduced chips, the MTIA 300 stands out as it is already in production. This chip is designed to handle ranking and recommendation tasks, similar to the MTIA 200. It is capable of delivering 1.2 petaflops of performance while processing data in the MX8 format and is equipped with 216 gigabytes of High Bandwidth Memory (HBM). The architecture of the MTIA 300 includes one compute chiplet and two network chiplets, along with several stacks of HBM. Each compute chiplet consists of a grid of processing elements (PEs), which are designed with redundancy to enhance manufacturing yield.
The other three chips in the lineup are intended for a wider array of applications, extending beyond ranking and recommendation to include generative AI software, such as large language models (LLMs). The MTIA 500, the most powerful chip in this new series, can achieve 10 petaflops of performance with MX8 data and supports a more efficient data format known as MX4. This advancement allows for a reduction in the data size that AI models must process, thereby accelerating response times.
The MTIA 500 is constructed with four logic chiplets and is surrounded by multiple HBM stacks, allowing it to store up to 516 gigabytes of data—double that of the MTIA 300. Additionally, it features a System on Chip (SoC) chiplet that facilitates data transfer to and from the host server, enhancing overall efficiency.
Expected to enter production in 2027, the MTIA 500 will be accompanied by the MTIA 450, a less advanced chip also optimized for generative AI inference workloads. Both chips incorporate circuits designed to speed up specific, resource-intensive tasks within the inference process, such as FlashAttention, a widely used implementation of the attention mechanism that LLMs utilize to interpret input data.
Meta's engineers emphasized that the MTIA 400, 450, and 500 share the same physical infrastructure, including chassis, racks, and network systems. This modular approach allows for seamless integration of new chip generations into existing setups, streamlining the transition from development to production deployment. Such reusable designs minimize the resources required for the development and deployment of multiple chip generations.
In addition to the hardware advancements, Meta employs custom compilers to optimize AI models specifically for its MTIA chips. The company also utilizes a software module called the Hoot Collective Communications Library, which manages data flow between processors. This library performs certain calculations using transistors located close to memory cells, which helps to reduce data travel times and enhance overall performance.
This announcement follows closely on the heels of Meta's recent agreement to purchase billions of dollars worth of processors from Nvidia Corp. and Advanced Micro Devices Inc. Furthermore, reports indicate that Meta plans to incorporate Google LLC’s Tensor Processing Units (TPUs) to support its LLM operations. This strategic move underscores Meta's commitment to bolstering its AI capabilities through both proprietary and third-party technologies.
Why it matters
The introduction of Meta's custom AI chips signifies a pivotal moment for the company as it seeks to enhance its AI infrastructure. By developing specialized hardware, Meta aims to optimize the performance of its AI models, which are integral to its business operations, from content recommendations to advertising strategies. The MTIA series of chips not only showcases Meta's engineering prowess but also reflects a broader trend in the tech industry towards custom silicon designed for specific workloads.
For creators and technologists, the implications of Meta's chip development are profound. The ability to process large amounts of data quickly and efficiently can lead to more sophisticated AI applications, enabling creators to develop richer, more engaging content. Moreover, as generative AI continues to evolve, the demand for powerful inference capabilities will only grow, making advancements like those seen in the MTIA chips crucial for staying competitive in the AI landscape.
As Meta continues to innovate in AI hardware, it sets a precedent for other companies in the industry. The integration of custom chips with existing technology infrastructures could inspire similar initiatives among competitors, ultimately accelerating the pace of AI development and deployment across various sectors.
Frequently asked questions
- What are Meta's new AI chips designed for?
- Meta's new AI chips are designed to enhance internal artificial intelligence workloads, specifically for tasks such as ranking and recommendation, as well as generative AI applications.
- When will the MTIA 500 chip enter production?
- The MTIA 500 chip is expected to enter production in 2027.
- What is the performance capability of the MTIA 500 chip?
- The MTIA 500 chip can provide 10 petaflops of performance when processing data in the MX8 format.
Related stories
AI & art news in your inbox, daily
The day's top stories, summarized. Free, no spam, unsubscribe anytime.
