ai · June 8, 2026

google/gemma-4-12B-it-qat-q4_0-gguf · Hugging Face

Huggingface.co · View original source

google/gemma-4-12B-it-qat-q4_0-gguf · Hugging Face

In a significant advancement in the field of artificial intelligence, Google DeepMind has unveiled new versions of its Gemma 4 family of models, optimized with Quantization-Aware Training (QAT). This innovative approach allows these models to maintain a quality comparable to bfloat16 while dramatically reducing the memory requirements necessary for loading the model. The release includes four distinct QAT checkpoints, expanding the accessibility and usability of these advanced AI tools.

The Gemma 4 models are multimodal, capable of processing both text and image inputs, with audio support available on specific versions (E2B, E4B, and 12B). They are designed to generate text outputs, making them versatile for various applications. This latest iteration includes open-weight models in both pre-trained and instruction-tuned formats, enhancing their adaptability for different tasks. Notably, Gemma 4 features an extensive context window of up to 256K tokens and supports over 140 languages, showcasing its multilingual capabilities.

Model Architecture and Capabilities

Gemma 4 models are built on a combination of Dense and Mixture-of-Experts (MoE) architectures, which makes them suitable for a wide range of tasks, including text generation, coding, and reasoning. The models come in five different sizes: E2B, E4B, 12B, 26B A4B, and 31B. This variety enables deployment across various environments, from high-end mobile devices to powerful servers, thereby democratizing access to cutting-edge AI technology.

Each model in the Gemma 4 family is engineered to be a highly capable reasoner, featuring configurable thinking modes that enhance their cognitive abilities. The extended multimodalities allow the models to process not only text and images but also video and audio, particularly in the E2B, E4B, and 12B versions. This broad capability set positions Gemma 4 as a robust tool for developers and researchers alike.

One of the standout features of the Gemma 4 models is their hybrid attention mechanism, which combines local sliding window attention with full global attention. This design ensures that the final processing layer retains a global awareness, crucial for handling complex tasks that require long-context understanding. Additionally, the models utilize unified Keys and Values in their global layers and apply Proportional RoPE (p-RoPE) to optimize memory usage for lengthy contexts.

The naming convention for the models includes the letter 'E' in E2B and E4B, which stands for 'effective' parameters. The smaller models leverage Per-Layer Embeddings (PLE) to enhance parameter efficiency for on-device deployments. Instead of increasing the number of layers or parameters, PLE allows each decoder layer to utilize a small embedding for every token, resulting in a smaller effective parameter count compared to the total parameter count.

The Gemma 4 12B Unified model distinguishes itself with an encoder-free architecture, streamlining the processing of multimodal data. By eliminating dedicated encoders, this model directly projects raw image patches and audio waveforms into the LLM's embedding space, thereby reducing latency and enabling more efficient fine-tuning.

Safety and Ethical Considerations

Safety and ethical considerations are paramount in the development of AI models, especially those intended for open use. The Gemma 4 models have undergone rigorous safety evaluations, similar to those applied to Google's proprietary Gemini models. Developed in collaboration with internal safety and responsible AI teams, these models have been subjected to both automated and human evaluations to enhance their safety features.

The improvements in content safety are notable, with Gemma 4 models outperforming their predecessors, Gemma 3 and 3n, in various safety testing categories. The evaluations indicated minimal policy violations across text-to-text and image-to-text tasks, demonstrating a commitment to responsible AI development. The models aim to prevent harmful content generation while maintaining low rates of unjustified refusals.

As multimodal models capable of processing vision, language, and audio, Gemma 4 opens up a plethora of potential applications across diverse industries. The development team has carefully considered a range of ethical concerns, ensuring that the model's capabilities align with responsible AI practices.

In conclusion, the release of the Gemma 4 family of models marks a pivotal moment in the evolution of open AI technologies. With their advanced capabilities, efficient architectures, and a strong emphasis on safety and ethics, these models are well-positioned to become integral tools for creators and technologists alike, fostering innovation in various fields.

Frequently asked questions

What is Quantization-Aware Training?
Quantization-Aware Training (QAT) is a technique that allows machine learning models to maintain performance while reducing the precision of the computations, thereby lowering memory and computational requirements.
What are the different sizes of the Gemma 4 models?
The Gemma 4 models are available in five sizes: E2B, E4B, 12B, 26B A4B, and 31B, catering to various deployment environments from mobile devices to powerful servers.
How does the hybrid attention mechanism in Gemma 4 work?
The hybrid attention mechanism combines local sliding window attention with full global attention, allowing the model to efficiently process complex tasks while maintaining a global context.

AI & art news in your inbox, daily

The day's top stories, summarized. Free, no spam, unsubscribe anytime.