ai · March 24, 2026

From experiment to production: A reliable architecture for version-controlled MLOps

Redhat.com · View original source

From experiment to production: A reliable architecture for version-controlled MLOps

In the rapidly evolving landscape of artificial intelligence and machine learning (AI/ML), managing the data that fuels these models presents a significant challenge. Practitioners often find themselves grappling with questions about data versions when model performance falters or when analyses yield inconsistent results. The dilemma of tracking which dataset was used for a particular training run or report can be daunting. However, Red Hat has introduced a promising solution aimed at alleviating these concerns with its new AI quickstart, which integrates the orchestration capabilities of Red Hat OpenShift AI with the data versioning functionality of lakeFS.

The Architecture Behind the Solution

At the core of this innovative quickstart is Red Hat OpenShift, a platform that provides a robust, enterprise-grade foundation for deploying applications. Unlike Kubernetes, which can often feel like a do-it-yourself project, OpenShift offers a more streamlined experience, handling scalability and security. This is crucial when transitioning AI models from local environments to production, ensuring that the underlying infrastructure can support the demands of real-world applications.

Building on this foundation is Red Hat OpenShift AI, where data scientists can perform the essential work of model development and deployment. This layer consolidates various tools into a unified dashboard, minimizing the need for users to navigate between disparate applications. It integrates Jupyter Notebooks for experimentation, model serving for deployment, and automated data pipelines, effectively bridging the gap between coding and delivering functional AI services.

The integration of lakeFS into this architecture is where data management truly excels. Traditionally, developers use Git for version control of code, but lakeFS extends this concept to data management. By acting as a control plane over object storage, lakeFS allows users to branch, commit, and revert datasets with the same ease as managing source code. This capability is particularly beneficial for ensuring the reproducibility of AI models; if unexpected behavior arises, users can roll back to the specific data state used during training.

The Benefits of lakeFS and OpenShift AI

One of the standout features of this architecture is its ability to handle multimodal data. lakeFS treats structured data, semi-structured JSON, and unstructured formats like images uniformly, providing a comprehensive version control system. This is especially important for industries that require strict compliance, such as finance and healthcare, as it allows organizations to maintain a verifiable chain of custody for training data.

The quickstart is not merely a basic demonstration; it is based on a real-world use case of fraud detection, providing a full lifecycle workflow. Users can learn how to effectively manage data as a dynamic entity, which is crucial for both auditing and debugging. By utilizing lakeFS as an AI data control plane on Red Hat OpenShift, organizations gain powerful tools to streamline their data management processes.

A significant advantage of this setup is the zero-copy branching feature, which enables teams to create copies of large datasets—such as a 1TB dataset for testing—almost instantaneously without duplicating the underlying data. This capability enhances data CI/CD practices, allowing teams to implement pre-merge hooks that validate data quality before it enters the training pipeline. If data ingestion issues arise, the system supports instant rollbacks, enabling users to revert to a previous state of the data repository swiftly.

The architecture also supports all AI data formats, ensuring that teams can apply version control across various types of data and their associated metadata. This comprehensive approach is vital for maintaining compliance and audit readiness, particularly in regulated sectors.

Why it matters

The implications of adopting this architecture for creators and technologists are profound. By leveraging lakeFS-based data versioning, machine learning teams can potentially deliver 2-3 times the number of models with fewer resources. This efficiency stems from the elimination of environment drift, rework, and dataset duplication, allowing teams to scale their output without necessitating an increase in personnel.

Faster experimentation cycles are another key benefit, as zero-copy branching enables teams to test new datasets and features without the delays associated with data duplication or infrastructure provisioning. This can reduce testing time by up to 80%, allowing organizations to innovate more rapidly.

Additionally, the architecture enhances risk management and operational stability. Immutable data commits create a clear chain of custody for training data, ensuring that every deployed model can be traced back to a specific dataset snapshot. This level of compliance is essential for meeting internal governance requirements and is particularly beneficial in regulated industries. In the event of faulty data entering the pipeline, the ability to roll back data quickly minimizes downtime and prevents costly retraining cycles.

The quickstart's design is also user-friendly, making it accessible for non-experts to begin their journey into professional AI operations. With straightforward deployment processes and clear workflows, teams can transition from manual data management to a reliable, version-controlled architecture with ease.

In conclusion, the Fraud-Detection-data-versioning-with-lakeFS quickstart represents a significant advancement in the management of AI data lifecycles, offering organizations a structured approach to enhance their AI operations and improve compliance, scalability, and efficiency. Those interested can explore the quickstart further through the official repository provided by Red Hat.

Frequently asked questions

What is lakeFS?
lakeFS is a data versioning tool that allows users to manage datasets similarly to how Git manages code, enabling branching, committing, and reverting data.
How does Red Hat OpenShift AI enhance AI model deployment?
Red Hat OpenShift AI provides a unified dashboard that consolidates tools for experimentation, deployment, and automation, streamlining the AI development process.
What industries can benefit from this architecture?
Highly regulated industries such as finance, healthcare, and telecommunications can benefit from the compliance and audit readiness features of this architecture.

AI & art news in your inbox, daily

The day's top stories, summarized. Free, no spam, unsubscribe anytime.