Enterprise AI should train where enterprise data lives

Kalyan Veeramachaneni
by Kalyan VeeramachaneniAugust 19, 2026

This principle has guided every major product decision we've made.

Recently, an interesting debate has emerged in the enterprise data space over the best way to access AI technology. Should enterprises access AI through an API —the primary deployment architecture offered by AI labs — which involves sending those labs their proprietary data? Or should they access AI by installing open source models (also called open weight models) on-premises and further training them on enterprise-specific data (also known as fine-tuning)? Several tech leaders, including Satya Nadella from Microsoft, are warning that when enterprises use AI models from labs like OpenAI and Anthropic, those labs gain ever-increasing access to those companies’ most sensitive business information. This debate is no longer just about privacy. It is increasingly about ownership — of enterprise knowledge, enterprise models, and long-term competitive advantage.

At DataCebo, we have long recognized that enterprise data is a strategic asset. We believe organizations should maintain full control over how their data is used, how AI models are trained and deployed, and ownership of the models and knowledge derived from their data.  In this article, we discuss how our commitment to enterprise AI has challenged our product, pricing, and packaging — and how we stay true to that commitment.

Training your own generative models

But first, some background about us. Our product is a software stack that allows enterprises to build generative models for their relational data stores. Unlike foundation models for language, these models can be built from scratch, locally with minimal compute and infrastructure needs. Our enterprise users typically follow this workflow: 

  1. Install our software stack within their environment; 
  2. They train generative models using their real data, and 
  3. They sample synthetic data from these models that is just like their real data.

They can then use that synthetic data to accelerate software testing, train machine learning models when real data is unavailable, and overcome data access barriers. Meanwhile, the trained models remain theirs, and their real data never leaves their environment. 

At DataCebo, we made a simple promise to our customers from day one: you can train your models on your data in your environment. Many synthetic data platforms have taken an opposite approach, and require their customers to upload real data to a hosted service. For us, this never made sense. We have talked to more than 300 enterprises about what they need from a synthetic data platform. Their number one requirement was simple: they did not want to upload their proprietary production data to a hosted service. The market response to hosted-service deployment architectures was brutal. Over the years, we've watched the industry steadily move toward customer-managed and on-premises deployments — the very direction our customers were asking us to take from day one.

AI models should train where your data lives 

This promise now affects every aspect of our strategy, including product development, packaging and pricing. Here's how they have influenced our day-to-day decisions in these three areas: 

Product. Because of our promise, we always knew that our software — used to develop generative models on an enterprise's data — must be installed and run within that enterprise's environment. This means we won't have access to the software once deployed — but if anything breaks, we still need to be able to help our customers, without accessing their data. Because of this, we defined our abstractions so that as the data flows through our software stack, it goes past several checkpoints, each representing the completion of a significant step in the modeling process. Each of these is documented, and users are given clear messages and guidance as they run their data through the software. 

Achieving this required a disciplined development process. It also made us parsimonious — we can't release bug-fixing patches every night, because that would disrupt the flow of the many users who installed the previous version within their environment. To make stable releases, we invented 20 bots that emulate our customers' usage under various scenarios. These bots run around the clock as our team is developing. 

We are proud to say that across several dozen product deployments, we have never needed to access a customer's data in order to debug an issue. That's because our promise — that you will be able to install and run our software in your environment — drives our decisions on a daily basis. 

Pricing. Many of the significant “deployment” decisions in the AI product world are actually pricing decisions. Companies build around the simplest pricing model: identify a unit of usage, and then price for that, so that the more the customer uses, the more they pay. A hosted service lets companies experiment with this type of pricing quickly and at scale.  It allows them to experiment with data collection strategies, pricing strategies. But it doesn't work for us, because it would mean breaking our promise. We refused to compromise our deployment model merely because it would simplify pricing. Ultimately, we came up with a pricing strategy that works well for us and our customers, and we are able to capture usage amounts and clearly communicate what is being sent to our servers. 

Packaging. Similarly, shipping software that can train generative models “on-prem” requires an entirely new way of thinking. We had to figure out how to authenticate users, and again, how to gather usage information for pricing. For security reasons, we can't reveal the techniques we eventually developed in this area — but trust that they inspired some unique and intense product discussions. 

Our final thoughts and takeaways

Over the next decade, enterprises won't merely own their data. They'll own the models trained on that data. Those models will become strategic assets in the same way databases, software, and intellectual property are today. Enterprise AI software should behave like databases. 

Enterprises should treat data as a strategic asset.  An enterprise's data encapsulates the business decisions and strategic intelligence that gives them their competitive advantage. Enterprises know this already, but they should double down, and put up guardrails to protect their data.

Enterprises should also treat their trained models as strategic assets. All models trained on enterprise data are also strategic assets, whether they're trained from scratch or fine-tuned (as with many open weights models). These models encode intelligence that is specific to that enterprise data, and will become competitive assets.  

Enterprises should demand on-prem, secure installation of any AI software stack. At DataCebo, we have demonstrated that it is possible to build an AI software stack, deliver it to enterprises, and let them run it on-prem within their environment. Don't let other vendors tell you this is impossible — if we can do it, they can too. 

Enterprises should work with startups and businesses that provide a localized deployment model. This article is meant to share the challenges a product company like ours faces when doing what's best for enterprises. Enterprises can help create a win-win situation by participating in  collaborations and partnerships, and by deliberately and systematically sharing how they use such products, so that we can make them better. Ultimately, this will help push the whole industry in the right direction.

Popular topics
Enterprise AI should train where enterprise data lives
Applications

Should enterprises access AI through an API which involves sending those labs their proprietary data? Or should they access AI by installing open source models on-premises and further training them on enterprise-specific data?

Kalyan VeeramachaneniAugust 19, 2026
How to generate synthetic survey responses for market research
Applications

Trained AI models do what LLMs cannot: Generate survey responses with the statistical variety and demographic accuracy your analysis depends on.

Neha PatkiJuly 29, 2026
Why synthetic data is not the same as data masking with ETL
Applications

Synthetic data does what ETL pipelines cannot: Create unlimited test data with low infrastructure, storage, and system complexity.

Neha PatkiJuly 28, 2026

Join the DataCebo Forum

Discuss SDV features, ask questions, and receive help.

Visit the DataCebo Forum

Explore our blog

Read our newest insights about synthetic data, updates on our products, and successful use cases.

Read our blog
Datacebo logo

Make synthetic data a reality

© 2026 DataCebo, Inc.