Introducing generative relational models for enterprise data
Generative AI has transformed the way we create and work with language, software, images, and video. But what about structured enterprise data?

Enterprise data is proprietary and a source of competitive advantage
Generative AI has dramatically accelerated how software is written and how enterprise teams work with language, images, and code. But the operational intelligence unique to a company, its customers, transactions, relationships, constraints and business rules lives inside its own complex databases and data warehouses. This is the ‘data moat’ that companies must guard and sustain.
The answer is not a general model trained elsewhere. It is a generative relational model trained directly on enterprise-owned structured data. This allows the enterprise to bring its own proprietary context to software, analytics and AI systems without repeatedly moving or exposing the underlying production records.
LLMs bring the world’s intelligence to the enterprise. Generative relational models bring the enterprise’s intelligence to AI.
From a representative subset to a reusable enterprise AI asset.
In 2024, we launched SDV Enterprise with a simple promise: provide a representative subset of your structured database. The software will train a generative relational model that learns the underlying statistical patterns and relationships within the database.
Unlike large language models, these models can often be trained in minutes to an hour, using only a subset of the data and, in many cases, on a standard CPU machine. The resulting model belongs to the enterprise and can faithfully generate entirely new, highly realistic structured data.
Minutes–hours
Typical training time
Any relational depth
Enterprise-scale schemas
Reusable
One model, many workflows

Statistically similar data is only the beginning
To become reusable AI assets, generative relational models must progressively understand the many layers of knowledge embedded within enterprise data.
Embedded context in data values
Formats and values carry business meaning. IDs, addresses, phone numbers, and other data types each have rules that shape how they relate to the rest of the database.
Generative relational models must preserve this hidden context—not just produce values that look statistically plausible.
Complex relationships across data
Composite keys, polymorphic relationships, reference tables, and databases split across teams create connections that are rarely captured in one clean schema.
Models trained independently must still recognize these connections and generate data that remains consistent across them.
Business logic
Application rules are deterministic—not database relationships and not statistical correlations. A customer may need a savings account before becoming eligible for a credit card.
Constraint-Augmented Generation lets the model preserve these rules while maintaining structure and statistical fidelity.
Representative training data
Enterprises may have terabytes of production data, but training on all of it is rarely feasible or necessary. The real challenge is selecting a representative subset without losing important context.
Training algorithms must understand that the subset represents a much larger, multi-table database.
Task-oriented generation
Random sampling may never produce the low-frequency event a user needs. Generation must respond directly to a task, condition, population, or scenario.
The model should generate complete, coherent records—even when only a few requested variables are specified.
Users gave the model the context it could not infer.
When SDV was redesigned as SDV 1.0, we deliberately built it to accept user input anticipating enterprise complexities we had not yet encountered. Users could program the behavior of the generative relational model through metadata, settings, data processing, constraints, and sampling instructions.

Enterprise scale changes the question.
Programmability works when users know the logic, formats, and relationships they need to specify. At enterprise scale, organizations often do not know all of that context and their data is notoriously underdocumented.


Four levels of enterprise readiness.
The levels do not measure the quality of the underlying generative relational model. They describe how much enterprise data complexity a platform can handle—and how much manual work users must do to make it work.
Level 1
Simple data platforms
Single-table and simple structures
These platforms support a single table or simple, common relationships. When the data becomes more complex, they often fall back to anonymization, noise, or column-by-column transformations.
More manual
Level 2
Pipeline-based platforms
Pre-processing and post-processing
The core AI models a single table, while ETL pipelines reshape complex data before training and reconstruct it afterward. The scripts and ongoing maintenance make this difficult to scale.
Level 3
Configurable enterprise platforms
Enterprise complexity, user configured
The platform can handle enterprise data, but users must define metadata, formats, relationships, constraints, and other settings required to address that complexity.
Level 4
Self-configuring enterprise platforms
Configuration learned from the data
Additional AI learns formats, business logic, keys, relationships, and embedded context directly from raw data—bringing enterprise generative AI closest to full automation.
More automated
Stop moving production data. Start working with the model.
Test data management workflows are already being reimagined and are particularly well suited for disruption. Instead of copying, masking, and transforming production databases through infrastructure-heavy workflows, teams train a model once and generate entirely new records whenever they need them. The same is happening with data sharing and optimization.


Your structured data already contains the knowledge.
Build the model that understands it.
Explore SDV Enterprise and create generative relational models that learn from the statistical patterns, relationships, business logic, embedded context, and intent within your data.