Synthetic data gets an AI-native upgrade with Perforce Delphix 

Synthetic data gets an AI-native upgrade with Perforce Delphix 

Perforce calls itself the DevOps tech stack company (which is presumably DevOps, all the way up the stack with governance), and the organisation is active again this month with the announcement of Delphix Synthetic Data. This is an AI-native software service designed to generate “realistic, scenario-specific test data” that maintains referential integrity across enterprise systems. So why does that matter?

The suggestion here is that there is a “critical enterprise bottleneck” in terms of securing realistic, compliant test data. By preserving cross-system referential integrity using safe AI (controls to keep related data matched and consistent across different databases, applications, or systems), teams can test complex, multi-database edge cases on demand without exposing sensitive production data or stalling CI/CD pipelines. 

Data structures, relationships & context 

The software automatically identifies data structures, relationships, and business context across sources, eliminating much of the manual configuration associated with legacy synthetic tools. Developers and AI agents get immediate access to safe, high-quality data for new applications, features, and agentic workflows, especially when production data is restricted, incomplete, or unavailable.

“Organisations need realistic test data that can be generated quickly, scale across complex environments, and meet data privacy requirements,” said Jim Mercer, program vice president, software development, DevOps, and DevSecOps, IDC. “Solutions like Delphix Synthetic Data that combine AI-driven automation with data quality and control are better positioned to address that need.”

What AI needs synthetic data 

Perforce suggests that synthetic data has “rapidly become essential” for enterprise software delivery. As organisations adopt AI-assisted and agentic development, synthetic data allows teams to generate data when production data doesn’t exist. They can use this data to help develop new applications, features, and test unique edge cases.

However, legacy synthetic tools require heavy manual configuration and often fail to deliver the realism and referential integrity enterprise teams need to deliver high-quality software.

Research from a 2026 Perforce survey of 518 enterprise technology leaders suggested that “a significant gap” exists between the needs of enterprise development teams and the capabilities of existing synthetic data solutions. 

What is an AI-native approach to synthetic data?

With an AI-native approach to synthetic data, the company saysy Delphix Synthetic Data uses AI to automatically scan metadata and schemas and configure synthetic data, and statistical analysis to understand data shape and distribution. Users can then describe the data they need and modify it using natural language prompts. Instead of waiting days or weeks for low-quality test data, developers, testers and agents can generate the data they need in minutes.

“Your masked production data only tells you what already happened,” said Ilker Taskaya, field CTO, Perforce Delphix. “Testing needs the cases that aren’t in production yet –  and it needs them to hold together across every database and file format in the environment. Delphix Synthetic Data generates data in the shape teams specify, with referential integrity intact, so agentic development isn’t waiting on test data.”

Delphix Synthetic Data is part of the Delphix DevOps Data Platform, which combines synthetic data generation with proven data masking, data delivery, and centralised governance. 

Together, Perforce claims these capabilities help organisations support a “broader range of development and testing scenarios” while maintaining consistent control over sensitive data. Developers and AI agents can access data through the user interface, APIs, and MCP-enabled workflows, enabling secure data access wherever development occurs. 

Who else ‘makes’ synthetic data?

Companies in this space include Gretel.ai (Gretel Navigator), Tonic.ai (Tonic Structural), MOSTLY AI (MOSTLY AI Platform), Hazy (Hazy Platform), Synthesized (Synthesized SDK), Synthesis AI (OpenSynthetics), Syntho (Syntho Engine), YData (YData Fabric), and Betterdata (Betterdata AI)… and of course SAS.

SAS Data Maker is a low-code/no-code synthetic data platform designed to generate high-fidelity information that mirrors real-world data sets. 

“As a leading synthetic data platform, SAS Data Maker enables you to augment existing data or create entirely new assets – reducing acquisition costs while accelerating AI development. SAS Data Maker synthetic data fuels your AI models without risking PII exposure,” notes the company.

IBM is big on also big on this subject. The company says that synthetic data can “act as a placeholder for test data” and is primarily used to train machine learning models, serving as a potential solution to the need for (yet short supply of) high-quality real-world training data for AI models. 

“Synthetic data is also gaining traction in sectors like finance and healthcare where data is in limited supply, time-consuming to obtain or difficult to access due to data privacy concerns and security requirements. In fact, research firm Gartner predicts that 75% of businesses will employ generative AI to create synthetic customer data by 2026,” stated IBM on its product pages.

Produced by none of the above, this instructional video explains synthetic data well.