Can enterprises protect data without making AI less reliable?
As organizations invest in AI, many are discovering a new bottleneck: obtaining data that is both protected and useful. Engineering The post Can enterprises protect data without making AI less reliable? appeared first on The New Stack .
Protecting data has become a growing concern for enterprises as they embrace artificial intelligence (AI). A survey conducted by the Perforce Delphix "2026 State of AI and Data Privacy Report" reveals that 26% of surveyed organizations find privacy controls hinder the acquisition of production-quality data, while 25% struggle to maintain relationships across data entities, and 51% encounter data quality challenges.
The report highlights that AI models trained on incomplete or distorted data produce less reliable outputs. Furthermore, when test environments contain unrealistic data, defects can escape into production, and poorly trained AI models can result in inaccurate analytics. These challenges underscore that data protection must not compromise the utility of data for AI and engineering teams.
Successful organizations recognize that compliance, quality, and speed can coexist. Protected data must remain realistic enough for testing and validation, representative for analytics, accessible for engineering teams, governed for regulatory requirements, and connected to preserve referential integrity. The key to overcoming these challenges lies in maintaining referential integrity between entities, as broken relationships can render protected data unrepresentative of reality.
For example, billing validation requires referential integrity when a customer has multiple products, charges, and invoices across several database tables. Lost relationships can lead to inaccurate billing, incomplete analytics, and flawed AI models. Many organizations focus solely on masking sensitive fields, but masked data that loses referential integrity presents a different kind of risk.
Engineering teams should measure privacy success beyond compliance metrics, evaluating data quality, realism, and referential integrity. Industry standards for enterprise-grade masking algorithms ensure consistent masking output across systems and environments, preserving relationships. By prioritizing these factors, enterprises can strike a balance between protecting sensitive data and maintaining its value for AI and engineering initiatives.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.