Data Governance and Privacy: Securing Enterprise Data with Automated Masking
Overview
A major Australian telecommunications provider identified significant risk surrounding the exposure of Personally Identifiable Information (PII) in its non-production environments. Innablr's data engineering and governance practice was engaged to undertake a comprehensive review and design an enterprise-grade Test Data Masking (TDM) capability. The objective was not merely to implement a tool, but to establish a robust data governance framework that protected sensitive customer data while maintaining the functional integrity required by engineering and quality assurance teams.
The Challenge
The organisation relied on a manual process for provisioning non-production databases. This practice inadvertently exposed live production data — including sensitive customer PII — across all lower environments, including development and testing. The risk materialised into a critical issue when test communications were inadvertently triggered to real customers from a non-production system.
This exposure spanned a complex, hybrid data estate comprising six key applications hosted across both on-premises infrastructure and the cloud, including critical systems such as BRM Retail and Siebel. Early attempts to mitigate the risk using manual scripting proved neither accurate nor scalable. The customer required a solution that could secure data without disrupting established BAU data refresh processes, and which could handle complex referential integrity requirements across interconnected operational systems and downstream analytical platforms, including the enterprise data warehouse.
The Innablr Approach
Innablr undertook a detailed discovery and architecture assessment to design an automated data compliance operating model. Following a thorough cost-benefit analysis of market-leading toolsets, Delphix was selected as the foundation for the masking capability. The Delphix Continuous Compliance Engine was chosen for its ability to integrate into DevOps architectures, its support for heterogeneous enterprise databases — including Oracle, PostgreSQL, SQL Server, DB2, and distributed platforms such as Hive — and its scalability for growing data volumes.
Establishing the Governance Framework
Before technical implementation began, Innablr worked closely with business stakeholders, database administrators, and security teams to establish a formal data governance framework. This included formal sign-off on PII scope, environment booking controls, and integration of masking activities into the existing BAU production-to-test refresh cycle. Strict environmental controls were designed to ensure that temporary access to unmasked data during the refresh window was locked down, requiring formal security exception approvals (TRA), secure VLAN configurations, and Netskope access controls.
Architectural Design and Integration
Innablr engineered an in-place masking architecture that aligned with the organisation's network topology and security requirements. The Delphix Continuous Compliance Engine was deployed as a dedicated appliance, connecting to source and target databases via JDBC over the database listener port, with administrator access secured over TCP 443. Data was transformed in memory and written back to the target database, ensuring no unmasked data persisted within the masking platform itself.
The solution integrated with the organisation's enterprise toolchain, including:
Integration Point
Technology
Target databases
Oracle (BRM Retail, Siebel)
Database connectivity
Identity and access
Microsoft Active Directory / LDAP
Phased Delivery Model
The implementation followed a rigorous, structured delivery model with formal governance gates between phases:
Design Phase — Innablr conducted infrastructure readiness assessments, confirmed PII scope with business owners, provisioned the Delphix appliance to an approved architecture design, and established environment booking and access controls. A formal go/no-go checkpoint was held before development commenced.
Development Phase — Masking jobs were built and validated in an isolated development environment (BRM DEV-4) using non-production SIT data. This segregated engineering pattern allowed the team to develop and test custom masking algorithms without contending with heavily utilised testing environments. Development followed an iterative approach: Iteration 1 leveraged out-of-the-box masking algorithms, while Iteration 2 introduced custom algorithms, external validations, and issue resolution.
Integration and Masking Phase — Once validated, masking jobs were applied to the Performance Testing (PT) environment. Following the BAU production-to-PT data refresh, the Delphix engine connected to the target database and applied masking rules in place, after which the environment was returned to the test team.
Acceptance and Handover Phase — Innablr ensured the organisation's QA capability was preserved post-masking by reusing existing sanity checks, automated regression suites, and integration test cases. A formal test summary report was produced and accepted by the platform owner before transition to BAU operations.
Preserving Referential Integrity Across the Data Estate
A critical requirement was maintaining referential integrity across a highly interconnected application landscape. Masking algorithms were engineered to ensure that obfuscated customer records in BRM Retail remained consistently linked to corresponding records in Siebel, WinCollect, and the downstream data warehouse. This allowed complex batch processes, end-of-month reporting, cross-system functional testing, and data warehouse reporting to proceed without disruption after masking was applied.
Outcomes
The customer now possesses a governed, automated capability to sanitise data in non-production environments through a centralised web interface. By replacing manual scripts with a repeatable, auditable masking process, the organisation has materially reduced its reputational and regulatory risk exposure — particularly relevant in the context of Australia's strengthened privacy penalties and heightened scrutiny following recent high-profile data breaches.
The operating model is designed for controlled rollout across additional environments with minimal support required for data model changes. As the engagement progresses, data virtualisation is being evaluated as the next phase of implementation to further accelerate secure test data provisioning across the broader application estate.