Overview
The institution ran an on-premises Cloudera Hadoop cluster as its enterprise data lake, feeding regulatory reporting, risk, and business intelligence across the organisation. After years of service the platform was straining against the demands placed on it:
The challenge
The institution ran an on-premises Cloudera Hadoop cluster as its enterprise data lake, feeding regulatory reporting, risk, and business intelligence across the organisation. After years of service the platform was straining against the demands placed on it:
- Cost. Hardware refresh cycles, the on-premises data centre footprint, and the Cloudera subscription were consuming a growing share of the technology budget, with little headroom to scale.
- Performance. Overnight batch windows were tight. Critical regulatory pipelines were taking [~8 hours] to run, leaving no margin for reprocessing or late data.
- Capability. Machine learning, advanced analytics, and self-service data access were difficult to deliver on the legacy stack. New data products took [weeks] to stand up.
- Governance. Lineage and sensitive-data classification were fragmented across the Hive metastore and downstream tools, creating friction against BCBS 239 risk data aggregation expectations and APRA CPS 234 and CPS 230 obligations.
- Talent. Deep Hadoop operations skills were increasingly scarce and expensive to retain.
A data centre exit deadline turned a strategic ambition into a delivery imperative.
The approach
Innablr led the migration end to end, from target-state architecture through to production cutover and operating model uplift.
- Discovery and dependency mapping. We catalogued every pipeline, table, and downstream consumer, using AI to accelerate analysis of a large and partly undocumented code base, then sequenced the migration by business criticality and risk so the most sensitive regulatory workloads moved under the most scrutiny.
- Target architecture. We designed a medallion lakehouse on Databricks, with Delta Lake for the bronze, silver, and gold layers and Databricks SQL serving enterprise BI, risk, and finance reporting.
- Spark-native migration. Because the Cloudera workloads were already Spark and Hive based, much of the estate moved to Databricks on a shared Spark lineage, lowering migration risk. Jobs were re-engineered for Delta and the lakehouse rather than lifted and shifted, and the Hive metastore was consolidated into Unity Catalog.
- AI-accelerated code transformation. We applied AI across the heavy lifting of migration: analysing legacy logic, transforming Spark, Hive, and SQL code to lakehouse-native patterns, and generating tests and documentation. Engineers reviewed and validated every output, so the speed gain came without ceding control of correctness or governance.
- Governance by design. Unity Catalog provided unified lineage, attribute-based access control, and automated PII tagging across the estate, giving risk and compliance teams a single, auditable view.
- Engineering discipline. Pipelines were deployed through Databricks Asset Bundles and Terraform, with CI/CD promotion across development, test, and production environments, replacing manual hand-offs with repeatable, governed releases.
- Parallel run and validation. Legacy and lakehouse pipelines ran side by side, with automated reconciliation, until outputs matched the institution's sign-off threshold. Cutover then carried no surprises.
The architecture
Legacy state
- On-premises Cloudera Hadoop data lake (Spark, Hive, Impala)
- Hive metastore for table definitions, with lineage spread across tools
- Manual orchestration and ageing on-premises hardware
Target state
- Databricks Lakehouse [on Google Cloud], medallion architecture on Delta Lake
- Unity Catalog for governance, lineage, and access control
- Databricks SQL serving enterprise BI, risk, and finance reporting
- MLflow for model development and tracking
- Databricks Asset Bundles and Terraform for CI/CD and infrastructure as code
The results
- [~45%] lower platform total cost of ownership after exiting the on-premises Hadoop estate and decommissioning legacy hardware
- Critical regulatory reporting pipeline runtime cut from [~8 hours to ~90 minutes], restoring margin in the batch window
- [~1,800] jobs and [~3 PB] of data migrated, with parallel-run reconciliation and zero unplanned downtime at cutover
- Time to deliver a new data product reduced from [weeks to days]
- Migration delivered [~40%] faster through AI-assisted analysis and code transformation, with engineer review at every step
- A unified, auditable governance layer spanning [~9,000] tables, strengthening BCBS 239 and CPS 230 alignment
- A foundation for machine learning use cases, including fraud, credit risk, and customer analytics, that were impractical on the legacy stack
"[ An ageing Hadoop estate was replaced with a single governed lakehouse, and in doing so freed our teams to build rather than maintain. The cost and performance gains were significant, but the bigger shift was confidence in our data.]" — [Head of Data Platforms, major financial services institution]
Capabilities applied
Data platform modernisation, lakehouse architecture, Databricks, Unity Catalog governance, AI-assisted code analysis and migration, Hadoop and Spark migration, DataOps and CI/CD, regulatory and compliance alignment.