Smart Data built an AI-assisted migration tool that converts legacy ETL workflows, orchestration scripts, and deployment configurations into production Databricks pipelines, with unit tests and infrastructure as code generated alongside every conversion.
Client
Fortune 100 grocery retailer
Industry
Retail
Timeline
AI-assisted legacy migration, proof of concept delivered in roughly three months

Project Overview
A Fortune 100 grocery retailer had committed to Databricks as its data platform. The commitment was the easy part. Standing between the decision and the outcome were hundreds of legacy ETL workflows and orchestration jobs, built up over years across enterprise domains, each one encoding business logic the company depends on every day.
Rather than staffing a large manual translation effort, Smart Data built an AI-assisted migration tool that converts legacy logic into Databricks pipelines on PySpark and generates the tests and deployment automation that make each conversion verifiable. A three-month proof of concept covering a real cross-section of the estate proved the approach holds at Fortune 100 scale.
Project Challenge
The conventional path through a migration this size is manual translation: engineers read each legacy workflow, understand what it does, and rebuild it on the new platform. At this scale, that math does not work. Four things stood in the way:
Hundreds of legacy ETL workflows and orchestration jobs spread across enterprise domains, each encoding business logic the company runs on daily
Manual translation scoped to multiple teams working for multiple years
Drift, fatigue, and inconsistent quality across a program measured in years rather than months
A platform decision already made, with migration cost threatening to stall it before the business case could be realized
The client did not need a bigger team. They needed the economics of the migration to change.
Project Approach
Smart Data built an AI-assisted migration tool using Claude and structured prompts. The tool converts legacy ETL logic, orchestration scripts, and deployment configurations into Databricks pipelines built on PySpark, generating unit tests and infrastructure as code alongside every converted pipeline.
Key Implementation Highlights:
AI-assisted conversion of legacy ETL logic and orchestration scripts into PySpark pipelines on Databricks
Unit tests generated with every conversion, so output arrives verifiable rather than merely plausible
Infrastructure as code through Terraform and GitHub Actions, generated alongside the pipelines
A proof of concept scoped across 31 workflows, 36 mappings, 60 orchestration jobs, and 50 shell scripts
Scope drawn from real production domains rather than a curated sample, to test the approach against the variety of a true enterprise codebase
Senior engineering effort redirected from writing translations by hand to validating output against known behavior
Generating tests and deployment automation with the pipelines is what makes the approach viable. Without them, AI conversion produces code a team still has to read line by line. With them, the platform team can review, run, and trust the output.
Project Results
The proof of concept answered the question the client actually needed answered, which was whether the approach would hold across the rest of the estate:
Proof of concept delivered in roughly three months with a small team, against a manual alternative scoped at multiple teams over multiple years
78% of the legacy transformation types in the estate already covered by the tool at the end of the proof of concept
95% confidence the approach generalizes across the remaining enterprise domains, based on the variety covered in the initial scope
Converted pipelines shipped with generated unit tests and infrastructure as code, so the client's platform team could verify the output rather than take it on faith
This engagement is also the foundation of Smart Data's Databricks partnership. The work proved, at Fortune 100 scale, that an enterprise can move from a legacy ETL estate to production Databricks pipelines on a timeline that keeps the business case intact.
Key Technologies
Databricks, PySpark, Claude, Terraform, and GitHub Actions.




