

A migration path from Segment Hightouch and Amplitude to Databricks CustomerLake
AUG. 10, 2026
6 Min Read
Contract renewal is the best moment to move CDP work from a legacy tool into the lakehouse you already govern.
That judgment comes down to cost, control, and timing. U.S. internet advertising revenue reached $258.6 billion in 2024, which means audience activation mistakes now sit next to very large media budgets. If you already run a governed lakehouse platform, renewal gives you a rare opening to move customer profiles, identity work, and campaign activation into infrastructure you already own. That path avoids another term of license fees plus the read, write, and compute costs that legacy CDPs add on top.
Key Takeaways
- 1. Contract renewal is the cleanest point to replace a legacy CDP because you can validate CustomerLake in parallel before budget and term are locked again.
- 2. The strongest business case comes from removing duplicate data movement, duplicated governance layers, and separate activation tooling when Databricks is already in place.
- 3. Migration success depends less on raw data transfer and more on rebuilding identity rules, audience logic, and rollout checkpoints with clear ownership.
Renewal timing sets the lowest risk migration window

Renewal timing lowers migration risk because you can stand up CustomerLake beside the existing CDP, validate data and audiences in parallel, and shift live campaigns only after results match. That sequencing keeps finance, security, and marketing aligned around one deadline instead of forcing a rushed cutover.
A practical pattern starts about 12 to 18 months before the contract end date. Your team maps the current audience logic, destination setup, identity rules, and campaign dependencies, then rebuilds the same outputs inside Databricks. A national retailer might keep paid media activation on the old platform for one holiday cycle while email suppression lists and loyalty segments move first.
This window matters because the business case is strongest before a new term is signed. Procurement already has the vendor under review, security doesn’t need to reopen a platform that is already approved inside Databricks, and marketing can compare campaign performance side by side. You’re making a contract choice using operating evidence and side-by-side performance results.
"Renewal timing lowers migration risk because you can stand up CustomerLake beside the existing CDP, validate data and audiences in parallel, and shift live campaigns only after results match."
CDP spend keeps rising after the license is signed
CDP cost rarely stops at the subscription because most platforms charge you to ingest data, process identity, and push audiences back out to channels. A $500,000 license often sits on top of duplicated storage, extra compute, and operational work that finance won’t see in the vendor quote.
A common pattern looks simple on paper and expensive in practice. Web events land in a warehouse, move into the CDP for profile stitching, then move again into ad, email, and analytics tools. Every pass adds I/O, pipeline support, failure monitoring, and reconciliation work. A media business with daily audience refreshes can end up paying for the same customer record several times before a single campaign launches.
This is why contract renewal becomes a finance and operating model issue. When the data platform is already Databricks, keeping a separate middleware CDP needs a stronger justification than “it already works.” If the same identity, segmentation, and activation outcomes can run where your governed data already lives, duplicated movement turns into a visible line item instead of accepted overhead.
CustomerLake shifts the comparison when Databricks is present
CustomerLake changes the comparison because it places the CDP inside the lakehouse instead of outside it. That means profile building, identity resolution, and campaign activation happen on governed tables your team already uses, with the same security posture and operating model already accepted for Databricks.
The practical difference shows up in two places. Profile work moves into the profile agent, which handles customer 360 records, identity logic, and optional enrichment through partners such as Acxiom, Epsilon, or Neustar. Activation work moves into the campaign agent, which takes approved audiences and pushes them into campaign execution. You keep the CDP function, but you stop paying for an extra layer that copies data in and out.
The comparison below shows where cost, governance, and migration effort diverge once the lakehouse already holds the customer data. It also clarifies where each platform adds operational overhead. Finance gets a cleaner TCO view. Platform teams get fewer systems to support.
| Platform | How spend usually grows | Where data movement happens | How governance is applied | What migration work matters most |
|---|---|---|---|---|
| Segment | License cost grows with event volume and destination usage | Data is collected into a proprietary schema and then pushed outward again | Governance is split between the warehouse and the CDP | Move schemas, audiences, and destinations into Databricks tables and activation flows |
| Hightouch | Spend rises with sync jobs, model maintenance, and destination complexity | Data starts in the warehouse but still relies on outbound sync layers | Governance is better aligned with the warehouse but activation remains separate | Consolidate activation inside CustomerLake and reduce reverse ETL dependency |
| Amplitude | Cost follows event tracking volume and analytics usage | Behavioral data often sits in a separate analysis path from activation | Governance spans analytics tooling and warehouse controls | Rebuild event analysis and audience logic in Databricks-native pipelines |
| CustomerLake | Spend is tied to compute and storage already governed inside Databricks | Customer data stays where the business already stores and manages it | Unity Catalog policies extend to profile and activation workflows | Focus shifts from tool integration to model quality, sequencing, and cutover testing |
Segment migrations start with a proprietary schema exit
Segment migration starts with data model extraction because audience logic and destination rules are built on top of a proprietary event and trait structure. You will move those definitions into Databricks tables first, then rebuild activation paths in CustomerLake against governed profile data.
A typical enterprise has years of hidden logic buried in tracking plans, custom traits, personas, and destination mappings. A subscription company might define “likely churn” from page views, support events, billing status, and email history, then send that segment to paid social, lifecycle email, and call center tools. That segment needs a line-by-line rewrite into documented tables, identity joins, and audience rules.
"Identity resolution should follow the business model you serve, because person-based graphs and account-based hierarchies solve different problems."
The hard part isn’t raw data movement. The hard part is finding business logic that grew inside the tool without clear ownership. Good migrations treat Segment as a requirements reference and rebuild the durable source of truth in governed tables. Teams that do this well inventory every active audience, retire unused destinations, and rebuild only the logic that still affects revenue, suppression, or compliance.
Hightouch migrations fold reverse ETL into Databricks workflows
Hightouch migration is usually more direct because the warehouse already holds the important data and business logic. The main job is moving activation from scheduled syncs into CustomerLake so segmentation and campaign delivery happen inside the same governed operating model.
A data team that already uses SQL models for lifecycle stages has a head start. Those models can remain in Databricks, then feed CustomerLake audiences instead of feeding outbound reverse ETL syncs. A bank, for instance, might keep product propensity scores in existing tables and swap the Hightouch sync layer for direct audience activation with the campaign agent.
This path often produces the cleanest cost story because the architecture is already close to warehouse-native. The team is starting from a warehouse-native setup, so the job centers on activation design instead of schema rescue. You’re deciding which syncs should remain, which destinations can move into CustomerLake first, and which service-level agreements matter for campaign timing. The result is less orchestration overhead and fewer moving parts for tech leaders to support.
Amplitude migrations keep event analysis inside warehouse pipelines
Amplitude migration succeeds when event tracking and behavior analysis are preserved inside Databricks instead of recreated as a disconnected reporting layer. You will keep event collection, session logic, funnel analysis, and audience triggers, but you’ll run them on warehouse tables that also feed activation.
A product-led business usually cares about pathing, retention cohorts, and feature adoption more than simple list building. Those patterns can be rebuilt with event tables, time-window logic, notebooks, and dashboards that sit next to the activation layer. Lumenalta often sees teams keep the existing event stream in place during validation while the same funnel definitions are recreated against Delta tables and checked for parity.
The important tradeoff is operational ownership. Marketing gets closer to governed customer data, but product analytics and data engineering need shared definitions for sessions, events, and user states. If that work is skipped, audience logic drifts. If it is done well, behavioral analysis stops living in a separate tool chain and starts feeding campaigns, personalization, and retention work from the same source.
Identity resolution must match the customer model

Identity resolution should follow the business model you serve, because person-based graphs and account-based hierarchies solve different problems. CustomerLake supports both paths, so your migration will work best when you decide early if the primary key is a household, an account, an individual, or an external identity spine.
A telco usually needs internal hierarchy resolution more than third-party graph expansion. One billing account can map to an address, a household, several devices, and multiple people, each with different permissions and marketing relevance. That structure matters because about 2.6 people live in the average U.S. household, which makes household relationships more than a niche edge case.
A media brand has a different need. Addressable advertising often depends on a recognized external spine from providers such as Acxiom, Epsilon, or Neustar. Migration work should keep those paths distinct. You’ll avoid a lot of rework if your team writes down which use cases need internal account resolution, which need external enrichment, and which must stay separated for privacy and measurement reasons.
Parallel rollout reduces campaign risk during contract decisions
Parallel rollout keeps campaigns live because it replaces a single cutover date with staged failover checkpoints. You can run profile creation, audience qualification, and destination activation in CustomerLake while the legacy CDP still supports active campaigns, then move traffic only after each step passes agreed tests.
A disciplined rollout usually starts with low-risk workloads such as suppression lists, seed audiences, or internal reporting segments. Paid media lookalikes, high-value lifecycle triggers, and account-based orchestration move later. Lumenalta typically sequences this work with finance, marketing ops, and platform teams in the same review cycle so rollback choices stay tied to revenue impact and contract dates, not tool preference.
The checkpoints below work because each one can be tested before cutover. They focus on parity, timing, and rollback. Marketing can read them quickly. Engineering can verify them with logs and counts.
- Audience counts match within an agreed tolerance across both systems.
- Destination delivery arrives on the same schedule as the current campaigns.
- Consent rules and suppression logic produce the same exclusions.
- Identity joins resolve the same households, accounts, or individuals.
- Rollback can restore the prior activation path within one campaign cycle.
That is why renewal is the right forcing function for a move into CustomerLake. You’re not betting the business on a big launch day. You’re making a controlled contract decision with proof that the governed lakehouse path can carry live marketing work. Lumenalta fits best here when you need the migration runbook to reflect revenue, risk, and operating cost at the same time.
Table of contents
- Renewal timing sets the lowest risk migration window
- CDP spend keeps rising after the license is signed
- CustomerLake shifts the comparison when Databricks is present
- Segment migrations start with a proprietary schema exit
- Hightouch migrations fold reverse ETL into Databricks workflows
- Amplitude migrations keep event analysis inside warehouse pipelines
- Identity resolution must match the customer model
- Parallel rollout reduces campaign risk during contract decisions
See how lakehouse-native customer activation simplifies cost and control.









