placeholder

Senior Databricks Data Engineer, Data Testing & Validation

Data Engineer
placeholder

About the Role

We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business.

The platform replaces an existing integration system that today connects the client's engineering tools and produces the reports their teams rely on. The rule is simple: the data in Databricks has to match what that system holds and what its reports say. It has to match on day one, and it has to keep matching every time new data arrives.

You will write the tests that prove it. You are a data engineer, not a manual tester. Your code checks that every table reconciles to the legacy system or its reports, that incremental loads add exactly what changed and nothing twice, and that duplicates never make it through. When a pipeline breaks one of those rules, your tests catch it before the client does.

What You'll Do

  • Build the automated test framework for the platform's pipelines, in PySpark and SQL, and run it on every deployment throughout all environments.
  • Write reconciliation tests that compare Databricks tables to the legacy system and its reports, with clear output on where they differ and why.
  • Test incremental loading: change data capture, appends, and batch refreshes. Prove that reruns are idempotent, that merges and upserts update the right rows, and that late or re-delivered files do not double-load.
  • Test for duplicates on composite keys and on the relationships that link records across source systems.
  • Write unit tests for transformation logic so that keying, joins, and business rules are verified in isolation before they run against real data.
  • Define data quality expectations in the pipelines themselves, so bad data is quarantined or fails the run instead of landing silently.
  • Test how pipelines behave when source files change shape: new columns, renamed columns, missing columns, and mixed types.
  • Wire the tests into Databricks Workflows and the CI/CD pipeline so they gate promotion between environments.
  • Test Unity Catalog access controls, row filters, column masks, and classification tags to confirm that export-controlled data is only visible to authorized users.
  • Plan and run user acceptance testing sessions with client stakeholders, capture feedback, and track defects to resolution.
  • Document test results and acceptance evidence for each deliverable, and report quality status in weekly reviews.

What We're Looking For

  • 5+ years in data engineering, including building test and validation frameworks for production pipelines, not only running tests written by others.
  • Strong PySpark and SQL, used to write reconciliation and assertion logic at scale.
  • Hands-on experience reconciling a new platform against a legacy system or its reports, and explaining the variances.
  • Deep understanding of incremental load correctness: CDC and append semantics, idempotency, MERGE behavior, watermarks, and late-arriving data.
  • Experience finding and resolving duplicates on composite keys across multiple sources.
  • Experience with a Python test framework such as pytest, and with data testing tools such as chispa, Great Expectations, or pipeline expectations in Delta Live Tables or Lakeflow.
  • Experience running tests from Databricks Workflows and CI/CD across multiple environments.
  • Organized, detail-oriented, and comfortable holding a delivery team and a client to a quality bar on a tight timeline.
  • Clear written communication - test evidence and variance reports are the core deliverables.
  • Ability to complete the client's background check and onboarding and to work on client-furnished equipment.
  • US citizenship.
  • Fluent English, both written and spoken.

Nice to Have

  • Hands-on experience with Databricks, Unity Catalog, or Delta Lake.
  • Python or PySpark for automating data validation.
  • Testing Unity Catalog access controls, row filters, and column masks.
  • Prior work in a defense, aerospace, or FedRAMP environment, or with CUI or ITAR-controlled data.
  • Familiarity with PLM or MBSE tools such as Teamcenter or Cameo, or with Oracle E-Business Suite data.
  • Experience with test management and defect tracking tools such as Jira or Xray.

Why Lumenalta is an amazing place to work at

At Lumenalta, you can expect that you will:

  • Be 100% dedicated to one project at a time so that you can innovate and grow.
  • Be a part of a team of talented and friendly senior-level developers.
  • Work on projects that allow you to use leading tech.

For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:

  • Delivery Partner Program qualification. Databricks co-delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
  • Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and the Champions already on our team mentor candidates through it.
  • Certifications covered. We pay the exam fee for any Databricks certification you want to take.

Our Process

A screening call, a technical interview, and a HackerRank assessment. The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through an Oracle integration you built, how you handled security and access constraints, how you modeled logistics or ERP data alongside engineering data, and what you would change. Relevant project experience matters more to us than certifications.

Location

This is a fully remote position. This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship. This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.

Application Deadline

Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

Traits of a Lumen

Radically Engaged

Strong performers, constant communication, quality work. We deliver impact.

Bright mindset

Ambitious, energized, kind. We tackle with optimism.

Lead the way

Professional, adaptable, thoughtful. We set the standard.

Lightspeed

Agile, collaborative, action-first. We move fast to deliver the best.

Join the bright side

Hiring Process

Learn more about how we interview and select candidates.

Career Opportunities

Find a role that best matches your skill set and career goals.