About the Role
We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business.
The platform depends on engineering data that today lives in Cameo system models, Teamcenter PLM, and Siemens Capital electrical designs. The upstream engineering teams export that data from the source tools and land it as files in S3. Those exports are large, change shape as the tools and models evolve, and were never designed with downstream consumers in mind.
You will build the ingestion layer that turns those exports into trusted Delta tables. You will stand up the S3 landing zone and ingestion framework, own the pipelines for all three engineering sources, and make every pipeline resilient to schema changes, so the digital thread team downstream can rely on what arrives.
What You'll Do
- Establish the S3 landing zone and ingestion framework used for all file-based engineering sources, including folder conventions, file validation, and failure handling.
- Build Auto Loader pipelines for incremental ingestion of Cameo, Teamcenter, and Capital exports landed in S3.
- Parse engineering export formats (CSV, XML, JSON, and similar) into well-typed Bronze and Silver tables with explicit schemas.
- Define and implement the incremental and batch refresh patterns applied across in-scope sources, including nightly batch loads.
- Configure schema evolution in Databricks declarative pipelines, including new column handling, rescued data columns, and controlled failure on renamed columns.
- Handle bad, late, and re-delivered files without corrupting or double-loading the tables.
- Implement baseline data quality checks and ensure lineage is captured for every pipeline.
- Orchestrate pipelines with Databricks Workflows.
- Work with the source system engineering SMEs to understand export contents, refresh cadence, and how model elements, requirements, and parts are represented in each file.
What We're Looking For
- 6+ years in data engineering, including production experience building ingestion pipelines on Databricks.
- Deep experience with Auto Loader and incremental file ingestion from S3.
- Experience parsing structured and semi-structured export formats (CSV, XML, JSON) into typed tables.
- Experience with Databricks declarative pipelines or Delta Live Tables, including schema evolution configuration.
- A track record of pipelines that handle bad, late, and re-delivered files safely.
- Strong PySpark and SQL, and working knowledge of Delta Lake internals such as managed vs. external tables and predictive optimization.
- Experience with Databricks Workflows for orchestration.
- Working fluency in PLM and MBSE vocabulary, enough to talk to engineering SMEs about parts, BOMs, revisions, and model elements.
- Strong PySpark and SQL, and working knowledge of Delta Lake internals such as managed vs. external tables and predictive optimization.
- Ability to complete the client's background check and onboarding and to work on client-furnished equipment.
- US citizenship.
- Fluent English, both written and spoken.
Nice to Have
- Hands-on experience with exports or data from Cameo / Teamwork Cloud, Teamcenter, or Siemens Capital.
- Experience ingesting from REST APIs into Databricks.
- Experience with Lakeflow Connect or other CDC-based ingestion tools.
- Prior work in a defense, aerospace, or FedRAMP environment, or with CUI or ITAR-controlled data.
- Databricks Data Engineer Professional certification.
Why Lumenalta is an amazing place to work at
At Lumenalta, you can expect that you will:
- Be 100% dedicated to one project at a time so that you can innovate and grow.
- Be a part of a team of talented and friendly senior-level developers.
- Work on projects that allow you to use leading tech.
For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:
- Delivery Partner Program qualification. Databricks co-delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
- Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and the Champions already on our team mentor candidates through it.
- Certifications covered. We pay the exam fee for any Databricks certification you want to take.
Our Process
A screening call, a technical interview, and a HackerRank assessment. The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through an ingestion pipeline you built from a complex source system, how you handled CDC and schema changes, how you validated the data, and what you would change. Relevant project experience matters more to us than certifications.
Location
This is a fully remote position. This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship. This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.
Application Deadline
Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

