About the Role
We partner with global enterprises to design and build large-scale data platforms that power the products and operations at the center of their business.
Not all of that data lives in enterprise systems. Manufacturing and supply chain teams maintain parts, supplier, and program data in Excel workbooks and flat-file repositories. These are authoritative sources, but they have no schema, no change log, and no owner-enforced structure. Each workbook is a small, undocumented system of its own.
You will bring those sources onto the platform. You will inventory them, build ingestion that turns inconsistent workbooks and files into governed Delta tables, and make the result join cleanly to parts and revisions from the engineering systems, so the digital thread includes the data that today lives only in spreadsheets.
What You’ll Do
- Inventory the Excel workbooks and file exports that business and engineering teams rely on, and document their structure, owners, and downstream use.
- Build Auto Loader pipelines to ingest workbooks and flat files landed in S3, with explicit schemas and typed Bronze and Silver tables.
- Parse multi-sheet, inconsistently structured workbooks at scale: shifting header rows, merged cells, embedded totals, mixed types in a column, and layouts that differ between versions of the same file.
- Configure schema evolution in Databricks declarative pipelines, including new column handling, rescued data columns, and controlled failure on renamed columns.
- Implement data quality expectations that catch the failure modes these sources produce, such as duplicate rows, blank keys, and silently changed formats.
- Conform ingested data to the platform's part number plus revision keying so it joins to BOM and engineering data from other sources.
- Ensure lineage is captured for every pipeline, orchestrate with Databricks Workflows, and document each source for handoff.
- Work with the business owners of each workbook to confirm meaning, business rules, and what "correct" looks like.
- Translate spreadsheet logic, such as formulas, lookups, and pivot tables, into maintainable SQL and documented business rules.
- Validate that platform outputs match the numbers users trust today, and explain differences when they do not.
- Document datasets, metrics, and definitions so reporting stays consistent after handoff.
What We’re Looking For
- Production experience building ingestion pipelines on Databricks.
- Hands-on experience ingesting Excel and flat-file sources into Spark or Databricks at scale and the judgment to know when each is appropriate.
- A track record of making human-maintained data reliable: shifting headers, merged cells, mixed types, and layout drift between versions.
- Experience with Auto Loader for incremental file ingestion from S3.
- Experience with Databricks declarative pipelines or Delta Live Tables, including schema evolution and data quality expectations.
- Experience conforming ingested data to shared keys so it joins with data from other systems.
- Strong PySpark and SQL, and working knowledge of Delta Lake and medallion architecture.
- Experience working directly with non-technical data owners to pin down definitions and business rules.
- Ability to complete the client's background check and onboarding and to work on client-furnished equipment.
- US citizenship.
- Fluent English, both written and spoken.
Nice to Have
- Experience extracting from Microsoft Access databases (ODBC, JDBC, or file-level tooling).
- Prior work with manufacturing or supply chain data, especially BOMs, parts, revisions, and supplier records.
- Familiarity with PLM or ERP data, such as Teamcenter or Oracle E-Business Suite.
- Prior work in a defense, aerospace, or FedRAMP environment, or with CUI or ITAR-controlled data.
- Databricks Data Engineer Professional certification.
Why Lumenalta is an amazing place to work at
At Lumenalta, you can expect that you will:
- Be 100% dedicated to one project at a time so that you can innovate and grow.
- Be a part of a team of talented and friendly senior-level developers.
- Work on projects that allow you to use leading tech.
For Databricks practitioners specifically, we are a Databricks partner with Champions and MVPs on our team, and that carries three concrete things:
- Delivery Partner Program qualification. Databricks co-delivers its Professional Services engagements through a short list of approved partner organizations, and practitioners have to be individually qualified before they can work on them. The qualification is granted once at the individual level and does not expire, so it stays with you. Partner organizations are the route in, so the qualified pool stays small. That scarcity is most of the value, and we sponsor it.
- Champion path. Databricks Champion is an individual recognition with its own badge, a dedicated Databricks solutions architect as your mentor, access to product teams, confidential roadmap briefings, and an invitation to the internal Tech Summit. Nomination has to come from a partner organization, which puts it out of reach for independents and most contractors. We nominate actively and regularly, and the Champions already on our team mentor candidates through it.
- Certifications covered. We pay the exam fee for any Databricks certification you want to take.
Our Process
A screening call, a technical interview, and a HackerRank assessment. The technical interview is a conversation rather than a whiteboard exam. We will ask you to walk through a governance design you have shipped, a federate-versus-ingest call you made, what you would change, and what you would expect security reviewers to push hardest on. We will also ask for a writing sample. Relevant project references matter more to us than certifications.
Location
This is a fully remote position. This position supports a U.S. Government contract that requires all personnel to possess U.S. Citizenship. This role is remote within the United States only. Candidates must reside in the United States while assigned to our client and be physically located within the United States when performing project work or accessing project systems or data. Availability to work overlapping Pacific, Central, or Eastern U.S. time zones.
Application Deadline
Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

