About the Role
As a Senior Computer Vision Engineer, you will lead the end to end development of an enterprise workplace safety and compliance detection pipeline. You will design, train, calibrate and deploy production computer vision models that turn raw client camera footage into operational insight that a reviewer can act on and a client can trust.
The work spans detection and multi-object tracking, region derivation from pose landmarks, instance segmentation for measurement, and temporal and zone aware rules. It also spans the parts that decide whether a model survives contact with a real site: honest evaluation against a contractual target, generalisation across camera viewpoints and locations, compute cost, and a lifecycle that keeps the model accurate after launch.
This role bridges hands on ML engineering with client facing leadership. You will architect pipelines that move from cloud batch execution to edge inference as a configuration change, work directly with global enterprise stakeholders, and transfer capability to client nominated technical staff.
What you'll do
- Build the core detection pipeline. Design and build detection on client camera footage: person and object detection, multi-object tracking, region derivation from pose landmarks, and classifiers that include an explicit "cannot determine" outcome rather than guessing.
- Extend beyond frame level detection. Take the pipeline into instance segmentation for area measurement, temporal reasoning for duration based rules, and zone aware logic where detection rules differ by location.
- Own the ground truth. Lead construction of a labelled dataset from raw footage, including labelling guidelines, adjudication process, class imbalance handling, split hygiene for video, and hard case mining.
- Train, calibrate and prove performance. Train, fine tune and calibrate models against contractual performance targets. Own the held out evaluation methodology end to end: sample selection, adjudication, and honest reporting of TP, FP, TN, FN, error rates and the confidence interval around them.
- Produce outputs reviewers can work with. Derive person level and event level outputs with duration, feeding operational dashboards and alerting where reviewers inspect the underlying frames.
- Design for generalisation. Build so the system holds up across camera viewpoints and sites, and so that adding a camera or a location is a bounded exercise rather than a new project.
- Make deployment location a configuration choice. Architect so inference location is a configuration change: cloud batch now, edge inference as camera counts scale.
- Engineer for compute cost. Consumption is a live client concern. Own frame sampling strategy, motion and zone gating, the CPU and GPU split, and batch against streaming trade offs.
- Own the model lifecycle. Versioning, drift detection, and a retraining loop fed by human reviewer decisions.
- Advise on capture quality. Write camera placement, positioning and image quality recommendations for the client's facilities team.
- Transfer the capability. Deliver recorded enablement sessions and work alongside client nominated technical staff so the client can run and extend what you build.
What we're looking for
- Experience. 5+ years in applied machine learning, with at least 3 years of computer vision shipped to production and operating on real world footage rather than curated datasets.
- Detection and tracking, hands on. Deep experience with object detection and multi-object tracking. You have trained, fine tuned and debugged these, not only run pretrained weights. Familiarity across detector families including YOLO, two stage detectors such as Faster R-CNN, and transformer based detectors such as RT-DETR or RF-DETR, with tracking experience such as SORT or DeepSORT.
- Segmentation and measurement. Instance segmentation experience, and the ability to convert pixel measurements into real world units through camera calibration.
- Pose and keypoints. Human pose estimation or keypoint based region localisation.
- Dataset construction. Demonstrated experience building labelled datasets from raw footage, including labelling strategy, handling severe class imbalance, and avoiding leakage in video splits.
- Evaluation fundamentals. Precision and recall trade offs, operating point selection, probability calibration, and the judgement to say clearly what has not been measured, including to non technical stakeholders.
- Domain shift. Practical understanding of domain shift. You have dealt with a model that worked on one camera and failed on another, and you know what to do about it.
- Distributed processing. Working knowledge of distributed processing. Your pipelines will run on Databricks and Spark. The platform team owns the environment, but you need to be effective in it rather than blocked by it.
- Core tooling. Proficient in Python and PyTorch or equivalent.
- Licensing awareness. Awareness of open source licensing implications for commercially delivered models, including copyleft terms attached to common vision frameworks.
- Data governance. Comfortable working under data governance constraints: client tenant only, no local copies of footage, imagery of identifiable individuals treated accordingly.
- Communication. Strong written and spoken English, and the confidence to present methodology and defend results directly to client stakeholders.
Nice to have
- PPE detection, workplace safety, or video surveillance analytics.
- Temporal action recognition or behaviour classification over video, as distinct from single frame classification.
- Multimodal work combining vision with sensor or time series data.
- Fine grained discrimination between visually similar classes.
- Edge inference deployment, for example NVIDIA Jetson or similar.
- Databricks ML tooling: MLflow and Unity Catalog governed data.
- Systems designed around a human in the loop review step.
- Animal or livestock monitoring.
- Prior consulting or client facing delivery experience.
Location
Fully remote. Must have availability to work overlapping U.S. Pacific, Central, or Eastern time zones.
Application Deadline
Applications will be accepted until October 4th, 2026. Candidates can expect feedback by October 12th, 2026.

