Requirements
Must have:
- We require a bachelors degree in Computer Science, Computer Engineering, Software Engineering, Electrical Engineering, or a related discipline, plus 4+ years of data engineering experience, or a masters degree with 2+ years of experience.
- We need strong proficiency in Python and SQL, with the ability to deliver production-grade data pipelines.
- We expect experience with cloud data platforms, preferably AWS services such as S3, Glue, Athena, and Redshift, or equivalent tools, along with infrastructure-as-code experience using Terraform or CloudFormation.
- We look for a solid grasp of data partitioning approaches and columnar storage formats such as Parquet or ORC.
- We need experience designing and operating data pipelines that handle time-series and binary data.
- We value the ability to assess open-source tools and adopt them when they are the right fit rather than building everything internally.
- We look for strong instincts around data quality, including monitoring, validation, and lineage tracking.
- Bonus: experience with autonomous vehicles, robotics, or other sensor-driven autonomous systems.
- Bonus: advanced experience with Foxglove or Rerun, including custom extensions or integration into structured log review or annotation QA workflows.
- Bonus: familiarity with the MCAP CLI and/or Python library, including converting MCAP data into columnar formats for querying and downstream processing.
- Bonus: experience with ML data curation, including diversity sampling, pseudo-labeling, and dataset versioning.
- U.S. citizenship is required due to access restrictions governed by U.S. law.
Responsibilities:
- We help design and organize our data lake, including schema definitions, partitioning strategy, and metadata indexing.
- We build and support end-to-end ingestion pipelines that move high-bandwidth sensor logs from vehicles into cloud storage, even with intermittent or unreliable connectivity.
- We implement validation and integrity checks to catch corrupted data, missing sensors, and calibration inconsistencies before downstream processing.
- We define retention, tiering, and lifecycle policies that balance storage costs with development value.
- We create tooling to query raw logs and produce curated datasets for training and evaluation.
- We build automation for cost-effective pseudo-labeling workflows at ingest scale.
- We develop data quality and model performance metrics that guide labeling toward the most valuable examples.
- We deploy and maintain visualization tools for log review, annotation QA, and autonomy debugging.
- We integrate visualization tools with the data lake so engineers can jump from a dataset record or model failure back to the source logs.
- We collaborate with autonomy engineers to define custom visualization panels and build metrics for analyzing unstructured operating environments.
- We create dashboards that show data coverage by terrain, operating environment, and geographic region.
- We establish and document data contracts between our data services and model training consumers.
- We partner with perception, planning, and embedded engineers across the full data lifecycle, from logging schema design and collection triggers to dataset interfaces for training and evaluation.
- We follow and help shape our data engineering standards, best practices, and tooling choices.
- We contribute to the data roadmap and communicate findings to senior technical leadership.
Company:
We are Torc, a leader in autonomous driving since 2007 and now part of the Daimler family. We focus exclusively on software for automated trucks, with the goal of transforming how freight moves. Our team is building the data infrastructure that powers our autonomy program, and this role sits on a lean, high-ownership team tackling large-scale sensor data problems with real impact.
We offer a collaborative, energetic, and team-focused culture, along with a competitive compensation package that includes bonus and stock options, 100% paid medical, dental, and vision premiums for full-time employees, a 401(k) with 6% employer match, flexible scheduling, generous paid vacation available immediately, company-wide holiday closures, and AD&D and life insurance. We are committed to a diverse and inclusive workplace and encourage candidates to apply even if they do not meet every qualification. The role includes a U.S.
pay range of $139,000 to $166,800.