A backend and infrastructure developer to support two infrastructure changes in the NDT platform:
Migrating the data layer from blob storage to a (semi-)structured database.
Moving model inference off Databricks onto the Kubernetes GPU cluster.
Alongside this, the developer helps identify resource and speed improvements across the pipeline. The role is collaborative: target designs are created with the existing team, the developer drives establishment of the required infrastructure, and the migration is completed together. The individual is expected to know the listed technologies well enough to make sound default choices without spending time investigating tools or options.
Expected Tools & Knowledge
Kubernetes: deployment, scaling, and resource/GPU scheduling
Specify the data and compute infrastructure for Vestas to provision; setup is handled by Vestas.
Drive migration of model inference from Databricks to Kubernetes GPU compute, specifying the cluster with the Vestas Kubernetes team (VKS).
Improve compute performance through better scaling and resource use.
Own infrastructure design decisions and drive each migration through with the team.
Potentially replace token-based authentication with OAuth.
Potentially establish continuous monitoring, such as Grafana dashboards for inference throughput, GPU utilization, and resource use.
Make and document sound default architectural decisions.
Deep Python expertise with a track record of deployed, maintainable, production-grade services.
Strong REST/FastAPI experience, including routing, dependency injection, and API design.
Works effectively within an existing team and hands off cleanly to subsequent maintainers.
Data Layer & Databases
Expertise in SQLAlchemy and relational modeling.
Hands-on experience with PostgreSQL and Snowflake, with the ability to justify the choice based on access pattern: transactional vs. analytical and structured vs. semistructured.
Experience migrating data from unstructured/blob storage into a structured store without disrupting a live pipeline.
Kubernetes & Compute
Expertise in Kubernetes deployments, scaling, and GPU workload scheduling.
Ability to move compute-heavy inference workloads onto a cluster and tune them for throughput and cost.
Ability to identify resource and speed improvements across a data and compute pipeline.
Collaboration & Judgment
Brings strong default decisions from prior experience, with minimal ramp-up and no time lost researching tooling.
Designs with the team and drives the infrastructure work through completion.
Documents architectural decisions clearly for the team that inherits the work.
Interesseret i denne stilling?