Skip to main content

Technical Program Manager – Infrastructure Engineering

Technology
GMI Cloud
臺北市, 台灣4天前截至 2026/11/3
全職

職缺描述

Technical Program Manager – Infrastructure Engineering

About the Role

GMI Cloud is building next-generation AI infrastructure designed for large-scale GPU training and inference workloads. Our platform supports high-density GPU clusters deployed in modern data centers across multiple regions. We combine infrastructure automation, cloud-native orchestration, and high-performance hardware to make GPU compute simple, reliable, and cost-effective.

As we scale, we’re hiring a Technical Program Manager to coordinate cross-functional engineering efforts and help deliver platform reliability, performance, and developer experience.

As an Infra Engineering TPM you will be the connective tissue between SRE, Platform Developers, Architects, and the broader engineering organization. You will own planning, execution, and communication for infrastructure programs—ensuring work aligns to product priorities and is delivered on time with high quality. You’ll support the Head of Infrastructure to turn strategic priorities into actionable roadmaps and to unblock teams operating in a fast-moving start-up environment.

Responsibilities

  • Program leadership
  • Lead end-to-end technical programs across infrastructure domains (GPU orchestration, cluster management, networking, storage, telemetry, provisioning pipelines).
  • Align program scope, milestones, dependencies, success metrics, and delivery timelines.
  • Drive program rituals: kickoff, sprint alignment, risk reviews, and postmortems.
  • Cross-team coordination & stakeholder management
  • Proactively support the Head of Infrastructure in setting priorities, sequencing work, meeting coordination, and driving execution across teams.
  • Serve as the coordinator within infrastructure organization: SRE, Platform Developers, Architects.
  • Proactively surface and manage dependencies, blockers, and risks across teams.
  • Communicate program status, trade-offs, and timelines to the Head of Infrastructure and other stakeholders.
  • Delivery and process
  • Manage roadmaps and release plans for multi-team infrastructure initiatives; work directly with the Head of Infrastructure to align execution with strategic goals and adjust priorities as needed.
  • Implement scalable program management practices (OKRs, milestones, RACI, risk logs, SLAs).
  • Ensure engineering deliverables meet performance, reliability, security, and cost targets.
  • Technical grounding & decision support
  • Maintain a strong technical understanding of GPU infrastructure, ensure architecture and program decisions align with NVIDIA NCP (NVIDIA Cloud Platform) reference architecture and best practices where applicable — including hardware configurations, networking/topology, telemetry, and security recommendations.
  • Facilitate architecture reviews and help translate architecture into executable work for platform engineers and SREs.
  • Reliability & operations enablement
  • Coordinate SRE-driven reliability initiatives (SLOs, incident response playbooks, SOP).
  • Help prioritize reliability work vs. feature work and plan operational runbooks, testing, and rollout strategies.
  • Metrics & continuous improvement
  • Define and track key KPIs (uptime, mean time to recovery, provisioning latency).
  • Use data to drive prioritization and continuous improvement cycles.

Requirements

Minimum Requirements

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.
  • 5 years of technical program management experience at an infrastructure- or platform-oriented organization (start-up experience strongly preferred).
  • Strong technical background with hands-on familiarity in at least one of: GPU compute stacks, Kubernetes, cluster orchestration, or cloud infrastructure.
  • Demonstrated experience coordinating across SRE, platform engineering, and architecture teams to deliver complex multi-quarter programs.
  • Proven track record of delivering projects with multiple engineering teams and external dependencies.
  • Excellent written and verbal communication—able to summarize technical trade-offs and program status for executive stakeholders.
  • Strong organizational skills: roadmaps, dependency mapping, risk management, and prioritization.
  • Data-driven: experience defining, monitoring, and improving KPIs and SLAs.
  • Comfortable with ambiguity and rapid change in a start-up environment.
  • Experience with incident management, postmortems, and on-call processes.
Nice-to-haves
  • Direct experience with GPU orchestration tools (Device Plugin, MIG, NCCL, CUDA drivers), Kubernetes GPU scheduling frameworks, or custom scheduler implementations.
  • Experience with on-prem/cloud deployments.
  • Familiarity with infrastructure-as-code, CI/CD, and observability tooling (Prometheus, Grafana, OpenTelemetry).
  • Prior experience at a GPU-focused company, ML infrastructure team, or HPC environment.
Keywords
monthsOfExperience: 60Plug-inOrchestrationCUDAOCamlGrafanaExecutableCloud computingCluster analysisCI / CDData clusterKubernetesCI/CD

對這個職缺感興趣嗎?