we are looking for a ML Engineer.Own the model plane end to end - vLLM inference, fine-tuning (GRPO/RL and LoRA), and the evaluation-and-promote pipeline that takes an owned model from checkpoint to production serving. We train and serve our own models on GPUs in our own cluster.Requirements: 2+ years of ML engineering or applied ML researchDeep familiarity with HuggingFace transformersExperience running inference servers (vLLM, TGI, or Triton)Python and CUDA fundamentalsUnderstanding of quantization (AWQ, GPTQ, GGUF)Bonus: GRPO/RLHF/DPO training, GPU workloads on KubernetesThis position is open to all candidates.
מתעניינים במשרה הזו?