Lead algorithm research, training fine-tuning, and engineering implementation of Large Language Models (LLMs) to enhance model performance and efficiency;
Lead the implementation of LLM compression, distributed training, and inference acceleration technologies (e.g., quantization, MoE, FlashAttention, etc.);
Design model optimization solutions tailored to business scenarios (e.g., dialogue systems, content generation, knowledge reasoning) to address challenges such as data sparsity and hallucination mitigation;
Track the latest advances in academia and industry to drive the translation of technological achievements;
Lead the delivery of technical solutions and collaborate with engineering teams to deploy high-performance services;
Participate in Agent architecture design and complete the end-to-end implementation of Agent projects.
Qualifications
Master’s degree or above, with 2 years of LLM R&D experience and 5 years of overall algorithm R&D experience;
Proficient in PyTorch/TensorFlow frameworks and familiar with distributed training tools;
Deep understanding of Transformer architecture and its derivative technologies, with extensive experience in LLM fine-tuning;
Well-versed in LLM training techniques such as fine-tuning, distillation, and reinforcement learning;
Proven track record in the full lifecycle experience of LLM training and deployment;
Proficient in Agent architecture;
Publications in top-tier conferences are preferred;
Passionate about embracing innovative technologies like LLMs;