ML Systems Engineer · LLM Inference · GPU Kernel Developer
Send a job offer directly to this candidate
ML Systems and GPU Kernel Engineer with deep hands-on experience in LLM inference optimization, custom CUDA/Triton kernel development, and transformer internals. Built a complete GPT-style inference engine from scratch — independent of HuggingFace - with FlashAttention-style tiling, fused kernels, KV-cache management, and INT8/FP16 quantization. Familiar with fine-tuning techniques including LoRA and RLHF paradigms. 2+ years of production C/C++ systems work at HCL Technologies.
Pursuing M.Tech (AI & ML) at IIITDM Jabalpur. Targeting applied GenAI Research and ML Systems roles where inference speed, model quality, and large-scale deployment intersect.
Software Engineer at HCL Technologies (2021-09 – 2023-10)
M.Tech in Computer Science (AI & ML) – IIITDM Jabalpur (2024-06 – 2027-06)
B.Tech in Information Technology – IGEC, Sagar (2017-01 – 2021-01)