LLM Evaluation, Coding Benchmarking & ML Research Reproducibility | AI Researcher
Send a job offer directly to this candidate
LLM evaluation and coding benchmark specialist with 2+ years of experience designing, implementing, and reviewing repository-level software-engineering and machine-learning evaluation tasks for AI agents. Authored and validated 50+ containerized benchmark tasks across Python, Go, Rust, C/C++, Ruby, and ML evaluation workflows. Experienced in adversarial task design, hidden-test development, rubric design, Docker-based reproducibility, automated verification, oracle validation, Harbor evaluation runs, deterministic scoring, and model failure analysis.
Also brings hands-on AI research experience in medical image segmentation, computer vision, biomedical signal processing, and time-series modeling through research at CURAJ, NIT Jamshedpur, IIT Kharagpur, and IIITDM Kurnool.
LLM Coding Benchmark Contributor & Quality Reviewer at Multiple AI Evaluation Platforms, including Snorkel AI Terminal-Bench (2024-01 – Present)
Research Intern – Oral Cancer Histopathological Image Analysis at National Institute of Technology Jamshedpur (2025-12 – Present)
Research Intern – EEG Signal Analysis for Autism Detection at Indian Institute of Technology Kharagpur (2025-08 – 2025-11)
Deep Learning Research Intern – Human Activity Recognition at Indian Institute of Information Technology, Design and Manufacturing Kurnool (2025-05 – 2025-07)
Integrated M.Sc. in Computer Science – Central University of Rajasthan (CURAJ) (2021-01 – 2026-01)
Higher Secondary in CBSE – Jawahar Navodaya Vidyalaya (2020-01 – 2020-12)