Job Description:
We are looking for an AI Eval Engineer to build and improve evaluation systems for AI-powered education products. This is a hands-on role focused on evaluating LLMs, developing Python-based evaluation tools, analyzing AI outputs, and enhancing model quality.
Key Responsibilities:
- Build and maintain evaluation pipelines for LLM-based applications.
- Develop Python scripts and tools for testing, analysis, and regression tracking.
- Create evaluation datasets, rubrics, and quality benchmarks.
- Analyze AI outputs to identify hallucinations, reasoning gaps, and quality issues.
- Collaborate with engineering and product teams to improve prompts, workflows, and AI performance.
Required Skills:
- 3+ years of experience in Software Development, Applied AI, AI Evaluation, QA Automation, or related fields.
- Strong proficiency in Python.
- Hands-on experience with LLM APIs, prompt engineering, JSON, and API-based workflows.
- Experience with AI evaluation, debugging, and quality analysis.
- Good understanding of RAG, structured outputs, and multi-step LLM workflows is a plus.
- Exposure to EdTech, curriculum, or educational AI products is an added advantage.