Sr. Machine Learning Engineer at HyperSoft Tech (2024-05 – Present)
- Serving as AI Team Lead, overseeing and managing a team of machine learning engineers and DevOps engineers.
- Managing AI projects by assigning tasks, providing technical support to the team, and coordinating API and load testing with the SQA team.
- Maintaining AI deployments on cloud platforms, collaborating with the DevOps team to monitor and resolve errors and bugs daily.
- Led efforts in developing and deploying Generative AI workflows (Diffusion Models) alongside LLaMA-based models for NLP and CV applications.
- Designed and developed a production-ready AI-powered Speech-to-Text, transcript analysis, and Spoken English coaching platform using OpenAI's open-source Whisper model with Faster-Whisper and an open-source Qwen LLM, delivering grammar correction, fluency assessment, vocabulary enhancement, and structured JSON responses.
- Implemented scalable, GPU-accelerated inference using FastAPI and the vLLM engine, optimizing concurrent request processing and low-latency AI serving without relying on external APIs.
- Built a real-time Speech-to-Speech conversational AI platform for English language learning using OpenAI's open-source Whisper model, an open-source Qwen Large Language Model (LLM), and the open-source Kokoro Text-to-Speech model, enabling GPU-accelerated, low-latency offline voice interactions.
- Developed a FLUX-based multi-image face swapping pipeline using ComfyUI, Hugging Face models, and FastAPI, enabling scalable deployment across cloud-based GPU infrastructures.
- Developed a Qwen-based image generation pipeline using ComfyUI, FastAPI, and WebSockets, enabling scalable deployment across server-based and serverless GPU infrastructures.
- Developed a real-time voice cloning and song generation pipeline using the RVC framework and Hugging Face models in Python, and deployed it via FastAPI for seamless voice-to-song conversion.
- Developed a music generation pipeline using Audiocraft, Hugging Face models, Python, and FastAPI.
- Implemented end-to-end RAG workflows using LangChain, combining custom LLaMA fine-tuned models, Chroma database, and Hugging Face text encoders for domain-specific applications.
- Led initiatives on fine-tuning LLaMA-based models using LoRA, QLoRA, and PEFT techniques on cloud infrastructure for domain-specific tasks.
- Developed and maintaing multiple video downloading and data extraction pipelines using Selenium, yt-dlp, BeautifulSoup, and FastAPI, enabling large-scale collection of video content and metadata from diverse platforms.
Machine Learning Engineer at HyperSoft Tech (2023-05 – 2024-04)
Developed and deployed AI-driven products by training and optimizing generative models and computer vision models.
- Developed and deployed AI-driven products by training and optimizing generative models (Stable Diffusion, Audiocraft, RVC, GANs) and computer vision models (YOLO, FastSAM, Detectron2), integrating them into production APIs using TensorFlow, PyTorch, FastAPI, and Flask.
- Performed training of generative and computer vision models by applying feature engineering workflows, such as dataset collection, image curation, and CVAT annotation.
- Performed model optimization to minimize GPU consumption and deployed AI solutions on cloud environments through Docker-based containerization.
- Implemented and deployed AI solutions across multiple cloud services, such as AWS (EC2, Lambda, SageMaker), Google Cloud Platform (Compute Engine, Cloud Run), Salad Cloud, and Modal, optimizing for scalability and performance.
- Assisted in MLOps and DevOps tasks for AI deployments, including CI/CD setup, Docker containerization, GitHub management, and Hugging Face model integration.
Machine Learning Engineer at 99technologies (2022-07 – 2023-04)
Led end-to-end development and production operations of Machine Learning and Computer Vision solutions for USA-based clients.
- Led end-to-end development and production operations of Machine Learning and Computer Vision solutions (Object Detection and OCR) for USA-based clients.
- Trained and optimized YOLO-based computer vision models using client-specific datasets collected and prepared for production.
- Led end-to-end dataset preparation, including collection and CVAT-based annotation of custom and client-provided datasets for training computer vision models.
- Developed and integrated OCR systems within the Audiono project, leveraging simulated key commands to streamline client workflows as part of end-to-end Machine Learning and Computer Vision solutions.
- Deployed ML/CV projects on client systems by integrating CI/CD pipelines with GitTea, implementing end-to-end MLOps for seamless model delivery and updates.
- Collaborated on web scraping projects to extract and process data from websites for US-based clients using Python, Selenium, and BeautifulSoup.
Artificial Intelligence Developer at SiParadigm Diagnostic Informatics (2021-10 – 2022-09)
Implemented Object Detection models for on-demand OCR solutions for banking sector clients.
- Implemented Object Detection models using TensorFlow and YOLO as part of on-demand OCR solutions for banking sector clients.
- Trained YOLO and TensorFlow Object Detection models on team lead-provided datasets to identify and detect text groups in images.
- Optimized trained models for edge deployment by converting to TensorFlow Lite and applying 16-8 bit quantization, enhancing inference speed and efficiency on mobile devices.
- Deployed TensorFlow Lite optimized models in an Android application, collaborating with the Android developer and guiding the installation of the TensorFlow library in Kotlin.
- Enhanced existing datasets for Object Detection by collecting additional images, labeling, and annotating them via CVAT, improving model accuracy and robustness.