Computer Vision Research Assistant at Center for Research in Computer Vision, UCF (2025-08 – 2025-11)
- Designed and implemented a knowledge distillation pipeline to transfer zero-shot image classification capabilities from CLIP's ViT-based visual encoder (e.g., FastViT) to a lightweight ResNet-50 backbone.
- Developing a pre-encoder token-pruning method for video MLLMs — run-length tokenization that merges temporally redundant patches across frames — to run long-video inference under a fixed compute budget with no loss in accuracy.
- Developing streaming video MLLMs that combine token compression and semantic-aware pruning to sustain high-FPS perception under fixed compute budgets, targeting latency-critical settings such as autonomous driving and robotics.
Research Assistant at Applied Multimedia Lab, NYUAD (2024-08 – 2025-05)
- Developed a emotion classification model and optimized a transformer-based ATCNet model to classify five emotional states from EEG signals, improving model accuracy from an initial 68% to a peak of 74.45%
- Engineered a deep learning pipeline for emotion recognition by implementing an ATCNet architecture that uses convolutional layers, multi-head self-attention, and temporal convolutional networks to effectively extract spatiotemporal features from raw EEG data
HCI Research Intern, Advised by Dr. Mark Billinghurst at Empathic Computing Lab (2024-06 – 2024-12)
- Fine-tuned and deployed a Llama-8B LLM using parameter-efficient QLoRA to accurately classify deictic language from live speech, enabling real-time, low-latency visual cues in a collaborative VR application.
- Engineered an end-to-end machine learning pipeline in Unity that integrated OpenAI's Whisper for speech-to-text with the custom-tuned LLM to seamlessly translate conversational intent into visual aids.