DIRECTOR – AI SYSTEM SOFTWARE at Cerebras Systems (2024-01 – Present)
- Directed kernel, runtime, graph-compiler, and inference-serving teams to architect and deliver a high-performance Multi-LoRA inference platform, enabling dynamic adapter composition with sub-millisecond switching latency and significantly improved accelerator utilization across large-scale deployments.
- Architected an automated, GPU-accelerated Data Factory pipeline that generated the largest open-community executable corpus for agentic RL, leveraging compiler-driven program synthesis and enabling scalable training of code-focused LLMs on heterogeneous compute.
- Led PyTorch framework soundness for the Cerebras system usability, ensuring correctness, coverage, and performance of operator lowering through PyTorch Lazy Tensor Core (LTC) and Torch-MLIR pipelines; closed 100+ lowering gaps and enabled near-complete model support on custom silicon.
- Designed and delivered distributed data-pipeline infrastructure optimized for heterogenous compute (CPU, GPU, and custom AI accelerator nodes) to support SOTA GenAI pretraining workloads (e.g., DeepSeek-V3 scale), improving throughput by 3× through kernel fusion, sharded I/O, and parallel scheduling.
- Directed large-scale data generation efforts producing more than 2 trillion curated math, code, and reasoning tokens, enabling high-quality pretraining and SFT for next-generation LLMs; optimized datatrove pipelines for bandwidth, memory locality, and cross-device parallelism.
SENIOR ENGINEERING MANAGER – AI KERNELS, COMPILER & PERFORMANCE at Intel (2020-01 – 2024-01)
- Led a global team delivering 50+ deep learning math kernels (attention, layernorm, GEMM variants, sparse kernels) for Intel GPU/Gaudi with 1.3×–4× speedups across GenAI, NLP, and vision workloads.
- Drove MLIR adoption, designing a new pattern-rewriting subsystem used by PyTorch for operator lowering; improved compile-time, graph fusion efficiency, and backend coverage across Intel accelerators.
- Delivered performance optimizations for models like LLaMA, DLRM, UNet3D, closing 70% of perf gaps vs hand-tuned baselines; improved end-to-end throughput for PyTorch workloads by 30–50% on Intel hardware.
- Piloted LoRA and code-generation experiments (e.g., StarCoder) on Intel's accelerator, demonstrating automated MLIR dialect synthesis and reducing operator-bring-up time by 40%.
- Contributed to MLPerf submissions, resolving CPU/GPU vector–matrix pipeline bottlenecks and achieving competitive throughput on Gaudi/Nervana hardware.
ENGINEERING MANAGER – FULL-STACK AI OPTIMIZATION at Intel (2018-01 – 2020-01)
- Led cross-domain teams (framework, bridge, kernels, compiler, algorithms) delivering best-roofline performance for BERT on Intel Nervana architecture; achieved 2× speedup through GEMM stacking + fusion.
- Built and scaled India's Deep Learning Kernel R&D org, growing team size by 100% and establishing technical leadership pipeline.
- Delivered full-stack optimizations across algorithmic, kernel, and runtime layers while maintaining zero regression across releases.
TECHNICAL LEAD – AUDIO DSP, FIRMWARE & DRIVER STACK at Intel (2012-01 – 2018-01)
- Designed and delivered portable audio kernel + firmware stack for automotive, IoTG, and client segments, enabling multi-OS deployments and reducing integration issues by 50%.
- Implemented secure bootloader for audio subsystem (authentication + sequencing), improving bring-up reliability across SoC generations.
- Integrated Google wake-on-voice, reduced keyword engine latency, and developed post-processing DSP algorithms (EQ, gapless playback).
TECHNICAL LEAD – AUDIO SYSTEMS & DSP at NXP / Trident (2007-01 – 2012-01)
- Led delivery of first Dolby MS10 decoder on TV SoCs; integrated dual-decode, AAC→AC3 transcode, and audio descriptor features.
- Owned audio system bring-up, customer-site tuning, and certification for Dolby/SRS.
SOFTWARE ENGINEER – AUDIO DSP at Philips (2004-01 – 2007-01)
- Developed peak/shelf filters, DSP optimizations, and UHAPI specification implementations for Philips media pipelines.