Generative AI & Data Engineer - Oaktree Capital Management - Los Angeles, CA
(2024-07)
Led GenAI and cloud data engineering for a global investment management firm — built RAG-powered analytics pipelines, LLM-integrated data products, anomaly detection systems, and modernised petabyte-scale AWS infrastructure.
- Designed and deployed end-to-end RAG pipelines integrating vector databases and embedding models to build intelligent document search and Q&A systems over investment datasets.
- Built Agentic AI workflows using LangChain and AWS SageMaker to automate multi-step data retrieval, summarization, and decision-support tasks across investment platforms.
- Implemented LLM-integrated data pipelines (Python, PySpark, dbt) to enrich structured financial datasets with AI-generated insights, classifications, and anomaly flags.
- Engineered semantic search solutions using vector embeddings and chunking strategies enabling natural language querying over petabyte-scale datasets in S3 and Redshift.
- Delivered ML-powered anomaly detection and predictive analytics (NLP, LSTM, SageMaker, Kubeflow) for real-time security KPI monitoring and performance forecasting.
- Built layered dbt models (staging to intermediate to marts) with Jinja macros and SCD Type 2 logic supporting AI-ready dimensional data marts for claims, providers, and encounters.
- Designed scalable ETL pipelines (AWS Glue, Redshift, PySpark); automated incremental loads with Apache Airflow; built serverless file processing with AWS Lambda.
- Enforced HIPAA and PCI-DSS compliance through Terraform-managed IAM roles, encryption at rest and in transit, and fully automated CI/CD via CodePipeline.
- Deployed Denodo data virtualization to unify structured and unstructured sources for AI/ML workloads; designed microservices for SaaS AI data products using Flask and Django.
Data Engineer - InMobi - Bangalore
(2021-03 - 2022-06)
Built distributed data solutions on Hadoop and AWS supporting ML model training, real-time streaming analytics, and cloud data infrastructure for manufacturing intelligence.
- Built and deployed ML models on AWS SageMaker for predictive analytics; developed in-time peak/valley alert monitoring using Python, SQL, and threshold-based auto-email triggers.
- Architected real-time streaming pipelines using AWS Kinesis Data Streams, Firehose, and Lambda to capture and store event data in S3 and DynamoDB for ML training datasets.
- Designed Hadoop-based distributed ETL solutions (HDFS, HIVE, Sqoop, Kinesis) and Spark/Kafka/HBase stacks for scalable feature engineering pipelines.
- Built CI/CD pipelines using Docker, Jenkins, and AWS ECS; used Terraform Cloud with S3 and Lambda for infrastructure deployment and EMR cluster management.
- Integrated Snowflake with S3 for nested JSON ingestion using Clone and Time Travel; built data quality and compliance checks using Airflow, SQL, and Cloud Functions.
- Followed Agile/SAFe delivery with Jira; tracked team velocity and sprint progress within Program Increment (PI) planning cycles.
Data Analyst / SQL Developer - AudLink - Bangalore, India
(2019-03 - 2021-03)
Delivered BI reporting, data warehousing, ETL development, and early ML prototyping for a quality and compliance-focused organization.
- Built ML algorithm prototypes using Python (Pandas, NumPy, Matplotlib) and R for predictive modeling and statistical analysis supporting business decision-making.
- Designed Power BI dashboards (bar, waterfall, gauge, tree map) with tabular cube models and data quality validation; built SSIS ETL packages using Lookup, SCD, Merge Join, and Pivot transformations.
- Developed conceptual, logical, and physical data models in Erwin for OLTP and reporting systems; administered and optimized SQL Server RDBMS production environments.