Senior Data Engineer at Wells Fargo International Solutions Pvt Ltd (2023-02 – Present)
Tech Stack: PySpark, BigQuery, GCP Dataproc, Cloud Composer, Airflow, Dataplex, Django
- Led end-to-end migration of on-premise SQL databases to GCP, onboarding 400+ tables into Cloud Storage, Cloud SQL, and BigQuery, and designed a three-layer Medallion Architecture (Raw → Business → Service) using Dataproc and PySpark.
- Applied Spark optimization techniques (caching, broadcast joins, smart partitioning) and BigQuery partitioning/clustering, cutting pipeline processing time by up to 83% and reducing compute costs.
- Automated daily data movement, error handling, and failure alerting via Airflow on Cloud Composer, reducing manual monitoring effort.
- Implemented data governance via Google Dataplex, including dataset tagging, access policies, and IAM-based role controls.
- Designed star/snowflake schema BigQuery models to support efficient business reporting and analytics.
- Architected a centralized REST API platform for self-service data access, eliminating redundant data duplication and cutting storage costs across business units, using Django ORM (SQL Server) and PyArrow for high-performance retrieval from Dremio.
- Led migration to OpenShift (OCP) with APIGEE integration for authentication, rate limiting, and monitoring — later adopted as the standard for other internal APIs.
- Led and mentored a team of 3 engineers, driving code reviews and testing standards that reduced production incidents; authored TDDs/BRDs and partnered with stakeholders on solution design.
- Engineered a GenAI agent with LangChain for automated data profiling at scale, applying stratified sampling on PySpark pipelines to surface statistical patterns and auto-generate data quality validation rules.
- Prototyped a GenAI-powered Spark Log Analyzer using LangChain and ChromaDB, chunking and embedding Spark job logs into a vector store to enable semantic search for root-cause debugging and LLM-generated performance optimization recommendations
Senior Software Engineer at DeepCompute Software (India) Pvt Ltd (2020-08 – 2023-02)
Tech Stack: AWS (S3, Glue, Lambda, Step Functions, Athena, OpenSearch, DMS, Transfer Family, Lake Formation), PySpark, PostgreSQL, Apache Parquet, Textract/Tika
- Architected an end-to-end serverless data lake on AWS, orchestrating hybrid ingestion of unstructured R&D research PDFs via AWS Transfer Family (SFTP) and structured clinical data from PostgreSQL via AWS DMS into Amazon S3.
- Engineered distributed PySpark ETL pipelines in AWS Glue to join relational PostgreSQL data with PDF document metadata, automating cataloging and reducing manual metadata processing effort by 90%.
- Automated document parsing and text extraction using AWS Lambda, Step Functions, and Amazon Textract/Apache Tika to convert unstructured research papers into structured, queryable data.
- Optimized analytical performance and cost by partitioning curated S3 data into Apache Parquet, improving Athena query speed by 60% and cutting query costs by 45%.
- Accelerated R&D discovery for 100+ research professors by integrating Athena catalogs with Amazon OpenSearch for full-text search, with direct secure PDF access embedded in reporting dashboards.
- Enforced HIPAA/GDPR compliance and reliability via AWS Lake Formation, fine-grained IAM policies, and automated CloudWatch monitoring, sustaining 99.9% pipeline availability.
Senior Software Engineer at Mieone Technologies Pvt Ltd (2017-03 – 2020-08)
Tech Stack: Django, REST API, ReactJS, Firestore, MySQL, Python, Barcode/POS Integration
- Built a barcode-based order return system for a leading Indian unicorn online retail platform, enabling users to scan items at POS terminals and instantly fetch product data from Firestore for real-time return processing, with statuses synced back to the main database to keep online seller inventory accurate.
- Reduced inventory return-to-seller turnaround time by 50-70% and improved return accuracy by eliminating manual data-entry errors.
- Designed backend services for Stockone, a multi-tenant SaaS warehouse management system supporting inbound, outbound, and inventory operations with dedicated customer portals and offline POS.
- Built high-concurrency REST APIs for real-time stock visibility, reducing manual reconciliation effort.
Senior Software Engineer at Headrun Technologies Pvt Ltd (2015-07 – 2017-03)
Tech Stack: Django, REST API, Python, Scrapy, MySQL
- Led backend development for a social fashion discovery mobile app, building Django REST APIs powering real-time trending brand tracking and in-app purchase flows.
- Built a large-scale Scrapy web scraping pipeline to collect runtime video availability and metadata from multiple OTT platforms, served through a unified API.
- Built a Django-based data comparison tool to validate content differences against reference sources like IMDB, improving data integrity across a multi-source content pipeline.
Software & QA Engineer at AccessAutomation Pvt Ltd & Headrun Technologies (2010-09 – 2015-07)
Tech Stack: Python, Scrapy, MySQL
- Developed automated test scripts for Dell India R&D's RAID storage management product (OMSS), improving test coverage and reducing manual regression testing effort.
- Built a social graph aggregation service to consolidate a user's friend networks from Google, Facebook, and Twitter into a single platform.