Lead Data & AI Platform Engineer at Cloudflare, Inc. (2025-10 – Present)
- Architected & built AI-driven Self-Healing Data Pipelines capable of automatically detecting, diagnosing, & remediating intra-DAG & inter-DAG (upstream/downstream) workflow failures across large-scale distributed data platforms.
- Developed a modular agentic remediation framework with centralized log collection, intelligent diagnosis, automated recovery workflows, Slack-based operational communication, and knowledge-base-driven pattern matching to enable autonomous RCA, improve pipeline reliability, and reduce manual operational overhead.
- Designed and developed MCP-compatible AI agent APIs and intelligent tool servers enabling autonomous agents to securely discover, access, and query distributed enterprise data platforms, orchestration systems, GitHub repositories, and operational metadata through standardized protocols & semantic context layers.
- Designed and developed a scalable cloud-native real-time revenue computation platform integrating Salesforce and enterprise data systems, leveraging API-driven architectures, PySpark, and SQL to automate complex business-rule-based financial calculations and synchronize computed metrics back into operational platforms.
- Led analytical modeling and exploratory data analysis across revenue, finance, and workforce domains, using BigQuery, SQL, Python, and Looker to generate predictive insights, identify business drivers, and influence executive-level planning.
Sr. Data Engineer at Trepp, Inc. (2023-10 – 2025-09)
- Design & developed data products/solution in datalake for Commercial Mortgage-Backed Securities (CMBS) & Commercial Real Estate (CRE) Product to measure & monitor securitized mortgage, commercial real estate property data.
- Source & migrate data through the data lake - using Spark/Python in AWS/EMR clusters & orchestrate the workflow through step functions/cloud watch.
- Integrated disparate data sources to deliver a 360-degree view of commercial real estate financial, owner, tenant, physical, economic unit's data.
- Configuring data quality checks programmatically in the production pipeline by enabling DQ rulesets.
- Ingested key CMBS data sets to the new Trepp data lake and worked to optimize CDC workflow for delivery via API framework to UI & external clients.
Sr. Data Engineer at Meta Platforms, Inc. (2021-09 – 2023-06)
Data Engineer in Enterprise Products team with experience from end-to-end data lifecycle starting from understanding system design, orchestration, application logging, data modeling, data warehousing, distributed data processing, data visualization – goal metrics/operational health of the product.
- Worked with XFN teams to create data engineering roadmap, followed by developing pipelines to create foundational data for single source of truth & create data visualization for product analytics & deep dive dashboard for product health.
- Developed Scalable Dataswarm data pipeline using Spark/Presto framework to process event logs from payment processing engine sourced from fusion/intern system to transform & load it in hive for near real time dashboard feed.
- Partnered with XFN team for org wide DE fixathon initiative to create deep dive dashboard to report key product metric & operational metrics with slice/dice of the data using key dimensions.
Big Data Engineer at Publicis Media (2016-05 – 2021-06)
- Developed & automated data pipeline using Apache Spark/Scala to orchestrate data processing from data ingestion to insight generation & integrate media logic computation process to optimize the performance of ad effectiveness.
- Design & developed scalable data pipeline using Apache Spark/Python to ingest/transform/analyze log level data from disparate source systems & load in Redshift for daily dashboard feed to measure audience interaction with media ads.
- Developed high velocity ETL data pipeline using kafka (MSK in Qubole) to stream media data in data lake & apply media logic to transform & load data in Hive metastore to support machine learning/data science projects.
- Design & develop ETL framework to extract media data from heterogeneous source system to apply business transformation to create monthly reports and load into downstream system EDW for dashboard reporting/data archival.
- Created Snowpipe to ingest unstructured raw data sourced from AWS S3 & move data to raw layer. Further transformed to a golden/presentation data layer for data analytics to create an aggregated table for metrics creation.
- Orchestrated the data pipeline using Apache Airflow by writing DAG scripts for scheduling & monitoring of the pipelines.
Sr.Data Analyst at Ernst & Young (2014-01 – 2016-05)
- Developed ETL workflow to extract files from the server & transform payroll data to ACA (Affordable care act) standard & load export ready files into upstream system.
- Advanced ad hoc data analysis for report amendment using Sql scripts.
- Designed email optimization method in loyalty portfolio to target customers based on demographics, consumer behavior using ETL framework.
- Devised methodical ranking for coupon offer system in EMJO application based on previous history & current business priorities using ETL framework.
- Charted daily reports to expert makers from various data sources to facilitate data analytics for reporting consumer buying patterns.
Data Analyst at Ericsson Inc. (2010-08 – 2013-12)
- Analyze resource forecast and demand to provide recommendations to drive resource strategy handshake between operations & engagement practice using MORE tools.
- Analyze demand management data for customer units (CU's) to raise actions items for operation to fill demand.
- Developed a prototype model for customer segmentation based on telecom numeric, effective date, digitized address which provides unmatched information to regulatory affairs for validation.
ETL Developer at Tata Consultancy Services (2006-07 – 2009-07)
Client: Nielsen, Equifax
- Developed a blended premium application which involves migration of data from progress database to oracle and extract data from oracle to create reports in.csv format based on business rules.
- Standardized 3rd party data to Equifax format (data profiling, data cleansing, and data validation), validate & load in DB2.