Data Engineer at Concentrix Catalyst (2026-04 – Present)
Designed and enhanced metadata-driven governance capabilities on the Databricks Lakehouse Platform to improve sensitive data discovery, secure enterprise data access, and automate governance reporting through scalable metadata engineering and policy-driven controls
- Designed and developed reusable metadata-driven SQL frameworks to analyze enterprise data assets and measure sensitive data classification, tagging coverage, masking policy adoption, and governance completeness across multiple business domains.
- Built governance intelligence solutions by transforming Unity Catalog metadata into actionable governance insights, enabling proactive identification of unclassified sensitive data, inconsistent tagging, duplicate classifications, and masking policy gaps.
- Solved complex masking implementation challenges by evaluating datatype compatibility, reusable masking strategies, policy conflicts, and alternative implementation approaches for heterogeneous enterprise datasets within Unity Catalog.
- Conducted engineering evaluations of platform capabilities by analyzing metadata behavior, automatic classification limitations, masking framework design, and architecture trade-offs to improve solution scalability, maintainability, and long-term adoption.
- Developed reusable governance metrics, reporting models, and auditing frameworks that provided continuous visibility into governance maturity, masking compliance, classification coverage, and metadata quality across enterprise data platforms.
- Partnered with architects and cross-functional engineering teams to design secure, metadata-driven governance solutions leveraging Unity Catalog, RBAC, dynamic data masking, and policy-based access controls aligned with enterprise security requirements.
- Produced technical design documentation, architectural recommendations, proof-of-concept findings, and platform evaluation reports, supporting engineering decisions and future evolution of the enterprise governance platform
Data Engineer at Concentrix Catalyst – Office of Decision Intelligence (2025-04 – 2026-03)
Designed and developed a centralized enterprise analytics platform by engineering scalable data pipelines, reusable transformation frameworks, and trusted analytical data products that enabled standardized reporting, executive decision-making, and AI-assisted analytics across multiple business domains.
- Designed and developed scalable ELT pipelines to integrate, cleanse, validate, and standardize enterprise operational, workforce, project, and financial data from multiple source systems into trusted analytical datasets.
- Engineered reusable transformation frameworks that centralized business rules and standardized data processing, improving consistency, maintainability, and scalability across enterprise reporting solutions.
- Developed dimensional data models and semantic datasets that standardized KPI calculations, utilization reporting, profitability analysis, forecasting, and operational performance metrics for enterprise-wide analytics.
- Built certified analytical data products that enabled executive dashboards, self-service analytics, and downstream reporting while reducing duplicated business logic and improving trust in enterprise data.
- Leveraged DOMO Magic ETL, SQL DataFlows, Beast Mode calculations, and Analyzer to standardize KPI logic and deliver reusable executive reporting assets.
- Optimized SQL transformations and analytical workflows to improve reporting performance, simplify downstream consumption, and deliver reliable, analytics-ready datasets.
- Integrated AI capabilities into the enterprise analytics platform using LangChain and DOMO AI, enabling conversational analytics, contextual insight generation, and intelligent exploration of trusted business data
- Partnered with business stakeholders and cross-functional teams to convert evolving analytical requirements into scalable, reusable data platform capabilities supporting strategic and operational decision-making.
Data Engineer at Valmont Valley Irrigation GenAI (2024-05 – 2025-04)
Designed and developed an event driven cloud processing platform that automated product data ingestion, image transformation, metadata enrichment, and AI-assisted content generation through scalable Azure workflows, enabling faster product onboarding and improving operational efficiency.
- Engineered event-driven processing pipelines using Azure Functions, Azure Logic Apps, Azure Blob Storage, and Python to automate ingestion, validation, transformation, and orchestration of enterprise product data and digital assets.
- Designed reusable processing workflows that standardized image preparation, automated background removal, and improved the quality and consistency of product assets consumed by downstream business applications.
- Integrated Generative AI capabilities into automated processing pipelines to generate product descriptions and enrich product metadata, reducing manual effort and accelerating product onboarding.
- Developed scalable orchestration workflows using Azure Logic Apps, Azure Functions, Azure Automation, and Azure Blob Storage to coordinate cloud services and improve workflow reliability, maintainability, and extensibility.
- Built modular processing components that promoted code reuse and configurable workflow execution while simplifying future enhancements across product categories.
- Implemented application monitoring and operational logging using Azure Application Insights and Azure Automation to support workflow visibility, troubleshooting, and operational reliability.
- Conducted unit testing and integration testing to validate workflow reliability, processing accuracy, and end-to-end automation behavior before deployment.
Big Data Engineer at Cigna (2023-11 – 2024-05)
Contributed to the modernization of a legacy enterprise data platform by engineering scalable ingestion, transformation, and migration pipelines that transitioned Hadoop-based data workloads to Databricks Delta Lake, enabling reliable cloud-native analytics and improving data processing efficiency.
- Designed and developed scalable PySpark ingestion pipelines supporting full-load and incremental processing, ingesting enterprise healthcare datasets staged in Amazon S3 into Databricks Delta Lake.
- Engineered distributed transformation pipelines that standardized, validated, and prepared large enterprise datasets for reliable downstream analytics and reporting.
- Developed Spark-based data processing workflows using Apache Spark (PySpark) to efficiently transform high-volume enterprise datasets while supporting scalable cloud-native data engineering practices.
- Implemented Delta Lake storage patterns that improved data reliability, simplified downstream consumption, and supported efficient analytical processing across curated datasets.
- Modernized legacy Hive-based processing workflows by contributing to the migration of Hadoop workloads into Databricks while preserving existing business logic and ensuring migration accuracy.
- Contributed to legacy Hadoop migration by supporting data movement across Teradata, HDFS, Amazon S3, and Databricks Delta Lake using Sqoop, Hive, DistCp, and PySpark while preserving existing business logic during platform modernization.
- Optimized Spark and SQL transformation logic using Delta Lake optimization techniques