Senior Data Engineer at Anblicks (2025-10 – Present)
Global Data Platform (GDP) — Snowflake medallion platform (Landing/Bronze/Silver/Gold/Platinum) on managed Apache Iceberg, with Reltio MDM, dbt projects executed via Airflow.
- Architected and delivered Phase 1 of a Snowflake medallion Global Data Platform (Landing/Bronze/Silver/Gold/Platinum) on managed Apache Iceberg with Reltio MDM, integrating 27 source systems end-to-end within a strict 6-month timeframe.
- Built Bronze as the system of record with full history via SCD-2 and soft-deletes (no physical deletes/transformation), Silver with domain-scoped canonicalization (pre-MDM core/extended, post-MDM mastered tables), and Gold delivering curated, domain-aligned data products and KPI aggregates — all as dbt projects executed via Airflow tasks (dbt run/build) across standardized ingest → validate → transform → publish → quality task groups.
- Categorized and integrated all 27 sources into domain-aligned canonical models spanning Finance, Sales, Opportunity, Lease, Property, Company, Contact, and Supplemental domains, bridging technical execution with cross-departmental financial and operational reporting needs.
- Enforced data quality using SODA CL checks at the Bronze-to-Silver gate and MDM eligibility scoring, and maintained Iceberg table health through compaction, snapshot retention, and orphan-file remediation.
- Enforced pipeline traceability by capturing execution timestamps via Airflow's execution context and auto-logging pipeline metadata into centralized audit tables, with query tagging and retry/backoff policies for observability, debugging, and governance compliance.
- Independently designed and deployed Snowflake-native AI accelerators (Cortex + agentic AI) that boosted team efficiency by 60-70%, turning manual source onboarding into reusable AI "skills" used across all 27 sources: AI Canonical Model Generator — automated canonical/Silver data modeling, cutting turnaround from days to minutes. AI dbt Code Generator — generated end-to-end dbt skeleton code for all 27 pipelines from mapping files, eliminating repetitive SQL authoring. DQ Accelerator — automated SODA CL data quality rule generation, replacing manual rule authoring.
Data Model Validator
Agent (Streamlit) — an AI co-reviewer that validates uploaded ERDs against industry data standards, profiles source data, and generates compliance checklists. Self-Service AI Skills — reusable skills for any team member to automate Source-to-Target Mapping (STTM), dbt code generation, and reverse-analysis of complex codebases.
- Engineered AI-generated root-cause and remediation summaries for data failures, and automated mapping of Critical Data Elements (CDEs) to the DQ Rule Catalog using Snowflake Cortex at >60% accuracy, with auto-generated business definitions.
- Built parameterized procedures and config-driven rollups that quantify data quality checks across Staging, Refined, Reporting, and Remediation layers.
Senior Data Engineer at Relanto (2025-06 – 2025-10)
- Delivered data engineering across Snowflake and cloud data pipelines, supporting warehousing, transformation, and analytics workloads.
Software Engineer (Sr. Developer) at Cognizant Technology Solutions (2021-10 – 2025-06)
Spectrum Product Development & Next Generation Spectrum 2.0
- Streamlined data retrieval by 30% via optimized Oracle database structures (normalized tables, indexed views, packages), enhancing reporting performance.
- Leveraged Snowflake features including zero-copy cloning, clustering keys, Streams, Tasks, and Snowpipe to build real-time and batch ingestion pipelines, improving data freshness by 30% and reducing query execution time for large datasets by 50%.
- Automated and scheduled Talend and Informatica ETL workflows using Autosys to load data from S3 to Snowflake, external DB to Oracle, and Snowflake to Oracle, reducing manual intervention by 30%.
- Performed end-to-end data validation between Oracle and Snowflake, implementing SQL-based checks to ensure 100% accuracy.
- Reduced system downtime by 40% through proactive monitoring, log analysis, and consistent code synchronization across environments.