AWS SILVER LINING SYSOPS FELLOW at AWS (2025-02 – 2025-07)
- Completed a 26-week intensive AWS training course, gaining hands-on cloud engineering skills from AWS field architects, training instructors, and customers.
- Configured 3-tier environments using CloudFormation with cost monitoring, deepening understanding of cloud-based infrastructure and data pipeline optimizations.
- Engaged in virtual classroom training, customer lectures, and interactive labs to develop practical expertise in AWS services such as EC2, RDS, S3, EBS, Lambda, and GenAI.
Data Engineer at Cognizant (2022-02 – 2022-12)
- Developed and optimized ETL pipelines using Python and PySpark to process large-scale structured and semi-structured datasets from real-time and batch sources.
- Wrote PySpark DataFrame and Spark SQL transformations for data cleansing, aggregation, and enrichment to support downstream analytics.
- Built reusable Python modules for data validation, logging, and error handling to improve pipeline reliability and maintainability.
- Integrated Kafka streams with PySpark for near real-time data ingestion and processing.
- Leveraged AWS Glue (PySpark-based) jobs to orchestrate scalable ETL workflows and load data into data warehouses.
- Used Git and Linux for version control, job scheduling, and deployment in cloud-based environments.
- Collaborated in Agile stand-ups with cross-functional teams to align data pipelines with application and business requirements.
Data Analyst at NLL Academy (2020-11 – 2021-08)
- Performed data analysis and exploratory data analysis (EDA) using Python (pandas, NumPy) to identify trends, anomalies, and key performance indicators.
- Built Python-based data visualizations to present insights to stakeholders, supporting business decision-making.
- Applied machine learning models in Python and R, including regression and classification techniques, to analyze large datasets.
- Wrote optimized SQL queries to extract data from Oracle databases and validated results through Python analysis.
- Created Tableau dashboards to translate analytical findings into clear, business-focused visual stories.
Jr. Data Engineer at Per Scholas / Cognizant (2019-08 – 2020-05)
- Designed and executed ETL workflows using Python and Apache Spark (PySpark) for healthcare datasets.
- Implemented Spark SQL and DataFrame operations to transform, filter, and aggregate high-volume data.
- Built Kafka producers and consumers to support streaming data ingestion and real-time processing.
- Loaded processed data into MongoDB and relational databases to support analytics and reporting use cases.
- Gained hands-on experience with Hadoop ecosystem tools through classroom labs and team-based projects.
Data Science Fellow at NYC Data Science Academy (2018-04 – 2019-04)
- Built and evaluated machine learning models using Python and R, including linear/logistic regression, decision trees, random forests, and clustering algorithms, to solve real-world analytical problems.
- Applied supervised and unsupervised learning techniques for classification, prediction, and pattern discovery on structured datasets.
- Performed feature engineering, data preprocessing, and model evaluation using metrics such as accuracy, precision, recall, and RMSE.
- Developed data visualizations and dashboards using R Shiny, Python, and plotting libraries to communicate model results and insights.
- Implemented natural language processing (NLP) techniques to analyze scraped IMDB user reviews, including text cleaning, tokenization, and sentiment visualization using WordClouds.
- Conducted end-to-end data science projects, from data collection and exploratory data analysis (EDA) to model building and result presentation.
Software Developer (Intern) at Consumer Priority Services (2017-01 – 2017-02)
- Designed and developed the company's website using ASP.NET, HTML, and Bootstrap.
- Cleaned data and maintained records for the manufacturer's database using Microsoft Visual Studio.
Software Programmer at Inteplast (2014-10 – 2015-05)
- Designed human-machine interfaces using VB.NET, PLC (Programming Logic Controller), and SQL Server to collect thickness data for all 24 zones.