Production Support Engineer - Cummins
(2024-03 - 2025-06)
Provided end-to-end production support for multiple business-critical applications, ensuring incidents were resolved within SLA timelines.
- Delivered detailed Root Cause Analysis (RCA) for recurring issues and collaborated with developers, QA, and business teams to implement permanent fixes.
- Supported and managed AWS cloud environments, including RDS upgrades, patching, EC2 instance management, CloudWatch logs, and DMS migrations.
- Monitored application performance with Dynatrace (dashboards, alerts, synthetic monitoring) to proactively detect issues and minimize downtime.
- Managed high-priority Jira tickets on Kanban board, tracking SLA adherence and providing clear communication to stakeholders.
- Led daily meetings with Customer Service team to review open issues, prioritize workloads, and align on resolutions.
- Created and maintained knowledge base documentation and SOPs for faster troubleshooting and knowledge sharing.
- Provided on-call and weekend production support, ensuring rapid recovery of critical systems during off-hours.
- Recommended process improvements in incident response and monitoring, helping reduce repeat issues and improve efficiency.
- Used Postman to validate APIs during incident troubleshooting, ensuring data accuracy and resolving integration-related issues quickly.
- Managed incident tickets on Jira Kanban board, prioritizing and resolving issues based on severity (Critical, High, Medium, Low) within SLA timelines.
Devops/Site Reliability Engineer - Tata Consultancy Services Ltd
Site Reliability Engineer - Kirkland's
(2021-09 - 2024-03)
Monitored and ensured 24/7 uptime of Kirklands.com, maintaining high customer satisfaction.
- Resolved escalated incidents and coordinated with vendors and internal teams to resolve outages quickly. Prepared RCA reports, and shared with business stakeholders.
- Conducted daily monitoring of application servers (CPU, memory, error rates, response times) using New Relic, Datadog, and AWS CloudWatch.
- Implemented preventive maintenance and coordinated system reboots and patching with infrastructure teams.
- Automated monitoring tasks through Jenkins jobs, reducing manual work and speeding up detection of external service issues.
- Partnered with development and offshore support teams via Jira to track fixes, improve workflows, and ensure timely incident resolution.
- Created and maintained SOPs for monitoring and incident response, improving consistency and knowledge transfer across teams.
- Reduced site downtime and improved customer experience by proactively identifying bottlenecks and resolving recurring issues before they escalate.
Build & Release Engineer - CGI Group Inc
(2016-03 - 2017-02)
Managed build and release activities across multiple applications using Jenkins, Git, Maven, and Ant.
- Developed automation scripts with Shell to reduce deployment errors and improve efficiency.
- Coordinated with project management, QA, and development teams to ensure smooth, error-free releases.
- Created and maintained build pipelines by merging source code from SCM for development, QA, and staging environments.
- Automated deployment steps using Bash shell scripting, reducing manual errors and increasing reliability.
- Configured Jenkins in master–slave architecture to improve build performance and supported migration from SVN to GitLab with identical configurations.
- Integrated static code analysis tools (Checkstyle, PMD, FindBugs, JUnit, DbUnit) into CI/CD pipelines for higher code quality.
- Documented build and deployment processes for knowledge transfer, reducing recurring support inquiries and onboarding time.