Service Delivery Manager at CoreCard Software Solution Pvt. Ltd. (2022-11 – 2026-07)
Managed end-to-end IT service delivery, ensuring services met agreed SLAs, OLAs, and KPIs. Acted as the primary point of contact for customers, conducting regular service review meetings and managing stakeholder expectations.
- Led Major Incident Management (P1/P2), coordinating cross-functional teams to restore critical business services within SLA targets.
- Drove Incident, Problem, Change, and Service Request Management processes in accordance with ITIL best practices.
- Performed Root Cause Analysis (RCA) for high-priority incidents and implemented corrective and preventive actions to reduce recurring issues.
- Monitored service performance using operational dashboards and prepared weekly/monthly service reports for leadership and customers.
- Ensured compliance with Service Level Agreements by tracking incident trends, service availability, response times, and resolution metrics.
- Coordinated planned changes through the Change Advisory Board (CAB), minimizing business impact and improving change success rates.
- Collaborated with Infrastructure, Application, Database, Network, Cloud, and Security teams to ensure seamless service delivery.
- Managed vendor performance and third-party service providers to ensure contractual SLA compliance.
- Identified opportunities for automation and continual service improvement (CSI), improving operational efficiency and customer satisfaction.
- Supported capacity planning, availability management, and disaster recovery readiness to ensure business continuity.
- Maintained ITSM documentation, knowledge articles, SOPs, and operational runbooks.
- Managed customer escalations and ensured timely resolution while maintaining high customer satisfaction (CSAT).
- Utilized ITSM tools such as ServiceNow, BMC Remedy, Jira, and monitoring tools including LogicMonitor, Dynatrace, Splunk, SolarWinds, and cloud platforms such as AWS/Azure (where applicable).
- Managed P1 and P2 incidents for critical banking applications, ensuring minimal business impact and swift service restoration within defined SLAs.
- Initiated incident notifications through relevant tools to promptly inform stakeholders, and quickly set up bridge calls to collaborate with technical teams for timely resolution.
- Followed escalation protocols as per the defined matrix to ensure proper communication, accountability, and SLA adherence at every stage.
- Conducted post-incident root cause identification, created preventive action tickets, and followed up on aging incidents to strengthen system reliability.
- Provided 24x7 on-call incident response support, including off-hours coverage.
- Managed PagerDuty for alerting — configured on-call rosters for technical and escalation teams, and monitored MTTA/MTTR compliance.
- Participated in client meetings (alternate days/weekly) to review performance, discuss incident gaps, and identify improvement opportunities.
- Prepared and shared daily, weekly, and monthly reports with internal management and the client team, summarizing incidents, actions, and improvements.
- Led end-to-end problem lifecycle management — from detection and logging to root cause analysis and permanent resolution.
- Conducted Root Cause Analysis (RCA) using structured methodologies including 5 Whys, Fishbone/Ishikawa, and Fault Tree Analysis.
- Maintained and continuously updated the Known Error Database (KEDB) with validated workarounds and permanent fixes.
- Performed trend analysis on recurring incidents to identify patterns and proactively prevent future disruptions.
- Facilitated Post-Incident Reviews (PIR) for major incidents, communicating findings and action items to stakeholders.
- Reduced MTTR significantly by building a robust KEDB and a standard workaround library.
- Managed end-to-end change lifecycle: RFC submission, impact assessment, CAB approval, scheduling, implementation, and post-implementation review (PIR).
- Chaired and participated in Change Advisory Board (CAB) reviews to evaluate risk, impact, and rollback plans for all change types.
- Classified changes (Standard / Normal / Emergency) and ensured appropriate approval workflows were followed.
- Coordinated with technical teams, business stakeholders, and vendors to plan and execute changes with zero or minimal downtime.
- Developed and maintained change documentation including test plans, rollback procedures, and implementation checklists.
- Tracked change success rates, failed changes, and unauthorized changes, reporting outcomes to leadership.
- Supported emergency change processes for critical security patches and P1/P2 incident resolutions.
- Coordinated end-to-end release planning and scheduling in alignment with change management and project delivery timelines.
- Collaborated with development, QA, and infrastructure teams to ensure smooth deployment of releases with minimal service disruption.
- Maintained the release calendar and communicated release windows to all relevant stakeholders.
- Ensured all release activities followed defined processes including pre-release checklists, go/no-go approvals, and rollback plans.
- Tracked and reported on release outcomes, including post-deployment validation and issue resolution.
Major Incident Management / Service Delivery at Capgemini (2021-11 – 2022-09)
- Upon ticket creation, promptly acknowledged issues, assessed priority and impact, and notified relevant stakeholders within SLA.
- Initiated troubleshooting bridge calls, coordinating with appropriate technology teams (Network, Wintel, VM, etc.) as required.
- Managed incident calls following the escalation matrix, ensuring clear communication throughout the troubleshooting process.
- For issues originating outside Capgemini infrastructure, shared logs with the Client Incident Manager and recommended escalation to the client's technology team with bridge details.
- Took full ownership of high-severity incidents, ensuring regular stakeholder updates until resolution and formal closure.
- Prepared detailed incident chronologies post-resolution and coordinated with the technical team to obtain RCA documents.
- Followed up daily on backlog incidents, coordinating with stakeholders for updates and managing the end-to-end incident lifecycle.
- Analyzed weekly SLA breach reports from EasyVista, facilitating discussions with technical teams to identify gaps and share corrective feedback.
- Hosted daily DUSTUM calls, engaging technology teams to address challenges and manage backlog tickets for timely closure.
- Identified recurring incidents from P1/P2 reports and collaborated with technical teams to raise problem tickets for root cause determination and RCA generation.
- Swiftly initiated problem tickets following MIM resolutions, coordinating RCA discussions with technical teams.
- Maintained a comprehensive tracker for all problem tickets and shared daily updates with management and clients.
- Verified CAB approvals, confirmed implementation dates (minimum two hours ahead), and validated CI details before assigning change tickets to the technical team.
- Maintained a tracker for all change tickets and provided weekly status updates to clients.
- Supported release scheduling by coordinating with change management and technical teams to align release windows with approved change cycles.
- Ensured pre-release validations were completed, including CAB sign-off, rollback readiness, and stakeholder communication.
- Tracked post-release stability, flagging any issues and ensuring timely incident or problem ticket creation when needed.
Major Incident Management at Genpact (2021-03 – 2021-11)
Provided global support across Genpact clients in India, UK, US, South Africa, and the Philippines, handling critical P1, P2, and P3 incidents.
- Upon ticket receipt from the Service Desk, acknowledged issues, set priority and impact, and coordinated with the SDM/requestor as needed — all within SLA.
- Released automated notifications via ServiceNow and SMS alerts through an SMS broadcast tool, including bridge details to relevant stakeholders within SLA timelines.
- Drove troubleshooting bridge calls, engaging Network, Wintel, VM, and other technology teams, business users, and ETOs as required.
- Followed the escalation matrix to escalate to the next level of support within defined timeframes to restore services promptly.
- Took full ownership of high-severity incidents, ensuring all stakeholders were continuously updated through to final closure.
- Prepared six-hourly reports and daily incident tracker reports, sharing them with relevant stakeholders.
- Worked on proactive alert tickets triggered from monitoring tools including Solarwinds, Logic Monitor, and CA Spect