Principal, Systems and Infrastructure Engineer - BCDR - WALMART - Bentonville, Arkansas
(2019-01 - 2026-12)
- Oversaw creation of business impact analyses across key sectors of company, identifying vulnerabilities and implementing corrective measures that improved system robustness.
- Collaborated with international partners to align global disaster recovery processes, establishing coherent and efficient approach to risk management and disaster recovery.
- Led team of analysts in conducting regular reviews of disaster recovery plans, resulting in continuous improvement of response procedures and communication strategies.
- Assisted with providing input on establishing clear strategy, policies, and procedures for DR governance.
- Designed and implemented DR exercise record workflows for E2E testing for three phases in repository tool.
- Developed ServiceNow change management record to autofill BCDR requirements for all production exercises.
- Implemented BCDR calendar to display each DR exercise from past to future dates with details via Tableau.
- Automated communications via Outlook to stakeholders for upcoming exercises.
- Created Jira boards for different services provided by BCDR to intake work, runbooks to failover and failback application traffic at GSLB level, and runbooks to scale down and scale up pods for application and services.
- Directed changes with BCDR repository tool, vendors implementing BCDR program, vendor contracts, and all DR exercises for critical business services (CBS).
- Carried out DR exercises for new applications during onboarding and before handing off to BCDR practitioners.
- Managed updates to crisis management documentation to reflect changes in technology, leadership, and procedures, keeping company compliant with evolving regulations.
- Assisted with testing cyber resilience exercise tools for critical business services.
- Worked with application teams to implement HA within region and then test capability.
- Collaborated with application teams on incidents they were part of and came to conclusion of automating their recovery as part of process.
- Oversaw current platform resiliency and established capability and automation as needed to mitigate the gaps.
- Developed and implemented disaster recovery plans, ensuring 99.9% uptime for critical applications.
- Collaborated with key stakeholders to define and prioritize RTO and RPO, as well as getting sign-off from one exercise to another on the scope.
- Coordinated with cross-functional teams to conduct regular disaster recovery drills, reducing recovery time by 80%.
- Utilized cloud technologies such as Microsoft Azure and Google Cloud to enhance disaster recovery capabilities and reduce downtime.
- Analyzed recovery exercises to refine and optimize disaster management plans.
- Analyzed recovery efforts post incident to identify areas for improvement.
- Coordinated with external agencies for comprehensive disaster response initiatives.
- Monitored and reported to senior management regarding disaster recovery metrics to track performance and compliance.
- Runbooks automation has saved 98% of time with failover / failback tasks.
Senior Disaster Recovery Lead - SUPERVALU INC. (UNFI) - Eden Prairie, Minnesota
(2016-01 - 2019-12)
- Conducted analysis of critical business applications to successfully execute business operations.
- Identified and mitigated risk through vulnerability scans and health check monitoring of critical dependencies at four data center locations.
- Utilized effective system to constantly monitor replication of 10-gig pipe to ensure replication of critical data to disaster recovery (DR) site.
- Identified short- and long-term solutions for DR through active collaboration with architects.
- Recovered critical data in shorter RPO and RTO through implementation and testing of tools.
- Developed DR project template to identify scope and objectives and measured objectives for success of DR test.
- Assessed DR, infrastructure, and software plans to recover and protect business IT infrastructure in event of disaster.
- Managed DR contracts for increases in compute and capacity for critical dependencies.
- Conducted DR tests on annual basis around Enterprise Freeze.
- Lowered cost by reducing hardware from DR contract to include virtual environment in production.
- Reviewed all phases of DR plan, including requirements gathering, project planning, timeline management, resource management, test exercises, training, and role swaps.
- Carried out DR tests in collaboration with infrastructure teams for open systems.
- Supervised pre- and post-test activities and executed DR test.
- Supported internal and external auditors by providing auditing requirements.
- Identified applications and ensured compliance with SOX, HIPAA, and PCI.
- Improved skills of individual team members by conducting training sessions for infrastructure team members in DR tests.
- Fostered congenial and professional relationships with infrastructure teams to reinforce importance of DR.
- Identified goals and roadmap by developing partnerships with leadership.
Continuity Services Manager (CSM) - Computer ScienceS Corporation - Lilburn, Georgia
(2010-01 - 2016-12)
Developed, managed, and tested disaster recovery plans for fast-growing organizations from government, finance, transport, and healthcare sectors in North American region. Coordinated all aspects of contingency planning, including assisting clients in creating and updating recovery plans, documenting processes and procedures, and training client staff in business recovery techniques. Collaborated with IT departments to establish recovery policies, processes, and procedures.
Conducted large disaster recovery testing events to ensure readiness. Developed and delivered business impact analysis to internal and client management.
Business Continuity Services Lead, India - CSC INDIA PVT. LTD - Hyderabad, India
(2008-01 - 2010-12)
Led strategic and tactical aspects of company's information services business continuity and pandemic programs. Ensured minimal downtime of critical business operations in event of crisis or interruption through effective oversight of all continuity preparations and procedures corporation wide. Developed and implemented recommendations and provided enterprise-wide support for business continuity, crisis management, and disaster recovery.
Developed, managed, and supported LDRPS tool. Managed invoice service and test calendar. Reduced recovery time by 150%, improved test failure rates by 100%, and established management reporting system.
Data Center Supervisor / Disaster Recovery Coordinator - CIBA VISION - Duluth, Georgia
(2004-01 - 2008-12)
Oversaw personnel and process management and monitoring, including backups and restores, disaster recovery integrity across systems, change control migrations, and AutoSys / Control-M schedulers. Supervised communication and documentation of corrective actions, follow-up, backup, feedback, and other issues related to system processing, job, and process results. Coordinated activities and documentation reflecting system processing and results, including infrastructure, facility and security issues, feedback, and recommendations.
Improved support systems' efficiency across all functions through effective project management.
System Administrator - SOLVAY PHARMACEUTICALS, INC. - Marietta, Georgia
(2000-01 - 2004-12)