Production Support Engineer – Application Support L2 - TD Bank - Toronto, ON
(2024-03)
Own senior L2 production support for 15+ banking applications across DEV, QA, UAT, Stage, and PROD — primary escalation point for P1/P2 incidents, managing ~150+ incidents quarterly during business hours and on-call rotations with consistent SLA adherence.
- Lead Sev1/Sev2 bridge calls end-to-end — coordinate dev, DBA, infra, and business stakeholders in real time; author structured post-incident RCA reports and drive all remediation items to verified production closure.
- Operate Datadog APM dashboards daily — tune alert thresholds, review distributed traces to isolate latency and error spikes, and maintain service-level views that measurably improved production visibility for the L2 support team.
- Perform Dynatrace observability analysis during slowdowns and outages — review transaction traces, host-level metrics, memory pressure, and GC activity to scope user impact and support engineering teams with root cause identification.
- Correlate logs across distributed banking services using Splunk and ELK Stack — trace multi-system failure chains and feed structured findings directly into RCA documentation and ServiceNow ticket audit trails.
- Monitor Kubernetes workloads using kubectl — validate pod health, stream container logs, and confirm rollout stability during incidents and post-release verification cycles.
- Write SQL Server queries to investigate failed transactions, data mismatches, and locking issues not visible at the application layer — primary resolution path for data-related P2 incidents; support Oracle data discrepancy analysis alongside DBA teams.
- Validate Azure-hosted application health via Azure Monitor and App Services — confirm deployment stability and service availability following releases and infrastructure changes.
- Support Jenkins, GitLab CI/CD, and UrbanCode Deploy release cycles — monitor deployment runs, execute smoke tests, and flag regression risks before production promotion.
- Manage full incident lifecycle in ServiceNow and Jira Service Management — structured intake, triage, escalation, resolution, and audit-trail closure with stakeholder updates at every stage.
- Maintain and expand runbooks and KB articles for recurring issue patterns — standardisation of resolution steps reduced mean time to resolve repeat P2s and shortened onboarding ramp-up for new team members.
Application Support Engineer – Production Systems - LifeLabs - Toronto, ON
(2022-01 - 2024-02)
Delivered 24/7 L2 application support for 10+ laboratory and healthcare systems across QA, UAT, and PROD — responsible for production stability of platforms supporting lab processing and clinical result reporting in a PHIPA-regulated environment.
- Monitored application and service health using Splunk, ELK Stack, and PagerDuty — responded to on-call alerts, triaged severity, and engaged resolver teams within SLA windows; handled ~120+ tickets per quarter across production and UAT environments.
- Managed full incident and service request lifecycle in ServiceNow and Jira — categorisation, triage, escalation, stakeholder updates, and SLA-compliant closure for healthcare-critical workflows.
- Ran SQL Server queries to validate data accuracy, investigate failed batch jobs, and trace interface workflow errors across lab reporting pipelines and healthcare data feed integrations.
- Monitored batch and scheduled jobs across overnight processing windows — early identification of job failures reduced SLA breach risk on critical overnight chains and improved morning shift handoff readiness.
- Wrote Python and SQL scripts to automate recurring data validation checks and post-change verification — reduced manual effort per change window and improved consistency of release sign-off procedures.
- Monitored Kafka-based messaging pipelines for lab result event feeds — validated message flow, tracked topic consumption and processing lag, and escalated delays to the middleware team to prevent downstream reporting impact.
- Supported Azure Monitor across cloud-hosted healthcare components — tracked service availability, application errors, and infrastructure health to maintain environment stability between change windows.
- Participated in CAB (Change Advisory Board) reviews — validated change risk, reviewed pre/post-change checklists, and coordinated team readiness for scheduled maintenance windows in a regulated healthcare environment.
- Maintained runbooks, SOPs, and KB articles aligned to PHIPA compliance requirements — supported audit readiness and consistent incident handling practices across the support team.
Production Support Analyst - Sonata Software - Hyderabad, India
(2018-03 - 2019-12)
Built foundational L1/L2 application support capability across enterprise systems — managed incident tickets in ServiceNow and Jira, resolving issues within SLA windows and developing core triage and escalation discipline.
- Investigated application issues across Windows Server and Linux environments — checked server health, service status, application logs, and data inconsistencies to identify root cause and drive resolution.
- Ran SQL Server queries to validate data integrity, investigate failed transactions, and support defect analysis — developed core data-layer troubleshooting skills applied throughout subsequent roles.
- Monitored application services and batch jobs — reviewed logs and alerting output to detect failures early and escalate before SLA breach; built operational awareness across production workload patterns.
- Supported UAT and production go-live windows — monitored real-time application behaviour, communicated status to stakeholders, and confirmed stability before business handoff.
- Performed environment readiness checks and release validation post-deployment — confirmed services were healthy and application behaviour was stable before marking releases as successful.
- Initiated team knowledge base — documented common incident types and resolution steps, establishing a runbook foundation that reduced repeat escalations and supported faster L1 resolution.