Skip to main content

Server Engineer - Cloud Infrastructure

Technology
Sourcebae
₹10,00,000 - ₹24,00,000 /year4 days agoUntil 26/11/2026

Job description

NOC Server Engineer The NOC Server Engineer provides advanced 24x7 operational support for cloud-hosted server infrastructure in AWS and Azure. This role acts as the shift escalation lead for complex and high-severity incidents, ensuring service stability and rapid restoration. Kubernetes (EKS), Terraform, and foundational cloud networking knowledge are required skill sets to support modern cloud workloads.

Key Responsibilities:

  • Provide deep troubleshooting for Linux and Windows servers hosted in AWS and Azure, including OS, services, performance, and capacity.
  • Analyze alerts and telemetry using CloudWatch, Azure Monitor, Splunk, and Grafana to validate root cause and recovery.
  • Coordinate with Operations Center (OC), Incident Managers, and platform SMEs; maintain clear ownership and escalation.
  • Execute safe, pre-approved operational actions and validate post-change health using documented procedures.
  • Ensure accurate incident timelines, updates, and handovers in ServiceNow/Jira.
  • Mentor engineers and own structured shift handovers to maintain operational continuity.

Required Skill Set:

  • Cloud Server Engineering: Strong hands-on experience supporting AWS and/or Azure compute services (EC2, Azure VMs/VMSS).
  • Linux & Windows Servers: Advanced OS-level administration, patching, log analysis, service troubleshooting, and performance tuning.
  • Kubernetes (EKS) & Container Orchestration: Experience triaging EKS node and pod issues, crash loops, scaling symptoms, and workload health.
  • Infrastructure-as-Code (Terraform): Ability to execute, review, and validate Terraform-based changes following operational guardrails.
  • Cloud Networking Fundamentals: Working knowledge of VPC/VNet, subnets, routing tables, security groups/NSGs, load balancers, DNS, and basic connectivity troubleshooting.
  • Observability: CloudWatch, Azure Monitor, Splunk, Grafana/Prometheus; alert tuning and dashboard interpretation.
  • ITSM & Operations: Strong incident documentation, communication, and SLA discipline using ServiceNow or Jira.

Required Experience:

  • 6 years of experience in Cloud Operations, NOC, SRE, or Infrastructure Support roles.
  • Proven experience leading or handling P1/P2 incidents in 24x7 production environments.
  • Hands-on experience working with runbooks, escalation matrices, and shift handovers.

Work Arrangement:

Work from Office with structured handovers and escalation ownership.

Keywords
OrchestrationGrafanaJiraCloud computingLinuxNode.jsNode

Interested in this role?