AWS Cloud Engineer / SRE - Tata Consultancy Services - Indore, IN
(2024-04)
Client: LSEG – World-Check One
- Managed AWS infrastructure for a mission-critical KYC platform with 100M+ API calls/day across 3 regions, maintaining 99.95% uptime.
- Automated infrastructure provisioning using Terraform + Terragrunt, reducing manual setup time by 70% and config drift by 90%.
- Implemented CI/CD pipelines with Git + Jenkins, cutting deployment time from 2 hours to 25 minutes.
- Built CloudWatch + Datadog dashboards and alarms for ECS, RDS, Lambda, reducing MTTD by 40% and MTTR by 30%.
- Reduced Lambda cold start p99 latency by 70% from 3s to <900ms by implementing Provisioned Concurrency and right-sizing memory from 512MB to 1GB based on CloudWatch metrics.
- Achieved 99.95% uptime by eliminating timeout errors from 50/day to 0/day through dynamic timeout tuning and memory optimization, cutting retry costs by 60%.
- Led incident response for P1/P2 outages; conducted blameless postmortems and implemented fixes that reduced repeat incidents by 50%.
- Optimized AWS costs using Savings Plans, Reserved Instances, and Lambda right-sizing.
- Proactive monitoring on ECS Fargate + Docker; reduced deployment failures by 60% via blue/green deployments.
- Supported 24x7 production for high-throughput transaction system handling hundreds of millions of events/day across multi-region AWS infrastructure. Maintained 99.99% uptime.
- Managed 100+ production incidents/month across Aurora MySQL clusters, SQS queues, and ECS microservices. Performed rapid triage and RCA on distributed systems with high availability requirements.
- Resolved Aurora MySQL InnoDB HLL alerts and replication lag on 6 MultiAZ clusters. Ran processlist diagnostics, terminated idle queries, and paused batch jobs with zero data loss.
- Managed SQS queue backlog, DLQ alerts, and manual redrives for 4+ queue types including OGS Android seNotification. Applied controlled throttling without message loss.
- Built observability dashboards in Datadog correlating ECS CPU/memory, SQS consumption, DB active sessions, latency, and error rates. Tracked SLIs/SLOs and error budgets proactively to prevent SLA breaches.
- Resolved ECS OutOfMemoryError and container restarts across microservices on Docker. Managed stateful workload recovery and health checks for distributed microservices architecture.
- Led incident management lifecycle in ServiceNow + BigPanda for alert correlation and noise reduction. Conducted weekly reviews across 100+ monthly incidents to drive systemic reliability improvements and reduce MTTR.
- Managed Active-Passive multi-region infrastructure with Global DB Replication across 6 Aurora MySQL, 3 Elasticsearch, and 6 Redis clusters.
Systems Engineer - Tata Consultancy Services - Indore, IN
(2023-01 - 2024-03)
Client: NSK Ltd. & MMC (Mitsubishi Motor Corporation)
- Provisioned and managed multi-user AWS Workspaces on production environments, ensuring high availability and security compliance.
- Implemented cost-saving initiatives including resource rightsizing and Reserved Instance adoption, reducing cloud spend measurably.
- Engineered AWS Transit Gateway configurations for scalable, secure network connectivity across VPCs and on-prem environments.
- Automated monthly security posture reporting using CloudFormation, tracking and remediating AWS Security Hub alerts at scale.
- Configured and managed RDS instances (setup, security, backups, performance monitoring) ensuring 99.9%+ database availability.
- Managed OS-level patching of EC2 instances via AWS Systems Manager Patch Manager, maintaining compliance across the fleet.
- Conducted AWS Well-Architected reviews using AWS Lens, identifying opportunities for cost, security, and performance improvements.
- Generated and deployed SSL/TLS certificates (CSR, ACM, private key management) for production workloads.
- Enforced security controls, IAM policies, and access management using AWS Organizations across multiple accounts.
- Administered jump server access controls and authentication protocols for secure remote access to production systems.
- Monitored and managed cloud costs using AWS Cost Explorer and Budget alerts to prevent overruns.
Systems Engineer - Tata Consultancy Services - India
(2022-06 - 2022-12)
Client: Microsoft – Azure DevOps Environment
- Administered Azure DevOps Server (TFS) at application and database levels, including version upgrades and instance cloning.
- Managed SQL Server high availability with Always-On and failover cluster instances, ensuring zero-downtime operations.
- Configured Azure Automation Accounts and Traffic Manager profiles for intelligent traffic routing and workload automation.
- Performed automated password rotation using Azure Key Vault, strengthening credential security posture.
Systems Engineer to Senior Systems Engineer – AWS Infra - Infosys Pvt. Ltd. - Indore, IN
(2020-04 - 2022-04)
Client: DHS Australia
- Led AWS infrastructure activities for a large government client, overseeing VPC security, IAM policies, and EC2 provisioning.
- Designed and implemented Auto Scaling policies (Target Tracking, Simple, Step) and Elastic Load Balancer configurations for high-traffic workloads.
- Built Lambda-based automation for server scheduling, CloudWatch alarm logging to S3 (JSON/SNS), and server inventory generation.
- Created CloudWatch Dashboards and alarms via shell scripts, enabling real-time monitoring of multi-instance environments.
- Performed Penetration Testing remediation: applied IAM policies per pentester recommendations, configured Kali Linux server, and managed AppLocker & Group Policies.
- Imported existing infrastructure into Terraform, enabling IaC-driven management and drift detection.
- Generated Monthly EC2 monitoring reports using Trusted Advisor with rightsizing recommendations, contributing to cost reduction.
- Managed Active Directory administration via PowerShell including bulk user access provisioning and Group Policy management.
- Provisioned and secured AWS VPCs, security groups, EC2 instances, EBS/EFS storage, and RDS (MS SQL Server) environments.
- Administered Linux (RHEL/Ubuntu) instances including Logical Volume Management (LVM) resizing on EC2.
- Implemented AWS Config rules to enforce compliance checks across the cloud environment.
- Optimized EC2 purchasing decisions (On-Demand, Reserved, Spot) based on client workload analysis, reducing infrastructure costs.
Systems Engineer Trainee - Infosys Pvt. Ltd. - Mysuru, IN
(2019-10 - 2020-03)
- Completed 3.5-month stream training in MS Technologies: C#, HTML, CSS, and Angular.
- Completed 2-month generic training in Python, MySQL, Data Structures & Algorithms, and OOP concepts.