The Identity Solutions program within the Services group works to prevent fraud by securing a user’s identity through their devices, identity information, and usage patterns such as passive Biometrics
Comprised of best-in-class solutions from Ekata (a Mastercard company), we use complex machine learning, combined with device, identity, and transaction information from billions of transactions to reduce user friction to enable highly secure real time transactions
The Identity Solutions – ‘Cloud Infrastructure - Edge’ team is looking for a Lead Site Reliability Engineer to help support, design, and build our highly available and scalable AWS Cloud infrastructure which our products are built on
The team builds infrastructure for Engineering projects, extends our platform for the future, and improves the availability and performance of our applications
This position also designs, develops systems, automation, and tools to help make it easier for Engineering teams to deploy services in a fast, automated and reliable fashion
Be a team and highly accountable project leader in a geo-diverse team. Mentoring the more junior members and being an SME
Plan, design, build, and scale our infrastructure
Work with tools such as Jenkins, Ansible, Argo CD, Terraform, CloudFormation, Resource Manager and many more to ensure that our stack is well represented as Infrastructure as Code
Manage, design, and improve security and availability monitoring and processes for all services, ensure defined security policies are consistently implemented across all environments
Deploy and supervise workloads to cloud environments, proven experience with all of the core services within AWS including instance management, IAM configuration, Database, Caching and general support/troubleshooting
Have a deep understanding of the core components required to run Kubernetes (EKS) and be able to build a cluster from scratch or scale out if needed
Have perfected load balancing, service mesh and always looking for ways to improve availability, response time, and uptime
Maintain, and manage quality documentation for systems owned by the Infrastructure team
Design, build, and improve monitoring tools to identify and resolve issues before they happen. Mainly with Prometheus
Help and advise other teams, troubleshoot and solve failures and performance problems, participate in on-call rotations
Have a basic passion for working with Go, Python, Rust or even Bash to build custom tools and improve system integration. Take code ownership to the next level and act as an advocate for writing code that aligns with industry best practice
Keywords
monthsOfExperience: 36UnixOrchestrationRustScalabilityCloud computingLinuxDevOpsCluster analysisMysqlPythonCI / CDData clusterMemory managementWindows Media PlayerAWSAnsibleDockerJenkinsKubernetes