As a Senior Site Reliability Engineer at our company 911, you'll own the infrastructure that keeps our platform reliable, scalable, and secure - work that directly supports mission-critical 911 systems used by public safety agencies. You'll drive infrastructure-as-code practices across AWS, lead observability efforts through Datadog, and bring modern AI-assisted engineering approaches into how the team builds and operates.What You'll DoOwn and evolve AWS infrastructure using Infrastructure-as-Code (Terraform / Terragrunt)Architect and scale AWS environmentsDeploy, scale, and manage containerized workloads using Kubernetes and Docker; contribute to HA/DR architecture and platform strategyLead deployment and release processes using Argo (reference JD also names Bitbucket, Jenkins as part of the CI/CD toolset).Define and enforce SLOs, SLIs, and error budgets; drive toil reduction across the platformDrive full utilization of Datadog for monitoring, dashboards, and alerting across the platform (reference JD also names Prometheus, Grafana as potential observability tooling)Build self-service internal developer platforms that empower teams to ship faster.Take end-to-end ownership of infrastructure projects - define success criteria, execute, and measure outcomes.Partner cross-functionally with engineering teams (e.g., network engineering, Dev owners) on long-term technical planning.Bring AI-assisted engineering practices (e.g., Claude, MCP integrations) into daily workflows to improve team efficiencyDocument work and provide cross-training to peers.Resolve JIRA tickets across Cloud, CI/CD, deployments, and monitoring.Requirements: At least 6 years of experience as a DevOps/SRE engineer in a cloud environmentHands-on, production-level AWS experience.Hands-on production experience with Kubernetes and containerizationExperience with Terraform/Terragrunt (or similar Infrastructure-as-Code tools) - requiredStrong Bash scripting skillsDeep understanding of SRE principles: SLOs, SLIs, error budgets, toil reduction, blameless post-mortemsStrong incident management / on-call experienceSolid understanding of APIs, microservices, and distributed systemsDemonstrated experience leading a project end-to-end, from defining success criteria through delivery and measurementCommunicates effectively across teams and can drive long-term technical planningPractical experience with AI-assisted engineering tools (e.g., Claude, Cursor) and MCP-style integrations is a strong plusExperience building AI/ML infrastructure (model deployment, inference pipelines)-plus.This position is open to all candidates.
מתעניינים במשרה הזו?