We are looking for a Senior Software Engineer who is focused on developing platform capabilities, automation, observability, and engineering tooling that improve scalability, reliability, and developer productivity while reducing operational complexity. This person will work closely with engineers across the In-Memory Database Platform, Core, and SRE teams to solve high-impact engineering problems and build scalable solutions.
Responsibilities
- Design and develop automation and platform tooling to operate and scale Valkey reliably and efficiently.
- Build and enhance observability capabilities, including metrics, dashboards, alerting, logging, and distributed tracing.
- Develop tooling and automation for cluster lifecycle management, diagnostics, and other operational workflows, reducing manual intervention and operational complexity.
- Investigate complex distributed-system and performance issues and implement solutions that improve platform reliability, scalability, and performance.
- Build and improve CI/CD automation and engineering tooling to enable safe and efficient development and deployment of platform services.
- Collaborate with Platform, Core, and SRE engineers on architecture, APIs, automation, and tooling, and contribute to the technical documentation needed to develop and maintain the platform.
Requirements
- Strong proficiency in Java and experience developing production backend systems.
- Experience with at least one scalable distributed data store such as Cassandra, Redis, Valkey, or MongoDB.
- Experience designing and building RESTful services and resilient, high-performance backend components.
- Experience with Kubernetes and developing or deploying applications in containerized environments.
- Experience building CI/CD automation using technologies such as Gradle, Jenkins, Spinnaker, and GitHub.
- Experience developing or integrating observability capabilities, including metrics, dashboards, alerting, logging, or tracing for distributed systems.
- Strong understanding of distributed-system principles, including high availability, scalability, fault tolerance, and performance.
- Ability to investigate complex system behavior and translate operational or scalability challenges into engineering solutions.
- Comfortable working independently on ambiguous engineering problems and collaborating across multiple teams.
- Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience
Nice to have
- Experience developing platforms or tooling for Redis, Valkey, or other large-scale distributed database systems.
- Experience building automation for database provisioning, configuration, upgrades, scaling, failover, or other cluster lifecycle operations.
- Experience developing internal platforms, control planes, or infrastructure automation used by other engineering teams.
- Experience with scalable messaging systems such as Kafka or RocketMQ.
We offer
- Opportunity to work on cutting-edge projects
- Work with a highly motivated and dedicated team
- Competitive salary
- Flexible schedule
- Benefits package - medical insurance, vision, dental, etc.
- Corporate social events
- Professional development opportunities
- Well-equipped office
- Please note that all onboarding must occur in person and you may be asked to travel to attend
About Us
Grid Dynamics (NASDAQ: GDYN) is a leading provider of technology consulting, platform and product engineering, AI, and advanced analytics services. Fusing technical vision with business acumen, we solve the most pressing technical challenges and enable positive business outcomes for enterprise companies undergoing business transformation. A key differentiator for Grid Dynamics is our 8 years of experience and leadership in enterprise AI, supported by profound expertise and ongoing investment in data, analytics, cloud & DevOps, application modernization and customer experience.
Founded in 2006, Grid Dynamics is headquartered in Silicon Valley with offices across the Americas, Europe, and India.