DevOps Engineer
Technology
広域東京エリア, 日本1ヶ月前まで 2026/9/17
正社員
職務内容
We are seeking an Operations Lead Engineer to drive the reliability, release automation, and continuous improvement of our business applications. This role balances day-to-day production stability with proactive SRE/DevOps initiatives in a highly collaborative, global environment.
Key Responsibilities
- Production Operations & Release Management
- L1/L2 Support: Lead incident response, troubleshooting, and root cause analysis (KEDB).
- Release & Batch Execution: Coordinate change management and deployment activities using Jenkins, GitHub, and Control-M.
- Pipeline Engineering: Design, implement, and optimize CI/CD pipelines (Jenkins, GitHub Actions).
- Toil Reduction: Automate repetitive tasks (deployments, data extraction, batch jobs) using scripting languages.
- Alerting & Metrics: Build and maintain telemetry dashboards using Prometheus, Grafana, Dynatrace, CloudWatch, and Splunk.
- MTTR Reduction: Optimize alerting logic to accelerate incident recovery and analyze performance trends.
- Hybrid Cloud Support: Manage application configurations on AWS and OpenShift.
- Infrastructure Coordination: Collaborate on network configs (DNS, Load Balancers, Firewalls) and SSL certificate renewals.
Required Skills & Experience
Must Have
- Experience: 3 years in IT Operations, SRE, or DevOps.
- OS & Cloud: Strong administration skills in Linux (RHEL) or Windows Server; basic knowledge of AWS and containerization (Docker, Kubernetes/OpenShift).
- CI/CD & Git: Solid experience with Jenkins and Git-flow (branch management, PRs).
- Scripting: Proficiency in at least one scripting language (Python, Shell/Bash, Groovy, PowerShell).
- Observability: Hands-on experience with modern monitoring tools (e.g., Grafana, Prometheus, Splunk).
- Agile & Docs: Experience working in Scrum/Kanban (Jira) and writing technical runbooks.
- Experience in the Financial or Insurance industries.
- Experience with enterprise job schedulers (like Control-M).
- Knowledge of Infrastructure as Code (IaC) via Terraform.
- Practical understanding of SRE concepts (SLOs, Error Budgets, Post-mortems).
Keywords
monthsOfExperience: 36KanbanGrafanaOpenShiftJiraCloud computingLinuxDevOpsGroovyPowershellPythonScrumCI / CDWindows ServerShell scriptAWSDockerGitGithubJenkins
この求人に興味がありますか?