We are looking for a highly experienced Senior Site Reliability Engineer (SRE) to design and operate a highly automated, multi-cloud provisioning platform. You will drive reliability, scalability, and automation while making independent architectural decisions and shaping SRE best practices across the team.
Senior Site Reliability Engineer
Your responsibilities:
- Architect and improve the reliability, scalability, security, and performance of a multi-cloud provisioning platform.
- Design and maintain production-grade cloud infrastructure, with a strong focus on AWS.
- Manage, scale, and provision Kubernetes environments.
- Design end-to-end automation and CI/CD pipelines to minimize manual operations and technical toil.
- Define, monitor, and continuously improve SLIs and SLOs, including latency, throughput, error rates, and capacity.
- Identify platform bottlenecks, single points of failure, and opportunities to simplify complex systems.
- Lead incident response, troubleshooting, Root Cause Analysis (RCA), and implementation of long-term preventive solutions.
- Collaborate with Product, Technical Leads, and engineering teams to embed reliability and security throughout the SDLC.
- Make independent architectural and technical decisions for production environments.
- Promote engineering best practices and mentor other engineers.
We are looking for you, if you have:
- 10+ years of hands-on experience in Software Engineering, DevOps, SRE, Platform Engineering, or Systems Infrastructure.
- 3+ years of advanced, hands-on AWS experience, including designing and troubleshooting production cloud environments.
- 3+ years of production-level Kubernetes experience, including managing, scaling, and provisioning clusters.
- Strong experience with Infrastructure as Code and infrastructure automation.
- Experience designing and maintaining automated CI/CD pipelines.
- Strong understanding of reliability engineering, including SLIs, SLOs, monitoring, capacity, and performance management.
- Proven experience handling critical production incidents and conducting RCA.
- Strong knowledge of cloud architecture, networking, security, and highly available systems.
- Proven ability to work with a high degree of autonomy and make sound architectural decisions.
- Strong problem-solving and troubleshooting skills.
- Excellent communication skills and experience collaborating with technical leads and cross-functional engineering teams.
- Ability to mentor engineers and promote engineering best practices.
Nice to have:
- Experience with Terraform.
- Experience with Azure and/or Google Cloud in addition to AWS.
- Strong scripting or programming skills in Python, Go, Bash, or similar languages.
- Experience building self-healing or highly automated infrastructure platforms.
- Experience with modern monitoring and observability solutions.
- DevSecOps and cloud security experience.
- Experience working with product-driven platform engineering teams.
We offer:
- Participation in interesting and demanding projects.
- Flexible working hours.
- A great, non-corporate atmosphere.
- Possibility to work remote or hybrid (2 days per week from the office).
- Opportunities for development and promotion.
- Attractive package of benefits.
We reserve the right to contact the selected candidates.