- We are looking for a skilled and motivated Lead Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and performance of our systems and services. You will work closely with development and operations teams to build and maintain robust infrastructure, automate processes, and drive engineering best practices
- Monitor, maintain, and improve the reliability and availability of production systems
- Respond to and resolve incidents, conducting thorough post-mortems to prevent recurrence
- Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
- Collaborate with development teams to build reliability into services from the ground up
- Design and implement automation to reduce toil and improve operational efficiency
- Participate in an on-call rotation to support critical systems
- Contribute to capacity planning and performance optimization efforts
- Document systems, processes, and runbooks to support the wider team
Benefits
- Comprehensive health coverage for employees and their families, at little or no cost to employees
- Free working lunch in the office Monday through Friday
- A social community involved in sports, charities, and in-office events
- Certification reimbursement for eligible expenses related to the CFA, IPM, CAIA, and FRM exams
- Generous PTO for personal, vacation, parental and medical leave
- Wellness discounts across gyms and other facilities
Strong understanding of core Kubernetes concepts including Pods, Deployments, Services, ConfigMaps, and IngressHands-on experience deploying, managing, and troubleshooting workloads in KubernetesMust be fluent in English both verbal and writtenUnderstanding of Kubernetes networking, storage, and security best practicesFamiliarity with Helm for application packaging and deploymentBachelors degree in computer science or relevant degreeExperience with Kubernetes cluster management and administrationWilling to work a hybrid modelCommitment to a blameless culture and continuous learningStrong problem-solving and analytical skills with a methodical approach to troubleshootingExcellent communication skills with the ability to collaborate across technical and non-technical teamsAbility to work effectively under pressure, particularly during incident responseA proactive mindset with a focus on automation and continuous improvementExperience contributing to open-source projectsFamiliarity with SRE principles as defined by the Google SRE handbookPrevious experience in a DevOps or Platform Engineering role
#J-18808-Ljbffr…
