Overview
In this Site Reliability Engineer II role within Corporate Risk Technology, you will help secure and optimize mission-critical platforms through code and cloud infrastructure. You’ll collaborate across teams to implement reliable deployment and monitoring practices and apply AI-assisted insights to improve incident response and operational signals. Your work enables scalable, available systems that support risk management. This is a chance to contribute to a growing, technology-driven environment that values curiosity and continuous improvement.
Responsibilities
- Guide peers on designing reliable systems and promote SRE best practices across the team
- Collaborate with software engineers to design, test, and implement CI/CD-based reliability approaches
- Implement infrastructure as code and network-as-code for scoped applications and platforms
- Partner with stakeholders to resolve complex problems using SLI/SLOs to prevent customer impact
- Identify roadblocks and propose technology-driven improvements to address business problems
- Use availability, reliability, and scalability principles to improve outcomes with partners
- Utilize enterprise AI capabilities to speed incident triage, troubleshooting, and post-incident analysis
- Leverage AI-assisted signals to identify reliability risks and prioritize repeatable improvements tied to SLOs
Key requirements
- Formal training or certification on site reliability engineering concepts
- Proficiency in site reliability culture and practices
- Proficiency in at least one programming language (Python, Java/Spring Boot, or .NET)
- Experience in observability (monitoring, alerting, telemetry)
- Knowledge of software applications and processes within a technical discipline (cloud, AI, mobile)
- Working knowledge of enterprise AI capabilities and data sensitivity awareness
- Ability to review and validate AI-assisted recommendations with security and data handling in mind
- collaboration
- curiosity
- problem-solving
- site reliability engineering concepts
- CI/CD tooling
- infrastructure as code
…
