Overview
In this role you will lead and shape the reliability strategy for critical applications within the Athena Core team. You will drive incident prevention, triage, and post-incident analysis, collaborating across cross-functional partners to meet service levels. You will mentor engineers, promote SRE culture, and steer AI-assisted reliability workflows to improve stability at scale. This is a high-impact position at a globally recognized financial services firm, offering a chance to influence how complex platforms stay resilient and secure.
Responsibilities
- Model and promote SRE culture, sharing knowledge through internal forums
- Lead reliability initiatives using data-driven analytics to improve service levels
- Define and align SLIs/SLOs with stakeholders and establish error budgets
- Leverage enterprise AI capabilities for major-incident triage and post-incident analysis
- Serve as main incident contact for applications and resolve issues to minimize financial impact
- Provide high-level technical guidance and mentorship across engineers
- Drive reuse of AI-assisted reliability workflows across SDLC/toolchain with proper controls
- Be part of the first line of support for the Athena Core team
Key requirements
- Formal training or certification in SRE concepts
- Bachelor in Computer Science
- Proven experience in reliability, scalability, performance, security, and enterprise architecture
- Programming: Python, Java/Spring Boot, .Net
- Experience using enterprise-authorized AI to improve SRE workflows with strong validation and data-sensitivity awareness
- Ability to evaluate AI recommendations for correctness and risk and define guardrails
- Linux expertise with strong observability skills (monitoring, SLAs, telemetry)
- CI/CD practices and tooling proficiency
- Containerization and orchestration expertise
- Networking troubleshooting experience
- Advanced knowledge with ongoing self-education in emerging technologies
- leadership
- mentorship
- collaboration
- Python
- Java/Spring Boot
- .Net
…
