Overview
As Lead Site Reliability Engineer at FactSet, you will ensure the reliability, scalability, and performance of our systems and services. You will work with development and operations to automate, build robust infrastructure, and promote engineering best practices. You will manage incidents, define SLOs/SLIs, and drive capacity planning and performance optimization. This role offers the chance to impact mission-critical financial platforms used by a global client base.
Responsibilities
- Monitor and improve production reliability and availability
- Respond to incidents and conduct post-mortems
- Define and track SLOs and SLIs
- Collaborate with development teams to bake reliability in from the start
- Design and implement automation to reduce toil
- Participate in on-call rotation
- Contribute to capacity planning and performance optimization
- Document systems, processes, and runbooks
Key requirements
- Kubernetes (Required)
- Bachelor in Computer Science or relevant degree
- Hands-on experience deploying, managing, and troubleshooting Kubernetes workloads
- Familiarity with Helm packaging
- Understanding of Kubernetes networking, storage, and security best practices
- Experience with cloud platforms (AWS, GCP, Azure)
- CI/CD tooling (GitHub Actions, ArgoCD, Harness)
- Monitoring/observability tools (Prometheus, Grafana, OpenTelemetry)
- Infrastructure as code (Terraform, Pulumi)
- Config management (Ansible, Puppet, Chef)
- Programming/scripting (Python, Go, Bash)
- English fluency (verbal and written)
- Hybrid work model
- Strong problem-solving and analytical approach
- Excellent communication across technical and non-technical teams
- Proactive with a focus on automation and continuous improvement
- Kubernetes
- Helm
- Cloud platforms
…
