Overview
As Engineer II on the TechOps SRE team, you’ll own the reliability of CrowdStrike’s commercial cloud and help drive scalable, automated solutions. You’ll work on large-scale distributed systems, ensuring availability, latency, and throughput while improving monitoring and incident response. You’ll collaborate with global teammates and continuously adopt new technologies to raise the team’s technical IQ. This role offers hands-on engineering with a strong impact on mission-critical services in a fast-paced, AI-enabled environment.
Pay / Benefits
- Market-leading compensation and equity
- Wellness programs (physical and mental)
- Generous vacation and holidays
- Parental and adoption leave
- Professional development opportunities
- Employee networks and volunteer opportunities
Responsibilities
- Manage availability, latency, throughput, monitoring, incident response, and capacity planning for mission-critical platforms
- Troubleshoot and resolve server hardware and platform issues across thousands of servers and VMs
- Engage in on-call rotations and lead incident analysis to drive systemic improvements
- Develop and advocate automation and tooling to improve operations and efficiency
- Stay current with new technologies and expand expertise across the architecture and processes
- Collaborate with global engineers to understand and optimize the overall process flow
- Mentor and uplift team technical capabilities and champion reliable deployment practices
Key requirements
- Bachelor’s degree in Computer Science or equivalent experience
- Minimum 5 years in a large-scale production environment
- At least 2 years of software engineering experience
- 2+ years in one or more of: C++, Java, Python, Go
- Experience with storage technologies (SAN, NAS, NFS, Object Storage, iSCSI)
- Experience with Infrastructure tech (Linux, Windows, VMware, Docker, Kubernetes)
- Experience writing technical documentation
- Configuration management with Puppet, Chef, Ansible
- Strong analytical skills with urgency, ownership, and drive
- Ability to work in a diverse, team-focused environment
- Experience leveraging AI technologies to enhance decision-making and workflows
- Analytical mindset
- Ownership and drive
- Team collaboration across distributed teams
- Linux engineering and administration for large-scale servers
- Monitoring/telemetry stacks: ELK, Prometheus, Grafana, Zabbix
- Performance tuning and fault finding through metrics
…
