Overview
In this role you will own the design, deployment and operation of streaming and storage infrastructure at scale within WebMD’s Data Platform. You’ll work across Kafka and Ceph-based systems on hybrid on-prem and cloud environments, shaping architecture, capacity planning and incident response. You’ll partner with multiple engineering teams to deliver reliable, observable platforms that handle billions of events daily. This is a chance to influence platform standards and tackle large-scale distributed systems challenges in a remote-friendly setting.
Pay / Benefits
- remote work possible
- large-scale infrastructure experience
- opportunity to define platform architecture
Responsibilities
- Design and optimize Kafka architecture, topic governance, partition strategy and throughput
- Operate Ceph storage, pool design and capacity planning
- Develop and implement operational automation to reduce manual work and speed incident response
- Maintain SQL Server backup and recovery pipelines and support basic clustering
- Build data tooling with self-service capabilities and observability for the team
Key requirements
- 5+ years operating distributed systems at scale in production
- Deep expertise in Kafka, Ceph, or similar distributed infrastructure
- Proven ability to design for scale and reliability
- Experience mentoring engineers and making technical decisions
- Willingness to work 9am-6pm ET U.S. hours
- Mentoring and technical leadership
- Ownership across system layers
- Ability to simplify complex systems
- Apache Kafka
- Ceph
- SQL Server backup and recovery
…
