Site Reliability Engineer, Big Data (Remote, International)

Company: PulsePoint
Apply for the Site Reliability Engineer, Big Data (Remote, International)
Location:
Job Description:

Overview

In this role you will own the design, deployment and operation of streaming and storage infrastructure at scale within WebMD’s Data Platform. You’ll work across Kafka and Ceph-based systems on hybrid on-prem and cloud environments, shaping architecture, capacity planning and incident response. You’ll partner with multiple engineering teams to deliver reliable, observable platforms that handle billions of events daily. This is a chance to influence platform standards and tackle large-scale distributed systems challenges in a remote-friendly setting.

Pay / Benefits

  • remote work possible
  • large-scale infrastructure experience
  • opportunity to define platform architecture

Responsibilities

  • Design and optimize Kafka architecture, topic governance, partition strategy and throughput
  • Operate Ceph storage, pool design and capacity planning
  • Develop and implement operational automation to reduce manual work and speed incident response
  • Maintain SQL Server backup and recovery pipelines and support basic clustering
  • Build data tooling with self-service capabilities and observability for the team

Key requirements

  • 5+ years operating distributed systems at scale in production
  • Deep expertise in Kafka, Ceph, or similar distributed infrastructure
  • Proven ability to design for scale and reliability
  • Experience mentoring engineers and making technical decisions
  • Willingness to work 9am-6pm ET U.S. hours
  • Mentoring and technical leadership
  • Ownership across system layers
  • Ability to simplify complex systems
  • Apache Kafka
  • Ceph
  • SQL Server backup and recovery

…

Posted: October 1st, 2026