Overview
In this role, you design and evolve systems that support Prime Video’s live OTT delivery, focusing on automation, monitoring, and reliability for 24/7 live operations. You will work cross-functionally to resolve live-event issues and scale our live Ops infrastructure, influencing customer experience at scale. Your work helps ensure smooth streaming during high-profile events while advancing the overall live monitoring ecosystem. This is a fast-paced, impact-driven opportunity to shape the reliability of Prime Video’s live services.
Responsibilities
- Design and implement systems supporting Prime Video’s Live OTT delivery
- Develop tooling and automation to scale Prime Video Live Operations
- Grow monitoring capabilities and systems reliability
- Develop operational metrics and maintain service level agreements
- Provide escalation support for incoming trouble tickets
- Collaborate with other teams to integrate into the broader Live Ops environment
Key requirements
- Experience programming with at least one modern language (C++, C#, Java, Python, Golang, PowerShell, Ruby)
- Experience using AWS cloud solutions in a DevOps environment
- Experience with CI/CD pipelines build processes
- Troubleshooting and debugging technical systems or operating highly available distributed systems
- Knowledge of system architecture, optimization, system dynamics, statistics, reliability analysis, and electronic system design
- Experience in network fundamentals (DNS, DHCP, TCP/IP, routing, switching, HTTP)
- collaboration
- ownership
- problem-solving
- C++, C#, Java, Python, Golang, PowerShell, Ruby
- AWS cloud solutions
- CI/CD pipelines
…
