Overview
In this role you will own the AIOps product roadmap and translate operational needs into actionable requirements. You’ll drive adoption of enterprise-scale AIOps capabilities and collaborate across IT Operations, SRE, and platform teams to improve reliability and reduce complexity. You’ll define success metrics and present progress to senior stakeholders, shaping the operating model with data-driven decisions. This role offers the chance to advance AI-driven observability and self-healing initiatives that impact global technology operations.
Pay / Benefits
- Health & Wellness
- Flexible Downtime
- Continuous Learning
- Invest in Your Future
- Family Friendly Perks
- Beyond the Basics
Responsibilities
- Own and execute the AIOps product roadmap aligned with reliability goals and adoption targets
- Translate operational needs into product requirements, user stories, and backlog items
- Own high-impact AIOps capabilities: noise reduction, event management, correlation, anomaly detection, root cause analysis, predictive alerting, self-healing remediation
- Collaborate with Operations, SRE, platform engineering, service management, application teams, security, and vendors to scale adoption
- Define and track success metrics: alert quality, incident reduction, MTTD/MTTR, service health, automation coverage, platform adoption
- Present roadmap progress, risks, decisions, and outcomes to senior stakeholders
Key requirements
- 10+ years in product management, IT operations, SRE, observability, or related enterprise technology roles
- Strong understanding of AIOps concepts: event correlation, anomaly detection, root cause analysis, noise reduction, predictive analytics, automated remediation
- Experience defining product roadmaps, managing requirements and backlog priorities, delivering platform capabilities across lifecycle
- Bachelor’s degree in Computer Science, Engineering, Information Systems, or related field, or equivalent practical experience
- Machine Learning expertise in time-series forecasting, clustering, and correlation
- Generative AI experience with LLM prompting, Retrieval-Augmented Generation (RAG), and LangChain
- Cloud infrastructure knowledge of AWS, Azure, or GCP components
- Analytical mindset
- Strong problem-solving
- Prioritization skills
- Time-series forecasting, clustering, correlation algorithms
- LLM prompting, Retrieval-Augmented Generation (RAG), LangChain
- AWS, Azure, or GCP infrastructure knowledge
…
