Site Reliability Engineer 2

Company: Oracle Corporation
Apply for the Site Reliability Engineer 2
Location:
Job Description:

Overview

In this role you will help Oracle Analytics remain reliable and scalable by diagnosing issues across distributed systems, building automation, and supporting 24/7 operations. You will partner with analytics developers and other teams to reduce downtime and improve incident response. You’ll shape runbooks, dashboards, and AI-assisted tooling to enable proactive, data-driven cloud operations. This is a hands-on SRE role in a fast-paced, customer-focused environment.

Pay / Benefits

  • competitive benefits
  • flexible medical
  • life insurance
  • retirement options
  • volunteer programs
  • accessibility accommodations

Responsibilities

  • Provide 24×7 operational support for Oracle Analytics services and release cycles
  • Respond to incidents, diagnose complex production issues, and drive mitigation and post-incident actions
  • Read and troubleshoot existing application/service code to identify root causes and safe remediation
  • Develop deep product expertise to prevent regressions and reduce recurring incidents
  • Build and maintain operational tooling, dashboards, monitoring, runbooks, and knowledge-base content
  • Develop AI-assisted automation tools to improve incident investigation, monitoring, and reporting
  • Create scripts/services for monitoring, telemetry collection, capacity analysis, patching, and remediation
  • Analyze service health, workloads, and capacity trends to identify reliability risks
  • Improve CI/CD processes, deployment practices, and operational readiness
  • Collaborate with Development, Support, Product Management and other teams on issue resolution
  • Share operational knowledge and continuously improve team processes
  • Follow security, compliance, change-management, and operational procedures

Key requirements

  • BS or MS in Computer Science, Engineering, or equivalent practical experience
  • Experience supporting cloud infrastructure and operational processes
  • Strong understanding of networking basics (DNS, HTTP/HTTPS, TLS, load balancing)
  • Linux/Unix administration experience
  • Experience with cloud services and large-scale distributed applications in production
  • Ability to troubleshoot complex issues by reading existing code and systems
  • Experience documenting runbooks and operational guides
  • Experience working in agile environments
  • Strong written and verbal communication, including remote team collaboration
  • Ability to work independently and participate in on-call and after-hours support
  • Strong communication skills
  • Collaborative mindset with remote teams
  • Analytical and methodical problem-solving approach
  • Networking and TCP/IP fundamentals
  • Linux/Unix system administration
  • Python, Bash, JavaScript/Node.js

…

Posted: October 1st, 2026