Overview
In this role you will help optimize OCI’s critical components and internal tools as part of the Cloud Performance Organization. You’ll design and build scalable, high-throughput cloud services from the ground up, balancing performance with reliability. You will work closely with cross-functional teams to set SLOs, implement observability, and drive improvements that reduce cloud costs and improve customer experience. This greenfield opportunity offers autonomy, mentorship, and a chance to influence how Oracle scales its AI-enabled cloud solutions. You will be part of a dynamic, inclusive culture that values bold ideas and practical impact.
Pay / Benefits
- flexible medical coverage
- life insurance
- retirement options
- volunteer programs
- inclusive culture
- accommodations for disabilities
Responsibilities
- Lead development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads
- Define scalability requirements, optimize performance, and leverage distributed state management and data-plane platforms
- Design fault-tolerant, highly available systems using redundancy, replication, failover, load shedding, throttling, and rate limiting
- Establish SLOs, KPIs, telemetry, dashboards, and alerts to ensure reliability and performance
- Design performance, load, fault-injection, and brownout testing, and implement replication and synchronization for correctness and availability
- Proactively diagnose production issues, guide incident response and root cause analysis, and ensure operational readiness
- Enable in-service maintenance and upgrades with minimal customer impact
- Mentor engineers in troubleshooting and operational practices
- Implement encryption, access controls, and security remediation for multi-tenant environments
- Ensure compliance with applicable standards and maintain required documentation
- Develop and maintain IaC and automation for cloud infrastructure
- Enable safe and repeatable patching, updates, and rollbacks through effective change-management practices
- Plan and manage moderately complex initiatives with technical oversight
- Collaborate across teams to align objectives and deliver solutions
- Promote inclusive collaboration and diverse perspectives
- Analyze and resolve moderately complex issues with clear recommendations
- Document and share effective problem-solving practices
- Stay current with industry trends and continuously develop technical skills
- Coach and mentor junior engineers and promote knowledge sharing
- Identify and implement improvements to processes, workflows, and team effectiveness
- Support talent development through candidate interviews and hiring recommendations
Key requirements
- collaborative mindset
- mentorship
- problem-solving
- distributed systems design
- scalability and performance optimization
- observability and telemetry
…
