Overview
In this role you lead the development and initial architecture of scalable distributed systems, focusing on elasticity, high throughput, and fault tolerance. You define scalability requirements, drive data-path optimization, and ensure reliability through redundancy and automated failover. You establish telemetry, dashboards, and SLO-aligned practices, while embedding security, IaC, and change-management into operations. You’ll mentor peers and contribute to production readiness in a leading AI and cloud solutions environment.
Pay / Benefits
- flexible medical
- life insurance
- retirement options
- volunteer programs
- disability accommodation support
Responsibilities
- Lead development and initial architecture of scalable distributed systems with horizontal/vertical scaling.
- Optimize code and data paths for high-throughput, hyper-scale workloads.
- Define scalability requirements and ensure design meets elasticity and performance targets.
- Design fault-tolerant systems with redundancy, replication, and automatic failover.
- Handle network partitions and unreliability with load-shedding, throttling, and rate-limiting.
- Define KPIs, telemetry, dashboards, and alerts; design validation, replication, and synchronization for correctness and durability.
- Diagnose production issues, mentor peers, and ensure operational readiness.
- Implement robust security controls, remediation, and multi-tenant compliance documentation.
- Develop automation and IaC to manage cloud infrastructure and automate patching/updating/rollback within change-management plans.
Key requirements
- collaboration and cross-functional partnership
- problem solving and critical thinking
- mentorship and knowledge sharing
- system design and architecture for distributed systems
- scalability engineering and elasticity
- high-throughput data processing
…
