Overview
In this role you ensure the stability of mission-critical Palantir workflows and go on-call to resolve incidents before customers are affected. You translate field learnings into product changes, tools, and improved processes to raise reliability at scale. You actively automate manual tasks and advocate for product enhancements based on real-world challenges. You also document best practices to uplift the wider team and organization. This is a hands-on, impact-driven position with a strong emphasis on continuous improvement.
Responsibilities
- Develop a deep understanding of Palantir’s products and operational processes
- Go on-call, respond quickly to mission-critical incidents
- Diagnose, resolve, and proactively prevent issues in the field
- Collaborate with internal stakeholders to improve scalability and reliability of Foundry workflows
- Identify recurring pain points and automate or streamline workflows
- Advocate for product enhancements based on field insights
- Create clear, actionable documentation and share best practices to elevate reliability
Key requirements
- Ability to work independently and with others to solve ambiguous technical challenges
- Excellent written and verbal communication with technical and non-technical stakeholders
- Proficiency in Python, Java, and SQL
- Familiarity with parallel data processing and Spark optimisation
- Strong organizational skills and attention to detail with prioritisation
- Resourcefulness and creativity in fast-paced environments
- Experience with root cause analysis and documenting solutions for broader impact
- Enthusiasm for hands-on problem-solving, continuous improvement, and knowledge sharing
- communication
- collaboration
- problem-solving
- Python
- Java
- SQL
…
