Overview
In this role, you will own the North agent harness—the execution layer that makes agents reliable and production-ready. You’ll drive the agent loop, context management, and tooling orchestration to support long, multi-step workflows in enterprise settings. You’ll collaborate closely with engineering and the Modeling team to validate capabilities and evolve the harness in step with model progress. You’ll translate real-world enterprise feedback into concrete product and model requirements, shaping a scalable, secure agent platform for impactful automation.
Pay / Benefits
- health and dental benefits
- 6 weeks vacation
- remote-friendly with global office presence and stipends
- parential leave top-up
- co-working stipend
- weekly lunch stipend and snacks
Responsibilities
- Define and own the North harness roadmap across agent loop, context engineering, tool orchestration, sandbox execution, and sub-agent delegation
- Serve as primary interface between North engineering and Cohere’s Modeling team to validate new harness capabilities before build
- Own North’s agentic evaluation framework ensuring compatibility with training infrastructure and serving as a bridge between product and research
- Engage enterprise customers to surface agentic failures and translate findings into product and model requirements
- Stay current with open-source and commercial agent ecosystems to guide adoption and architecture alignment
Key requirements
- 5+ years of product management experience in agentic AI systems, developer infrastructure, or applied ML products
- Deep understanding of modern LLM agent architectures (multi-agent systems, tool-augmented reasoning, memory and retrieval, programmatic orchestration, RAG, long-horizon execution)
- Strong grasp of agentic evaluation design (measuring task completion, failure recovery, long-horizon reliability) and diagnosing model vs. scaffolding gaps
- Technical depth to contribute to architecture decisions at implementation level (design docs, async execution, filesystem design, sandboxed environments)
- Ability to switch between ML research discussions and engineering architecture conversations
- Track record of shipping platform-layer products with measurable impact on reliability, performance, or capability
- Cross-functional collaboration
- Strategic thinking and roadmap ownership
- Effective communication with engineering and research teams
- LLM agent architectures
- Multi-agent systems and tool-augmented reasoning
- Memory and retrieval mechanisms
…
