Overview
In this role you will design and iterate on a next-generation video generation foundation, collaborating with Lightspeed Tech Center to push state-of-the-art across games. You will tackle long-context modeling, efficient tokenization, and scalable pre-training at massive scales. You’ll drive research, share findings through papers, and influence the technical roadmap for AI-enabled video creation. This is a mission-driven opportunity to shape cutting-edge video tech for multi-device gaming at scale.
Responsibilities
- Design and iterate on the next-generation video generation foundation architecture
- Explore core technologies for long video generation, including long-context attention, KV Cache compression, and memory mechanisms
- Research high-compression-ratio video tokenizers and unified modeling for multi-resolution/multi-frame-rate videos
- Lead video pre-training at large token scales, defining data mixtures and curriculum learning strategies
- Explore Scaling Laws and scientific scale-up paths, improving model capabilities
- Track industry state-of-the-art, run comparative experiments, guide roadmap decisions, and publish academic papers
Key requirements
- Ph.D. in AI-related fields with first-author papers at top-tier conferences
- Proficiency in diffusion models and autoregressive generation principles and engineering
- Experience training video/image generation models from scratch; strong PyTorch and large-scale distributed training skills
- Deep understanding of Video VAE/Tokenizers design trade-offs; 3+ years of relevant research experience
- Publications related to video generation or diffusion models
- Experience in core industry product R&D or hands-on work with latest technologies is preferred
- Experience leading end-to-end video foundation model projects is preferred
- leadership
- collaboration
- communication
- diffusion models
- autoregressive generation
- PyTorch
…
