Abstract: In this technical talk, Paras and Ajay Jain, brothers and co-founders of AI video generation research lab Genmo, will pull back the curtain on the current landscape of AI video generation technology.
This talk is a unique opportunity for developers to learn about diffusion model architecture from one of its co-inventors: Ajay, co-author of the DDPM paper. Today, almost all major image and video generation models are based on Ajay's work. Ajay and his co-researchers were also the first to develop a text-to-3D model, “DreamFusion” capable of producing high-fidelity content.
Paras and Ajay's session will provide an in-depth exploration of the methodologies used and challenges faced while building Genmo's newest model, launching this September. The challenges they’ll discuss solutions for include:
Optimizing GPU infrastructure to improve latency speeds and scale with growing user demand
Solving the ‘incoherence’ challenge of long-context video generation with prompt design and AI video workflows
Paras and Ajay will also discuss methods for increasing AI-generated video quality. For example, improving frames-per-second speed above the Hollywood standard of 24 FPS, preserving motion quality and boosting photorealism.
The speakers will also share samples of video generation models in action to highlight the current landscape of model capabilities and deficiencies.
Paras and Ajay will conclude with a glimpse into future developments in video generation. Model builders will leave this session with a deep understanding of the technical intricacies involved with building state-of-the-art video generation models, the challenges that still need solving, and where the next generation of innovators can begin.
Bio: Ajay Jain is the CTO and co-founder of Genmo, a research lab building frontier models for video generation. Prior to founding Genmo, Ajay earned his Ph.D. at UC Berkeley, where his research pioneered the application of diffusion models for image generation and 3D modeling. Ajay is the co-inventor of modern diffusion architecture with the DDPM paper. He and co-researchers also developed the first text-to-3D model capable of producing high-fidelity content ("DreamFusion," 2022). Ajay's research won the Outstanding Paper Award at ICLR 2023 and his work was also presented at the 2023 Davos Summit.



