Building Video Understanding Agents and models that can reason, interact and act on long form videos
Video is the richest and least searchable format of information we produce. Feeding a 2-3-hour video directly to an LLM is impractical, as the cost and latency make it a non-starter.
This talk covers building an agent-powered system that ingests long-form video and makes it fully queryable like a long form text document.
The talk also covers my journey of building video understanding models and interaction models for live video.
About the speaker
Logesh Kumar Umapathi
Machine learning Engineer at Blackbox.ai
Logesh is a Machine Learning Engineer at “Blackbox.ai” (https://www.blackbox.ai/), where his work focuses on building agentic systems and models that automate software development and improve developer productivity. His interests include code-generation LLMs, synthetic data generation with LLMs, and aligning code LLMs with human preferences. He has contributed to notable Code LLM research and open-source projects, including “StarCoder” (https://arxiv.org/abs/2305.06161), “SantaCoder” (https://arxiv.org/abs/2301.03988), and the “BigCode Evaluation Harness” (https://github.com/bigcode-project/bigcode-evaluation-harness). He was also recognized with the AV Luminary Award for Best AI Scientist by Analytics Vidhya. Previously, he was a Lead Machine Learning Engineer at “Saama Technologies” (https://www.saama.com/), where he led machine learning efforts for a product focused on accelerating clinical trials and reducing the time to market for new drugs. Outside of his work, Logesh enjoys speaking at machine learning events, reading books, and photography.