This course explores modern approaches to learning and inference in computer vision and 3D scene understanding. Students will study four complementary paradigms that shape the field today: implicit neural representations (INRs) for encoding signals and geometry in continuous form; analysis by synthesis with differentiable rendering, illustrated through neural radiance fields (NeRF); amortized inference for fast 3D reconstruction, with recent feed-forward models such as DUSt3R and VGGT; and generative video models together with vision-language-action models (VLAs) for decision-making in dynamic environments.

The course combines theoretical foundations with hands-on experimentation. Teaching is organized around four lectures (3 hours each) introducing the core concepts and recent literature, two practical lab sessions (3 hours each) in which students implement and analyze key methods, and a final project (two 3-hour sessions) where small groups investigate a topic of their choice, building on the techniques covered in class. The project concludes with an oral presentation and a short written report.

By the end of the course, students will be familiar with state-of-the-art approaches to neural scene representation and generative perception, and will have gained practical experience implementing, training, and evaluating such models.