Concevoir des modèles multimodaux

Build Multimodal Models

NEOPOLIS AKADEMY

Build Multimodal Models

Practical exploration of multimodal models for vision, audio, video and image generation. Focuses on integrating multimodal data flows, compatible architectures, open-source libraries and concrete steps to assemble models and pipelines.

View this course on Akademy ↗

Enrolment and practical details are available on Neopolis Akademy.

Build Multimodal Models

What you will explore

Survey of multimodal model types and concrete application scenarios.

Processing and fusion of vision, audio and video signals to feed unified models.

Techniques for conditional image generation and output post-processing.

Pipeline architecture: preprocessing, multimodal encoding, fusion and decoding.

Examples of practical tools and components for experimentation and prototyping.

STEP BY STEP

Course programme

01Build Multimodal Models

Build multimodal models by leveraging available models and datasets, preprocessing steps per modality, and approaches for classification, VQA and image editing. The module covers vision, audio, text and multimodal generation.

Accessing resources: how to navigate models and datasets, identify popular text‑to‑image models and preprocess multimodal data.

Unimodal models: practical approaches for vision, audio and text tasks such as image classification, object detection and background removal.

Multimodal classification: zero‑shot image classification with CLIP, automatic caption quality evaluation and multimodal sentiment analysis.

Multimodal generation: Visual Question Answering using ViLT and document VQA with LayoutLM, plus image editing using diffusion models.

Explore this module on Akademy ↗

Programme source: Neopolis Akademy. Original course page