NEOPOLIS AKADEMY
Multi-Modal Systems with the OpenAI API
Design of multimodal systems using the OpenAI API: audio processing, moderation and customer-support flows. The course describes multimodal ingestion, transcription, automated moderation and integration into support pipelines.
View this course on Akademy ↗Enrolment and practical details are available on Neopolis Akademy.

What you will explore
Outlines architectures for ingesting audio and other media and extracting usable text or metadata.
Explains transcription workflows, cleaning and time-alignment for audio recordings.
Introduces automated moderation approaches and content filtering suitable for customer flows.
Shows how to combine transcripts and models to orchestrate responses and routing in support systems.
STEP BY STEP
Course programme
01Multi-Modal Systems with the OpenAI API
Explores multimodal features of the OpenAI API: speech processing, transcription and translation. The course demonstrates creating podcast and call transcriptions, handling non‑English languages, and translating transcriptions into English or Portuguese.
Covers multimodal use cases of the OpenAI API focused on speech‑to‑text and translation.
Provides concrete workflows: transcribing a podcast, processing customer calls and detecting speaker language.
Describes how to translate transcriptions (e.g. Portuguese → English) for integration into customer‑support workflows.
Explore this module on Akademy ↗Programme source: Neopolis Akademy. Original course page