Fondamentaux des grands modèles de langage

Large Language Model Foundations

NEOPOLIS AKADEMY

Large Language Model Foundations

Exploration of large language model fundamentals: Transformer architecture, tokenization, dataset preparation and training workflows. A technical approach focused on understanding components and training stages.

View this course on Akademy ↗

Enrolment and practical details are available on Neopolis Akademy.

Large Language Model Foundations

What you will explore

Transformer architecture: attention mechanisms, feed-forward layers and optimization principles.

Tokenizers and sequence representation: splitting methods, encoding and handling context-length limits.

Datasets: collection, cleaning, sampling and preparation for supervised or staged training.

Training workflows: fine-tuning pipelines, evaluation, hyperparameter management and training monitoring.

STEP BY STEP

Course programme

01Large Language Model Foundations

Foundations of LLMs and the ecosystem: NLP/Transformers concepts, tokenizers, pipelines, fine‑tuning, datasets, the Hub, Gradio/Argilla tools, training loops, evaluation and RL methods for LLMs. Also covers debugging and model sharing.

Introduces NLP and Transformer basics, how they work and how pipelines solve tasks.

Covers tokenizers, sequence handling, data processing, training loops, fine‑tuning (Trainer API, SFTTrainer, LoRA) and learning curve analysis.

Presents using the Hub for models and datasets, dataset creation/annotation, debugging training pipelines, Gradio for demos and advanced RL techniques applied to LLMs.

Explore this module on Akademy ↗

Programme source: Neopolis Akademy. Original course page