Library / Artificial Intelligence

Modern Computer Vision AI with Vision Transformers and LLMs

On Udemy

About this course

Modern Computer Vision AI with Vision Transformers and LLMsModern Computer Vision has evolved far beyond traditional image classification. Today, Vision AI systems can recognize objects, detect multiple objects, segment images, understand natural language, reason about visual scenes, analyze videos, generate realistic images, and intelligently edit existing images. These capabilities are transforming industries such as healthcare, autonomous driving, robotics, manufacturing, retail, surveillance, and smart automation.

This course provides a structured learning journey through the world of Modern Computer Vision, Vision Transformers, Vision-Language Models, and Large Language Models (LLMs). Rather than learning individual models in isolation, you'll understand how they connect together to build intelligent Vision AI systems capable of solving real-world problems.

Your learning journey includes:

Building a strong foundation in Modern Computer Vision and Vision AIUnderstanding how Computer Vision has evolved from CNNs to Vision Transformers and Foundation ModelsLearning the core Computer Vision tasks:

Image Recognition

Object Detection

Image Segmentation

Vision-Language ModelsVision Reasoning

Depth & Pose Estimation

Video Intelligence

Image Generation

Image EditingExploring state-of-the-art Vision AI models, including:

ResNet-50Vision Transformer (ViT)YOLODETRDINOGrounding DINOSegment Anything Model (SAM)CLIPBLIPTimeSformerDiffusion ModelsBuilding practical Vision AI applications using Python, PyTorch, Hugging Face Transformers, OpenCV, Ultr

Ready to start? Continue on Udemy to enroll.

Start learning on Udemy (opens in a new tab)

Prices, discounts and availability are set by Udemy. We may earn a commission when you purchase through links on this site.