Udemy
Computer Vision : OCR using Python - GenAI with LLM & RAG
Artificial Intelligence · Data Science · Development
Library / Artificial Intelligence
On Udemy
Learn how to build practical AI-powered Computer Vision applications using YOLO26, Gemini, and AI Vision Agents.
In this course, you’ll learn YOLO26 for object detection, segmentation, pose estimation, classification, and OBB, build AI Vision Agents with Gemini and YOLO, and create real-world applications for OCR, Document AI, invoice processing, and food & nutrition analysis.
The course focuses on practical projects that combine Computer Vision, Generative AI, and Vision Agents.
What You Will LearnYOLO26 & Computer VisionYOLOv8, YOLO11 & YOLO26Object Detection, Segmentation, Pose Estimation, Classification & OBBYOLO26 Custom Object Detection: Dataset Creation & Custom Model Training
How to Annotate / Label a Custom Dataset Using Roboflow
Train YOLO26 on a Custom Pothole Dataset for Pothole Detection
Building a Vehicle Intensity Heatmap from YOLO26 Detections
Real-Time Bird’s Eye View (BEV) System using YOLO26 and OpenCVAI Vision Agents for Computer Vision
Build a Minimal Vision Agent using LLM + YOLOObject Counting with YOLO and Streamlit
Autonomous Vision Agents with Gemini & YOLO
Build a Fully Autonomous AI Vision Agent with Gemini & YOLOAdd Segmentation and OCR tools Document & Invoice Processing
Build a Document & Invoice Processing Agent Extract and analyze information from documents
Build a Food & Nutrition AgentCombine Computer Vision with n
Ready to start? Continue on Udemy to enroll.
Start learning on Udemy (opens in a new tab)Prices, discounts and availability are set by Udemy. We may earn a commission when you purchase through links on this site.
0 courses
Udemy: 2026-09-27 · Coursera: 2026-09-27
Prices and discounts are shown on each provider's site.
Try fewer words or clear your filters.