Library / Artificial Intelligence

Testing LLMs : With DeepEval, Promptfoo, RAG & CI/CD

On Udemy

About this course

AI/LLM Testing Mastery: From DeepEval to Production CI/CDTraditional tests can be green while your chatbot is still wrong, biased, or jailbroken. This course teaches you how to test LLM applications the way production teams actually have to: with metrics, judges, adversarial checks, and quality gates.

You will learn to evaluate correctness, relevance, faithfulness, hallucination, toxicity, bias, and security. You will test RAG pipelines and AI agents, run prompt-injection and jailbreak cases, and wire the whole suite into GitHub Actions so a bad answer can block a deploy.

We use DeepEval and Promptfoo, local models with Ollama (no API key required to learn), and finish with a portfolio-ready capstone: a support chatbot, an evaluation suite, and a CI pipeline.

You will learn how to:

  • Explain why unit, integration, and E2E tests are not enough for generative AIScore free-form answers with LLM-as-judge instead of exact string matches
  • Test RAG end to end — retrieval quality and grounded generation
  • Test tool selection, multi-step agents, and memory
  • Red-team chatbots for injection, jailbreaks, and data leaks
  • Put evaluations in CI/CD with thresholds, quality gates, and cost control
  • Debug a failing eval: was the chatbot wrong, or was the judge wrong?
  • Includes hands-on labs, homework after each module, and a GitHub project.

Ready to start? Continue on Udemy to enroll.

Start learning on Udemy (opens in a new tab)

Prices, discounts and availability are set by Udemy. We may earn a commission when you purchase through links on this site.