|
所在平台: Udemy |
课程主页: https://www.udemy.com/course/ai-testing-deepeval-ragas-ollama/
课程评论:没有评论
课程名称:使用Ollama测试AI与大型语言模型(LLM)应用程序:DeepEval、RAGAs及更多 课程概述: 本课程旨在教授如何测试和评估AI应用程序,特别是大型语言模型(LLM)。通过实操练习,您将掌握关键技能,帮助质量保证(QA)人员、AI QA、开发人员、数据科学家及AI从业者使用最前沿的技术来评估AI性能、识别偏见,并确保应用程序的稳健开发。 课程涵盖主题: 第一部分:AI应用程序测试基础 - LLM测试简介 - AI应用类型 - 评估指标 - LLM评估库 第二部分:使用Ollama部署本地LLM - 本地LLM部署 - AI模型 - 本地运行LLM - Ollama实现与GUI/CLI - 将Ollama设置为API 第三部分:环境设置 - 使用Jupyter Notebook进行测试 - 设置Confident AI 第四部分:DeepEval基础 - 传统LLM测试 - DeepEval代码实例:AnswerRelevance与Context Precision - 在Confident AI中评估 - 本地LLM测试 - 理解LLMTestCases和Goldens 第五部分:高级LLM评估 - 使用LangChain进行LLM评估 - 评估答案相关性与上下文精准度 - 偏见检测 - 使用GEval建立自定义标准 - 高级偏见测试 第六部分:使用DeepEval进行RAG测试 - RAG简介 - 理解RAG应用程序 - 演示与创建GEval - 测试简洁性与完整性 第七部分:高级RAG测试 - 创建多组测试数据 - 在Confident AI中的Goldens - 实际输出与检索上下文 - 从数据集中生成LLMTestCases - 进行RAG评估 第八部分:测试AI代理和工具调用 - 理解AI代理 - 与代理合作 - 在无实际系统的情况下测试代理 - 使用多个数据集进行测试 第九部分:使用RAGAS评估LLMs - RAGAS简介 - 上下文回忆 - 噪音敏感性 - MultiTurnSample - 概述和有害性的通用指标 第十部分:使用RAGAS测试RAG应用程序 - 引言与设置 - 创建检索器和向量存储 - RAG的MultiTurnSample数据集 - 使用RAGAS评估RAG 此课程为希望深入了解AI与LLM测试的从业人员提供了全面的学习体验。
Testing AI & LLM App with DeepEval, RAGAs & more using Ollama and Local Large Language Models (LLMs)Master the essential skills for testing and evaluating AI applications, particularly Large Language Models (LLMs). This hands-on course equips QA, AI QA, Developers, data scientists, and AI practitioners with cutting-edge techniques to assess AI performance, identify biases, and ensure robust application development.Topics Covered:Section 1: Foundations of AI Application Testing (Introduction to LLM testing, AI application types, evaluation metrics, LLM evaluation libraries).Section 2: Local LLM Deployment with Ollama (Local LLM deployment, AI models, running LLMs locally, Ollama implementation, GUI/CLI, setting up Ollama as API).Section 3: Environment Setup (Jupyter Notebook for tests, setting up Confident AI).Section 4: DeepEval Basics (Traditional LLM testing, first DeepEval code for AnswerRelevance, Context Precision, evaluating in Confident AI, testing with local LLM, understanding LLMTestCases and Goldens).Section 5: Advanced LLM Evaluation (LangChain for LLMs, evaluating Answer Relevancy, Context Precision, bias detection, custom criteria with GEval, advanced bias testing).Section 6: RAG Testing with DeepEval (Introduction to RAG, understanding RAG apps, demo, creating GEval for RAG, testing for conciseness & completeness).Section 7: Advanced RAG Testing with DeepEval (Creating multiple test data, Goldens in Confident AI, actual output and retrieval context, LLMTestCases from dataset, running evaluation for RAG).Section 8: Testing AI Agents and Tool Callings (Understanding AI Agents, working with agents, testing agents with and without actual systems, testing with multiple datasets).Section 9: Evaluating LLMs using RAGAS (Introduction to RAGAS, Context Recall, Noise Sensitivity, MultiTurnSample, general purpose metrics for summaries and harmfulness).Section 10: Testing RAG applications with RAGAS (Introduction and setup, creating retrievers and vector stores, MultiTurnSample dataset for RAG, evaluating RAG with RAGAS).