Vollständiges Transkript anzeigen (5.262 Wörter)
Hi everyone. My name is Nick, and I'm here with Akshay to give a talk about evals. We are from we're from Lyft, and we've been building Lyft customer support AI agent for a year two now, and gave a lot a lot of thoughts about how to build eval that actually matters and scale our AI agent multi AI multi AI agent system. Uh, just a little bit of quick introductions. My name is Nick. I'm a data science manager. I've been at live for six years, a long time lifter. Uh, really excited to talk to you a little bit more about evals. I'll hand it to Akshay. Hi everyone. Uh, I'm Akshay. I'm in Nick's team and we've been working together uh on customer support agents for Lyft. Improving the harness, improving the evals, things like that. And I've been at Lyft for almost four years now, and I'm very excited to be here and talk about uh building evals that actually matter. Super excited to be here, and super honored to be, to be on the online track for AI Engineer Warfare. Uh, yeah, and let's dive in. For the agenda of today, uh our talk will primarily focus on eval. We will start by sharing how we think about the end-to-end pipeline for our evaluation for building support, customer support AI agent system. Uh, we'll go into deep dive into each component, uh more deeply as we go. Uh, we'll start by talking about offline evaluations, online evaluations, eval harness, as well as what we are planning to build going forward. Alright, let's dive in. Um, I want to quickly explain the, you know, high-level system of how we think about evaluation system for AI agents. Uh, so here you can see we have a development phase and the production phase. Uh, so during development, if you're building agents, you should be very familiar with, uh, you know, managing context, building web pipeline to give your agents educational contents, uh, defining your tool, building your agentic graph, as well as writing a system prompt. Uh, so once all of that engine, agent engineering process is done, you have an, you have an AI agent. Uh, the way we think about this is before we launch this AI agents to productions, we want to go through a rigorous offline evaluation process to make sure that this agent, uh, actually has sufficient performance before we launch this to a live, uh, users. So coming from, you know, a data science and machine learning background, uh, we've been building model machine learning model for, for, for, for a while. And I think the way that we think about agent development is very similar to building machine learning models as well. If we are running, uh, offline evaluations for our machine learning model before that goes to production, I think we should do the same for AI agents, as well as any other agentic platform, uh, agentic applications. Uh, but what I think, you know, offline evaluation, how that, how that is different than, uh, traditional machine learning model is that, you know, we typically we are building specifically for customer support AI use case, we'll building an agent that's multi-turn. Uh, so for offline evaluation, there will be a component of simulated conversations. So you want, you typically want to have a datasets, synthetic data set that's representative of your production traffic, uh, have a user LLM that plays, plays out the complete multi-turn simulated conversation and as well as having a grader, such as LLM as a judge, uh, to be able to evaluate how good that interactions was. And then we have a launch gate, right? We want to make sure that we have certain criteria on our offline eval and we're meeting that criteria before we decide to launch this AI agent to productions. And so the, the real imperative here really is that we don't want to use our live user as, you know, test data for our AI agents. And I think in any, any cases that, that is not, not good practice, so we really want to emphasize the importance of having an offline evaluation process. So once the agent is in production, we also have an online evaluation pipeline as well. We have our own favorite tracing tools to, uh, trace all the execution and context that the AI agent used to respond to a real user in in production's environment. We have our online grader, as well, that grades how well our AI agent is doing in productions, as well as having a human in the loop pipeline to do error analysis, identify failure mode and feed that, that insights to the development teams to continuously improve our AI agents. I want to quickly go over I think three of the most common reasons why we think, uh, evaluations typically fail in, uh, for for different teams. So the first reason is that, you know, the the grader that we create, the scores that we create needs to be meaningfully gating something. Uh, this is what we really emphasize on in the previous slide that we need to have a launch gate. If your LLM as a judge is just floating out there, there there's a score but no one is really using that score as a meaningful gate, uh, for your development and productions environment, then that LLM as a judge is not, not valuable. Uh, we have also seen a lot of a lot of mishap people have when they are creating their LLM as a judge. Uh, there's a lot of different opinion out there on in terms of how do you create uh a good LLM as a judge. And typically, uh and and unfortunately also very early on in our journey, the LLM is, the LLM judge that we created are, um, very noisy, too generic. Uh, it will output a score and but people don't really believe in, uh, in what the LLM judge is doing, or they don't think the LLM judge insight is actually actionable. And finally, I think when something regresses in productions, we need to have clear mechanism to be able to to to catch that regression as well as identify clear owners to be able to take actions on, uh, the insights of our graders and uh regression gates. Very cool. I want to sequence into talking about our offline evaluation system. And as we as I touched on early on, I think this is the most critical piece uh going from development cycle to to productions. We really want to have a robust offline evaluation system to be able to get more confidence in uh in the AI agents that we are shipping to productions. We took a lot of inspiration from, um, this paper called Tallybench, uh, which is developed by the wonderful people at Sierra AI, uh, and this is specifically for customer support AI agent, but we also think this is applicable for any user-facing agentic applications, uh, agentic applications. Uh, so yes, you can see this is a offline, offline simulations where you have the AI agent as well as a user LLM that are interacting with each other to produce the multi-turn traces, uh, you have agent domain policy, which is essentially what the, you know, for each customer support use case there is instructions policy on how to handle different a different customer support issues. So, taking inspiration from Talibench, this is sort of a high-level approach that we have in creating our offline simulator. Uh, as we mentioned earlier, we have our Langgraph agents that we built and we have defined, uh, an instructions for our our user LLM to run. So for simulation lies, we define the user intent, we define, you know, what this intent is supposed to to represent and we define the user data point or the world state of the of our user. For example, you know, the driver that are coming to us might be a luxury driver that might been driving for us for for a couple years and so and so forth. And finally, to define the user, user behavior or user personas as we would typically see with our real, uh, you know, real Lyft user as well. We also created these different persona for our user. One example here is that this can be a loyal long-time Lyft customer but they are frustrated with uh with Lyft earning systems. So in our offline simulator, we have this Langgraph agent that are, you know, interacting with our user LLM model and generating this multi-turn trajectories uh, of uh, of this multi-turn agentic uh, agentic trajectory, and we also build our offline, offline grader or offline evaluator. Uh, LLM judge is a big component of that. We'll direct dive a lot more deeper into how to build a grade, LLM as a judge. Uh, and apart from, you know, LLM as a judge, we also have more deterministic, uh, evaluator as well, and these usually looks like a code assertion as you see in traditional unit test. Uh, and for example here, some some of the deterministic criteria that we have created so far, uh, looks something like this, right? If, you know, if in this in this specific interactions, the AI agent is supposed to grant concession, we will write like a rules like this whether or not the AI agent has indeed grant concessions and we will compare the agent agent tool calls with the expected outcome to make sure that uh, we have we can measure the accuracy of the agent instruction following behavior. So, one of the big gotcha that we faced with running our offline evaluation is the creating synthetic data is is one of the key challenges in making sure our offline data set is representative of our production's data. So, uh, here we can, you know, um, here we have a meme here. What ideally what you don't want to be doing is just to prompt an LLM model to generate 50 different test query for your offline data sets. Uh, so here's a couple suggestions where you can approach this, you know, much more much more realistically and assemble a data set that that closely resemble to your production data. So, this is also something that we we try to do for customer support AI agent as well. We take uh, we take some sample from our production data. Uh, so a Lyft user has been reaching out to support for a long time, we take a sample of our real production example and supplement our offline data set with that. And you and the second thing that we can we we do is we mutate, uh, different criteria from our offline data sets to be able to cover different golden path and edge cases. So, one problem with our offline evaluator is that as as we mentioned in the previous slide, if you were simply uh sampling different 50 50 test use test, uh, test query from uh by using a LLM model, uh, another thing an another big gotcha that we we face in in building this offline simulator was that we were using a frontier lab LLM, uh, model to roleplay Lyft user in our offline evaluations. Uh, and for the most part, you know, our frontier lab model are trained to be helpful, Assistant, rather than a uh Lyft user that might not, uh that might not sound as nice as, as always. So, in our first pass at running our offline evaluation, what we noticed is that our LLM user sounds almost too nice. And as you can see on, on the right here, these verbatims are very, very very, very complete. The user, the LLM user are very patiently explaining the issues that they are facing in productions and our first attempt at our offline evaluation gave us 90 plus pass rate or accuracy rate, right. Uh, this almost sounds too good to be true and I think it indeed is the, too good to be true. So in reality, these are the real user verbatims that we get in productions. Uh, as you can see here, most user, they are they're impatient, they're already frustrated. So the verbatims they, they they don't want to explain their issues like a AL LLM user will. Uh so typically what we see in productions is is something like this, uh and in in some in reality these makes, you know, AI agents much more difficult to evaluate. So, what can we do to improve our LLM user LLM? Uh, simulate this real life user much closely, right? Uh so what we, what we, how we approach this is we fine-tuned a LLM model with Lyft user verbatim. So instead of speaking like this very verbose and very nicely impatiently explaining their issues, uh our LLM user will produce verbatim that's resemble this much closely. And the benefit of this is while, you know, we did see our evaluation score goes down after we fine-tune a LLM model that speaks more like our Lyft user and therefore making our evaluation more difficult, uh, but in reality this is really what you want uh when you're building this user simulator, right. If you have an eval that's too easy, that doesn't give you any real, uh production insights into how your AI agent is actually going to perform. Uh and this also give you a lot more room to be able to tweak your AI agents to deal with, quote unquote, difficult user. And as I as I gave a little bit of sneak peak earlier as well, we also take a lot of work to define, Lyft user personas, uh, this can really ground our LLM user to adopt a specific Lyft user persona, and therefore, uh simulate our real life user much more closely. Uh, a couple of different user persona that we define here, uh are, you know, Bypasser, user who just want to escalate to agent, regardless of and not giving AI a chance, Refund seeker, uh AI skeptics. Alright, I'll hand it off to Akshay to talk more about our LLM judge. All right, uh, hi everyone. So I'm gonna take from here. I'll talk about second problem, which we usually face when we are doing evals with LLM as a judge. And here you can see, uh, this is how pretty much everyone is using LLM as a judge to evaluate their agents. The focus can be slightly different based on the use case. Someone can focus more on safety, someone can focus more on cost and latency, someone can focus more on quality, but people are going to measure these sort of metrics more or less. So uh we want to detect leaks, we want to detect safety issues, um and and things like that. So some parts are deterministic, which can be evaluated by code, and some are not, which are evaluated by the LLM. Now, the problem with this approach is uh that these metrics are too generic and not actionable. So for example, we also started with our eval journey using pre-built metrics from DeepEval framework, which were uh which were measuring tool usage appropriateness, response helpfulness, conversation naturalness, completeness, and things like that. And we did see those metrics. But the problem was these metrics were not actionable. They were not giving us any actionable insights. If, let's say, response helpfulness is 0.5, then what do we do with it? So, uh things like and other other scores like toxicity score, bias, fairness, conciseness, all these are kind of relevant, but if the metrics are just uh scores, we we don't know what to do with them. So we can use these pre-built eval metrics as a baseline, but we shouldn't use them as our core eval metrics, because we want eval metrics to be actionable and tied to the business outcome or the product which we are focusing on. Uh, we want to define, okay so, what LLM as a judge should be. Uh, we collaborate very closely with domain experts and utilize their insights. So evals should be framed around a task success or failure. And a binary outcome is very easy to calibrate and train um LLM judge that can consistently score your agentic trajectory. When we partner with domain experts and data scientists to build metrics which are actionable and aligned with business goals, we can see much more meaningful results, uh, and actionable insights. Not only this is more consistent, uh, but when an agent fails an interaction, we can systematically analyze the error pattern and have and actually know what what we can do to fix this. Here is an example of an uh actionable metric for our use case, uh which is called education rubric. Here the we defined the metric how it should be and we define success and fail criteria for LLM judge. So for example, the AI agent tries too many times to educate the user, uh if it could have uh escalated the the issue. Or if it's escalating too soon without giving any chance to educate the user. So things like that comes under failure category, so we we mark this as fail, uh but if it's an expected behavior, we mark it as as pass. Uh, okay, so with this, how do we validate if our uh LLM judge is working as expected? First thing is we need to treat it as a classifier. So how we train our classification models in machine, traditional machine learning, we can also treat our evaluation judges as as those traditional ML classifiers with binary outputs. Once we have binary outputs for every metric uh or a task based on our business goals and functional requirements, we can hand label uh around 100 examples with pass fail labels and then split the data into uh train, dev and validation sets like how we used to do with machine learning models. Uh, then we score precision and recall uh for our judge based on human labeled ground truths, which will give us an actual report on how good our uh judge is performing. Uh, and how do we split our data to do to calculate precision and recall is this. So, we split the data similar to how we used to do in training machine learning models, but the difference is the percentage of splits uh and here we are not actually training model weights, so we are just using the data to inform judges prompt. So the percentages are a little bit different. For example, in training data, we we pick a few short examples from that from this set uh for the judges prompt. And then we iterate the prompt against the dev set uh and improve our harness or our prompt, and then finally we validate against the test set uh to see that we didn't overfit on the dev examples. This is the practical uh way of like splitting the data and then calculating precision and recall scores for your judge to actually know that the judge is working as as expected. All right, the next light. Okay, so this is another thing which uh most of us ignore when we are doing evaluations, uh which is criteria drift and validating the validators. The key idea is that we actually discover what our evaluation criteria is by looking at the data and grading our outputs. Uh our and our sense of quality will also evolve evolve with new data we see and more uh examples we we grade. So, the evaluation should not be decoupled from model observations. In fact, they should be developed uh they should be co-developed with the model when we are testing the evaluator and calculating the precision and recall scores. So, uh the the there is always a gap when we, when we talk about LLM as a judge. We cannot define the criteria beforehand and then evaluate agents against them. Uh our criteria should also be evolving as and when we see more examples, and then we should refine our metrics uh for our judges and then evaluate judges on top of those metrics. All right, then next step. One of the things which which can make the numbers we are reporting more meaningful is to add some statistical rigor to it, right? Uh we report alignment rates as bare point estimates. So, if we add confidence intervals, we do some calibration and proper sampling, uh the same numbers can become more meaningful. We should definitely reserve the expensive rigor for the moments, uh number actually gate something, or like like a shipping decision or we are reporting numbers to uh company leaders. But depending on the use case, we should definitely have uh confidence intervals for the for the numbers we report. Because every score needs an interval. This is just a small example to give you uh an insight on what this actually means. Uh let's say we have two evaluators and one scores 84 percent and one and the other scores 88 percent and the number of samples we number of traces which we have uh used is let's say 50. So, to show that like this is this is a very small gain and we we need much more than 50 examples to actually uh show that this gain gain is real. Uh with just 50 examples and only 4 percentage point percentage gains, uh gain, we we don't actually know if this gain is real or not. So bigger gains and bigger gains and paired designs need far less uh and we can reserve the rigor statistical rigor for for things which matter the most. Okay, so this is a non-exhaustive list of observed eval anti-patterns. I'm not going to read all of them, but uh these definitely contain some low-hanging fruits. We need to put in time and effort if we if we need meaningful evals. So we cannot rely on LLMs for everything yet. And what I can say is ignoring the data is one of the most important uh things which we shouldn't do, uh and which we sometimes don't focus on due to lack of time or resourcing, but it acts as the foundation for meaningful eval evaluations. If you don't look at the data, you won't be able to create meaningful criteria, uh or labels. And if you don't have labels, you won't be able to evaluate your judges, and if you are not evaluating your judges, you don't know if your agentic pipeline is working as is expected. So this acts as a base, so uh one of the most important things we should not ignore. Okay, uh we we said that we want to make metrics more actionable and standardize the pipeline, but how actually we should do it. So this gives you like a template to to do an error analysis loop, and it is important to know that this loop is something which runs continuously. It's not uh not an one-off audit. Uh, so we deep dive into raw traces. So once we have logged our traces for our agentic flows, multi-agent systems or whatever we have in our uh use case, then we pinpoint failure modes. So basically we try to identify what exactly is failing. And we only keep the metrics that change a decision. We remove all the noise, we only prioritize on the metrics which are tied to business use cases, functional, uh requirements, and actually something which changes a decision. Then we form a fresh premise to re-evaluate. And then we repeat it. So we can have like a regular cadence of doing this pipeline. It can be weekly, it can be bi-weekly, but this is something which needs to run uh continuously, and it's it's not a one-off audit. Okay, so Tracing. Tracing as we said is uh one of the most important things and everything kind of depends on it. For diving deep into raw traces, we definitely need to log them first. So we can use tools like LangSmith, LangFuse, etc. to log and view the traces. Uh and here each trace captures the full graph execution, which nodes ran, what LLM saw, which tools were called, uh what was the token usage, what was the latency for every call and things like that. We can also enrich traces with metadata if we want which gives you more insights than uh than the actual data. You can also end this traces with metadata if we want which gives you more insight than uh than the actual data. You can also end this traces with metadata if we want, which gives you more insight than uh than the actual data. Okay, and we also have annotation queues. Annotation queues are nothing but an interface, which is very helpful for domain experts to label or give feedback to evaluators and in an easy to understand UI. So they don't have to look at the raw traces, uh JSONs and stuff like that to figure out what to focus on. They can use this annotation queue and they can give feedback uh or label examples easily. We can then add these traces to datasets for offline evaluation, or we can use this for uh calculating our precision and recall for our judges. So this forms the basis to validate the evaluation with ground truth labels. Okay, I now I'll hand it over to Nick to to close the evaluation loop. Thank you Akshay. Uh and I think the the goal of having eval is to be able to feed uh our evaluation insights back into improving the model performance or the agent's performance. So here, uh I would introduced a couple way that we think about continual learning for our AI agent and closing the evaluation loop. Um and and here you have model learning, context learning, and harness learning. A model learning is really about, you know, post-training, updating an underlying model weights, and training a custom LLM model. Uh, context learning and harness learning is really more about improving uh everything else other than the model. Uh, so context is improving what information the agent actually sees, this can be document, this can be uh the stored memories of the user, tool outputs and so and so forth. Harnest, uh that means updating the model system prompt, and tool schemas, control flow, routing, retries, and and so and so forth. So, I think from, you know, the error analysis that Akshay has shared earlier, I think that has really helped us to understand how can we improve uh our agent prompt, as well as updating, uh updating our knowledge base and tune the context management strategy that we have for our AI agent, and the identify failure mode really help us uh be able to feed that insights into actual improvement for the agent AI agents. So, I want to quickly talk about what's next for us, uh in our journey of building customer support AI agent at Lyft. Uh we have a we we have a we have some level of ability to run our offline simulator, but in fact, I think, you know, it's not repeatable. These are currently stored as scattered script across different notebooks and different analysis repo. Uh, I think one thing that we are looking really, really looking into investing is a systematic eval harness, uh, and then having a a a harness system that can help run our offline evaluation in in a systematic and standardized manner and allowed a different people to contribute to uh our evaluation suite with very, we with we with pre-defined uh primitive and config based uh workflow. And another thing that we've been thinking a lot about is uh post-training. Uh as we mentioned earlier, you know, identify model failure mode has really helped us tune the agent context and as well as the agent harness. Uh but over the years, we have gathered a lot of real user signals on our agent performance as well. Uh so we we are really starting to think about how do we fine-tune uh model that does different tasks for our customer support AI agent, uh as well as framing a reward modeling problem for to to enable reinforcement learning. I want to quickly share a little bit about the work that we're doing with eval harness and how we think about uh building uh building, building eval harness for uh agent facing user-facing agentic applications. Uh so again, eval is really that scaffolding that you that we need to be able to run eval efficiently in a standardized format across uh all the different agents and sub-agents that we have for customer support AI agents. And and as you can see from my offline simulator slide earlier, uh our eval harness is config-driven, and these are typically store as YAML file that's easily editable uh by different contributors and not just by engineers. Uh analysts and data scientists can can contribute to this evaluation suite as well. And you know, with thousands if not tens of thousands of examples, uh in our evaluation suite, we need, you know, parallelisms and throughput to be able to run our offline evaluation in a reasonable amount of times. And to enable uh user to be able to different user to be able to contribute to our eval evaluation suite, we also define, you know, primitive around our eval eval harness. Uh these are high-level things like tasks, data sets, personas, LLM adapter, and evaluator. And after all, you know, the benefit of having a eval harness is that you can define the config ones and run these eval indefinitely many times across different at different touch points or different gates of your agent development process. This can be, you know, locally when you're developing this agent, uh at any point when you tune a prompt, you can run the evaluation suite and get a immediate in in, uh immediate feedback on how your agent is doing compared to the previous versions. We can run these at pre-commit cook, uh to make sure that our performance doesn't degrade before we push a change uh to our agent service. Uh another area that we are looking at is also at CI/CD, and how we can use our eval harness to uh build our regression test suite, uh acceptance acceptance test suite as well. So, that wraps up our presentation today. Uh we've gone through a lot of a lot of a lot of different topics uh for evaluations. Uh really, I think this is an end-to-end journey uh for building a evaluation pipeline that works for customer support AI agent or any user-facing agentic applications. Uh, Akshay and I, we are very interested to hear about, you know, what you all have been working on and share any learnings that you have for eval. So, feel free to contact us, uh if you have any questions or you just want to share or brainstorm about how to improve your evaluation system. All right, thank you and I hope you all enjoyed the AI engineer Warfare. Thank you. Thanks all.