Gain trust in your Coded Agent before moving to production with UiPath

GeeksforGeeks · Intermediate ·🤖 AI Agents & Automation ·7mo ago

Key Takeaways

Demonstrates how to gain trust in a coded agent using UiPath before moving to production

Full Transcript

Okay, we're live. Yeah. So, hello everyone. Welcome back to the finale of this trilogy of workshop along with UiPath. I am Samund Ratan Jen, your host, the director of data science at Geeks for Geeks. In today's session, we will look at how the world is moving towards agentic AI future and how UiPath platform enables AI agents, robots, people and even models to work together harmoniously to revolutionize industries and enhance human potential. Remember, AI agents are autonomous software that can perceive environments in real time with objectives and execute complex multi-step task with minimal human intervention. They are delivering adaptable scalable workflows. UiPath is delivering adaptable scalable workflows that drive efficiency, decision making and innovation. And now for the final time, Geeks for Geeks is proud to host Isubu Jakhan, Adrien Thomas and Radu Mukau. Isu is a product builder and engineer at heart with a simple belief that our purpose is to improve people's lives. Bu he builds meaningful products and he thinks that is one way of doing that. is now working on AI and agents to deliver on that mandate at scale. Adrien Adrien is a principal product manager at UiPath I think. Okay. He is he focuses on building reliable intelligent agentic systems that empower businesses to accelerate productivity with confidence. Radu is a software engineer at UiPath focused on designing and orchestring reliable high impact systems. As one of the main contributors to UiPath's opensource libraries for coded agents, he is deeply committed to improving the developer experience, making it easier and more intuitive to build AI powered solutions. Now on to the main now onto the main people in this particular session. Yeah. Cool. Thank you for uh the introduction indeed. This is the third and final workshop on how to build and deploy trust decoded agents with UiPath. So in the previous sessions we went over the fundamentals. We touched on what UiPath stands for. Uh what is the UiPath mandate and where agents and in particular code agents or agentic workflows partic in particular fit in in the bigger picture and we highly recommend you go over the previous workshops and at the beginning of each section you will see an introduction for UiPath and where the code agents fit into the bigger picture. We will skip that for this session. because we want to allocate more time on uh for the demo and for demonstrating the latest and greatest and the trust principles. Now uh be besides that uh we also touched on how one can use coded agent frameworks like llama index and launch chain and integrate them natively as first class citizens within uh UiPath and deploy them to the UiPath platform to be orchestrated and securely governed and we even went beyond uh beyond that. In the second workshop, we took you through the typical iterative uh emergent if I may journey of building and evolving a coded agent to map a real use case scenario. We started with a simple agentic workflow addressing a very simple maybe risk-free scenario and we evolved that to a more complex agentic workflow with multiple steps multiple flows uh large functionality coverage blend of deterministic and undeterministic checks and tools and we end up with an escalation mechanism. Now the UiPath mandate as we mentioned previously is to provide the best-in-class platform and tools for agentic uh automation which sounds uh nice uh aspirational even but let me give you the let's say the true meaning of it. what is the real embodiment of that statement of that mand mandate and how does the as we say rubber meet the road >> [clears throat] >> uh for that uh and well it's clearly not by uh chasing the hype and not chasing any buzzwords and not chasing the next shiny thing around the corner is a totally different approach. It's about being ground grounded in knowing where the true value is and where the true challenges uh really are and I know to gain um uh and to to go straight to tackle them. And this workshop is definitely all all about that is uh about tackling the biggest challenges. And today uh we will tackle the challenge of trusting the agent before going to uh production because developing a functionally ready uh aentic automation that aligns with its intended uh logic is clearly commendable. It's nice maybe not even easy to do yet it represents only the midpoint only maybe a fraction of uh the entire path to production. The remaining journey lies in earning the trust in agentic workflow. Uh because trust is and will remain the currency of production. So in what follows we will tell the same story. Um uh and you might already know a part of that story and the story is about how one can cut through the noise and go beyond the hype and join the as we say the few the 5% the 1% of those who deploy dependable agents into production with undisputed value. So trusting these agents making critical decision over real uh data. So it's the same story. We will tell the same story but now we will open a new chapter of the story. Let me u sh pinpointing exactly what we will uh be showcasing today. uh as topics we will go uh with evolves for coded agents the studio app and solution integration and we'll also touch on local dev traces and debug um experience now as we introduce the topic still uh the reason for UiPath coded evolves is that with this uh great power which is agentic workflows comes great responsibility so we'll have to have an equal force of the same nature to control and to gain trust in the agents. So we move from uh simple agents react agents one LM calling tools in a loop up until it reaches a conclusion a resolution and it returns a result to the new world which is the one we evolve in and not only us but in general where we'll have multiple steps or multiple flows. So that's an agentic workflow. So as code agents are know supercharged with a fine grain level of control and flexibility that helps us push the boundaries of what an agent can do. The same way eval for code agent are supercharged with a fine grain control and flexibility to match that force. This is the the principle here. The way we will showcase it today will be via um real use case one that we not only imagine. We've seen this and we've seen this from multiple um customers having the same kind of problem and taking the same approach. We will take a snippet of that and show it to you and we will apply the events over them to showcase why they are important in a meaningful uh manner. So uh without further ado, the the key that we want to show is a logistic agentic workflow. And I know that this image might scare some of you like it's cumbersome. It's complicated. But fear not, we will break this down into simple steps and it will very easy to comprehend. But let's go through through the workflow and let me uh show what the workflow would actually look like. And remember whenever we're talking about monogentic workflow the final ultimate goal is to have one that is uh autonomous meaning that it can take the steps from start to finish without human intervention unless it's highly necessary or it's as a form of escalation. So imagine that we have know a sample email like this. It's an unstructured input saying that we want to arrange transport transportation for uh some components. These are the shipment details pretty much [clears throat] standard but the way they we receive them is not standardized. The type of data might be but not its structure. And this would be a typical uh email requesting from our company which is a logistic company to make them an offer and to confirm with that uh close the email part. Uh if you start the flow the next phase would be a validation phase. We want to see if the customer requesting it is actually an existing customer. What is the history of the customer? Do we have it the database? What is the history of the payments with that customer? Do you want to encourage them to make other orders or maybe even have blacklisted customer that should not receive a code for some reason? So this is a very important validation phase which is some a part we do validate but imagine at this step we can add other validations as well. Next one would be to understand the content to summarize it to label it in a certain flow. In this case, the type is for instance new order. What rather will showcase is that we can have multiple pets. It can be a new order. It can be a complaint. It can be requesting some details which is not directly related to an order or even a general category which is not yet handled with that. The next part would be to calculate the discount. So we do have the list price but for every customer we calculate a discount which is based on seasonality might be based on uh history that we have with that customer and other factors and we will have to feed in these details to the agent such that based on multiple factors it will calculate a proper uh discount and then calculate the base cost that we have and uh calculate the final offer. The cost will be based on uh distance will be based on u the the travel is it internal or external how many days the weight and so on. It's a lot of factor that are considered here in order to calculate the cost. This is an agentic work to combine to aggregate the data from multiple sources and to have a final offer. Uh this is a human review which is an form of escalation. Meaning that if the confidence of the agent is below below a certain threshold that the offer is a coherent compelling one then instead of discarding the entire flow we can keep all the details that we've gathered up until this point give them to someone in the company to review maybe adjust some of it and then continue with the work. For instance, if the confidence here is below 80%, it would say this is the offer. Do you agree? Uh do another check. And let's say [clears throat] that the human approved the flows. So reuses all the steps. Uh the agent continues with that approval com composes the response. uh so uh PDF with a breakdown of whatever is needed a code and an email that is nicely formulated and very importantly of of course it will send the confirmation to the client but not this for be to be a meaningful agent it should be integrated into the company's uh business uh process which is uh at the last step step is to start some internal processes uh to actually make the uh this offer uh a concrete one within the company and uh schedule uh the transport, schedule the drivers uh and all the necessary steps that are needed in order to actually um act on uh on this offer and provide this the final service to the customer. [clears throat] Now going back to how this would look like uh in code um and I will just go through a flow uh we have multiple ones but for instance the the flow that you've seen is the same here with the mermaid diagram where we have the input this one which is the um I can zoom in a bit is the email structure one the next one will be to validate the company as you've seen we had a validation step and say that we do have a company in our database case and it's a valid one after that will be understand uh the request in that case that I've shown you was a new order so we'll go through this path to calculate the discount based on the history with that uh customer and other factors and then we'll go through to handle the order it can be a restriction so we can have some exceptions here but let's say that we are on a typical path where we do calculate the final uh order and we assemble the cost and we can propose a final code to the one that was requesting it. Um, and for some reason, let's say that we have the confidence below the score that we wanted, which should be a rare event, but it's important to have it. the human would approve it and then we will finalize the input and create the processes that follow uh the the request which is our internal processes to actually provide the service. So this would be the context and now as we will go through we will see how we can gain trust in this agent. So functionally this is ready. You have the pad that we need and now the question that we you will see again and again and again is how can we trust each step and each flow and how and what UiPath provides as tools and frameworks to give you uh the confidence in this agent to go into production and to trust them with real data which is codes and uh with real decisions along the way. So with that I will hand this to uh Radu who is the I would say uh engineering driving force uh behind uh evaluators. So I'm absolutely glad and uh that we can hear directly from the engineer that are actually building this and are working this day in and day out at UiPath. So with that Radu please. >> Okay thanks for setting up the scene. So as you said the code is ready the implementation is done we know it works on the happy pads we'll also test this now and some key aspects about this agent it's fairly complex it uses a data fabric for data storage two indexes for performing a rag a queue for audit and for the processing an escalation app and even an MCP server and all of those components are available through the open source SDKs okay we won't go into details here because we don't want to waste time on this. So, uh let's start and test this agent. Make sure it does what it's supposed to do. The fourth step is to authenticate. We do that using the CLI, the UiPad CLI, and we authenticate against the staging environment for this demo. Okay, we wait for it to to face the authentication details and the authentication was successful. As with any UiPath agent, you need to initialize it. This command basically is taking care of everything. It creates configuration files. It creates bindings. We'll touch that later. So, basically everything needed for this agent to run both locally in cloud and for evaluations to run both locally and incloud is handled pretty much by this command. We can see here it created a bunch of files. It created the mermaid diagram, the one that I was showed before. So this is not the mermaid diagram is not created by the developers themselves. We just uh analyze the code and basically this is created out of the box. So you have a visual way of representing this agent where the nodes in here are actually functions into the Python code. Great. So let's continue. The development is done. Let's test this. We [snorts] have this very nice uh dev terminal. To use it, you just need to add the UiPad dev dependency. It's already installed on my machine. And the next step is to run the UiPad dev C cli command. What this command is doing, it's opening an interactive terminal application. And we can use this application to run the agent, see realtime logs, see the traces, see all the steps it took. So this is very very useful in the development process. This is how it looks like. It it uh inferred from the code that as input it needs an email address, an email content and optionally a confidence threshold. The default value it's 80. We can play with that. I have here uh a sample input and this is the input order. It's the one that talked about. So basically we receive an email from from a company and we should validate this company. The domain it's correct and we have a contract with this company and then this is the email content. They want to uh arrange transportation from their warehouse in London to their warehouse in Oslo. And we have a bunch of of data here. Okay. To run this, we just copy paste this as JSON format. I have it here. And we can copy paste it into the dev technaliz input. And click on run. Now real time in the logs panel. We are going to see the logs of this agent. What it does. We will see that it successfully identified the company using the domain. It's the Nordic wholesale group. It fetched the ID from data fabric. Using this ID, it's looking for the previous orders and it decided to apply a 5% discount. And we can also look at the justification and because they had four or in the last seven days, they are qualified for a 5% discount. This all was handled by the agent using uh UPET data fabric. Also in the traces panel, we can see traces as they come. So those are the steps that the agent took. Even though in the logs we can print here just what we want like API calls. This comes out of the box or certain certain audit stuff. In the traces panel we can see the entire agent flow every step it took. We can see that it forex went through the understand request node. It understood that this is a new order. Then it went to the calculate discount node. Then it went to the handle new order. And in here we can see every tool that the agent invoked. It validated the shipment capacity. It used the shipment retriever tool. In here is the company data, the actual stuff it uses to make sure it adheres to the company policies. And at the very end, it finished. And in the details panel, voila, we can see the output. No error. And we have this email. The discount was applied. This is the estimated price based on the agent findings. Okay. So we covered the handle uh or scenario. There's one more I want to handle until we move on to eval. We also have a brand new CLI command which is the debug command. Let's let's have a quick look at it. I maybe will zoom in a bit. Okay. So this CLI command you just say UIP debug and then this is the name of the agent. You can provide an input file and an output file if you don't write to want to write it in the terminal. Uh this let's run this command and as you can see now the input is a complaint one. We are not on the same workflow anymore and the agent should take different approaches. Let's have a look at the input complaint. So this is the input complaint. So they are very very uh unhappy with our services because the delivery was late because they had a huge loss because of this. And let's see what our agent is doing. Now remember we ran the debug commandment not the run one we showed before and this means that we can place dynamic break points on this agent locally. Let's have a look at the diagram once again and see where we would like to place a break point. So we go to the autogenerated agent mark made. Let's view it as a beautiful chart. Let me close some of the windows. Okay. And remember this is a complaint email, right? So it should route the request to this branch and at some point we'll have the collect output node. Okay. To place a break point, it's very easy. We just write to have to write B from break point and then write the node name. So in this case would be collect output break point set and now we can write the agent. Let's run it. And here we'll see step by step what is the current agent state. So at each node we see okay uh this is let's see this is the validate company node. We see what the agent did. It validated the company. It extracted the company name and the company ID. Okay, here I invoked an AI model. We can see the prompts it gave it. We can see the tools. You can see everything and request. It classified this as uh complaint. So everything is working as expected. But there's one thing I want you to pay attention to. Oh, and also we hit a break point. Now the execution stopped and we can resume it. But there's one thing as input. We gave this time a 99 confidence threshold. This confidence threshold as you said before basically notifies the agent that if its confidence is below this number, it should escalate this to a human. It should not take any action. Let a human review this. And this is what makes agents reliable. So the agent gave it a pretty pretty good confidence score which is still below 99. And let's see what's happening. Let's continue because we hit the break point and now if you look at the logs we see the job got suspended and a polling mechanism started. Why is that? Well, as I said the confidence was not good enough and the agent actually created a human in the loop task for someone to approve it. And the main uh takeaway here is that you can do this fully from your uh local environment. Now we navigate to the UIP action center. If we if we create this, we should see a pending task here. It was created 25 seconds ago and it's already assigned to me because I instructed the agent to do that. And uh we can either approve or reject it. Here the agent uh pasted information that is uh important for me. So they said okay this is a complaint. This is the original email and here is the proposed email response and I can have a look at it and I can either approve it or decre if if everything looks good, I approve it and the workflow goes ahead. But let's let's imagine something is not right. So we we analyze this, we look over our business case and we we figure out well the refund is not big enough. This is a valuable customer. We want to refund them more. So we can give feedback, realtime feedback to the agent. I'll say no, please refound uh this amount, not this one. And I'm going to reject. Now if we go back to the development ID without doing anything, you'll see that the execution resumes because the task was completed. Now the action the agent can resume. So this way you can really test real case scenario without uh leaving your ID or leaving it just to uh approve or deny human escalations. Now we can see here in the state that it got our input. we didn't approve it and we wanted it to take another route and we hit a break point again because now the agent reanalyzes what it should do based on our uh based on our input and it created an action again. This means I can go back to action center. I should have another pending task was created 10 seconds ago and let's see if the agent did what was expected from him. Let's see if it understood that we wanted to change something. And if we have a look here. Okay. Now, it decided to refound the number that I gave it. This is the power that human need provides. Okay. So, everything works as expected. I can approve this. Everything is fine now. And if we look back, okay, the execution finishes. So, let's recap. Where are we? We developed this pretty complex agent. Everything is working fine but uh I would say can we deploy to production just yet? What do you think? >> Uh I would say absolutely no. Wait. So is the confidence of not deploying to production now would be 100%. So [laughter] let's change that uh from 100% no to production to 100% go to production. But before that since you mentioned the dev experience >> maybe it's it's worth considering here something. Let's imagine that things didn't work as you expected. You've went through the human in the loop. You said change from uh this amount of refund to another one and at the second run. Let's say imagine that it didn't change the sum. How could you use the dev experience to investigate where the problem is? >> Well, in the dev experience, I have basically all the traces or all the logs. I can see step by step why the LLM decided to do that or not do what was expected from him. As I s as I showed you, we also have the prompts that are passed into the LLMs. This complex workflow is using three agents. That means three LLMs with tools and I can see what's happening and then and then correct that. But again, this is happening only on the happy path. Those are the scenarios I tested. We are not covering all the scenarios just running the agent locally. This is where eval come into play. >> Correct. And you'll see that at least for this I said they would have this bug on on a path that we know what is the expected behavior and we can go we can have a complimentary debug experience and dev experience where you can see full traces identify at which step it took the the wrong turn and maybe change it. So we are we should be in a state before moving to evolve we should be a state where we are satisfied with the functionality and [clears throat] once we are satisfied with its functionality once everything looks good once that email looks like okay it's perfect job done it's weekend now we are halfway there and let's see what are the next steps >> yeah exactly so to run the UiPath evaluations for coded agents you can do them both locally and studio web for the sake of this demo we'll show both approaches. So at first let's create a new project in studio web. What is studio web? Studio web it's our cloud native web ID. I have uh where is it? I have here a solution that I created before logistics agent uh gigs for gigs in it. I have an IPA workflow. I can have anything. But I want to create a new project. For this I click on the plus button. I select agent. Let's say it's logistics uh gigs for gigs. Let's create it. And now I have to choose what kind of agent this will be. Uh let's wait for a second for it to load. Okay. So I have these three options and I want to publish a coded agent. Let's choose the coded agent. And it's very very simple to publish an agent to studio web. You just need to copy this environment variable and then run the UiPad push command. That's it. Let's do that. What this what this command is doing it's first of all let's add this to the uh M file and then run UiPad push. This command is not only pushing the source code to studio web but also the evals. So if you prefer a more deed approach, you don't want to leave your favorite ID that's perfectly fine. You can work with evals, develop them locally, run them locally, but also if you are in a collaborative environment or you want to use the the more beautiful studio canvas, you can also do that. So as you can see, it's uploading the files that this agent is required to run and also it's uploading some evaluation files that we created uh before and I'll show you those. And it's it's doing some work stuff. It's importing some resources. We'll also touch that a bit later. Okay, let's go back to Studio Web and see what it looks like. And this is it. We push the code. Here on the main page, we see the entire mermaidmate. We also see some metadata like the publisher name, when it was last pushed, and was the current project version. Now, in here inside the source code, I have all the files needed for this agent to run. So, once a project is pushed into studio web, it get fully run on cloud. we don't need if we don't want we don't need to touch the the local ID. Okay. And in here you'll also notice that I have this breakpoints list. So I can place break points and debug this fully in UiPad cloud. Let's place a break point. Let's say I want to have a break point for the understand request note. We can see on the main diagram that it's seted a break point there. And then we can debug this agent clicking here. Now we uh the the canvas uh identified that this is the input of our agent. And here we provide the input as you do locally. But there's another very cool thing I'll show you. For let me just copy paste. Let's test with the same order which is the most most complex workflow. So this is the sender. This is the email content. We leave the confidence threshold as uh let's say 80. And we have this environment variables uh field. Basically, we're not pushing from your local environment any secrets, any environment variables for obvious reasons. But the agent might need environment variables for it to work. And in this case uh our agent is built in such a way that if you provide it with an environment variable that's specifying an MCP server URL it will consume also an MCP server and uses tools and let's do that. I'll click on save on debug and now the execution is starting in cloud. to see the actual job that's executing. We can go to the orchestrator and in our personal workspace, we'll see that a serverless job started. That's the job that's running our agent and it will it's reporting to to studio web. So in my workspace, if I go to automations jobs, I should see here a debugging job in progress. This one started 13 seconds ago. Let's check it. This is the input I gave it here is the same the experience that we have in the dev terminal. We can see the logs. We can see traces. And we see that okay, we hit a break point. Let's head back to studio web. The execution is paused here because we hit a break point. We can also see traces here. So we hit a break point after the validate company note. Okay, let's continue with the execution. And the execution resumes. It should go till the end because I don't have any any more break points and the confidence threshold is low enough. But what I wanted to show you here is that remember we set the MCP server URL environment variable that MCP server it's placed inside the shed folder and it's an MCP server fully available and uh run inside UIPad. It's a Google maps MCP. We can see it is running because the agent requested it and it's very easy to set it up. It takes like five seconds taps I'd say. Let's see how this was configured. We just said, "Okay, we want you to run this command with the server that's already available from Google." And then we had to provide optionally for the server a Google maps API key that we wrote as an secret asset. And that's it. Once you click create, once you create the MCP server, it is available and pluggable into any coded agent using the MCP protocol. And of course you can see traces for the MCP server as well. If we click on this MCP server and we look at the traces panel you can see that our agent forex listed the tools to see what tools it can use from this Google maps MCP server and then it decided to use the maps distance metrics because it's it wanted to find out how long would it take to drive from London to Oslo to deliver the package. If we have a look at the input we see that okay it called this tool with source as London with destination as Oslo and then the tool said okay it takes approximately 21 hours and this uh will uh will be understanded by the agent and the execution finished it was successful we can see the output here everything works as expected okay so we debug gau agent even in in studio web. Let's get now to the fun part. Uh let me refphrase this project briefly and let's get to the evals. So beforeoting it to production as you said we need to make sure the agent is reliable. We need to gain confidence. We can't we can't promote it just yet because it works with some simple data. We need to make sure that no matter what the case is, our agent will will work as expected. Okay. So on the left panel you'll see here evaluation set and evaluators. If we go to the evaluators we see we already have some evaluators. These are the one I created locally and pushed to studio web. But let's see what evaluators we have available in in iPad. So basically you have 11 evaluators available and basically those are uh should work with any coded agent that you have out of the box. We have from output validators to LLM based evaluators to to call validators. Let's take for example a LLM based one to create a validator is very very easy. You just select the desired one. You have some configuration to do like what model should this LLM use. This LLM is looking at the evaluator trajectory and then right here okay so this agent should do those stuff and this evaluator will evaluate if that was the case or not. Since you mentioned the Google Maps um MCP, which is by the way a crazy thing if you think the alternative to that is to have a bespoke system that would implement those APIs and would get the data or even worse a [clears throat] human actually in going to the web and inputting the uh start and the destination and retrieving the the final uh distance. So given that part of our agent is that we have an external tool code that we did not program. We don't have access to the implementation of it. We just enter that very brief command to bring those packages MCP packages within UiPath and host them within UiPath. So then out of these 11 evaluators can we evaluate that even if we have external tool cost and what would be generally speaking a good evaluator for for that case. >> Yeah. So we do have exactly this we have tool call validation. So those evaluators here are evaluating all the tool the agent is calling and not only that it evaluates the that they are providing the right arguments to the tools because a tool may be incredible but if your agent is hallucinating it and then it's not calling it the right way. It doesn't matter. It's evaluating how many times it calls any tool. It's evaluating in which order it's calling the tools. We also have this evaluator for this example. and it's evaluating even the output of that tool. >> Okay, so we we can gain trust in an agent even if it uses tools that are not within UiPad and [clears throat] are external and you have a bunch of ways to tackle them and to see if the proper tool was called with the proper arguments and if the output was as as expected. >> Yeah. Yeah. Exactly. That's very well said. Okay. And now let's go. We had a look at the evaluators. Let's go to the evaluation set. I created this before. So we have three evaluations in this evaluation set. Let's take a look at the first one. We can run them in cloud or globally. We can run one at a time. All doesn't matter. Let's have a look at the first one. So this is an evaluation here. It has a name and an input. we say okay for this input which is the one we tested so far uh we are applying those evaluators here you see that we don't need to apply all evaluators we just need to configure those that make sense and we have two evaluators for this data point we have a LLM judge output this LLM power evaluator will look at the expected uh output which is this one and we'll look at the actual output and provide a score how much uh how close those two those two strings are and also we have a detonistic trajectory evaluator which is actually a tool order. So as you as we said before uh actually this is this is a very key point. It's not enough for an evaluator to evaluate the output. You can build an amazing agent. You can build an output evaluator and that output may actually be correct but if it came like that for the wrong reasons is not good enough. We need to make sure the output is okay but also the agent took the right steps to achieve that output. And let's see here we have those three tool calls. We say okay first of all I want the agent to call the understand request tool to understand the request and how to route it. Then this is very important. I want the shipment retriever tool which is an ECS index. It's the the company data that we have for delivery. I want this tool to be called before the valid achievement capacity tool. Why? Because the agent may look at the email content and say okay actually I'll I'll explain this while I I run this this uh evaluation. So why this is very important. Okay evaluation is is running. We can see in here in the runs tab that it's in progress and to monitor the evaluation job. We can see uh here in orchestrator in personal workspace a job that uh started not this one sorry let me refresh this this one studio debugging and here we can see the command it ran so is the eval command. What does that mean that we can also run this command locally but let me get back to the tool caller. So we need to make sure for this agent that it first calls the tool that fetches company data and after only then validates the shipment category because the agent may look at the email and say okay they want to transport 3,000 kilograms. This is for sure a high capacity but what if that's not it shouldn't be for the agent to decide should be our own company policy data. So this is why we needed to call that tool before sorry after it fetches the data from our company. So the run finished. Let's refresh it. And something is not good. As you can see we get an 85 score for the output evaluator which is a pretty good score but we get a zero score for the evaluator to lower that. So we are looking here at a false positive. The output is correct but not for the for the correct reasons. Let's have a look at it. Why was this uh 0% score? Well, it was because we expected to call understand request she retriever and then validate. But no, it understood the request and then called the valid sheet for so the agent the LM actually hallucinated and categorized this capacity as it wanted. it doesn't it was not grounded into the context we gave it right then that's not good and this way we can catch bugs like that that can have major major implications in production >> yeah and I think this is a beautiful showcase on why when we mention and we do have some sobering news and sobering posts around the internet like only 10% of the project provide on the value only 5% of them are trusted in production and this is a beautiful example on why that happens because everything looks good from even from a human perspective. The final output is beautiful. Uh the response would be 10 out of 10 but it misses entirely and hence it cannot be trusted and without proper evalance as Raj said you cannot catch this. So one important aspect here is that these are we do see them because we uh uh we designed in such a way to emphasize them but in a normal typical workflow you would not have the full view of this and you would not be able to catch this at design debug time. [clears throat] This is beyond that and hence the importance and the difference that this evolves make in delivering on the overall agentic promise. >> Yeah. Okay. So we we identified the problem, right? It's not taking the steps you should take. Let's head back to the dev diagonal. And uh here we we can see the problem. It has this tool that's defined like this in code. And it said okay if validation fails then you should stop processing immediately. The problem with this is that the agent hallucinates and okay it it believes that this tool should be called forex but it shouldn't. So we can actually uh play a little prompt engineering here and I'll I'll use this little script to replace this prompt here with with this one in our prompt.py file. Let's do that. Okay, so we updated the new prompt. Now we clearly state that validation human capacity should be called only after we get output from the shipment retrieval tool as I I discussed before. And also let's do something different this time. Let's not push the code again and run the evals in cloud. Let's run them locally and see what's happen. So to run evaluations, it's that easy. You just can have this command uipal the name of the agent and then output file that's that's optional if you want to see the output of all evaluators now this command here will run all the evaluators and we'll see there it's running the agent is running everything is starting but we can't really get feedback from that what's happening well actually we can do that because this run even though it's local it's reporting real time to studio If we go to the runs tab, we can see run is in progress is this one. This this right here is actually the local run. So even though I'm developing locally for better visibility, for better tracing, I also have the this studio web interface that I can I can use because those evaluations are as I said real time reporting what they did to studio web. So we can see now what happened. the agent uh the the first eval was completed and let's see the score this time. The LLM judge got a point8 as the first time but the tool color got a perfect score. This is because we changed the prompt and now everything works as expected. The dev terinal here I can also see the the logs for this specific run. And if I go to to studio web again let's click refresh. This was the eval run new successful. Let's see. Evaluator tool got a perfect score this time because we changed the prompt and now is the correct time. Cool. So this is a very very nice example of how we can combine output evaluators with trajectory evaluators to to reach that confidence threshold that you need to promote this agent to production. But that's not all. We also have two more uh use cases that are very nice. Okay. So view let me ask you something here. So this agent when I describe to you what this agent does and based on the evaluators that we have available here how would you evaluate the discount part this discount logic because that's a custom uh calculation based on the business case of my company. me as a company, I decide if I want to apply a 10% discount, a 7% or more. What evaluator I can use from here to to make sure the discount is uh correctly calculated by the agent. >> Oh, it's custom logic, right? So there will not no [clears throat] out ofthebox evaluator would be able to provide that. No matter how large is the menu of evaluators that you have, there will be no one that that will be specific for my use case. So I would guess none of them. >> Yeah, that's correct. But that's that's also okay because we do have custom evaluator text. So basically we we give you out of the box 11 evaluators, but if you want to implement some custom logic, you can totally do that. Those are fully customizable. Let's see how you can do that. So to add a new evaluator for discount in this case, you have two CLI commands. You have the UiPad add evaluator command and you give it a name. This will create a a template for the evaluator. Then you go ahead and implement the code. And after that you have the register evaluator. This is analyzing the evaluator you created, extracting the tile from there. So we make sure this evaluator will work with UIP coded agents both locally and in cloud. And let's see how this evaluator looks like that I implemented before. You can find it here under evaluations evaluators custom and we have discount.py. Okay, this is how the evaluator looks like. It's it's highly customizable. You have a lot of a lot of stuff to configure here. But the most important part is that you need to implement this method. That's it. You just need to implement this method and based on the input you get you you you get to decide what the output should be and we basically give you here everything you give we give you the agent input the agent output the trace so you can do whatever evaluation you want using the agent trajectory the agent behavior and simulation instructions if applicable. So let's quickly have a look at this agent what it does. It checks in all the traces for the one span we care about which is the calculate discount one and then we get the discount percentage that the agent gave uh to this order and then we get how many orders we had in the last seven days from the previous nodes and then we have this this logic here which is very simple as I showed you in the prompt. If you have more than seven or decks, it's 15. If not 10, three, five, or nine. And that's pretty much it. This is a custom evaluator. And this evaluator works for my specific use case. >> Now, a deterministic oneistic evaluator for the discount or is it a blend of deterministic and undeterministic? >> Well, mine is a deterministic one because I I can't really play with this. I can let an agent hallucinate or even an LLM that's evaluating another LLM hallucinate and say okay it's fine for you to give a 99% discount that shouldn't be the case so in my personal case for a logistics company this will be a deterministic evaluator >> but that's actually a good point because here it's code I can write anything I can have a mix of deterministic LLMs anything I want to do to evaluate this agent and decide what score should I give it >> yeah because as we discussed with customers and partners and as they develop use case. One frequent question that surfaces is how can an LLM judge judge an LM? Wouldn't that just continue with their issue? then the the judge would not judge correctly. And this is a again a beautiful example on how we can with UiPath eval blend a deterministic evaluators with in this case a custom deterministic one for uh something that is heavily high risk right the discount can uh end up in in something that we heavily influences the outcome the business outcome of this agent in the whole picture. >> Yeah. Yeah, that's a very good point. Well, in some scenarios, it may be beneficial to have this scenario, an LLM judging another LLM because the temperature is different, the model, the prompts, the tools they may use, that's fine. But yeah, in those critical cases, we we should rely on custom custom evaluators or deterministic ones even if we can reuse something. So I while while we are explaining our CBI this eval set we can see the discount evaluator here this is a custom one doesn't exist anywhere else but in our project because we created it and we get a perfect score and here is the output of the agent it ran. We can also have a look in uh studio web at the runs. Uh let's see uh we have this one and it got a perfect score. Why is that? Well, I also get the justification here because I wrote it in the evaluators. In the last seven days, we got four orders for this company and we expected a 5% discount and the actual discount is 5%. So, this is a perfect score. Using this evaluator, we can be sure that the agent is compliant to the company company policy that we have. >> Okay. So then one one challenge before we move to the next one is I think it would be fair to say that discount is not a static set it once in stone and never look back at it uh thing right. Uh so if someone wants to change it and the business owner would like to have direct influence and confident that so one he can change the discount based on several factors without changing the agent right and [clears throat] secondly to have the confidence that once it changes then the agent would pick up the new information and would integrate it uh within its logic. How would you go about this? Well, that's another feature that I planned for the end of this presentation, but we can touch it a bit now. Uh, you can achieve that using bindings. So, binding, it's another concept from UiPath that lets you drop to choose drop in replacements for your agent without modifying the code. So, in my code, I can have hardcoded. Okay, at this step and this is actually what's happening in this agent. I have hardcoded in code. Let's let's have a look actually because this is a good question. Uh let's have a look where the the main agent code code actually resides and let's see here. Okay. And I have some retrievers that I'm using to perform my some context grand retrievers and uh let's see not pre this one. This is hard code right because this is code I say okay I want you to look in orchestrator for the index called shipment index dev in the shipment index f in the shared folder and this is hard code it once the code is running and it's running good and the evaluations are giving it a high score I don't want to touch the code again right because it's already done it's I I got trust but then I don't want to use the shipment index dev I want to use the shipment index production while when I promote this agent to production. Well, bindings, UiPad bindings let you do that. You can at runtime you can choose what what resources will be consumed by this agent even if they are hardcoded in the code. This is a very powerful feature and I'll show this at at the end of the presentation. >> Cool. Cool. So, similarly we could reference the discount policy for Q2 or for Q3 dynamically without uh making a single change in the code. Yes, because an agent or an agentic workflow in this case it's good enough if it has the edges correctly. If the logic is good, if the prompts are good, if the the tools the internal tools are good, everything else should be evaluated in another way because if our way agent works as we expect, it should work good with index A, index B, index C with the many types of discounts. It can be fully customizable. Okay, cool. Okay. So I have just one more evaluation I want to show you which is the let's say the last very nice feature that we we introduced with this evaluation release. So if you remember at the beginning of the presentation I showed you when we ran iPad debug I showed you that we said okay you have a confidence threshold of 99. Everything below 99 you should escalate. Let a human be sure you took the right decision. But how how is that working? The human in the loop. How is human in the loop working with evals? Because you can't really have you may have an eval set that has 100 evaluation points. You can't put someone each time you evaluate the agent go and respond to the action created in action center and said resume reject. You can't interrupt the agent when evaluating it. And here comes the third and and final point from this presentation which is mocking. >> Yeah. And we have limitation here, right? We one thing that we cannot do and we will not be able to do is to evaluate a human. That that's out of scope. [laughter] >> Well, I mean, yeah, I guess I I wouldn't be I wouldn't completely agree with this statement, but uh [clears throat] let's let's suppose that's the case. Okay. So here we can mock the function that's because we this is a scenario we want to test. We want to test that if the agent is not confident enough it escalates the the the task to a human and the human does something and the agent continues. This is a valid scenario we need to test before promoting to production but we don't want to have the to deal with the human in the loop part and this is extremely easy to do with UiPad. You just have this import from UIP inval mock import mockable. This is a decorator and then we use this decorator to mock to apply to any function we need to be mocked within the evas. So we said for the create approval task that's actually interrupting the graph and waits for a human response. We want this to be mockable. So if I want in a data point I can say this should be mocked for this input and this output. This is for this input. This is the output you have to to provide. And how is that configured? Very easy. We have here the evaluation set. And the last evaluation point I have to show you, it's I think it's this one. New York successful. It's yes compliant complaint with human approval. This is the first input I showed you. When you we run the agent, the confidence threshold is very high. It's 99. So the agent almost every time will escalate this. But we define a mocking strategy for this valid point. And we say okay for the function create approval task that we decorated as mockable. So we know it can be mocked given any two arguments that will be passed to this function. I want you to return that is approved and this is the output from the human response looks professional and addresses the customer constraints appropriately. That's all I need to do and let's see if it works. Let's go to the default evaluation set. Search for this eval point. I think it was complaint with human approval. Let's see. Complaint with human review. This one uh before running it. Uh I want to show you another combination of uh of trajectory and uh output based. I have here an exact match evaluator that comes out of the box. Here I just say that okay I want I want you to look at the error part of the agent. It should you can consider this successfully if it has no error and I have a complaint trajectory evaluator and I I write it here and I tell you what it should do. The agent should classify the email as a complaint in natural language and this will look at the agent trace not the agent object. Okay, let's save and let's run this. Now remember for this eval point the confidence threshold is 99. So if we are not mocking that function it will not work right the agent will escalate the the problem to a human someone is going going to uh need to reject or approve but we don't want that. So in here in orchestrator as always we can see what this job is doing. Let's see if it started. It did. We can look at the the this is very nice. When evaluating at the agent, you don't care only about the evaluators, but you also care about the agent trajectory. So you know what you are evaluating and using this discover discoverability tools. You can see here in very easy uh what this what this agent is doing. Okay. So let's see let's have a look at the traces in here. human review node. This is the node that should escalate the problem. Let's see details about it. So let's look at the output of this one. And the output is what we mocked inside the evaluation. So this was not actually called. But when the agent when the evaluators are evaluating this, they're going to get this trace. They're going to get okay. So the agent was approved by the agent task was approved by a human. And if we go back to to UiPad Studio, let's see here the responses of of those runs. And we see that we get a perfect score for both evaluators that we applied. The LM J trajectory gave it 100%. Because the agent process the email as complaint as we asked it it to do that it look over the traces and the exact match gave it 100% because we have no error. So that's pretty much it with evaluators. Of course, now you can do I mean as we showed you almost sky is the limit here. If you don't have the perfect built-in evaluator to use, you can have a custom evaluator. If you don't want to call in the evaluator external services that may that may not be important and they may change data, you don't need to do that. You mock everything you don't want to be called like you do in a regular unit test. Okay. So, we gain trust. Can we promote it now to production? What do you see? >> Yeah. [snorts] Um, now I think uh we are pretty much there. >> Okay. So, we showed you in the last two uh two workshops that you can push the agent from within the the terminal. You can run your iPad pack, your iPad publish. But once an agent reaches studio web, you can do with it everything else you can do within studio web for other project types. So we can actually click on publish here and this will this is very nice. This is where this studio integration really plays a key role because when we publish this solution this this becomes a self-contained solution that we can take and publish it to another organization to another tenant and it should work out of the box because we do have here the resources used by the agent the resource that we can override as I said and I'll show this in in a moment after we publish this and also we have here an active process we can have five more agents 10 more agents doesn't matter once Studio web, they will all be contained in to the same solution and that should work out of the box in any environment, not only where you tested it when you have those those resources. Okay, but you can find more about this in in our uh in our studio documentation. Uh let's see what the publish job is doing. uh I should be notified once those those agents that actually the agent and IP job are properly packaged and deployed. Uh let's see hopefully did not fail. We can let's check actually in the solutions tab. So in here uh in my personal workspace I have this solutions and in here I should see the published solutions from uh from within uh UiPad. Okay. And we see that this is the one logistic G for G but it is it is in a draft state because most likely I need to uh oh no never mind it was successful. I just needed to give some time. Okay so it was successful. The solution was deployed. We can see it here. This is the current version of the solution. And we can navigate to the solution folder. Okay. So this is the solution folder. We have here a couple of of processes. This is the RPA workflow. This is the approval app because as I said now it works as a fully contained solution. We we have this app inside the solution once we import it. And we have the agent here. And let's let's edit this agent and and see a bit here in the package requirements. Those are the bindings I just told I was yapping about before. So in here you can see the resources used by the agent and what's actually used at runtime. So we see we have an index here which is the company policy index. If we want we can override it but for this one let's keep it the company policy. We have an app here and we can we can choose another app that we wanted to use. Let's choose this one. And for index, and this is a key part, we have here the shipment index dev, but we just promoted this agent to production. We want to test it. We want, sorry, we want it to run with a realtime index that contains uh realtime data. So instead of the shipment index dev, let's choose the shipment index production. Let's update the agent and let's run it. And with that we can conclude. Since we promote the agent to production, we need to be highly confident that it it took the the right decision. So let's say 99 confidence threshold. And as always, let's use the input we we tested before. So remember the code is the same. I didn'tify I didn't modify anything in the code. The agent is the same. The input is the same but there was a there was a runtime configuration that I did before. Instead of test let this be because now we are in a real world scenario. This is John do and the email content is the same. [clears throat] Okay. And let's type this agent. Now why was was very important that I chose the production index not the development index. So remember the email is the same right? The shipment should go to London to Oslo just as before. But if we go here in the shed to the storage buckets where we have the actual data ingested by the indexes and we have a look at shipan data production. This will contain a key difference. Let's download this file and have a look at it. Let me put it on the right screen. Uh, one sec. It's this one. Where is it? Okay, in the production index from London to Oslo, we actually have a temporary uh um restriction, right? We cannot perform shipments from this location due to current international restrictions. And uh this is key. So actually even if you don't modify the code, the agent can operate with any sorts of data. And let's see what it did. So before when running with the same the same input and the same uh code it said okay it's fine I applied the 5% discount this is the sum go ahead and we can see here the agent is suspended because I gave it a 99 confidence threshold. We have the exume conditions tab and we see that okay the resume condition for this agent is actually uh uh uh an action created. We can have a look at the traces [clears throat] and let's see how this section looks like in in action center for the the grand finale. Okay, so this is the task it created and because I used a different index and that index contains restrictions with the same code that I gain confidence in with the same agent I'm going to get different outputs and in here we can see that the proposed email response is that aggregating form there is currently a restriction on the requested delivery route. So this is where bindings do all the magic. Once an agent is is done, once you finish developing it, that's all. No need to touch the code. You evaluate it, you promote it to production, and you pick what resources that agent uses. Let's approve this. The job should pick up from where it left off. The state switched to resume. And we should see in in a second the the final output in here of this job. So yeah, I think this is pretty much it. We can we can recap at the end what we did. So we developed an agent. We made sure it works on our test scenarios. Then we evaluated not only the output but also the trajectory and the custom logic that we did. In order to achieve that, we mocked some of the functions that we needed and the last step we package the project. We publish it and we feed it the production resources, not the development resources anymore. And in here, the output of the agent is that I will inform you there is currently a restriction on that route. >> Yeah, thank you R. I think this is a beautiful showcase on where we stand as UiPath and what is our mandate and if you remember at the beginning we stated that we are not chasing the buzzwords and not chasing the hype we we are very keen on understanding where the true challenges are where the true value is and tackle those problem so what we've shown here is that we understand that there is no silver bullet in this agentic space there is a lot of promises but there is no one prompt or one LM model that would do everything but rather we've shown that from build to deploy we do need the proper tools and the proper platform in order to actually deliver on on this promise and you you've [clears throat] seen that as Radu um did a recap so in the build we had the UiPath Python SDK to have the access bucket human in the loop to create that data fabric and all that to bring that into the uh dev experience we've provided uh dev have a debug tool that we've showcased to uh know reach that functional ready state for the agent and after that all the tools for the eval like we can eval trajectory tool calls [clears throat] end to end flow and we can have uh very powerful tools like having a mocking and uh the custom evaluators along the way and we even provide the dual eval experience you can develop the eval in code in your IDE and then synchronize it with studio or even use the uh UIX canvas to build those same evalu capabilities in terms of eval. So we have a full eval experience but that's not the end of it. You need to provide a good deploy experience where you can have it packed in the solution. Ability to securely orchestrate and govern and for instance have possibility to have different bindings that you can control the behavior of uh the agent and the proper context that you give to the agent. And one final part that we did not touch on is the what happens after deployment because the story continues even after that while you can monitor it. You can provide feedback to that agent, incorporate that feedback and watch if not only you've gained trust but next one is to see that it works as expected and you can mate maintain trust and maybe that is a story for another time but most of the story we've told already and I hope this created some uh a good feeling on what is the power and the need for this kind of tools. So thanks and with that I think we can open for Q&A if we still have some time for that. So uh the questions I gleaned from the session today was the first one was obviously uh where can we get the the code for this uh the if if that's possible is that something that we can provide. Uh as of now no this is not a public repo but uh typically what we do we refine the code and we then publish them as samples in our uh repo. >> Yeah. So for sure very soon you are going to see this as a sample in our UI launching repo as this was a launching agent. But uh as I said before in the forex workshop a very good starting point for anyone that wants to use and develop such agent is our sample. So in our documentation you can find the link to our examples and there then uh we have a lot of samples very easy samples not this complex and each of them shows you can use rack how you can use an MCP how you can use an index and you combine all of those and you can end up with with agent such as this but will be this this also will be available once we refine it a bit. >> Okay. So uh is it okay if I ask some questions? >> Sure. >> Yeah. So uh you you talked about the business logic basically the the operational flow in the beginning right? Uh can I kind of understand this uh but can you explain to the learners to the beginners why that is necessary from a technical perspective as well? Uh why why is the business case needed from a a technical perspective? Is that the question? >> Yeah. So why do technical people need to understand the business use case at the end of the day? >> Mhm. Okay. Yeah. Well, as you see as you seen in this agent, the the diagram of it was was pretty large and this was this agent was specifically developed for the gigs for gigs demo. we may reuse it in the future but in a real real world scenario it may have like 10 times that amount of nodes. So a developer good engineer should uh understand the business requirements and based on that you can develop your agent. This is this is the the main advantage of workflows. We we call this an agent but it's actually an agentic workflow that uses utilizes three sub aents if you if you may. So formally you would have just one react agent. You have an LLM. We give it some context and we we pass it it we throw at it some tools and you they would get the human request and they will do stuff. But now with with workflows what what coded agents offer is that deterministic approach that blends it perfectly with LLMs when you can have a deterministic choice like for discount if you have more than three work days in the same day five discount you should take that deterministic choice because you are sure that the agent will will take the right decisions when you can do that when you are working with structure data when they should take actions when they should call tools and stuff like that for sure use an LLM. So this this is the golden path here. You understand the business case where you can do something deterministic you do it but when you can do it you let the LLM do their magic but constraining them with those indexes with this context grounding with those tools that you first validate. So that's very important. So yeah that's that's really uh a keen insight uh Radu. So uh another question that I saw in the comments is u are these agents they are they going to be ultimately placed in like humanoid robots or something like that? Is that the future that we are looking at? Yeah, in a sense um I think what matters are the principles behind this. Some of them would be applied to a wide range of use cases. Um and whatever we presented we think that they hold through as principles as approaches as need and the uh implementation for a host of uh use cases and uh the only new ones that we will see for other use cases is in the customization part and here is where the coded agentic workflows come into play where we can integrate with bespoke systems and have as mentioned before have the possibility to insert deterministic code which would be custom and predictable alongside with undeterministic uh ones. So yeah the same principle hold true and coded workflows coded agentic workflows do provide the kind of means the control and flexibility that you would need to enable those uh scenarios both from a developing build perspective and evaluation perspective. >> Okay. So u how okay so another question that I saw was that for such an agentic workflow how much evaluation should be done like I know the answer there's no there's no such thing as too much evaluation right at the end of the day right but but what should be an appropriate amount of time let's say we built the agent in 3 months then how much more uh efforts and time should we put in evalu situation. >> No, perfect. That that's a very good point and that's another challenge that we we somewhat touched upon obviously with evas but we did not discuss about evas coverage what means to to have a proper and did you evaluate it fully the the workflow means meaning that the potential flows are evaluated such that whenever it goes to a certain path you had the proper evaluator the relevant evaluator along that path and for that indeed that's an a an uh an issue a challenge on top of what we've discussed today on how you can generate or build manually those general way of thinking about this is to think in flows and for that flows to think of possible paths and for those possible paths to add the evaluators add the steps or tools uh such that the main decision points are covered. That is the the typical approach. What we see in the future is that uh it and something that we are working on is to have a tool that would understand the workflow. Something that is beyond a power of a single developer to understand all the flows all the potential paths and once you understand or system understands the potential paths to generate on top of that the eval coverage, right? Understand the flows and proposing the proper evals along the paths. And for the developer would just be to understand what are those and uh tweak it here and there with some more relevant inputs. uh but this is what we I think this is a very relevant question and something that um should be addressed and a solution that should be uh provided and things that we are actually uh now working on but that uh if you do release it I would love to have another workshop on this because that would be quite revolutionary uh full test coverage of the entire code is like the holy grail of programming since what the 1990s I But yeah [laughter] so so that would be absolutely revolutionary. Uh >> okay [laughter] so uh another question is that I'll just rephrase it. Aush another question is that we know that there are some parameters to optimize for you know creativity or or maybe abstraction of the output right but what are some parameters that we can consider for making it more friendly in tone or or making it more robust uh in sense of in the sense of market research and market friendliness. So what would be some parameters that you would uh modify or or you would change? Mhm. Well, uh there are yeah some hip parameters that you can tweak uh when working with LLMs, but our those frameworks that are very popular nowadays and we are have integrations with them like longchain they let you choose what should be there is a clear distinction about the system prompt and the user prompt. So, and each time you create an agent, you can provide a system prompt for it. And it's that easy. The system prompt can contain how it should respond to questions, how technical it should be, how helpful or whatever. And no matter what the user prompt is, it's very hard to do sort of an uh let's say an injection when it comes to the tone the agent should have and stuff like that because the system prompt will always have priority over the user chrome, >> right? So basically it it you can control the persona of the LLM with this system prompt. >> Yes. Yes. Also in the example show just showed you any every agent that was using tools they had a system come behind and they said okay you are a discount expert you know how to calculate those using that you are very very uh you pay attention to the details. So you know that was for the first one. the second one, the one that classified the request, it had an entire different system. So those should be fine- tuned depending on the the thing the agent is trying to achieve. >> So uh I think that's it. If there are no more questions, >> uh I think there was another one. I can see it here in the chat. Somebody asked if there is an premium option of the UiPath action center. And the question to that is yes. The answer sorry is yes. We do have the community plan. You can go ahead and create an account in UiPad platform today and you can use uh some resources and play with them and see if everything works fine before going commercial. I think you already have like out of the box like please correct me if I'm wrong. You do have like 25 2050 LLM calls free per day and and access to second resources and stuff like that. Yeah, we do have entitlements enough so that you can experiment with um all agentic features and they do those entitlements are a common pool and you can use that for whatever services including these ones. >> Okay. Uh another great question uh what are where uh have so one real world example maybe where UiPath or UiPath platform helped in uh security operations. Uh I mean >> yeah by the way uh one one of the common things with agents and with every automation is that they save time and what we found out u lately and I think it would be interesting to to just mention that the true value with agents is way beyond time. Um and I I found this for a lot of examples with customers where time was like a given and we've seen this in um trivial things like uh ticket uh classification where it saved a lot of time because it removed the the input and analysis of human in assisted cases because the summary of those use cases were provided to human and just for review and then it would complete the the use case. Um for instance in the logistics the one that we've shown this is a demo but we've seen this with our customers the time improvement for u logistic and uh that use case in transportation was uh I think from like 1 hour per transaction to like 5 minutes or so approximation. Um but generally uh now the time is commoditized as a value and true value is that is is the less human intervention on decision- making. That is the the currency here. The true the true value. >> True. True. Uh another great question. I think you answered that in the uh in the code where you were showing the code as well. Can agents be paused? >> Yes. Yes, they do. We do have the suspend and resume mechanism and we showed you in this example a human in the loop where a new human comes and resolves the task but we also offer agent in the loop. So imagine like you don't even need a human that's responding to the task. You have an agent. You can have multiple agents talking with each other. One suspends, one resumes, then it gets the output and and goes on. So yes, those agents are are developed in such a way that a human can intervene at any point you decide it's it's necessary. >> Okay. Okay. So uh guys uh any any closing remarks that uh any of you would like to make? >> Yeah. And as always we one thing that we love to do is to see what you've built and please do reach out to us and we are willing to help. Please check our documentation and uh the community offerings and [clears throat] uh yeah uh the community forum. We again eager to see what you've built and we are in that phase where we uh always listen very carefully to uh whatever challenges and uh you know bottlenecks or friction you would encounter and ready to uh help with those and uh increase the speed and the reliability of the agents. >> Yeah. And we also welcome contributions to our open source repository. So >> yeah double down on that. double down. >> Yeah. Okay. Uh Okay. So, first of all, I would just like to appreciate the UiPath team, right? So, they they showcase some a particular use case that I have never seen, right? And and this basically shows the amount of efforts they put in. I would just like to thank you guys. U and thanks for having us. >> Yeah, thank you. >> And to the audience and the participants, I hope you had fun. I hope you learned a lot and I hope you will be building amazing projects using the uh EUPath platform and you will be showcasing it to us as well. So that's that's pretty much it for today. Thank you for your time. I have been Samunatan. See you next time. Thank you guys. >> Thank you.

Original Description

Register here to be eligible for the Certificate: https://www.geeksforgeeks.org/event/UiPath-Llamalndex-Agent-UiPath-gfg This is the third and final workshop of the series. You need to register on the above link to be eligible to get the participation certificate. The links of the first 2 workshops are given below for you to watch and learn. As the world moves into an agentic future—the UiPath Platform™ enables AI agents, robots, people, and models to work together harmoniously to revolutionize industries and enhance human potential. AI agents are autonomous software systems that can perceive environments, reason in real time toward objectives, and execute complex, multi-step tasks with minimal human intervention. Fueled by large language models (LLMs), they are rapidly transforming business operations by delivering adaptive, scalable workflows that drive efficiency, decision making, and innovation. In this session you will learn to build, integrate and deploy enterprise grade Coded Agents with the UiPath platform using the LangGraph Framework. “We’ve seen a key part of this be AI observability, and we’re excited to integrate LangSmith with UiPath to help even more builders ship agents with confidence.” Harrison Chase CEO, LangChain Register now: https://www.geeksforgeeks.org/event/UiPath-Llamalndex-Agent-UiPath-gfg Watch the 1st Workshop: https://www.youtube.com/live/XqxO-1Bwv44?si=kNDSHS9XrIATh12Q Watch the 2nd Workshop: https://youtube.com/live/6lAic15ji10 #aiagent #uipath #agenticai
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Playlist

Uploads from GeeksforGeeks · GeeksforGeeks · 0 of 60

← Previous Next →
1 How I got into Walmart | Shailesh Sharma
How I got into Walmart | Shailesh Sharma
GeeksforGeeks
2 Upgrade yourself In 29 Days | GeeksforGeeks
Upgrade yourself In 29 Days | GeeksforGeeks
GeeksforGeeks
3 Learn AWS Fundamentals For Free
Learn AWS Fundamentals For Free
GeeksforGeeks
4 Conversation With Young Achievers | Meet the winners of Bi-Wizard Coding Contest | GeeksforGeeks
Conversation With Young Achievers | Meet the winners of Bi-Wizard Coding Contest | GeeksforGeeks
GeeksforGeeks
5 Meet The Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
Meet The Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
GeeksforGeeks
6 Interview Prep Strategies | PayPal
Interview Prep Strategies | PayPal
GeeksforGeeks
7 OLX Interview Preparation Strategies | Hukam Singh
OLX Interview Preparation Strategies | Hukam Singh
GeeksforGeeks
8 Meet Some More Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
Meet Some More Winners Of Bi-Wizard Coding Contests | GeeksforGeeks
GeeksforGeeks
9 Live Mock DSA
Live Mock DSA
GeeksforGeeks
10 Microsoft Azure For Absolute Beginners
Microsoft Azure For Absolute Beginners
GeeksforGeeks
11 Python for Data Science | Data Science Master Bootcamp | Arpit Jain
Python for Data Science | Data Science Master Bootcamp | Arpit Jain
GeeksforGeeks
12 Getting Started with Data Analysis | Data Science Master Bootcamp | Ashish Jangra
Getting Started with Data Analysis | Data Science Master Bootcamp | Ashish Jangra
GeeksforGeeks
13 How to prepare theory subjects for SDE interviews | Geeks Summer Carnival 2022
How to prepare theory subjects for SDE interviews | Geeks Summer Carnival 2022
GeeksforGeeks
14 Get Your Tickets To The Geeks Summer Carnival | GeeksforGeeks
Get Your Tickets To The Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
15 TED Talk Data Analysis Project | Data Science Master Bootcamp | Ashish Jangra
TED Talk Data Analysis Project | Data Science Master Bootcamp | Ashish Jangra
GeeksforGeeks
16 How I Secured AIR 9 in GATE'22 |  Tushar
How I Secured AIR 9 in GATE'22 | Tushar
GeeksforGeeks
17 Learn Java Backend Development | Geeks Summer Carnival | GeeksforGeeks
Learn Java Backend Development | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
18 How to Recognize which Data Structure to use in a question | Geeks Summer Carnival | GeeksforGeeks
How to Recognize which Data Structure to use in a question | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
19 Learn Data Structures and Algorithms | GeeksforGeeks
Learn Data Structures and Algorithms | GeeksforGeeks
GeeksforGeeks
20 Interview experience at Flipkart | GeeksforGeeks
Interview experience at Flipkart | GeeksforGeeks
GeeksforGeeks
21 Lets Prepare for GATE'23 the Right Way | Sakshi Singhal | GeekSummerCarnival
Lets Prepare for GATE'23 the Right Way | Sakshi Singhal | GeekSummerCarnival
GeeksforGeeks
22 Highest Paying Jobs in 2022 | Ishan Sharma | Geeks Summer Carnival 2022 | GeeksforGeeks
Highest Paying Jobs in 2022 | Ishan Sharma | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
23 Geeks Summer Carnival 2022 | 5th April- 11th April | GeeksforGeeks
Geeks Summer Carnival 2022 | 5th April- 11th April | GeeksforGeeks
GeeksforGeeks
24 Preparing for SDE interviews | Soham Mukherjee | Geeks Summer Carnival 2022 | GeeksforGeeks
Preparing for SDE interviews | Soham Mukherjee | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
25 Full Stack Development with React & Node | Utkarsh Malik | Geeks Summer Carnival | GeeksforGeeks
Full Stack Development with React & Node | Utkarsh Malik | Geeks Summer Carnival | GeeksforGeeks
GeeksforGeeks
26 Introduction to Open Source and Roadmap to GSOC 2022 | Geeks Summer Carnival 2022 | GeeksforGeeks
Introduction to Open Source and Roadmap to GSOC 2022 | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
27 Web Scraping in Action | Geeks Summer Carnival 2022 | GeeksforGeeks
Web Scraping in Action | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
28 Getting Hired at BITCS via GfG Job Portal | Get Hired With GeeksforGeeks
Getting Hired at BITCS via GfG Job Portal | Get Hired With GeeksforGeeks
GeeksforGeeks
29 How to build a faster landing Page | Geeks Summer Carnival 2022 | GeeksforGeeks
How to build a faster landing Page | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
30 Geeks Summer Carnival | 5th To 11th April, 2022 | GeeksforGeeks
Geeks Summer Carnival | 5th To 11th April, 2022 | GeeksforGeeks
GeeksforGeeks
31 How to get ideas for Startup | Geeks Summer Carnival 2022 | GeeksforGeeks
How to get ideas for Startup | Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
32 Journey from Tier 3 to JusPay | GeeksforGeeks
Journey from Tier 3 to JusPay | GeeksforGeeks
GeeksforGeeks
33 Geeks Summer Carnival 2022 | GeeksforGeeks
Geeks Summer Carnival 2022 | GeeksforGeeks
GeeksforGeeks
34 Dispelling Myths and Pre conceptions of Programming Languages
Dispelling Myths and Pre conceptions of Programming Languages
GeeksforGeeks
35 Must Do System Design Questions
Must Do System Design Questions
GeeksforGeeks
36 Understanding Sorting Techniques in an hour | Keerti Purswani | Geeks Summer Carnival
Understanding Sorting Techniques in an hour | Keerti Purswani | Geeks Summer Carnival
GeeksforGeeks
37 Get Hired at NEC | Job-A-Thon 8
Get Hired at NEC | Job-A-Thon 8
GeeksforGeeks
38 Journey from Tier 3 college to Microsoft | GeeksforGeeks
Journey from Tier 3 college to Microsoft | GeeksforGeeks
GeeksforGeeks
39 Get Hired with GeeksforGeeks at SuperK | Job A Thon 8
Get Hired with GeeksforGeeks at SuperK | Job A Thon 8
GeeksforGeeks
40 GeeksforGeeks: Redesigned
GeeksforGeeks: Redesigned
GeeksforGeeks
41 From Tier 3 to cracking multiple interviews | GeeksforGeeks
From Tier 3 to cracking multiple interviews | GeeksforGeeks
GeeksforGeeks
42 Live Mock DSA
Live Mock DSA
GeeksforGeeks
43 Youtube Data Analysis | Ashish Jangra | GeeksforGeeks
Youtube Data Analysis | Ashish Jangra | GeeksforGeeks
GeeksforGeeks
44 DSA Self-Paced Course Preview | Sandeep Jain | GeeksforGeeks
DSA Self-Paced Course Preview | Sandeep Jain | GeeksforGeeks
GeeksforGeeks
45 GATE Live Classes | Prepare for GATE CS 2023 | GeeksforGeeks
GATE Live Classes | Prepare for GATE CS 2023 | GeeksforGeeks
GeeksforGeeks
46 Journey from JIIT to Adobe
Journey from JIIT to Adobe
GeeksforGeeks
47 Life Is Unfair Ft. Shonty badmash | LIVE Discord Session | A GeeksforGeeks Exclusive
Life Is Unfair Ft. Shonty badmash | LIVE Discord Session | A GeeksforGeeks Exclusive
GeeksforGeeks
48 Interview Experience at Google | Tech Dose
Interview Experience at Google | Tech Dose
GeeksforGeeks
49 Live Mock DSA
Live Mock DSA
GeeksforGeeks
50 Interview Experience @ Amazon | GeeksforGeeks
Interview Experience @ Amazon | GeeksforGeeks
GeeksforGeeks
51 My journey through the tech world from India to US | Vidushi | GeeksforGeeks
My journey through the tech world from India to US | Vidushi | GeeksforGeeks
GeeksforGeeks
52 Complete Interview Preparation Course | GeeksforGeeks
Complete Interview Preparation Course | GeeksforGeeks
GeeksforGeeks
53 Live Mock DSA
Live Mock DSA
GeeksforGeeks
54 Getting Hired at FiftyFive Technologies | Job-a-thon 9.0
Getting Hired at FiftyFive Technologies | Job-a-thon 9.0
GeeksforGeeks
55 GFG Karlo, Ho Jayega | GeeksforGeeks ft. Khaleel Ahmed
GFG Karlo, Ho Jayega | GeeksforGeeks ft. Khaleel Ahmed
GeeksforGeeks
56 How I got job offers from 2 big companies : Arcesium & Microsoft | GeeksforGeeks
How I got job offers from 2 big companies : Arcesium & Microsoft | GeeksforGeeks
GeeksforGeeks
57 LINUX for Beginners | GFG x Itversity
LINUX for Beginners | GFG x Itversity
GeeksforGeeks
58 My interview experience at Walmart | GeeksforGeeks
My interview experience at Walmart | GeeksforGeeks
GeeksforGeeks
59 Get Hired at Speckyfox
Get Hired at Speckyfox
GeeksforGeeks
60 Live Mock DSA
Live Mock DSA
GeeksforGeeks

Related Reads

📰
The Fall of Static Audits: Analyzing AI-Driven Supply Chain Risk Management
Learn how AI-driven supply chain risk management can help mitigate disruptions and minimize losses, making traditional static audits obsolete.
Dev.to AI
📰
Building a multilingual voice agent: lessons from FR/EN/DE: field controls that hold
Learn how to build a multilingual voice agent by applying lessons from French, English, and German language support
Dev.to AI
📰
How I run an AI agent 24/7 on a Raspberry Pi (and don't lose its memory)
Run an AI agent 24/7 on a Raspberry Pi without losing its memory, a cost-effective and secure solution
Dev.to AI
📰
How MADDPG Combines Deep Learning with Multi-Agent Strategies
Learn how MADDPG combines deep learning with multi-agent strategies to enable AI agents to learn, collaborate, and compete in complex environments
Medium · Deep Learning
Up next
Best AI Agent Community to Accelerate Your Learning of AI (James Dooley Chats with Julian Goldie)
James Dooley
Watch →