Skip to main content

Write a comment

PREreview of When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems

Published
DOI
10.5281/zenodo.22922694
License
CC0 1.0

I run invoicing, payments and support for my own small businesses on agents like the ones in this paper.

So let's start, for example, that each business has its own leaks. so it's money leaks. It's invoices leaks. It's some customer support leaks.

Because if you're using, for example, AI, you understand that it can be wrong. Sometimes it can hallucinate, it can broke, it can be broken. And you need to sometimes repair it, or you want to like, do something new, or you want to add some information.

This paper counts it on 1,980 test cases, twelve error types. One failure they show: the agent skips the tool and answers from its head.

what I also find for myself that AI you know AI makes a lot of mistake a lot. It's like hallucinate. It can create something by himself and then say sorry yeah I was wrong.

In their example the model states a payment amount without querying the database. It looks like success.

This client, for example, asked for some kind of stock, or he wanted an invoice. So again, I have another AI agent who can create an invoice and send it by mail to the client. Client is paying your invoice. Next AI agent, he's managing, if really we received Stripe, PayPal, Zelle, wherever he want to pay. Then warehouse. Then shipping.

In a chain like that a wrong amount does not stop at one step.

anything is not okay it can hallucinate it can turn off. You ask him hey please change this number nine to number eight and he can change everything for you and you're like oh my god who is the best computer in the world like me or you?

The paper does not describe an approval step. My dental lab system works like this. So it's never sent anything to the lab by itself. So anything that leaves clinic needs a human approval or the system, if the system is not sure. And in this case, it stops.

So, if we're talking about money, the first thing that AI do, for example, if we're talking about invoices or, like, client's payment, it's not changing your system. So, you will have, like, the same view. You will have the same display.

And when I wrote to FDA, my point was simple. Big companies, they have engineers. So, they have their own people full-time that watch their AI. But a small clinic has one doctor. So, the rule must be simple for them. The machine prepares AI checks, human approves. And when the machine is not sure, it stops.

let's say CL code is doing something like for five seven eight hours and then Codex came and say hey here is a mistake here is the mistake and he's like oh thanks yeah let's rebuild it.

So just one time in the morning I open my computer I check that everything is working. So you're like you have some kind of orchestra and you just need on a top level to see that everything is okay.

Competing interests

I run Negodiuk LLC. I build AI agent systems for small businesses, including the dental lab ordering system in this review. I have no connection to the authors.

Use of Artificial Intelligence (AI)

The author declares that they used generative AI to come up with new ideas for their review.

You can write a comment on this PREreview of When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now