Skip to main content

Write a PREreview

Self-Correcting Multimodal AI Agents for Reliable Autonomous Decision-Making

Posted
Server
Preprints.org
DOI
10.20944/preprints202609.1427.v1

Recent advances in large language models and multimodal foundation models have enabled artificial intelligence agents to jointly process text, images, and other modalities while performing multi-step reasoning and interacting with external tools. Despite this progress, autonomous agents remain prone to compounding errors that originate in visual perception, language reasoning, tool invocation, or the interaction among these components. This paper proposes a self-correcting multimodal agent architecture that couples task decomposition and planning with an explicit verification and revision loop, allowing the agent to detect inconsistencies in intermediate reasoning steps and in final outputs before a task is considered complete. The architecture combines cross-modal evidence aggregation, secondary-model critique, rule-based consistency checks, and tool-grounded verification to identifycandidate errors, and it uses a bounded iterative revision procedure to correct them without unnecessary recomputation. We formalize the verification-and-correction process as a constrained optimization over an evolving action-reasoning trace and describe a concrete algorithmic instantiation suitable for visual question answering, document understanding, and tool-augmented reasoning tasks. We further describe an evaluation protocol spanning accuracy, factual reliability, task completion, and robustness to injected multimodal errors, together with an explicit accounting of the additional inference-time cost introduced by iterative self-correction. The framework is intended to give practitioners a concrete, reproducible template for building and evaluating multimodal agents that can detect and recover from their own mistakes prior to acting autonomously.

You can write a PREreview of Self-Correcting Multimodal AI Agents for Reliable Autonomous Decision-Making. A PREreview is a review of a preprint and can vary from a few sentences to a lengthy report, similar to a journal-organized peer-review report.

Before you start

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Start now