,
Department of Information Technology, Sri Sivasubramaniya Nadar College of Engineering, Chennai, Tamil Nadu, India
,
Department of Information Technology, Sri Sivasubramaniya Nadar College of Engineering, Chennai, Tamil Nadu, India
,
Associate Professor, Department of Information Technology, Sri Sivasubramaniya Nadar College of Engineering, Chennai, Tamil Nadu, India
Senior Product Manager in Tech, Amazon, USA
This study focuses on designing a self-service kiosk system called Bank and Loan Assistant Kiosk, which will be a breakthrough idea by using Multi-Agent System and machine learning technologies. Although considerable development has been made in online banking, there is a major flaw in the availability of voice-first and localized financial advice services in rural areas with low literacy rates. Existing systems of online banking depend mainly on the use of text input interface, cloud computing architecture, and universal language models, which are not suitable for rural India. The main objective is to increase the efficiency of the kiosk by offering real-time customized analysis for loans offered by the government welfare programs. This paper attempts to fill this gap by designing a system that uses three specialized agents - the User Agent responsible for intent detection and dialogue management using a voice interface; the RAG Product Agent for scheme retrieval via semantic vector search through a knowledge base; and finally, the deterministic rule-based Eligibility Agent used for verifying the scheme specific qualifications of applicants. All agents coordinate their actions according to the JSON-RPC 2.0 Agent-to-Agent (A2A) communication protocol. The fundamental principles of the system presented here have solid foundations in existing scientific literature. Indeed, Multi-Agent Systems have proven to be a suitable means for tackling the problem of unstructured financial data extraction and processing since unstructured financial reports were able to be turned into structured data using an average accuracy of roughly 95%. The system uses a fine-tuned OpenAI's Whisper model for voice recognition, which was trained on 1,200 Tamil audio samples and reached WER and CER of 10.77% and 3.37%, respectively. Additionally, end-to-end latency was observed to be 1.72 seconds, falling below the two seconds threshold required for real-time kiosk operation.
This is an open access article distributed under the Creative Commons Attribution Non-Commercial License (CC BY-NC) License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
The statements, opinions and data contained in the journal are solely those of the individual authors and contributors and not of the publisher and the editor(s). We stay neutral with regard to jurisdictional claims in published maps and institutional affiliations.