,
Assistant Professor, Department of Electronics and Computer Science, Fr. Conceicao Rodrigues College of Engineering, Bandstand, Bandra West, Mumbai, India, Department of Computer Science and Engineering, Koneru Lakshmaiah Education Foundation, Guntur, Andhra Pradesh, India
,
Professor, Department of Computer Science and Engineering, Koneru Lakshmaiah Education Foundation, Guntur, Andhra Pradesh, India
Professor, Department of Computer Science, Tunghai University, Taichung, Taiwan
The financial institution uses different types of documents, including bank statements, cheques, income tax forms, salary slips, and electricity bills. Current approaches deal with either the preprocessing stage, text and feature extraction, or even classification separately and fail to perform efficiently on low-quality, handwritten, and visually similar documents. The proposed method, HAPODCNN, which stands for Hybrid Top-Down Attention Pyramidal Atomic Orbital Dual-Path Convolutional Neural Network, is used for digitization and classification of financial documents in an end-to-end approach. The architecture includes MGSWF-GS as a preprocessing technique for image improvement, JA-ViT-UNet as a text and feature extraction technique, and a bidirectional pyramidal dual-path CNN along with spatial and channel attention for classification. Atomic Orbital Search is employed for model parameter tuning. The experiments were performed on 414 real financial documents in India with 5 categories. First, the 414 original images were partitioned into three sets: 70% of the images for training, 10% for validation, and 20% for testing. The data augmentation process was then performed on the training set only to increase the number of training data from 414 to 1,170. The suggested HAPODCNN yields 98.93% classification accuracy with 97.34% precision, recall, and F1 score, and consumes 18.7 million parameters and 3.42 GFLOPs. The model also reports 98.30% ROC-AUC, 3.2 ± 0.1% character error rate, 6.5 ± 0.2% word error rate, and 1.12 s latency. The suggested model ensures 96.1% accuracy even under severe document degradations, which implies its high robustness. As seen from the obtained results, the combination of degradation-aware preprocessing, transformer-based feature extraction, bidirectional attention-based classification, and AOS allows improving recognition and classification accuracy, which makes the suggested framework applicable to automated financial document processing tasks.
This is an open access article distributed under the Creative Commons Attribution Non-Commercial License (CC BY-NC) License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
The statements, opinions and data contained in the journal are solely those of the individual authors and contributors and not of the publisher and the editor(s). We stay neutral with regard to jurisdictional claims in published maps and institutional affiliations.