×
Home Current Archive Editorial board
Instructions for papers
For Authors Aim & Scope Contact
Original scientific article

ENHANCING SPATIAL-TEMPORAL FEATURE FUSION WITH CONVOLUTIONAL GRUS FOR BREAST TUMOR CLASSIFICATION

By
A. Aarthi Orcid logo ,
A. Aarthi

Assistant Professor, Department of Electronics and Communication Engineering, Paavai Engineering College, Namakkal, Tamil Nadu, India, Research Scholar, Department of Electronics and Communication Engineering, Ponnaiyah Ramajayam Institute of Science and Technology (PRIST), Deemed to be University, Thanjavur, Tamil Nadu, India , Thanjavur, Tamil Nadu , India

Smitha Elsa Peter Orcid logo
Smitha Elsa Peter
Contact Smitha Elsa Peter

Professor, Department of Electronics and Communication Engineering, Ponnaiyah Ramajayam Institute of Science and Technology (PRIST), Deemed to be University, Thanjavur, Tamil Nadu, India

Abstract

Even today, one of the main reasons for mortality among females is breast cancer, whose effective treatment needs earlier detection. In this work, a new methodology based on spatial-temporal Deep Learning for detecting and classifying breast tumors through the use of Convolutional Neural Networks (CNN) and Convolutional Gated Recurrent Unit (ConvGRU) is presented. The spatial part of the network extracts spatial features such as tumor texture, tumor shape from each slice individually, whereas the temporal part of the CNN identifies changes between different slices. The model presented in this study is capable of increasing accuracy, decreasing false positive rates, and improving robustness for tumors with irregular morphology by exploiting both spatial and temporal aspects together. For evaluation of the model, the publicly available BUSI dataset was used, and 0.9867 area under the curve (AUC), 96.88% recall, 88.57% precision, 95.40% accuracy, and 92.54% F1-score were obtained. Results have indicated the effectiveness of CNNs and ConvGRUs for automating breast cancer detection. Proposed methodology significantly surpasses baseline methods such as EfficientNetB7, DenseNet121, ConvNeXtTiny, XGBoost Ensemble, Soft Voting Ensemble in terms of classification accuracy and discriminative power. These findings support the model's applicability in clinical settings and highlight future directions in multi-class tumor classification and model optimization for deployment in resource-constrained clinical environments.

References

1.
Tran D, Bourdev L, Fergus R, Torresani L, Paluri M. Learning Spatiotemporal Features with 3D Convolutional Networks. 2015 IEEE International Conference on Computer Vision (ICCV). IEEE; 2015. p. 4489–97.
2.
Feichtenhofer C, Fan H, Malik J, He K. SlowFast Networks for Video Recognition. 2019 IEEE/CVF International Conference on Computer Vision (ICCV). IEEE; 2019. p. 6201–10.
3.
Tran D, Wang H, Torresani L, Ray J, LeCun Y, Paluri M. A Closer Look at Spatiotemporal Convolutions for Action Recognition. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE; 2018. p. 6450–9.
4.
Zhou B, Andonian A, Oliva A, Torralba A. Temporal Relational Reasoning in Videos. Lecture Notes in Computer Science. Springer International Publishing; 2018. p. 831–46.
5.
Bertasius G, Wang H, Torresani L. Is space-time attention all you need for video understanding?. In Icml 2021 Jul 18 (Vol. 2, No. 3, p. 4).

Citation

This is an open access article distributed under the  Creative Commons Attribution Non-Commercial License (CC BY-NC) License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. 

Article metrics

Google scholar: See link

Issue image
Issue 36, 2026
See full issue

Citations

Crossref Logo

0

The statements, opinions and data contained in the journal are solely those of the individual authors and contributors and not of the publisher and the editor(s). We stay neutral with regard to jurisdictional claims in published maps and institutional affiliations.