Gallbladder Cancer Classification using Parallel Transfer Learning with Multi-model Feature Fusion and LSTM
Hybrid parallel transfer learning + LSTM framework for gallbladder cancer classification, achieving 99.37% accuracy.
This thesis presents a deep learning model for gallbladder cancer classification from ultrasound images, developed at KUET under the supervision of Dr. Mostafa Zaman Chowdhury, Professor, Dept. of EEE, KUET (February 2024).
Gallbladder cancer is frequently diagnosed at an advanced stage with poor prognosis. Ultrasound imaging — the primary diagnostic modality — suffers from speckle noise and low contrast, making accurate classification difficult. This work proposes a parallel transfer learning architecture integrating VGG16, VGG19, XceptionNet, and ResNet50 with two LSTM layers for multi-model feature fusion and sequential pattern recognition.
Dataset
The publicly available GBCU (Gallbladder Cancer Ultrasound) dataset was used — 1,255 images across three classes: Normal (432), Benign (558), and Malignant (265). Images were split 70% training, 20% validation, 10% testing.
Preprocessing Pipeline
A custom preprocessing pipeline was developed to enhance ultrasound image quality before model input:
- Image Resizing — standardized to 512×512 pixels
- CLAHE (Contrast Limited Adaptive Histogram Equalization) — enhances local contrast, improving visibility of tissue boundaries concealed by speckle noise
- Laplacian Sharpening — emphasizes edges and fine structural detail critical for differentiating benign from malignant tissue
Model Architecture
Features are extracted in parallel from four frozen pre-trained CNN backbones (ImageNet weights):
| Backbone | Feature Extraction Layer | Output Shape |
|---|---|---|
| VGG16 | block5_conv3 | 8×8×512 |
| VGG19 | block5_conv3 | 8×8×512 |
| ResNet50 | conv5_block3_out | 4×4×2048 |
| XceptionNet | block14_sepconv2_act | 4×4×2048 |
Flattened features from VGG16+XceptionNet and VGG19+ResNet50 are concatenated into two separate streams, reshaped into sequences of length 8, and each passed through an LSTM layer (256 units). The two LSTM outputs are concatenated, then passed through Dense (256) → Dropout → Dense (128) → Dropout → Dense (64) → Dense (3, softmax).
Hyperparameters: Input 128×128×3 · Optimizer: Adamax · Loss: Sparse Categorical Cross-Entropy · Batch size 32 · Early stopping (patience 10) · 50 epochs
Results
5-fold cross-validation was used for robust evaluation on the GBCU dataset:
| Metric | Mean | Std Dev |
|---|---|---|
| Accuracy | 99.37% | 0.35% |
| F1-Score | 99.52% | 0.49% |
| Sensitivity | 99.64% | 0.33% |
| Specificity | 99.69% | 0.69% |
| Cohen's Kappa | 99.22% | 0.71% |
| ROC AUC | 100.00% | 0.00% |
Comparison with State-of-the-Art (GBCU Dataset)
| Method | Accuracy | Sensitivity | Specificity |
|---|---|---|---|
| Radiologist A | 70.0% | 70.7% | 87.3% |
| Radiologist B | 68.3% | 73.2% | 81.1% |
| GBCNet (CVPR 2022) | 87.7% | 91.9% | 96.7% |
| RadFormer (2023) | 90.2% | 92.9% | 90.0% |
| This Work | 99.37% | 99.64% | 99.69% |
Related Publication
As part of this thesis, an ensemble study exploring average combinations of VGG16, VGG19, XceptionNet, and ResNet50 was published at ICEEICT 2024 (best result: VGG19+XceptionNet, 85.44%). This motivated the full LSTM-based architecture presented in the thesis.