Hybridtransformer: Multi-Feature Token Fusion of Deep Cnn Features and Handcrafted Descriptors for White Blood Cell Classification

dc.contributor.authorHoşavcı, Reyhan
dc.contributor.authorDik, Sümeyye Zülal
dc.contributor.authorAkçelik, Zeliha Kaya
dc.contributor.authorArar, Mahmud Esad
dc.contributor.authorAram, Kadir
dc.contributor.authorKaya, Samet
dc.contributor.authorKuş, Zeki
dc.contributor.authorAydın, Musa
dc.date.accessioned2026-08-21T14:20:27Z
dc.date.issued2026
dc.departmentFSM Vakıf Üniversitesi, Mühendislik Fakültesi, Bilgisayar Mühendisliği Bölümü
dc.departmentFSM Vakıf Üniversitesi, Mühendislik Fakültesi, Elektrik-Elektronik Mühendisliği Bölümü
dc.departmentFSM Vakıf Üniversitesi, Mühendislik Fakültesi, Yapay Zeka ve Veri Mühendisliği Bölümü
dc.description.abstractAccurate white blood cell (WBC) classification is essential for diagnosing hematological diseases, yet it remains a challenging fine-grained visual recognition problem due to high intra-class variability, inter-class similarity among morphologically adjacent subtypes, and staining variability across imaging conditions. Existing CNN-based approaches, while effective, rely predominantly on deep learned features. Although prior hybrid studies combine them with handcrafted descriptors, such integration has largely been limited to feature concatenation or classifier-level fusion rather than tokenlevel cross-modal attention. This paper proposes HybridTransformer, a multi-modal Transformer-based framework that fuses deep CNN features with handcrafted descriptors within a unified token sequence. A pretrained EfficientNetV2 backbone extracts a 1280-dimensional feature vector, which is spatially partitioned into four sub-tokens. Three handcrafted descriptors, Local Binary Pattern (LBP) histograms, HSV color histograms, and Gabor filter statistics, are computed as complementary tokens encoding microtexture, staining color distribution, and multi-scale structural patterns, respectively. All tokens are projected into a shared embedding space and processed by a Transformer encoder, where multi-head selfattention enables data-driven cross-modal interaction. The framework is evaluated on the MLL23 dataset, a challenging 18-class peripheral blood cell benchmark comprising 41,621 expert-annotated images. HybridTransformer achieves 95.3% accuracy, 95.2% F1-score, 95.2% precision, and 95.2% recall, outperforming all five standalone CNN baselines and a Vision Transformer baseline. Systematic ablation studies confirm that deep CNN features are the dominant contributor, while each handcrafted descriptor provides consistent incremental gains. The proposed framework demonstrates that token-level multi-modal fusion within a Transformer architecture is an effective strategy for fine-grained hematological image classification.
dc.identifier.citationHOŞAVCI, Reyhan, Sümeyye Zülal DİK, Zeliha Kaya AKÇELİK, Mahmud Esad ARAR, Kadir ARAM, Samet KAYA, Zeki KUŞ & Musa AYDIN. “Hybridtransformer: Multi-Feature Token Fusion of Deep Cnn Features and Handcrafted Descriptors for White Blood Cell Classification”. Multimedia Systems, 32.7 (2026): 1-21.
dc.identifier.doi10.1007/s00530-026-02571-9
dc.identifier.endpage21
dc.identifier.issue7
dc.identifier.startpage1
dc.identifier.urihttps://hdl.handle.net/11352/6241
dc.identifier.volume32
dc.identifier.wosWOS:001832733000002
dc.identifier.wosqualityQ2
dc.indekslendigikaynakWeb of Science
dc.language.isoen
dc.publisherSpringer
dc.relation.ispartofMultimedia Systems
dc.relation.publicationcategoryMakale - Uluslararası Hakemli Dergi - Kurum Öğretim Elemanı
dc.rightsinfo:eu-repo/semantics/embargoedAccess
dc.subjectPeripheral Blood Cell Classification
dc.subjectTransformer Encoder
dc.subjectHandcrafted Descriptors
dc.subjectFine-Grained Visual Recognition
dc.subjectHematological Image Analysis
dc.subjectToken-Level Fusion
dc.titleHybridtransformer: Multi-Feature Token Fusion of Deep Cnn Features and Handcrafted Descriptors for White Blood Cell Classification
dc.typeArticle

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
Hoşavcı.pdf
Boyut:
2.45 MB
Biçim:
Adobe Portable Document Format

Lisans paketi

Listeleniyor 1 - 1 / 1
Yükleniyor...
Küçük Resim
İsim:
license.txt
Boyut:
1.17 KB
Biçim:
Item-specific license agreed upon to submission
Açıklama: