Low-data image classification remains a fundamental challenge in real-world machine learning, particularly for finegrained categories with high visual similarity and limited labeled data. We study this problem through a controlled benchmark on multi-class crop recognition from unconstrained field images, serving as a realistic testbed for low-data visual learning. We evaluate a diverse set of architectures, including classical convolutional neural networks (ResNet-18/50, MobileNetV2, EfficientNetB0, DenseNet-121), a modern convolutional design (ConvNeXtTiny), and transformer-based models (ViT-B/16, Swin-Tiny), under a unified training and evaluation protocol with multi-seed analysis. Beyond top-line accuracy, we incorporate macro-F1, confusion-matrix analysis, dominant misclassification pairs, and deployment-oriented metrics such as parameter count, FLOPs, and inference latency. Our results show that modern architectures substantially outperform earlier CNN baselines in this low-data regime, with ConvNeXt-Tiny achieving the strongest overall performance, while transformer-based models remain competitive and demonstrate strong transfer capability. At the same time, lightweight models such as ResNet-18 remain attractive under strict latency constraints. These findings provide practical guidance for model selection in low-data fine-grained recognition and highlight the importance of multi-seed, error-aware, and efficiency-aware evaluation in applied machine learning.

Low-Data Fine-Grained Crop Classification: A Multi-Seed Benchmark of CNN and Transformer Models

Imran, Faisal
Writing – Original Draft Preparation
;
Albarelli, Andrea
Supervision
;
Torsello, Andrea
Supervision
;
Gasparetto, Andrea
Validation
;
2026

Abstract

Low-data image classification remains a fundamental challenge in real-world machine learning, particularly for finegrained categories with high visual similarity and limited labeled data. We study this problem through a controlled benchmark on multi-class crop recognition from unconstrained field images, serving as a realistic testbed for low-data visual learning. We evaluate a diverse set of architectures, including classical convolutional neural networks (ResNet-18/50, MobileNetV2, EfficientNetB0, DenseNet-121), a modern convolutional design (ConvNeXtTiny), and transformer-based models (ViT-B/16, Swin-Tiny), under a unified training and evaluation protocol with multi-seed analysis. Beyond top-line accuracy, we incorporate macro-F1, confusion-matrix analysis, dominant misclassification pairs, and deployment-oriented metrics such as parameter count, FLOPs, and inference latency. Our results show that modern architectures substantially outperform earlier CNN baselines in this low-data regime, with ConvNeXt-Tiny achieving the strongest overall performance, while transformer-based models remain competitive and demonstrate strong transfer capability. At the same time, lightweight models such as ResNet-18 remain attractive under strict latency constraints. These findings provide practical guidance for model selection in low-data fine-grained recognition and highlight the importance of multi-seed, error-aware, and efficiency-aware evaluation in applied machine learning.
2026
2026 11th International Conference on Machine Learning Technologies (ICMLT)
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/5126773
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact