Low-data image classification remains a fundamental challenge in real-world machine learning, particularly for finegrained categories with high visual similarity and limited labeled data. We study this problem through a controlled benchmark on multi-class crop recognition from unconstrained field images, serving as a realistic testbed for low-data visual learning. We evaluate a diverse set of architectures, including classical convolutional neural networks (ResNet-18/50, MobileNetV2, EfficientNetB0, DenseNet-121), a modern convolutional design (ConvNeXtTiny), and transformer-based models (ViT-B/16, Swin-Tiny), under a unified training and evaluation protocol with multi-seed analysis. Beyond top-line accuracy, we incorporate macro-F1, confusion-matrix analysis, dominant misclassification pairs, and deployment-oriented metrics such as parameter count, FLOPs, and inference latency. Our results show that modern architectures substantially outperform earlier CNN baselines in this low-data regime, with ConvNeXt-Tiny achieving the strongest overall performance, while transformer-based models remain competitive and demonstrate strong transfer capability. At the same time, lightweight models such as ResNet-18 remain attractive under strict latency constraints. These findings provide practical guidance for model selection in low-data fine-grained recognition and highlight the importance of multi-seed, error-aware, and efficiency-aware evaluation in applied machine learning.
Low-Data Fine-Grained Crop Classification: A Multi-Seed Benchmark of CNN and Transformer Models
Imran, Faisal
Writing – Original Draft Preparation
;Albarelli, AndreaSupervision
;Torsello, AndreaSupervision
;Gasparetto, AndreaValidation
;
2026
Abstract
Low-data image classification remains a fundamental challenge in real-world machine learning, particularly for finegrained categories with high visual similarity and limited labeled data. We study this problem through a controlled benchmark on multi-class crop recognition from unconstrained field images, serving as a realistic testbed for low-data visual learning. We evaluate a diverse set of architectures, including classical convolutional neural networks (ResNet-18/50, MobileNetV2, EfficientNetB0, DenseNet-121), a modern convolutional design (ConvNeXtTiny), and transformer-based models (ViT-B/16, Swin-Tiny), under a unified training and evaluation protocol with multi-seed analysis. Beyond top-line accuracy, we incorporate macro-F1, confusion-matrix analysis, dominant misclassification pairs, and deployment-oriented metrics such as parameter count, FLOPs, and inference latency. Our results show that modern architectures substantially outperform earlier CNN baselines in this low-data regime, with ConvNeXt-Tiny achieving the strongest overall performance, while transformer-based models remain competitive and demonstrate strong transfer capability. At the same time, lightweight models such as ResNet-18 remain attractive under strict latency constraints. These findings provide practical guidance for model selection in low-data fine-grained recognition and highlight the importance of multi-seed, error-aware, and efficiency-aware evaluation in applied machine learning.I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



