When Pixels Lie: The Hidden Toll of Concept Drift

Authors

  • Gurgen Hovakimyan NOVA IMS-Information Management School, New University of Lisbon, 1070-312, Lisbon, Portugal , Centro de Investigação em Gestão de Informação (MagIC), 1070-312, Lisbon, Portugal Author https://orcid.org/0009-0002-4292-0166
  • Jorge Miguel Bravo NOVA IMS‐Information Management School, New University of Lisbon, 1070‐312, Lisbon, Portugal , Centro de Investigação em Gestão de Informação (MagIC), 1070‐312, Lisbon, Portugal Author https://orcid.org/0000-0002-7389-5103

DOI:

https://doi.org/10.31181/jopi41202670

Keywords:

Concept drift, Image classification, Model explainability, Grad-cam, Statistical drift detection

Abstract

Concept drift poses a major challenge to the reliability and transparency of computer vision systems. This study provides a systematic analysis of how four drift types, such as noise, blur, rotation, and brightness, affect both model performance and Grad-CAM/Grad-CAM++ explainability across CIFAR-10 and Tiny ImageNet datasets using three CNN architectures (ResNet-18, DenseNet-121, ShuffleNet v2). Following recent robustness literature, we treat these controlled test-time transformations as synthetic covariate-shift proxies for the input-distribution component of drift rather than as temporal concept drift in the strict sense. For CIFAR-10, accuracy declines from a baseline of 79–86% to as low as 24% under strong noise drift, while SSIM and IoU between baseline and drifted heatmaps fall from nearly 100% to as low as 18% and 25%, respectively. On Tiny ImageNet, the degradation is even more pronounced: top-1 accuracy collapses from 40–60% to below 4% under the strongest noise drift, with SSIM and IoU dropping to as low as 10% (per-architecture values are reported in the Results and Discussion section). Rotation and brightness prove markedly milder than noise and blur, indicating that high-frequency corruptions are the principal threat to both accuracy and attribution stability. Image-level paired tests (Wilcoxon signed-rank and paired t-test) with effect sizes and bootstrap confidence intervals, complemented by pixel-level distributional diagnostics (KS, Anderson-Darling, Wilcoxon rank-sum), confirm significant shifts in heatmap distributions, revealing that drift alters not only model outputs but also internal attribution patterns. The results demonstrate that different drift mechanisms degrade interpretability in distinct ways and that explainability metrics are often more sensitive to drift than accuracy itself. This work provides, to the best of our knowledge, one of the first unified, quantitative evaluations of explainability degradation under controlled visual drift, highlighting the need for adaptive, drift-aware XAI pipelines in dynamic real-world environments.

Downloads

Download data is not yet available.

References

Gama, J., Zliobaite, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), Article 44. https://doi.org/10.1145/2523813

Widmer, G., & Kubat, M. (1996). Learning in the presence of concept drift and hidden contexts. Machine Learning, 23(1), 69–101. https://doi.org/10.1007/BF00116900

Lu, J., Liu, A., Dong, F., Gu, F., Gama, J., & Zhang, G. (2018). Learning under concept drift: A review. IEEE Transactions on Knowledge and Data Engineering, 31(12), 2346–2363. https://doi.org/10.1109/TKDE.2018.2876857

Eliwa, E. H. I., & Abd El-Hafeez, T. (2025). Deep learning for sustainable agriculture: Automating rice and paddy ripeness classification for enhanced food security. Egyptian Informatics Journal, 26, 100785. https://doi.org/10.1016/j.eij.2025.100785

Garcia-Rubio, A. M., Heck, L., Schweitzer, K. M., Bateman, R. M., Alaeddini, A., & Castillo-Villar, K. K. (2026). Detecting concept drift in object detection models: A collaborative AI-human approach for defense applications. IEEE Access. Advance online publication. https://doi.org/10.1109/ACCESS.2026.3652106

Liu, C., Tang, K., Qin, Y., & Lei, Q. (2025). Bridging distribution shift and AI safety: Conceptual and methodological synergies. arXiv preprint arXiv:2505.22829. https://arxiv.org/abs/2505.22829

Brzezinski, D., & Stefanowski, J. (2014). Combining block-based and online methods in learning ensembles from concept drifting data streams. Information Sciences, 265, 50–67. https://doi.org/10.1016/j.ins.2013.12.011

Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. (2017). Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (pp. 618–626). https://doi.org/10.1109/ICCV.2017.74

Chattopadhay, A., Sarkar, A., Howlader, P., & Balasubramanian, V. N. (2018). Grad-CAM++: Generalized gradient-based visual explanations for deep convolutional networks. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) (pp. 839–847). IEEE. https://doi.org/10.1109/WACV.2018.00097

Omeiza, D., Speakman, S., Cintas, C., & Weldermariam, K. (2019). Smooth Grad-CAM++: An enhanced inference level visualization technique for deep convolutional neural network models. arXiv preprint arXiv:1908.01224. https://arxiv.org/abs/1908.01224

Hinder, F., Vaquet, V., Brinkrolf, J., & Hammer, B. (2023). Model-based explanations of concept drift. Neurocomputing, 555, 126640. https://doi.org/10.1016/j.neucom.2023.126640

Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., & Kim, B. (2018). Sanity checks for saliency maps. In Advances in Neural Information Processing Systems 31.

Ross, G. J., Tasoulis, D. K., & Adams, N. M. (2011). Nonparametric monitoring of data streams for changes in location and scale. Technometrics, 53(4), 379–389. https://doi.org/10.1198/TECH.2011.10069

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778). https://doi.org/10.1109/CVPR.2016.90

Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 4700–4708). https://doi.org/10.1109/CVPR.2017.243

Ma, N., Zhang, X., Zheng, H.-T., & Sun, J. (2018). ShuffleNet V2: Practical guidelines for efficient CNN architecture design. In Proceedings of the European Conference on Computer Vision (ECCV) (pp. 116–131). https://doi.org/10.1007/978-3-030-01264-9_8

Hovakimyan, G., & Bravo, J. M. (2024). Evolving strategies in machine learning: A systematic review of concept drift detection. Information, 15(12), 786. https://doi.org/10.3390/info15120786

Baena-García, M., del Campo-Ávila, J., Fidalgo, R., Bifet, A., Gavalda, R., & Morales-Bueno, R. (2006). Early drift detection method. In Fourth International Workshop on Knowledge Discovery from Data Streams.

Bifet, A., & Gavalda, R. (2007). Learning from time-changing data with adaptive windowing. In Proceedings of the 2007 SIAM International Conference on Data Mining. https://doi.org/10.1137/1.9781611972771.42

Raab, C., Heusinger, M., & Schleif, F.-M. (2020). Reactive soft prototype computing for concept drift streams. Neurocomputing, 416, 340–351. https://doi.org/10.1016/j.neucom.2019.11.111

Greco, S., Vacchetti, B., Apiletti, D., & Cerquitelli, T. (2025). Unsupervised concept drift detection from deep learning representations in real-time. IEEE Transactions on Knowledge and Data Engineering. Advance online publication. https://doi.org/10.1109/TKDE.2025.3593123

Wiles, O., Gowal, S., Stimberg, F., Alvise-Rebuffi, S., Ktena, I., Dvijotham, K., & Cemgil, T. (2021). A fine-grained analysis on distribution shift. arXiv preprint arXiv:2110.11328. https://arxiv.org/abs/2110.11328

Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., & Brendel, W. (2018). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International Conference on Learning Representations.

Liu, Z., Lu, J., Xuan, J., & Zhang, G. (2025). Learning latent and changing dynamics in real non-stationary environments. IEEE Transactions on Knowledge and Data Engineering. Advance online publication. https://doi.org/10.1109/TKDE.2025.3535961

Abdullahi, M., Alhussian, H., Aziz, N., Abdulkadir, S. J., Baashar, Y., Alashhab, A. A., & Afrin, A. (2025). A systematic literature review of concept drift mitigation in time-series applications. IEEE Access. Advance online publication. https://doi.org/10.1109/ACCESS.2025.3587231

Dong, S., Wang, P., & Abbas, K. (2021). A survey on deep learning and its applications. Computer Science Review, 40, 100379. https://doi.org/10.1016/j.cosrev.2021.100379

O'Shea, K., & Nash, R. (2015). An introduction to convolutional neural networks. arXiv preprint arXiv:1511.08458. https://arxiv.org/abs/1511.08458

Smilkov, D., Thorat, N., Kim, B., Viegas, F., & Wattenberg, M. (2017). SmoothGrad: Removing noise by adding noise. arXiv preprint arXiv:1706.03825. https://arxiv.org/abs/1706.03825

Marmolejo-Saucedo, J. A., & Kose, U. (2024). Numerical Grad-CAM based explainable convolutional neural network for brain tumor diagnosis. Mobile Networks and Applications, 29(1), 1–14. https://doi.org/10.1007/s11036-022-02021-6

Teixeira, B., Pinto, T., & Vale, Z. (2025). Detecting concept drift with SHapley additive exPlanations for intelligent model retraining in energy generation forecasting. In World Conference on Explainable Artificial Intelligence (pp. 1–17). https://doi.org/10.1007/978-3-032-08324-1_7

Beemelmanns, T., Sharifi, S., Mehrotra, M., Choudhuri, A., & Eckstein, L. (2026). Towards trustworthy and explainable AI for perception models: From concept to prototype vehicle deployment. arXiv preprint arXiv:2605.16087. https://arxiv.org/abs/2605.16087

Jang, M., Han, S.-H., Choi, H.-J., & An, B.-W. (2025). Improving load forecasting with feature selection via XAI and expert-guided prompting. IEEE Access. Advance online publication. https://doi.org/10.1109/ACCESS.2025.3637813

Michaelis, C., Mitzkus, B., Geirhos, R., Rusak, E., Bringmann, O., Ecker, A. S., Bethge, M., & Brendel, W. (2019). Benchmarking robustness in object detection: Autonomous driving when winter is coming. arXiv preprint arXiv:1907.07484. https://arxiv.org/abs/1907.07484

Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet? In Proceedings of the International Conference on Machine Learning (pp. 5389–5400). PMLR. https://arxiv.org/abs/1902.10811

Krizhevsky, A., & Hinton, G. (2009). Learning multiple layers of features from tiny images. University of Toronto.

Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems 32.

Duchi, J., Hazan, E., & Singer, Y. (2011). Adaptive subgradient methods for online learning and stochastic optimization. Journal of Machine Learning Research, 12, 2121–2159.

Ruder, S. (2016). An overview of gradient descent optimization algorithms. arXiv preprint arXiv:1609.04747. https://arxiv.org/abs/1609.04747

Harshit, N., & Mounvik, K. (2025). Improving real-time concept drift detection using a hybrid transformer-autoencoder framework. arXiv preprint arXiv:2508.07085. https://doi.org/10.36227/techrxiv.175393691.16483231/v1

Piaseczny, A., Shisher, M. K. C., Wang, S., & Brinton, C. (2026). RCCDA: Adaptive model updates in the presence of concept drift under a constrained resource budget. In Advances in Neural Information Processing Systems. https://doi.org/10.52202/085713-5183

Published

2026-09-21

How to Cite

Hovakimyan, G., & Bravo, J. M. (2026). When Pixels Lie: The Hidden Toll of Concept Drift. Journal of Operations Intelligence, 4(1), 65-94. https://doi.org/10.31181/jopi41202670