Intermediate layer configurations for image similarity: evaluating multi-level feature extraction across CNN and transformer architectures

Authors

  • Volodymyr Kubytskyi Taras Shevchenko National University of Kyiv
  • Artem Bozhok Taras Shevchenko National University of Kyiv https://orcid.org/0009-0009-9572-6501

DOI:

https://doi.org/10.17721/1812-5409.2026/1.28

Keywords:

multi-level feature aggregation, image similarity, convolutional neural networks, vision transformers, intermediate representations, near-duplicate detection, architectural comparison, ResNet-50

Abstract

Evaluating visual similarity between images requires vector representations that preserve information across multiple levels of abstraction. Multi-level feature aggregation from intermediate CNN layers has shown good results in near-duplicate detection tasks, yet no structured analysis exists of how architectural properties of different backbones affect the quality of such multi-level representations. This study addresses the question of which structural characteristics make an architecture suitable for intermediate layer aggregation.

We formalize the multi-level aggregation pipeline as Global Average Pooling applied to intermediate activations, followed by concatenation and L2-normalization, producing a fixed-length descriptor. A compact MLP classifier is trained on the resulting pairwise representations. We evaluate six backbone architectures – VGG-16, Inception v3, DenseNet-121, MobileNetV2, ResNet-50, and ViT-B/16 – analyzing their stage structure, spatial resolution progression, intermediate feature dimensionality, and information flow properties. Quantitative evaluation is performed on the INRIA Holidays dataset using F1-score for pairwise similarity classification. ResNet-50 with layers C2+C3+C5 achieves F1 = 0.77, outperforming all other configurations. VGG-16 multi-level aggregation yields F1 = 0.72 but suffers from parameter inefficiency (138M vs. 23M). DenseNet-121 reaches F1 = 0.74 despite feature redundancy from dense connections. ViT-B/16 intermediate blocks produce F1 = 0.70, limited by homogeneous block structure without clear spatial hierarchy. MobileNetV2 and Inception v3 show F1 = 0.68 and F1 = 0.71 respectively.

Clear stage hierarchy with residual connections, progressive spatial-to-semantic transition, and moderate intermediate dimensionality are the key architectural requirements for effective multi-level feature aggregation.

Pages of the article in the issue: 213 - 217

Language of the article: English

Author Biographies

  • Volodymyr Kubytskyi, Taras Shevchenko National University of Kyiv

    PhD Student

  • Artem Bozhok, Taras Shevchenko National University of Kyiv

    PhD Student

References

Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16×16 words: Transformers for image recognition at scale. In Proceedings of ICLR 2021. arXiv:2010.11929

Gkelios, S., Boutalis, Y., & Chatzichristofis, S. A. (2021). Investigating the vision transformer model for image retrieval tasks. In Proceedings of IEEE MMSP 2021. https://doi.org/10.1109/MMSP53017.2021.9733553

Hariharan, B., Arbeláez, P., Girshick, R., & Malik, J. (2015). Hypercolumns for object localization and fine-grained localization. In Proceedings of CVPR 2015 (pp. 447–456).

He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of CVPR 2016 (pp. 770–778).

Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of CVPR 2017 (pp. 4700–4708).

Jégou, H., Douze, M., & Schmid, C. (2008). Hamming embedding and weak geometric consistency for large scale image search. In ECCV 2008 (pp. 304–317). Springer.

Kordopatis-Zilos, G., Papadopoulos, S., Patras, I., & Kompatsiaris, Y. (2017). Near-duplicate video retrieval by aggregating intermediate CNN layers. In Multimedia Modeling. Springer.

Kubytskyi, V., & Panchenko, T. (2023). Enriched image embeddings as a combined outputs from different layers of CNN for various image similarity problems. In Lecture Notes on Data Engineering and Communications Technologies (Vol. 180, pp. 321–333). Springer. https://doi.org/10.1007/978-3-031-36115-9_30

Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of CVPR 2017 (pp. 2117–2125).

Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of ICCV 2021 (pp. 10012–10022).

Panchenko, T., Bozhok, A., & Kubytskyi, V. (2026). Multi-level CNN feature fusion from ResNet50 for near-duplicate image detection in real estate imagery. Informatica, 50(9). https://doi.org/10.31449/inf.v50i9.12111

Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of ICML 2021 (Vol. 139, pp. 8748–8763). PMLR.

Razavian, A. S., Azizpour, H., Sullivan, J., & Carlsson, S. (2014). CNN features off-the-shelf: An astounding baseline for recognition. In Proceedings of CVPRW 2014 (pp. 806–813).

Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L.-C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of CVPR 2018 (pp. 4510–4520).

Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.

Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., & Rabinovich, A. (2015). Going deeper with convolutions. In Proceedings of CVPR 2015 (pp. 1–9).

Thyagharajan, K. K., & Kalaiarasi, G. A. (2021). A review on near-duplicate detection of images using computer vision techniques. Archives of Computational Methods in Engineering, 28(3), 897–916. https://doi.org/10.1007/s11831-020-09400-w

Downloads

Published

2026-06-05

Issue

Section

Computer Science and Informatics

How to Cite

Kubytskyi, V., & Bozhok, A. (2026). Intermediate layer configurations for image similarity: evaluating multi-level feature extraction across CNN and transformer architectures. Bulletin of Taras Shevchenko National University of Kyiv. Physics and Mathematics, 82(1), 213-217. https://doi.org/10.17721/1812-5409.2026/1.28