Intermediate layer configurations for image similarity: evaluating multi-level feature extraction across CNN and transformer architectures
DOI:
https://doi.org/10.17721/1812-5409.2026/1.28Keywords:
multi-level feature aggregation, image similarity, convolutional neural networks, vision transformers, intermediate representations, near-duplicate detection, architectural comparison, ResNet-50Abstract
Evaluating visual similarity between images requires vector representations that preserve information across multiple levels of abstraction. Multi-level feature aggregation from intermediate CNN layers has shown good results in near-duplicate detection tasks, yet no structured analysis exists of how architectural properties of different backbones affect the quality of such multi-level representations. This study addresses the question of which structural characteristics make an architecture suitable for intermediate layer aggregation.
We formalize the multi-level aggregation pipeline as Global Average Pooling applied to intermediate activations, followed by concatenation and L2-normalization, producing a fixed-length descriptor. A compact MLP classifier is trained on the resulting pairwise representations. We evaluate six backbone architectures – VGG-16, Inception v3, DenseNet-121, MobileNetV2, ResNet-50, and ViT-B/16 – analyzing their stage structure, spatial resolution progression, intermediate feature dimensionality, and information flow properties. Quantitative evaluation is performed on the INRIA Holidays dataset using F1-score for pairwise similarity classification. ResNet-50 with layers C2+C3+C5 achieves F1 = 0.77, outperforming all other configurations. VGG-16 multi-level aggregation yields F1 = 0.72 but suffers from parameter inefficiency (138M vs. 23M). DenseNet-121 reaches F1 = 0.74 despite feature redundancy from dense connections. ViT-B/16 intermediate blocks produce F1 = 0.70, limited by homogeneous block structure without clear spatial hierarchy. MobileNetV2 and Inception v3 show F1 = 0.68 and F1 = 0.71 respectively.
Clear stage hierarchy with residual connections, progressive spatial-to-semantic transition, and moderate intermediate dimensionality are the key architectural requirements for effective multi-level feature aggregation.
Pages of the article in the issue: 213 - 217
Language of the article: English
References
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An image is worth 16×16 words: Transformers for image recognition at scale. In Proceedings of ICLR 2021. arXiv:2010.11929
Gkelios, S., Boutalis, Y., & Chatzichristofis, S. A. (2021). Investigating the vision transformer model for image retrieval tasks. In Proceedings of IEEE MMSP 2021. https://doi.org/10.1109/MMSP53017.2021.9733553
Hariharan, B., Arbeláez, P., Girshick, R., & Malik, J. (2015). Hypercolumns for object localization and fine-grained localization. In Proceedings of CVPR 2015 (pp. 447–456).
He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of CVPR 2016 (pp. 770–778).
Huang, G., Liu, Z., Van Der Maaten, L., & Weinberger, K. Q. (2017). Densely connected convolutional networks. In Proceedings of CVPR 2017 (pp. 4700–4708).
Jégou, H., Douze, M., & Schmid, C. (2008). Hamming embedding and weak geometric consistency for large scale image search. In ECCV 2008 (pp. 304–317). Springer.
Kordopatis-Zilos, G., Papadopoulos, S., Patras, I., & Kompatsiaris, Y. (2017). Near-duplicate video retrieval by aggregating intermediate CNN layers. In Multimedia Modeling. Springer.
Kubytskyi, V., & Panchenko, T. (2023). Enriched image embeddings as a combined outputs from different layers of CNN for various image similarity problems. In Lecture Notes on Data Engineering and Communications Technologies (Vol. 180, pp. 321–333). Springer. https://doi.org/10.1007/978-3-031-36115-9_30
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of CVPR 2017 (pp. 2117–2125).
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., & Guo, B. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of ICCV 2021 (pp. 10012–10022).
Panchenko, T., Bozhok, A., & Kubytskyi, V. (2026). Multi-level CNN feature fusion from ResNet50 for near-duplicate image detection in real estate imagery. Informatica, 50(9). https://doi.org/10.31449/inf.v50i9.12111
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., & Sutskever, I. (2021). Learning transferable visual models from natural language supervision. In Proceedings of ICML 2021 (Vol. 139, pp. 8748–8763). PMLR.
Razavian, A. S., Azizpour, H., Sullivan, J., & Carlsson, S. (2014). CNN features off-the-shelf: An astounding baseline for recognition. In Proceedings of CVPRW 2014 (pp. 806–813).
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., & Chen, L.-C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of CVPR 2018 (pp. 4510–4520).
Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., & Rabinovich, A. (2015). Going deeper with convolutions. In Proceedings of CVPR 2015 (pp. 1–9).
Thyagharajan, K. K., & Kalaiarasi, G. A. (2021). A review on near-duplicate detection of images using computer vision techniques. Archives of Computational Methods in Engineering, 28(3), 897–916. https://doi.org/10.1007/s11831-020-09400-w
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Volodymyr Kubytskyi, Artem Bozhok

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).
