Maqolalar

CROSS-INTERACTION-BASED MULTIMODAL FEATURE COMPARISON FOR MOVING OBJECT IDENTIFICATION IN CROWDED VIDEO SCENES

Jild 5 son 17 (2026) 57-63

2026-05-22 Maqolalar CC BY 4.0 Open Access

Mualliflar

  • Shohruh Begmatov Tashkent University of Information Technologies named after Muhammad al-Khwarizmi Doctor of Philosophy (PhD) in Technical Sciences, Doctoral (DSc) student
  • Mukhriddin Arabboev Tashkent University of Information Technologies named after Muhammad al-Khwarizmi Doctor of Philosophy (PhD) in Technical Sciences, Doctoral (DSc) student
  • Akhram Nishanov Tashkent University of Information Technologies named after Muhammad al-Khwarizmi Doctor of Science in Technical Sciences, Professor

Annotatsiya

Identifying moving objects in crowded video scenes is difficult because appearance information alone may be unreliable. Different people or objects may have similar visual appearances, while the same object may appear differently due to pose variation, scale changes, partial occlusion, illumination variation, or low visibility. To address this problem, this paper presents a cross-interaction-based multimodal feature comparison method for moving object identification. The proposed method represents each moving object using several complementary modalities, including appearance, geometry, spatial position, context, reliability, and clothing-color features. These heterogeneous features are projected into a common latent space before comparison. For two candidate detections, modality-wise comparison features are constructed using element-wise multiplication and absolute difference. Then, a cross-interaction function learns relationships between modalities, and an MLP estimates the final similarity probability. The proposed method is especially useful in difficult cases such as occlusion, lost track recovery, candidate ambiguity, and object reappearance. Compared with simple feature concatenation, the cross-interaction approach enables the model to learn conditional relationships across modalities and improves the reliability of moving-object identification in crowded scenes.

Kalit soʻzlar:

Iqtiboslar

N. Wojke, A. Bewley, and D. Paulus, “Simple online and real-time tracking with a deep association metric,” in Proceedings of the IEEE International Conference on Image Processing, pp. 3645–3649, 2017.

M. Ye, J. Shen, G. Lin, T. Xiang, L. Shao, and S. C. H. Hoi, “Deep learning for person re-identification: A survey and outlook,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 6, pp. 2872–2893, 2022.

H. Luo, Y. Gu, X. Liao, S. Lai, and W. Jiang, “Bag of tricks and a strong baseline for deep person re-identification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019.

X. Zheng, J. Zhu, Y. Sun, and Z. Zheng, "Multimodal person re-identification based on transformer relation regularisation," Information Fusion, vol. 104, article 102128, 2024.

K. Jiang, T. Zhang, X. Liu, B. Qian, Y. Zhang, and F. Wu, “Cross-modality transformer for visible-infrared person re-identification,” in Proceedings of the European Conference on Computer Vision, pp. 480–496, 2022.

Oʻquvchilar geografiyasi

5 Koʻrishlar
0 PDF yuklab olishlar
1 Davlatlar

    Yuklab olishlar

    Nashr qilingan

    2026-05-22

    Son

    Boʻlim

    Maqolalar

    Iqtibos keltirish tartibi

    Shohruh, B., Mukhriddin, A., & Akhram, N. (2026). CROSS-INTERACTION-BASED MULTIMODAL FEATURE COMPARISON FOR MOVING OBJECT IDENTIFICATION IN CROWDED VIDEO SCENES. Zamonaviy Dunyoda Innovatsion Tadqiqotlar, 5(17), 57-63. https://www.in-academy.uz/index.php/ZDIT/article/view/49858
    Innovative Academy RSC
    Article metrics Views and PDF downloads
    5 Views
    0 Downloads