TY - GEN
T1 - Token-Level Contrastive Learning for Open-World Weakly-Supervised Object Localization
AU - Li, Rouyi
AU - Luo, Zhaochuan
AU - Zhuo, Wei
AU - Shen, Linlin
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Open-World Weakly-Supervised Object Localization (OWSOL) targets the recognition and localization of both known and novel categories in open-world situations. The pioneering work primarily tackles the problem by proposing a generalized representation learning paradigm, which lacks targeted optimization for Vision Transformers (ViT). We argue that long-range visual dependencies of ViT are specialized in complete perception of objects. To better adapt ViT into OWSOL, in this work, we propose a Token-level Contrastive Learning (ToCL) framework. It mainly contains supervised contrastive learning on labeled data and, semantics-driven token-level contrastive learning on labeled and unlabeled ones. Specifically, contrastive learning is performed on both class and patch tokens, which learns complementarily for fine-grained semantics. Besides, tokens of foreground and background are learned to distribute apart by contrast. The above operations potentially enable self-attentions of ViT to accurately and completely focus on the target object regions. Extensive experiments on ImageNet-1K, iNatLoc500, and OpenImages150 datasets show that our method outperforms the state-of-the-art methods by a large margin. Code will be released.
AB - Open-World Weakly-Supervised Object Localization (OWSOL) targets the recognition and localization of both known and novel categories in open-world situations. The pioneering work primarily tackles the problem by proposing a generalized representation learning paradigm, which lacks targeted optimization for Vision Transformers (ViT). We argue that long-range visual dependencies of ViT are specialized in complete perception of objects. To better adapt ViT into OWSOL, in this work, we propose a Token-level Contrastive Learning (ToCL) framework. It mainly contains supervised contrastive learning on labeled data and, semantics-driven token-level contrastive learning on labeled and unlabeled ones. Specifically, contrastive learning is performed on both class and patch tokens, which learns complementarily for fine-grained semantics. Besides, tokens of foreground and background are learned to distribute apart by contrast. The above operations potentially enable self-attentions of ViT to accurately and completely focus on the target object regions. Extensive experiments on ImageNet-1K, iNatLoc500, and OpenImages150 datasets show that our method outperforms the state-of-the-art methods by a large margin. Code will be released.
UR - https://www.scopus.com/pages/publications/105028429004
U2 - 10.1007/978-981-95-5755-4_22
DO - 10.1007/978-981-95-5755-4_22
M3 - Conference contribution
AN - SCOPUS:105028429004
SN - 9789819557547
T3 - Lecture Notes in Computer Science
SP - 321
EP - 335
BT - Pattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
A2 - Kittler, Josef
A2 - Xiong, Hongkai
A2 - Lin, Weiyao
A2 - Yang, Jian
A2 - Chen, Xilin
A2 - Lu, Jiwen
A2 - Yu, Jingyi
A2 - Zheng, Weishi
PB - Springer Science and Business Media Deutschland GmbH
T2 - 8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Y2 - 15 October 2025 through 18 October 2025
ER -