TY - GEN
T1 - ReELFA
T2 - 2019 International Conference on Document Analysis and Recognition Workshops, ICDARW 2019
AU - Wang, Qingqing
AU - Jia, Wenjing
AU - He, Xiangjian
AU - Lu, Yue
AU - Blumenstein, Michael
AU - Huang, Ye
AU - Lyu, Shujing
N1 - Publisher Copyright:
© 2019 IEEE.
PY - 2019
Y1 - 2019
N2 - LSTM and attention mechanism have been widely used for scene text recognition. However, the existing LSTM-based recognizers usually convert 2D feature maps into 1D space by flattening or pooling operations, resulting in the neglect of spatial information of text images. Additionally, the attention drift problem, where models fail to align targets at proper feature regions, has a serious impact on the recognition performance of existing models. To tackle the above problems, in this paper, we propose a scene text Recognizer with Encoded Location and Focused Attention, i.e., ReELFA. Our ReELFA utilizes one-hot encoded coordinates to indicate the spatial relationship of pixels and character center masks to help focus attention on the right feature areas. Experiments conducted on the benchmarking datasets IIIT5K, SVT, CUTE and IC15 demonstrate that the proposed method achieves comparable performance on the regular, low-resolution and noisy text images, and outperforms state-of-the-art approaches on the more challenging curved text images.
AB - LSTM and attention mechanism have been widely used for scene text recognition. However, the existing LSTM-based recognizers usually convert 2D feature maps into 1D space by flattening or pooling operations, resulting in the neglect of spatial information of text images. Additionally, the attention drift problem, where models fail to align targets at proper feature regions, has a serious impact on the recognition performance of existing models. To tackle the above problems, in this paper, we propose a scene text Recognizer with Encoded Location and Focused Attention, i.e., ReELFA. Our ReELFA utilizes one-hot encoded coordinates to indicate the spatial relationship of pixels and character center masks to help focus attention on the right feature areas. Experiments conducted on the benchmarking datasets IIIT5K, SVT, CUTE and IC15 demonstrate that the proposed method achieves comparable performance on the regular, low-resolution and noisy text images, and outperforms state-of-the-art approaches on the more challenging curved text images.
KW - Attention drift
KW - Attention LSTM
KW - Center masks
KW - Encoded location
UR - https://www.scopus.com/pages/publications/85095363245
U2 - 10.1109/ICDARW.2019.40084
DO - 10.1109/ICDARW.2019.40084
M3 - Conference contribution
AN - SCOPUS:85095363245
T3 - 2019 International Conference on Document Analysis and Recognition Workshops, ICDARW 2019
SP - 71
EP - 76
BT - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 21 September 2019 through 22 September 2019
ER -