Abstract
Scene text detection and recognition have attracted increasing research attention recently, especially for texts of arbitrary shapes. In most of text spotting methods, text feature alignment is a key component to connect the detector and the recognizer for end-to-end training. Existing alignment methods can be roughly categorized into those based on global consistent transformations and based on character-level classification. However, these methods either are unreliable for heavily deformed text or ignore contextual information in recognition. In this paper, we propose a novel text spotter named TextTriangle, which detects and recognizes the arbitrary-shaped text in an end-to-end manner without character-level annotations. In TextTriangle, a text instance is described as a sequence of ordered triangles attached to each other. Based on this representation, a new PiecewiseAlign layer is designed to accurately extract features of the text instance with arbitrary shapes, which is the key to make the framework end-to-end trainable. Compared with the methods based on global consistent transformations, PiecewiseAlign adopts piecewise linear transformation for feature calculation. Experiments show that PiecewiseAlign is superior to TPS-based method in the text alignment, and TextTriangle achieves competitive performance on standard scene text benchmarks.
| Original language | English |
|---|---|
| Journal | International Journal on Document Analysis and Recognition |
| DOIs | |
| Publication status | Accepted/In press - 2025 |
Keywords
- End-to-end training
- Scene text detection
- Scene text recognition
- Scene text spotting
- Text feature alignment
Fingerprint
Dive into the research topics of 'TextTriangle: an end-to-end textspotter with piecewise linear alignment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver