Abstract
Glass segmentation is a challenging task due to its unique visual properties, such as transparency and reflectivity, which often result in incomplete segmentation and blurred boundaries. Existing methods typically trade off accuracy and efficiency, with high-accuracy models often too slow to meet the demands of real-time applications. To address this issue, we propose RTGlassNet, a novel end-to-end network designed specifically for real-time, high-fidelity glass segmentation. Our architecture employs a collaborative decoder consisting of two cooperating streams to effectively model semantic context and refine spatial details. The decoder is built upon two novel lightweight modules: the Asymmetric Context Integration Module (ACIM), which efficiently captures multi-scale context by revisiting a classic cascade architecture with depthwise separable convolutions, and the Edge Enhancement Unit (EEU), a parameter-free mechanism that adaptively enhances weak boundary signals within the feature space. At the foundation of our approach is a lightweight Differentiable Conditional Random Field (DiffCRF) module that innovatively applies separable convolutions to the message passing step, restoring pixel-perfect details using the original image as guidance for a final global optimization. Extensive experiments on benchmark datasets show that RTGlassNet achieves state-of-the-art accuracy and significantly outperforms previous methods while running at over 80 FPS on a single consumer-grade GPU, effectively resolving the persistent trade-off between accuracy and speed.
| Original language | English |
|---|---|
| Article number | 114181 |
| Journal | Pattern Recognition |
| Volume | 180 |
| Early online date | 6 Jun 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 6 Jun 2026 |
Keywords
- Boundary segmentation
- Differentiable Conditional Random Field (CRF)
- Glass segmentation
- Real-time segmentation
- Semantic segmentation
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver