Skip to main navigation Skip to search Skip to main content

ArchMap: Arch-Flattening and Knowledge-Guided Vision Language Model for Tooth Counting and Structured Dental Understanding

  • Bohan Zhang
  • , Yiyi Miao
  • , Taoyu Wu
  • , Tong Chen
  • , Ji Jiang
  • , Zhuoxiao Li
  • , Zhe Tang
  • , Limin Yu*
  • , Jionglong Su*
  • *Corresponding author for this work
  • Xi'an Jiaotong-Liverpool University
  • University of Liverpool
  • The Hong Kong University of Science and Technology (Guangzhou)
  • Zhejiang University of Technology

Research output: Chapter in Book or Report/Conference proceedingConference Proceedingpeer-review

Abstract

Digital orthodontics leverages 3D intraoral scanning and computational analysis to enable precise diagnosis, treatment planning, and outcome evaluation in a data-driven workflow. A structured understanding of intraoral 3D scans is essential for digital orthodontics. Such structured understanding converts raw geometric data into clinically interpretable representations, forming the foundation for reliable automated analysis throughout the orthodontic pipeline. However, existing deep-learning approaches rely heavily on modality-specific training, large annotated datasets, and controlled scanning conditions, which limit generalization across devices and hinder deployment in real clinical workflows. Moreover, raw intraoral meshes exhibit substantial variation in arch pose, incomplete geometry caused by occlusion or tooth contact, and a lack of texture cues, making unified semantic interpretation highly challenging. To address these limitations, we propose ArchMap, a training-free and knowledge-guided framework for robust structured dental understanding. ArchMap first introduces a geometry-aware archflattening module that standardizes raw 3D meshes into spatially aligned, continuity-preserving multi-view projections. We then construct a Dental Knowledge Base (DKB) encoding hierarchical tooth ontology, dentition-stage policies, and clinical semantics to constrain the symbolic reasoning space. Leveraging on this ontology, a schema-constrained vision-language inference pipeline transforms general-purpose VLMs into deterministic, contractcompliant structured predictors. We validate ArchMap on 1060 pre-/post-orthodontic cases, demonstrating robust performance in tooth counting, anatomical partitioning, dentition-stage classification, and the identification of clinical conditions such as crowding, missing teeth, prosthetics, and caries. Compared with supervised pipelines and prompted VLM baselines, ArchMap achieves higher accuracy, reduced semantic drift, and superior stability under sparse or artifact-prone conditions. As a fully training-free system, ArchMap demonstrates that combining geometric normalization with ontology-guided multimodal reasoning offers a practical and scalable solution for the structured analysis of 3D intraoral scans in modern digital orthodontics.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE International Conference on Big Data, BigData 2025
EditorsCheng-Zhong Xu, Leong Hou U, Xueqi Cheng, Jing Gao, Giuseppe Polese, Hong Mei, Paul Boniol, Michiaki Tatsubori, Chen Zhao, Dawei Zhou, Xiaohua Hu
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages7529-7538
Number of pages10
Edition2025
ISBN (Electronic)9798331594473
DOIs
Publication statusPublished - 2025
Event2025 IEEE International Conference on Big Data, BigData 2025 - Macau, China
Duration: 8 Dec 202511 Dec 2025

Conference

Conference2025 IEEE International Conference on Big Data, BigData 2025
Country/TerritoryChina
CityMacau
Period8/12/2511/12/25

Keywords

  • Dental Understanding
  • Tooth Counting
  • Visionlanguage Model

Cite this