TY - JOUR
T1 - Diff-Oracle
T2 - Learning Styles and Contents to Augment Realistic Oracle Characters in Diffusion Model
AU - Li, Jing
AU - Wang, Qiufeng
AU - Wang, Siyuan
AU - Zhang, Rui
AU - Huang, Kaizhu
AU - Cambria, Erik
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/21
Y1 - 2026/6/21
N2 - Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at https://github.com/JJJingLi/Diff-Oracle.
AB - Recognizing oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains because of the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel multi-modal conditional diffusion model that generates a diverse range of controllable oracle characters by inputting random combinations of references. Given the challenge of accurately describing oracle character styles using natural language, Diff-Oracle departs from traditional diffusion models that rely primarily on text prompts by introducing a style encoder. This encoder extracts style prompts from existing oracle character images, where style details are converted into a text embedding format via a pre-trained language-vision model. Additionally, given the lack of explicit content information for oracle characters, ensuring that generated characters accurately represent the intended glyphs is challenging. Therefore, we pre-generate pixel-level paired oracle character images (i.e., style and content images) by an image-to-image translation model, providing content information for the generation process. Meanwhile, Diff-Oracle integrates a content encoder designed to capture specific content details from content reference images. Extensive experiments on Oracle-241 and OBC306 datasets demonstrate that Diff-Oracle significantly outperforms existing generative methods in image quality and diversity. Moreover, Diff-Oracle substantially benefits downstream recognition tasks, outperforming all existing state-of-the-art methods by a large margin. In particular, on the challenging OBC306 dataset, Diff-Oracle achieves a 7.70% accuracy gain in the zero-shot setting and reaches 84.62% accuracy for unseen oracle characters, setting a new benchmark for oracle character recognition. The code is available at https://github.com/JJJingLi/Diff-Oracle.
KW - diffusion models
KW - Oracle character generation
KW - oracle character recognition
UR - https://www.scopus.com/pages/publications/105043586791
U2 - 10.1145/3806389
DO - 10.1145/3806389
M3 - Article
AN - SCOPUS:105043586791
SN - 1551-6857
VL - 22
JO - ACM Transactions on Multimedia Computing, Communications and Applications
JF - ACM Transactions on Multimedia Computing, Communications and Applications
IS - 6
M1 - 152
ER -