TY - JOUR
T1 - ON-DEMAND VEHICULAR SERVICE DEPLOYMENT IN 6G
T2 - A Collaborative Large-Small LM Architecture
AU - Ghebreziabiher, Amine Kidane
AU - Boateng, Gordon Owusu
AU - Ayepah-Mensah, Daniel
AU - Mizouni, Rabeb
AU - Mourad, Azzam
AU - Otrok, Hadi
AU - Bentahar, Jamal
AU - Muhaidat, Sami
N1 - Publisher Copyright:
© 2005-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Large language models (LLMs) offer significant potential for enabling on-demand service deployment in intelligent transportation systems (ITSs). However, their burgeoning size and high computational requirements limit their feasibility in 6G vehicular networks. To address these challenges, this article presents L2S-LM, a novel collaborative LLM-small language model (SLM) architecture for on-demand vehicular service deployment in 6G networks. The architecture follows a modular perception–prediction–placement pipeline, distributing tasks across the cloud, edge, and end layers. For perception, a cloud-based multimodal LLM (MLLM) generates semantic scene representations as model inferences, which a lightweight edge-deployed SLM interprets to predict appropriate on-demand services for vehicular traffic events. To enhance efficiency, a memory augmentation mechanism retrieves relevant historical predictions, reducing redundant perception computations. Finally, the placement module employs LM-assisted optimization to select the best node and allocate resources for service deployment. Experimental results demonstrate that L2S-LM achieves performance comparable to cloud-based LLM-only architecture while minimizing resource consumption.
AB - Large language models (LLMs) offer significant potential for enabling on-demand service deployment in intelligent transportation systems (ITSs). However, their burgeoning size and high computational requirements limit their feasibility in 6G vehicular networks. To address these challenges, this article presents L2S-LM, a novel collaborative LLM-small language model (SLM) architecture for on-demand vehicular service deployment in 6G networks. The architecture follows a modular perception–prediction–placement pipeline, distributing tasks across the cloud, edge, and end layers. For perception, a cloud-based multimodal LLM (MLLM) generates semantic scene representations as model inferences, which a lightweight edge-deployed SLM interprets to predict appropriate on-demand services for vehicular traffic events. To enhance efficiency, a memory augmentation mechanism retrieves relevant historical predictions, reducing redundant perception computations. Finally, the placement module employs LM-assisted optimization to select the best node and allocate resources for service deployment. Experimental results demonstrate that L2S-LM achieves performance comparable to cloud-based LLM-only architecture while minimizing resource consumption.
UR - https://www.scopus.com/pages/publications/105038968093
U2 - 10.1109/MVT.2026.3683168
DO - 10.1109/MVT.2026.3683168
M3 - Article
AN - SCOPUS:105038968093
SN - 1556-6072
JO - IEEE Vehicular Technology Magazine
JF - IEEE Vehicular Technology Magazine
ER -