Skip to main navigation Skip to search Skip to main content

平行智能范式视角下的视觉-语言-动作模型发展现状与展望

Translated title of the contribution: Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives
  • Bai Li*
  • , Jindi Hao
  • , Yueshuo Sun
  • , Yuqing Meng
  • , Jun Huang
  • , Yonglin Tian
  • , Zhengbing He
  • *Corresponding author for this work

Research output: Journal PublicationArticlepeer-review

3 Citations (Scopus)

Abstract

Vision-language-action (VLA) models are a comprehensive modeling approach for embodied intelligence that integrates visual perception, natural language understanding, and action execution within a unified framework, aiming to establish a continuous loop from environmental perception to task planning and action control. Their operational logic corresponds closely to the paradigm of parallel intelligence articulated in the early 21st century. That paradigm comprises artificial systems, computational experiments, and parallel execution, emphasizing virtual modeling, reproducible inference, and closed-loop interaction between the virtual and the real. The initial stage of VLA development, driven by multimodal deep learning, can be regarded as prototypical work within artificial systems, the subsequent stage characterized by large-scale models and cross-domain training expanded the scope of computational experiments, and the more recent focus on hierarchical control and virtual-real closed loops reflects the feedback correction and normative guidance emphasized in parallel execution. VLA models exhibit deep coupling between semantics and action, iterative cycles linking simulation with reality, and steadily improving verifiability. Nonetheless, challenges remain in generalization, semantic alignment, safety and interpretability, and deployment efficiency. Addressing these issues calls for contract-based task semantics, repairable long-horizon hierarchical planning, engineering-oriented use of world models, multi-level feedback and safety governance, and cross-platform transfer with human-machine collaboration. Examining VLA through the lens of parallel intelligence clarifies its developmental logic and provides methodological support for advancing toward trustworthy real-world applications.

Translated title of the contributionVision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives
Original languageChinese (Traditional)
Pages (from-to)290-303
Number of pages14
JournalChinese Journal of Intelligent Science and Technology
Volume7
Issue number3
DOIs
Publication statusPublished - Sept 2025
Externally publishedYes

ASJC Scopus subject areas

  • Computer Science (miscellaneous)
  • Computer Science Applications
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives'. Together they form a unique fingerprint.

Cite this