Abstract
Vision-language-action (VLA) models are a comprehensive modeling approach for embodied intelligence that integrates visual perception, natural language understanding, and action execution within a unified framework, aiming to establish a continuous loop from environmental perception to task planning and action control. Their operational logic corresponds closely to the paradigm of parallel intelligence articulated in the early 21st century. That paradigm comprises artificial systems, computational experiments, and parallel execution, emphasizing virtual modeling, reproducible inference, and closed-loop interaction between the virtual and the real. The initial stage of VLA development, driven by multimodal deep learning, can be regarded as prototypical work within artificial systems, the subsequent stage characterized by large-scale models and cross-domain training expanded the scope of computational experiments, and the more recent focus on hierarchical control and virtual-real closed loops reflects the feedback correction and normative guidance emphasized in parallel execution. VLA models exhibit deep coupling between semantics and action, iterative cycles linking simulation with reality, and steadily improving verifiability. Nonetheless, challenges remain in generalization, semantic alignment, safety and interpretability, and deployment efficiency. Addressing these issues calls for contract-based task semantics, repairable long-horizon hierarchical planning, engineering-oriented use of world models, multi-level feedback and safety governance, and cross-platform transfer with human-machine collaboration. Examining VLA through the lens of parallel intelligence clarifies its developmental logic and provides methodological support for advancing toward trustworthy real-world applications.
| Translated title of the contribution | Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives |
|---|---|
| Original language | Chinese (Traditional) |
| Pages (from-to) | 290-303 |
| Number of pages | 14 |
| Journal | Chinese Journal of Intelligent Science and Technology |
| Volume | 7 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - Sept 2025 |
| Externally published | Yes |
ASJC Scopus subject areas
- Computer Science (miscellaneous)
- Computer Science Applications
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Vision-language-action models under parallel intelligence paradigm: the state of the art and future perspectives'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver