A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration

A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration

Xinyu Liu, Qiqi Dong, Boya Jia, Yi Zhang, Binbin Lian · N/A · 2026

Framework

N/A

License

N/A

Stars

N/A

Summary

A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a conf...

Abstract Summary

A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.

Key Points

  • A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human inten...
  • This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motio...
  • It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bi...
  • A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is f...
  • A physical platform based on the UR3 collaborative robot is built for experimental validation

A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration

|Authors: Xinyu Liu, Qiqi Dong, Boya Jia, Yi Zhang, Binbin Lian

|Venue: arXiv preprint | Year: 2026

|arXiv: 2609.10339v1

Abstract

A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.

Key Contributions

  • A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human inten…
  • This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motio…
  • It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bi…
  • A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is f…
  • A physical platform based on the UR3 collaborative robot is built for experimental validation

Topics

  • human-robot-interaction

Code & Data

No code repository linked in paper metadata.

BibTeX

@article{Liu2026_260910339v1,
  title     = {A Confidence-Aware Multimodal Fusion Framework for Industrial Human-Robot Collaboration},
  author    = {Xinyu Liu and Qiqi Dong and Boya Jia and Yi Zhang and Binbin Lian},
  year      = {2026},
  eprint    = {2609.10339v1},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url       = {https://arxiv.org/abs/2609.10339v1}
}
Share

Related Papers

A New Human-Likeness and Comfort Index for Robot Movements Along Prescribed Paths

A New Human-Likeness and Comfort Index for Robot Movements Along Prescribed Paths

Rosanna Coccaro, Enrico Ferrentino, Antonio Parziale et al. · arXiv preprint · Jul 2026

As human-robot interaction rapidly spreads in numerous fields, the subject of robot acceptance gains increasing importance. Visual similarity to the human body, as occurs for humanoids, is generally not enough to ensure acceptance in physical interaction, as acceptance directly links to comfort a...

vision control human-robot-interaction
PDF Intermediate
No code repo Jul 2026
Assessing Physical Frailty and Fall-Risk Indicators with Social Robots: An in situ Evaluation with Older Adults

Assessing Physical Frailty and Fall-Risk Indicators with Social Robots: An in situ Evaluation with Older Adults

Aniol Civit, Antonio Andriella, Alba Martínez et al. · arXiv preprint · Jul 2026

Frailty assessments are crucial to evaluate the risk of adverse events and the health and social care needs of older adults, yet their administration remains resource-intensive and typically relies on coarse clinical outcomes, such as task completion times, which may overlook biomechanical indica...

vision reinforcement-learning human-robot-interaction benchmark
PDF Intermediate
No code repo Jul 2026
Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling

Catch, Throw, Repeat: Planning for Human-Robot Partner Juggling

Jonathan Rainer Lippert, Kai Ploeger, Abir Chowdhury et al. · arXiv preprint · Jul 2026

Dynamic object exchange between humans and robots remains a challenging problem due to uncertainty in perception, timing, and contact-rich interaction. Human-robot juggling represents a particularly demanding instance of this problem, requiring precise real-time coordination, predictive motion pl...

vision reinforcement-learning planning control human-robot-interaction
PDF Intermediate
No code repo Jul 2026
Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

Development of a Humanoid Robot Prototype for Multimodal Human-Robot Interaction

Thang Tran Viet, Thanh Nguyen Canh, Huy Uong Gia et al. · arXiv preprint · Sep 2026

Human-robot interaction (HRI) enables intuitive and intelligent collaboration between humans and robots in real-world environments. This paper introduces a humanoid robot prototype designed as a flexible testbed for developing and integrating artificial intelligence (AI) modules in HRI tasks. The...

manipulation vision vla reinforcement-learning control human-robot-interaction
Code PDF Intermediate
GitHub ★ — Sep 2026