关注学生:在线学习中用于自动化参与度预测的行为与上下文线索
文章背景与核心概要
本文针对在线辅导视频中自动化预测学生参与度这一复杂挑战展开了研究。由于参与度是一个包含行为、情感和认知状态的多维度结构,且受到高度的个体差异和标注主观性的影响,作者提出了一种新颖的多模态框架。
该方法通过 Perceiver IO 潜空间瓶颈(latent bottleneck),将预训练视频、音频和图像编码器的隐式时空特征,与结构化行为模态(头部姿态、视线、面部动作单元、情绪以及基于小波的音频特征)相融合。此外,学生和教师的个性被建模为变分后验,以实现参与者之间的部分参数共享(partial pooling),并采用具备不确定性感知的预测头来进行强健的风险量化。
Summary
This paper addresses the complex challenge of automated student engagement prediction from online tutoring videos. Because engagement is a multidimensional construct encompassing behavioral, emotional, and cognitive states—and is further complicated by high inter-person variability and annotation subjectivity—the authors propose a novel multimodal framework.
The approach integrates implicit spatiotemporal features from pretrained video, audio, and image encoders with structured behavioral modalities (head pose, gaze, facial action units, emotion, and wavelet-based audio features) using a Perceiver IO latent bottleneck. Furthermore, student and instructor personalities are modeled as variational posteriors to enable partial pooling across participants, and uncertainty-aware prediction heads are employed for robust risk quantification.
Metadata
- arXiv Identifier:
arXiv:2608.24340[cs.CV]- Authors: Alperen Kantarci, Visvanathan Ramesh, Gemma Roig
- Submitted: August 25, 2026
- Accepted Venue: ICMI 2026 (International Conference on Multimodal Interaction), October 5–9, 2026, Napoli, Italy
- Subjects: Computer Vision and Pattern Recognition (
cs.CV); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)- DOI: 10.48550/arXiv.2608.24340
- Related DOI: 10.1145/3776574.3832485
Abstract
从在线辅导视频中预测学生的参与度非常困难,因为参与度是一个包含不同行为、情感和认知状态的多维度结构。可靠的预测需要将不同类型的行为信号以及表现性线索结合起来。
The prediction of student engagement from online tutoring videos is difficult because engagement is a multidimensional construct comprising distinct behavioral, emotional, and cognitive states. A reliable prediction requires bringing together different types of behavioral signals as well as expressive cues.
通过对 CASED 数据集的分析,显而易见的是,由于高度的个体间差异以及参与度标注的主观性,参与度预测变得更加困难。为了应对这些挑战,我们开发了一个多模态框架,该框架集成了从预训练视频、音频和图像编码器中提取的隐式时空特征,以及头部姿态、视线、面部动作单元、情绪和基于小波的音频特征等结构化行为模态。
Through our analysis of the CASED dataset, it is clear that engagement prediction gets even harder due to the high inter-person variability as well as the subjectivity of the engagement annotation. To tackle these challenges, we develop a multimodal framework that integrates the implicit spatiotemporal features extracted from pretrained video, audio, and image encoders along with structured behavioral modalities like head pose, gaze, facial action units, emotion, and wavelet-based audio features.
我们通过 Perceiver IO 潜空间瓶颈整合了这些模态。此外,学生和教师的个性被建模为可学习嵌入上的变分后验,以实现参与者之间的部分参数共享。我们采用证据回归(evidential regression)和谱归一化高斯过程分类(spectral-normalized Gaussian process classification)预测头进行不确定性感知预测,以进一步提高鲁棒性和校准性。
We integrate these modalities via a Perceiver IO latent bottleneck. Moreover, student and instructor personalities are modeled as variational posteriors over learnable embeddings to enable partial pooling across participants. We employ evidential regression and spectral-normalized Gaussian process classification heads for uncertainty-aware prediction to further improve robustness and calibration.
对 CASED 挑战赛测试集的基准测试表明,所有参与的方法都收敛于随机猜测的性能水平,这揭示了该数据集的难度。在这个高度模糊的区间中,我们的框架取得了具有竞争力的性能,同时独特地提供了校准良好的不确定性指标,证明了可靠的风险量化是将参与度模型部署到现实世界教育工具中的基本先决条件。
Benchmark on the CASED challenge test set shows that all participating methods converge near random-chance performance, revealing the difficulty of the dataset. In this highly ambiguous regime, our framework achieves competitive performance while uniquely offering well-calibrated uncertainty metrics, demonstrating that reliable risk-quantification is an essential prerequisite for deploying engagement models in real-world educational tools.
Access Links & Resources
- Full-Text PDF: View PDF
- HTML Version: arXiv HTML (Experimental)
- TeX Source: arXiv Source Archive
- License: Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
External References
- Citations & Tools: NASA ADS | Google Scholar | Semantic Scholar
