TY - GEN
T1 - Toward Fair and Scalable Assessment of Socially Shared Regulation of Learning with Large Language Models
AU - Jiang, Yang
AU - Song, Yi
AU - Roll, Ido
AU - Hao, Jiangang
AU - Ruan, Chunyi
AU - Liu, Lei
N1 - Publisher Copyright: © The Author(s), under exclusive license to Springer Nature Switzerland AG 2027.
PY - 2027
Y1 - 2027
N2 - Effective collaborative learning depends on learners’ ability to collectively regulate their learning processes through socially shared regulation of learning (SSRL). Despite its central role in collaboration quality and learning outcomes, SSRL remains challenging to assess due to its multidimensional, dynamic, and interactional nature. Recent advances in large language models (LLMs) offer promising opportunities to automate discourse-based assessment, yet questions remain regarding their validity, robustness, and fairness across student populations. This study investigates the feasibility of using LLMs, specifically GPT-4o and GPT-5, to automatically code SSRL behaviors in collaborative discourse from a scenario-based science task, comprising 5830 student chat turns. Using a theory-driven coding framework, we compare LLM-based coding across prompt designs and model types, evaluate the consistency of LLM performance across underserved and non-underserved school contexts, and examine how SSRL behaviors differ between these student populations. Results show that LLM-based coding achieves strong agreement with human annotation, approaching human–human reliability, with context-enriched prompting yielding improved performance than general prompting. Importantly, LLM performance is stable across school contexts. Analyses further reveal meaningful differences in SSRL behaviors, with teams from an underserved school context demonstrating more cognitive information sharing behaviors, while teams from a non-underserved context engage more frequently in metacognitive planning and affective regulation. These findings highlight the potential of LLMs for scalable and equitable assessment of collaborative regulation and inform the design of responsible, adaptive AI systems that support collaborative learning for diverse learners.
AB - Effective collaborative learning depends on learners’ ability to collectively regulate their learning processes through socially shared regulation of learning (SSRL). Despite its central role in collaboration quality and learning outcomes, SSRL remains challenging to assess due to its multidimensional, dynamic, and interactional nature. Recent advances in large language models (LLMs) offer promising opportunities to automate discourse-based assessment, yet questions remain regarding their validity, robustness, and fairness across student populations. This study investigates the feasibility of using LLMs, specifically GPT-4o and GPT-5, to automatically code SSRL behaviors in collaborative discourse from a scenario-based science task, comprising 5830 student chat turns. Using a theory-driven coding framework, we compare LLM-based coding across prompt designs and model types, evaluate the consistency of LLM performance across underserved and non-underserved school contexts, and examine how SSRL behaviors differ between these student populations. Results show that LLM-based coding achieves strong agreement with human annotation, approaching human–human reliability, with context-enriched prompting yielding improved performance than general prompting. Importantly, LLM performance is stable across school contexts. Analyses further reveal meaningful differences in SSRL behaviors, with teams from an underserved school context demonstrating more cognitive information sharing behaviors, while teams from a non-underserved context engage more frequently in metacognitive planning and affective regulation. These findings highlight the potential of LLMs for scalable and equitable assessment of collaborative regulation and inform the design of responsible, adaptive AI systems that support collaborative learning for diverse learners.
KW - Assessment
KW - ChatGPT
KW - Collaboration
KW - Context
KW - Fairness
KW - Generative AI
KW - LLM-Based Coding
KW - Regulation
KW - Socially Shared Regulation of Learning
UR - https://www.scopus.com/pages/publications/105043986318
U2 - 10.1007/978-3-032-29773-0_14
DO - 10.1007/978-3-032-29773-0_14
M3 - Conference contribution
SN - 9783032297723
T3 - Lecture Notes in Computer Science
SP - 194
EP - 209
BT - Artificial Intelligence in Education - 27th International Conference, AIED 2026, Proceedings
A2 - Blanchard, Emmanuel G.
A2 - Chen, Guanliang
A2 - Chi, Min
A2 - Isotani, Seiji
PB - Springer Science and Business Media Deutschland GmbH
T2 - 27th International Conference on Artificial Intelligence in Education, AIED 2026
Y2 - 27 June 2026 through 3 July 2026
ER -