Measuring cognitive engagement in AI tutor conversations requires moving beyond traditional behavioral metrics like conversation length.
Using the ICAP framework, the authors developed a scalable, reliable labeling method to classify student engagement in GenAI-based tutor conversations as Passive, Active, or Constructive. Two human raters independently coded nearly 200 STEM-focused conversations, achieving high inter-rater reliability.
The authors then trained an LLM-as-a-judge, which closely matched human labels and enabled large-scale automation. They machine-labeled 50,000 conversations between learners and a GenAI-based tutor, all situated in a mathematics or science context.
Initial findings showed that most student interactions were Passive, suggesting a need for GenAI tutor interactions that promote deeper engagement, improved AI tutor design, and stronger student learning insights.