Group Engagement Recognition Dashboard

20 short classroom video clips were drawn from a published engagement-recognition dataset, in which the source paper's human coders had already labeled each clip's group engagement as either High or Medium, and that paper-assigned label is treated here as the Human Coded rating. Each clip was then independently re-analyzed, second by second, by a Gemini 3.1 Pro computer-vision (CV) pipeline, which produced its own engagement rating with no access to the paper's labels. This dashboard lines up the CV pipeline's ratings against those human-coded labels to see how well an off-the-shelf multimodal LLM reproduces expert human coding of classroom engagement.

Key

All Videos
High Engagement Rating by Human Coders
Medium Engagement Rating by Human Coders

At a glance

Videos: AI rating vs. Human Coded

Outer ring: AI overall rating

Inner ring: Human Coded label

Engagement Level

High Engagement
Medium Engagement
Low Engagement

Interactive Group Summaries

Seconds of footage by AI-coded engagement level

Hover a bar for both the group total and the per-video average

AI overall rating distribution

AI vs. Human Coded: Classification Accuracy

"Human Coded" = the label from the source paper's video set

Agreement rate (AI overall rating = human-coded label)

Confusion matrix, rows: human coded, columns: AI

Video Explorer