The problem
Students increasingly work with AI tools that hand them answers. That invites passive consumption instead of active thinking, and it carries real risks: cognitive offloading, weakened reasoning, reduced autonomy and self-confidence, and pressure on academic integrity. If we aim to cultivate critical thinkers, we must critically evaluate the tools we use, and how we use them.
The position: scaffold, do not substitute
My co-authors and I argue that AI in education should be judged by whether it develops human reasoning rather than replaces it. In a position paper at the ACM AI Leadership Summit 2026, From Substitution to Scaffolding: Breaking the Self-Reinforcing Harm Cycle of AI in Education (and Beyond), we propose this as a guiding principle for human-centered AI. Two companion papers develop the argument: Do AI Tutors Empower or Enslave Learners? (opening talk of the GenAIHEd workshop, AIED 2025) and AI in Education Beyond Learning Outcomes: Cognition, Agency, Emotion, and Ethics, which proposes a four-dimensional framework for evaluating educational AI beyond test scores.
MAIKE: the system
MAIKE is an open-source Socratic tutoring system I created and lead. Instead of answering, it analyzes a student’s argumentative essay through argument mining and responds with adaptive critical questions, prompting the learner to reflect, articulate, and revise their reasoning through meaningful conversation. It is designed to support, not supplant, student thinking.
The concept was first presented at the CAIHu Bridge at AAAI 2024; the full system at the Tools for Thought workshop at ACM CHI 2026. MAIKE runs on small open-source language models, which keeps it inexpensive, transparent, and deployable in real schools.
The evidence: a classroom study
Empirical studies have recently been conducted with secondary-school students:
- Ethics-approved classroom study
- ~60 secondary-school students
- Between-subjects design
- MAIKE vs. a conventional chatbot
- Pre/post argumentative essays
- Full dialogue logs for interaction analysis
The study explores how students interact with MAIKE and the impact on students’ written argumentation and critical thinking, relative to a conventional chatbot tutor.
The NLP underpinnings
To study reasoning at scale, I build tools that automatically read and assess students’ arguments, using small open-source language models:
- Argument mining: identifying, classifying, and assessing argument components in student essays (CMNA 2025, AI4ED @ AAAI 2025).
- Critical-question generation: the system that ranked 1st of 13 international teams in the CQs-Gen shared task at ACL 2025, with a two-stage pipeline combining creativity, analysis, and critical evaluation.
- Essay scoring: trait-level, ordinal scoring of argumentative essays rather than holistic averages (arXiv:2602.04604), plus a critical scoping review of LLM-based essay assessment.
Where this goes
After the PhD, my program is empirical: studying how AI systems affect human cognition, learning, and agency, with the study-design, measurement, and NLP machinery above, and with real participants. The four-dimensional evaluation framework (cognition, agency, emotion, ethics) would be the agenda; classroom-scale experiments are the method.
A first step is already underway: from March to September 2026 I am a visiting doctoral researcher at the ML4ED lab, EPFL (Prof. Tanja Käser), funded by a competitive ELIAS mobility grant. There I presented the classroom-study design and continued the collaboration that runs through all my PhD work.
Supervision, service, and funding
Student supervision.
Marta Serrador: M.Sc. thesis (2024–2025), Software Development for Mobile Devices, Universidad de Alicante: “Design and Implementation of an Interface for a Socratic Educational Chatbot” (Code); Nuria Riera: B.Sc. thesis (2024–2025), then research intern (Jun–Jul 2025), Multimedia Engineering, Universidad de Alicante; Daniel Frases: Research intern (May–Jun 2025), Advanced Technician in Cross-Platform Application Development, CFP Alfonso X El Sabio, Madrid; co-author of the ACL 2025 shared-task-winning paper.
Organization. Co-organizer, Collaborative AI and Modeling of Humans (CAIHu) Bridge, AAAI 2025, Philadelphia (Feb 2025).
Peer review. Educational Psychology (2026) · ACM Transactions on Intelligent Systems and Technology (2025–2026) · Acta Psychologica (2025) · Collaborative AI and Modeling of Humans, AAAI Bridge program (2025) · ACM Interactive, Mobile, Wearable and Ubiquitous Technologies (2024).
Competitive mobility funding. ELIAS mobility grant: ELLIS Research Exchange 2026, EPFL, Switzerland (Horizon Europe network of excellence); ELSA mobility grant: ELLIS Doctoral Symposium 2024, Paris, France (Horizon Europe network of excellence); ELISE mobility grant: AAAI 2024 Bridge Program, Vancouver, Canada (H2020 ICT-48 ELLIS network).