doi:10.2196/105029
Keywords
Talwalkar and colleagues [] address an important problem in medical education: faculty often observe learner performance but struggle to translate those observations into written feedback. Their randomized evaluation of an ambient AI scribe workflow is timely and practically relevant. The study is also valuable because it examines a real educational workflow, includes instructor review of AI-generated notes, and explicitly evaluates unedited AI summaries for mischaracterization and hallucination. These features make the study an important contribution to the emerging literature on AI-assisted feedback documentation. However, there are concerns regarding the interpretation of the results.
The central concern is that a well-documented AI-assisted feedback note should be distinguished from feedback that is educationally effective for learners. Their findings show that AI assistance improved the quality of written feedback documentation, as measured by the Evaluation of Feedback Captured Tool (EFeCT) score []. The EFeCT assesses whether specific feedback elements are present in the written note. It does not assess whether students understood, accepted, trusted, or used the feedback. This distinction matters because feedback becomes educationally meaningful only when learners can engage with it and use it to guide future action. Students’ use of feedback depends on self-regulation, beliefs, emotions, and their ability to identify next steps []. Molloy and colleagues [] similarly caution against treating feedback as a simple input separated from learner involvement and effects beyond the immediate task. Thus, higher EFeCT scores should be understood as evidence of improved documentation quality unless learner uptake and subsequent performance are also examined.
A second concern is that narrative length may have influenced the score difference. As Talwalkar and colleagues [] also note, human-only narratives were much shorter than AI-assisted outputs. Because the EFeCT awards one point for each feedback element present, longer notes have more opportunity to contain those scored elements and may, therefore, receive higher scores even when they are no more useful to learners. This matters for implementation. If institutions adopt AI tools because they improve documentation metrics, they may produce feedback records that receive higher scores while being more cognitively demanding and harder for learners to process. Shute’s [] review supports feedback that is specific and clear but also emphasizes that elaborated feedback should remain manageable for learners.
Related evidence supports this interpretation. Kondo and colleagues [] found that AI-generated feedback on clerkship logs was longer and more consistent, whereas supervisor feedback drew on clinical context and professional judgment. This suggests that AI may improve structure and consistency, while human supervisors may provide contextual educational judgment. These are not interchangeable qualities, but they may be complementary.
Taken together, these considerations highlight the need to distinguish documentation quality from educational usefulness in future research. In addition to documentation metrics, evaluations should include learner understanding, perceived actionability, trust, feedback use, subsequent performance, and the burden associated with reading longer feedback. These outcomes would clarify whether AI-assisted feedback truly improves learning.
Conflicts of Interest
None declared.
Editorial Notice
The corresponding author of “Ambient AI Scribes to Create Educational Feedback Notes for Medical Students: Randomized Trial” declined to respond to this letter.
References
- Talwalkar JS, Chartash D, Zhang L, et al. Ambient AI scribes to create educational feedback notes for medical students: randomized trial. JMIR Med Educ. May 28, 2026;12:e89996. [CrossRef] [Medline]
- Spooner M, Larkin J, Liew SC, Jaafar MH, McConkey S, Pawlikowska T. “Tell me what is ‘better’!” How medical students experience feedback, through the lens of self-regulatory learning. BMC Med Educ. Nov 22, 2023;23(1):895. [CrossRef] [Medline]
- Molloy E, Ajjawi R, Bearman M, Noble C, Rudland J, Ryan A. Challenging feedback myths: values, learner involvement and promoting effects beyond the immediate task. Med Educ. Jan 2020;54(1):33-39. [CrossRef] [Medline]
- Shute VJ. Focus on formative feedback. Rev Educ Res. Mar 2008;78(1):153-189. [CrossRef]
- Kondo T, Donkers J, Nishigori H, Rovers S, Heeneman S. AI-generated versus human supervisor feedback on medical students’ clinical clerkship logs: cross-sectional convergent mixed methods study. JMIR Med Educ. Jun 16, 2026;12:e90064. [CrossRef] [Medline]
Abbreviations
| EFeCT: Evaluation of Feedback Captured Tool |
Edited by Alicia Stone; This is a non–peer-reviewed article. submitted 18.Jun.2026; accepted 15.Jul.2026; published 26.Aug.2026.
Copyright© Amane Endo, Takeshi Kimura, Yuki Kataoka. Originally published in JMIR Medical Education (https://mededu.jmir.org), 26.Aug.2026.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Medical Education, is properly cited. The complete bibliographic information, a link to the original publication on https://mededu.jmir.org/, as well as this copyright and license information must be included.

