When Students Ask for AI Feedback, They Reflect Deeper: Results from a Large-Scale Study
An Innosuisse project with the Institute for Digital Technology Management (HAIS Lab) at Bern University of Applied Sciences (BFH), Fall Semester 2025
Key findings:
- Students who engaged with Rflect's AI feedback ("Go Deeper") went on to significantly extend their reflections afterward — a 44.1% increase in word count (p < .001).
- Students who engaged with AI feedback reflected substantially more often than peers in the same group who didn't — averaging 21.5 reflections over the semester compared to 13.25.
- Among students who used AI feedback, 79.1% went on to revise their reflection afterward.
- There were early, promising signs that using Rflect may help boost domain self-efficacy specifically for students who started the semester with lower self-efficacy.
- Students using the AI-supported version of Rflect reported modestly higher engagement with the reflection practice than those without AI feedback.
- While trust decreased in both groups over time, the decline was significantly more pronounced in the group using Rflect without AI features (the no-AI group).

Executive summary
During the fall semester of 2025, Rflect partnered with the Institute for Digital Technology Management (HAIS Lab) at Bern University of Applied Sciences on a large-scale longitudinal study, funded through an Innosuisse innovation project. The study set out to answer the following questions:
- How do students perceive the trustworthiness of Rflect's AI elements?
- And how do students' reflective writing skills develop over the course of a semester?
The study ran across 8 institutions in Switzerland and Germany, 21 lecturers, and 28 courses, comparing a control group (no reflection tool), a group using Rflect without AI features, and a group using Rflect with AI features. The clearest result: when students actively requested AI feedback, they went on to substantially extend their reflections — a 44% increase in word count. Also, the more students reflect, the longer their reflections become. While trust decreased and distrust increased in both groups over time, the decline was significantly more pronounced in the group using Rflect without AI features (the no-AI group). There were early signs that using Rflect and AI feedback may help boost self-efficacy for students who start out lower in confidence. At the same time, the study surfaced some friction: high survey drop-out rates and qualitative feedback pointing to specific product gaps. As with our previous studies, we see this as a promising but early-stage result that should be read with appropriate caution — and one that gives us a clear roadmap of what to improve.
About the study
The research team — Prof. Dr. Roman Rietsche, Prof. Dr. Thiemo Wambsganss, Léane Wettstein, and Katja Pott — designed a mixed-subject experiment with repeated measures. Courses were allocated to one of three conditions:
- Control group: students completed only a pre- and post-survey, with no interaction with any reflection tool.
- Experimental group 1 (No AI in Rflect): students used the Rflect app to reflect throughout the semester, without AI-powered features.
- Experimental group 2 (Full AI in Rflect): students used Rflect with AI feedback and an AI-powered reflection dashboard available to them.
All experimental-group students completed a pre-survey, reflected regularly in Rflect over the semester, and completed a post-survey; the control group completed only the surveys.
Participants: 654 students took the pre-survey across the three groups (326 in the Full AI condition, 97 in No AI, and 231 in the control group). As is common in semester-long, multi-institution studies, the final sample with complete data (pre-survey, post-survey, and interaction logs) was considerably smaller: 67 participants in the Full AI group, 51 in No AI, and 56 in the control group — a drop-out rate of 53%. Of the Full AI group's final sample, 32 students used Rflect's "Go Deeper" AI-feedback feature at least once.

Figure 1: The two AI functionalities implemented in the AI group, including "Go Deeper" and the AI-powered student dashboard.
Key findings
1. AI feedback is linked to longer, more developed reflections
The clearest quantitative result of the study: when students used the "Go Deeper" AI-feedback feature, their reflections grew substantially afterward. Median reflection length went from 85 words before AI feedback to 122.5 words after — a 37.5-word, 44.1% increase (p < .001). This was held across the full log dataset of 456 users and thousands of individual reflections logged over the semester (5,150 reflections from 312 AI-group users; 2,554 from 144 No-AI-group users).
Looking more closely at how AI feedback was used: among the 312 users in the AI group, 92 (29.5%) used Go Deeper at least once, and 220 (70.5%) never did. The students who did use it engaged more overall — averaging 21.52 reflections each, compared to 13.25 for those who didn't use the feature. Of the reflections where students received AI feedback, 79.1% showed an actual change in the text afterward, suggesting most students who engaged with the feedback used it to revise, not just to read.
What this means for educators: AI feedback doesn't reach every student — many of the AI group never used it in this study — but for the roughly one-third of students who do engage with it, the feedback loop meaningfully deepens the reflection. This points toward AI feedback as an opt-in tool best paired with prompts or lecturer encouragement that invite students to use it, rather than an assumed default behavior.

2. Trust and distrust evolved differently depending on AI support
Both trust and distrust in Rflect shifted over the course of the semester, and the pattern differed meaningfully between the AI and No-AI groups.
Trust and distrust were defined as follows:
- Trust = belief that someone or something is reliable, honest, or capable, so that you feel confident depending on them in situations of uncertainty and vulnerability. Example items: "I am confident in the Rflect app.", "The Rflect app is reliable."
- Distrust = belief that one should question motives and view actions with suspicion driven by expectations that the someone or something may be unreliable, harmful, or otherwise untrustworthy. Example items: "I am suspicious of the Rflect app's intent, action, or outputs.", "I am wary of the Rflect app."
- Distrust increased significantly across the semester in the No AI group (p < .0001), while the AI group showed no significant change (p = .66). The interaction effect between group and time was significant (F(1,208) = 9.3, p = .003).
- Trust decreased in both groups over the semester, but significantly more so in the No AI group (p < .0001) than in the AI group (p = .003); the interaction effect was marginally significant (F(1,208) = 4.04, p = .05). It's worth noting that trust ratings remained relatively high in both groups throughout, on a 7-point scale.
What this means for educators: students using the AI-supported version of Rflect held steadier views of the tool over time, while those without AI feedback grew more skeptical. This suggests the AI feedback loop itself may play a role in sustaining students' confidence in the tool across a semester — though both groups still saw some erosion in trust. It is worth noting that blind trust in any tool is not per se a good thing, so we take this one on the chin with a clear idea of what to improve.


3. Early signals on self-efficacy — promising, but not yet conclusive
Overall, there was no significant difference in general self-efficacy (F(2,152) = 1.63, p = .2) or domain self-efficacy (W = 1189.5, p = .39) between the three groups by the end of the semester. However, a more granular analysis revealed a marginally significant three-way interaction between group, students' baseline general self-efficacy, and time (F(1,86) = 3.60, p = .06):
- Among students who started with low general self-efficacy, there was no significant difference between groups, but a marginal positive effect was observed specifically in the AI group (p = .06).
- Among students who started the semester with high general self-efficacy, those in the No AI group saw a significant decrease in domain self-efficacy (p = .03), while the AI group showed no significant change (p = .18).
What this means for educators: these results are suggestive rather than conclusive (several land just outside conventional significance thresholds), but they point in a consistent direction: regular reflection, supported by AI feedback, may help support self-efficacy, particularly for students who don't already feel confident in the domain. This is a finding we plan to investigate further in future studies with larger samples.
4. Engagement showed a marginal uplift with AI
Students in the AI-supported group reported somewhat higher engagement with the reflection practice than those in the No AI group (W = 968, p = .055, r = .19) — a small effect that falls just short of conventional statistical significance, but consistent with the broader pattern of AI feedback modestly strengthening students' relationship with the reflection process.

What this means for educators: while personalized feedback expectedly boosts engagement, future studies could focus on how Rflect can support educators in delivering targeted feedback — with or without AI — to further enhance student engagement and learning outcomes.
5. What students said they liked, disliked, and wanted changed
Qualitative feedback added important context to the numbers. Students appreciated the simplicity of the Rflect tool, the design and UI, and the reflection practice itself. At the same time, they raised concerns about deadlines and the time required to complete reflections, the quality of the feedback and reflection questions, and — notably — uncertainty about data confidentiality, alongside requests for more actionable tips.
When asked what they'd like to see changed, students asked for more flexible deadlines for reflection tasks, clearer communication on data security, more personalized feedback and advice, push notifications and reminders, a voice-input option, and more varied, class-specific reflection questions. Many of these improvements have already been implemented in Rflect since.
What this means for educators: the qualitative feedback is a candid signal that, while the underlying reflection practice and AI feedback show real promise, the product experience at the time of the study — deadlines, feedback quality, data transparency — needs continued work to fully win students over, and student input has directly shaped several recent improvements to the tool.
Open questions and limitations
We want to be transparent about the parts of this study that limit how far these results can be generalized:
- High drop-out: 53% of students who started the pre-survey did not complete the study with valid interaction logs and post-survey data, shrinking the final analytical sample considerably (down to 174 students across all groups for attitudinal analyses). This is common in semester-long field studies but reduces statistical power and may bias the sample toward more engaged students.
- Marginal significance on several key results: many of the more encouraging findings — the self-efficacy interaction, the engagement effect — sit at or near conventional significance thresholds (p = .05–.06) rather than comfortably below them. We treat these as promising directions for further research, not settled conclusions.
- Uneven course and institution distribution: the final Full AI sample came from 2 institutions and 4 courses, while the No AI sample came from 3 institutions and 5 courses — a reminder that results may be shaped by course-specific factors as much as by the tool itself.
- Lecturer and onboarding variability: the research team noted that student motivation to engage with the study was closely tied to individual lecturers' own engagement, and that ensuring consistent onboarding messaging across many lecturers and courses was genuinely difficult.
- UI and prompt design matter: the team noted that the way the interface is set up, and the specific type of reflection prompts used, have an effect on how (and whether) students engage with AI feedback — variables this study could not fully isolate.
Where this leaves us
This study reinforces a pattern from Rflect's earlier efficacy work: reflection — and reflection supported by AI feedback in particular — shows real, measurable promise, but the effects are often concentrated among the subset of students who actively engage with the feature, and the product experience still has room to improve. The most encouraging signal is the strong, highly significant link between AI feedback and reflection depth.
For lecturers and institutions considering how to integrate structured reflection into their courses, the practical takeaways are: build in explicit encouragement or prompts for students to engage with AI feedback rather than assuming it will be discovered organically; be transparent with students about how their reflection data is handled; and expect that reflection tools, like most pedagogical interventions, will show their clearest effects among the students who most consistently use them.
As with our previous efficacy studies, we see this as one step in an ongoing process of building Rflect's evidence base. We're grateful to the 21 lecturers and 8 partner institutions who made this study possible, and to the HAIS Lab research team for their rigorous, honest analysis.
About Rflect
Rflect provides universities with the infrastructure to authentically teach and assess the human skills AI cannot replace. Learn more at rflect.ch. Have questions about this study? Get in touch: info@rflect.ch.