Back
Back to blog
Analysis

18 months of an AI tutor: what works and what does not

August 26, 2026
Hourglass
6 min
Kai Hoffmann
Founder's Associate

Questions at 11 p.m., two weeks before the exam. Explaining what a compiler is for the third time in a week. Building quiz questions even though the ideas run out after five minutes.

At more than 40 universities, tasks like these are now handled by an AI tutor. What holds up and what does not was reviewed by Alexander Pretschner in a webinar in May 2026: Professor of Software & Systems Engineering at the Technical University of Munich, co-founder of OneTutor, eighteen months of operation in his own lectures with up to 1,200 first-semester students.

The talk deliberately stays close to its subject. It is not about assessment in the age of AI, not primarily about the research, and not about detecting machine-generated text, but about a simpler question: what a tool changes in learning. That is exactly where two problems meet. The older one: a lecture runs in one direction, rarely fits everyone in the hall, and there is no money for tutorials. The newer one: whoever outsources their thinking learns nothing, a chatbot without course material invents answers, and whoever cannot pay for a licence is left behind.

The observations in this article come from a webinar held in May 2026 with Alexander Pretschner. The robust figures come from the accompanying research by the bidt, which studied the Bavaria-wide deployment in the 2025 summer semester. To make the insights discussed there accessible beyond the webinar itself, the key points are summarised here as a blog post. To make sure you don't miss future webinars, you can subscribe to our newsletter here.

What works

Two kinds of gaps, two tools

Some gaps are known: anyone who notices that a topic or a problem is not yet understood can ask a specific question. A chat is enough for that, as long as it is based on the course material rather than on half the internet and backs its answer with a reference to the relevant slide.

Other gaps stay invisible. Whoever does not suspect that something is missing does not formulate a question about it either. This is exactly where a quiz helps, one that probes the material from the outside and makes gaps in understanding visible before they show up in the exam.

This distinction explains why two tools are needed rather than one: the chat waits until someone comes, while the quiz actively goes looking.

For teaching that means: a chat window alone covers only half the job.

Questions nobody asks in the lecture hall

In a hall with 1,200 people, hardly anyone raises a hand to ask about the basics. In the chat, exactly that happens. Lecturers cannot see who asked, and that distance lowers the threshold.

The questions do not appear without lecturers

One prompt and 50 finished exam questions are ready? That is not how the system works. The lecturer provides topics and keywords, from which multiple-choice and free-text questions with model answers are generated, and then the work begins. Around 80 percent of the questions are directly usable in Pretschner's assessment. The rest are imprecise, ambiguous or simply wrong.

Curation is therefore not a convenience feature but a condition. Skipping this step risks distributing faulty questions to an entire cohort.

A second finding from practice is less obvious. The best upload material is not the slides but the transcripts of the recordings. Spoken language repeats itself, and from that redundancy the system recognises what carried weight in a session and what was only touched on.

For lecturers that means: anyone who records lectures anyway already has the best source on hand.

The feedback channel that did not exist before

Until now, what had not landed showed up at the end of the semester, namely in the exam. Now it is in the analysis on Monday: which questions were asked during the week, which topics fail in the quiz, where the error rate rises.

Alexander Pretschner has turned this into a routine. Half an hour before the lecture, a look at the analysis, then the session starts with whatever did not land during the week. The reasoning behind it is unusual: the explanation is usually to blame, not the students' understanding.

Two side effects stand out. Some questions run alarmingly far ahead of the actual material, which allows the pace to be corrected. And in the tutorial groups a whole category of questions disappears: what used to be clarified in the tutorial no longer arrives there.

For the university that means: the benefit does not sit with students alone. Lecturers and students work on the same platform, and out of that comes feedback that simply was not available before.

From what course size is this worth it? The chat pays off from around ten participants, because basic questions go unasked in small groups too. The aggregated analysis only pays off with large cohorts. Good when a lecture with 1,200 first-semester students and 60 teaching assistants would otherwise stay opaque. Also useful in a seminar with 20 people, except that the aggregated analysis of problems adds little there, since it is obvious even without a tool where the difficulties are.

What happens to student data? Lecturers see aggregated analyses, students always remain pseudonymous, and there is no attribution of questions or quiz results to individuals. Use requires consent and was not mandatory in any of the projects. According to the vendor's own account the service is hosted in Germany, with access via single sign-on and LTI 1.3. For IT that means: reviewing the data processing agreement remains part of procurement all the same.

How a semester runs

In practice one sequence has proven itself, and it starts before the semester: the course is set up, existing material moves into the system, and the first quiz questions are generated and curated, so that the tutor is usable on the first day of lectures. The difference is then made by introducing it in the lecture, because without that signal from the lecturer students do not get on board. After that a simple rhythm is enough: new quiz questions on a topic go live when the corresponding materials appear in the learning management system. That is not mandatory.

The price for this is 15 to 20 minutes per 90-minute lecture, reusable the following year. What that looks like in a specific course is shown in the article on the didactic use at TH Rosenheim.

The accompanying research supports this part. 84 percent of the 31 lecturers surveyed would use the tool again in the same course, and 86 percent of chat function users were fairly or very satisfied with it.

What does not work

A tool does not change an attitude

The clearest rejection in the talk is aimed at the hope that technology solves a motivation problem. Expectations of a degree programme vary widely. Some are interested in the subject, some in the qualification, most in both in shifting proportions. Anyone who only wants to pass does not suddenly learn differently with a tutor.

Good when a tool reaches those who ask questions anyway. Less good as an answer to a cohort with no interest in the subject.

No uploads, no tutorial

A system that starts with lecturers demands work from the people with the least immediate benefit from it. The question is left openly in the room during the talk: what is in it for me as a lecturer? Pretschner considers that attitude wrong, but he encounters it across subjects and types of institution.

The accompanying research supplies the number. Of 93 registered courses, 84 were active. In nine courses no material ever landed in the system, even though participation had been confirmed. Without material and without curated questions a course stays empty, no matter how well the technology works.

Usage piles up before the exam

Across the semester activity stays flatter. Then it rises steeply, shortly before the exam, and a second time before the resit. The most frequently named benefit is time saved.

Learning partly consists of working through difficulty. Time saved is therefore not a reliable indication of understanding, but at first only time saved. In the accompanying research, 86 percent consider the tool helpful for exam preparation, practically level with lecture scripts and course materials at 87 percent. So the tutor does not compete with the script, it sits next to it.

Does multiple choice encourage shallow learning? The talk explicitly concedes this risk. A pool of multiple-choice questions trains recognition, not argument. That is exactly why multiple choice is not the only question mode: free-text questions require students to formulate an answer themselves, which is then compared against a model answer. The next stage goes beyond that, as a tutor that asks rather than answers, following up on a concept from the session. The thinking behind it is set out in the article on the difference between an answering and a questioning AI. For lecturers that means: whatever the question pool does not train, the assessment itself has to demand.

Not everyone is reached

A fifth of the students surveyed did not use the offer. The reasons are revealing. A fundamental rejection of generative AI plays hardly any role, and neither do data protection concerns. What comes first are other learning methods that are felt to be more effective, especially those involving one's own research and one's own processing. A quarter consider generative AI unsuitable in their own subject area, and 8 percent were held back by access problems.

Does it work beyond STEM subjects? One of the larger deployments runs in philosophy. What stands out is the attitude, not the outcome. In the humanities and social sciences scepticism is greater, above all towards questions that have to allow for nuance. The accompanying research cannot test this, because a comparison between subject groups is not currently possible with the available case numbers.

For teaching that means: four out of five is not everyone. The remaining fifth is not a communication problem but didactic information.

What stays open

How could effectiveness be measured?

Whether students take up an offer is the easy question. The hard one is whether it works, and every recommendation to a university hangs on that.

The talk sorts through the options and leaves them open. Better grades, more engagement, fewer dropouts, more enjoyment of the subject. Plus a criterion that rarely appears in tenders: whether getting back in after two weeks of illness becomes easier. And the caveat that better grades are not the same as better understanding.

What the accompanying research shows and what it does not

Ten Bavarian universities, three years, scientifically led by the bidt. The first bidt report has been available since March 2026, covering the first empirical results on the use of OneTutor in the 2025 summer semester.

The result is more sober than the satisfaction figures suggest. Between frequent users, light users and non-users there is no difference in subjectively perceived learning success. The design barely allows such statements: not experimental, usage by individual choice, and exam results could not be linked in for data protection reasons. The authors do not consider a generalisation possible at this stage.

One set of figures from the talk, by contrast, looks like an effect. In an introductory lecture with around 1,200 first-semester students, 34 percent failed, 28 percent among users, and 13 percent among active users with at least 20 interactions. The same slide notes alongside them that there is no correlation between usage frequency and grades, and calls the findings preliminary.

What is documented, then, is acceptance. Impact is open. For procurement that means: a decision can today rest on usage and satisfaction, not on demonstrated learning gains.

Continuous assessment remains unsolved

The obvious next step would be to turn the conversation into evidence. Whoever can explain a concept in dialogue has understood it, and ongoing conversations could take pressure off exams. The talk describes this route and discards it in the same breath, at least for now. The reason is banal: an agent of your own can be placed in front of the conversation. So far there is no solution for that.

Can the chat be switched off before assessments? Yes, chat and quizzes separately, for take-home exams for instance. It does not solve the question of how independent work is evidenced.

What could change about the role of lecturers

Three developments are explicitly left open in the talk.

Economics. If a very similar introduction to computer science is delivered around thirty times in Bavaria, the question arises why it has to stay thirty.

Campus exodus. Students come to campus less often, especially in technical subjects. Whether that is good for learning is an open question.

One aside in the talk fits here. The lecture format is not treated as a mistake there. 90 minutes of coherent narrative have a value of their own, the flipped classroom depends on culture and cohort size and does not scale at 1,500 participants. What is meant is assistance, not replacement. The tool is explicitly not intended for homework, because homework is allowed to be hard.

What remains after eighteen months

Back to the beginning: the question at 11 p.m., the third attempt to explain a problem, the empty question catalogue. For these three things there is now a tool, and it is being taken up. Four out of five students use it when it sits in the course, and lecturers would predominantly use it again.

What follows from that for learning outcomes is open. The most robust study to date finds no measurable difference, and usage piles up before the exam instead of across the semester. What remains is the sentence the webinar ends on: what is needed is an understanding of where such tools work and where they do not. Only from that does the place emerge where benefit arises.

TRY ONETUTOR FREE

Try OneTutor for free

Up and running in minutes — no IT setup.

Test it with your own course materials.

Generate quiz questions from your materials automatically.

40+ universities already use OneTutor.

⚡ Ready in 2 minutes

Email

First name

Last name

Institution (optional)

Thanks — we're processing your request and will be in touch shortly.

Something went wrong. Please try again.

We use your details only to set up your test account.