Skip to main content
School of Biological Sciences School of Biological Sciences

Biology Teaching Professor Develops Cutting-Edge AI Tutor

Schema Study, which can be applied to multiple disciplines, is designed as a tool to help students learn, rather than provide quick answers.

Keefe Reuther.

Keefe Reuther

Four years ago School of Biological Sciences Assistant Teaching Professor Keefe Reuther was at a crossroads. Like many instructors facing a future of education with the emergence of AI, Reuther couldn’t sit back and watch — he had do to do something. When students seek help, Reuther wants them to actually learn, rather than being spoonfed quick answers from  AI chatbots such as ChatGPT.

He began working on a new tool that took advantage of AI’s enormous instructional potential, yet avoided spoon-feeding responses that would provide a short-term answer but hamper long-term course content digestion.

The product of his efforts became Schema Study, a custom-built, affordable, AI-based app grounded in course content rather than unchecked internet content (details in “Schema Study: A Large Language Model (LLM) Application for Asynchronous Student Learning and Inquiry,” published in the journal CourseSource). In this Q&A, Reuther delves into the app’s development, motivations and potential to become the future for biology instruction (and other disciplines) in the rising AI world.

How did you get started on Schema Study? What were the app’s origins?

Schema Study

Schema Study logo

November 2022, when ChatGPT became widely available, was an uncomfortable moment to be launching a new course. (Biological Sciences Assistant Teaching Professor) Liam Mueller and I had just piloted Data Analysis and Design for Biologists, an introductory course on the scientific method, statistics and programming for biology majors. It was immediately clear that a technology capable of generating fluent code and explanations on demand was going to change what teaching that material meant.

My reaction was genuinely mixed: something between childlike wonder and existential horror. But disengaging felt irresponsible. If generative AI was going to reshape how students interact with information and evidence, someone had to figure out how to use it deliberately, to accentuate what's genuinely useful and mitigate what isn't.

The first issue that grabbed my attention was one I knew I could solve on my own: a student stuck at 11 pm the night before a midterm, no instructor available, no study partner online. Flashcards, recorded lectures and practice quizzes can all deliver information, but none of them can respond to where a student actually is in their thinking. That gap is where Schema Study started.

What motivated you to develop this tool?

The core problem is scale. A single instructor, even with a small teaching team, cannot provide individualized feedback to hundreds of students at once. Educational research has developed real solutions to this: flipped classrooms, active learning, peer feedback. These work. But they all require other people to be present and engaged at the same moment as the student.

What I kept returning to was the Socratic ideal: one mentor, one student, a back-and-forth conversation where the instructor meets the learner exactly where they are and helps them reason their way forward. That kind of dialogue is extraordinarily effective and extraordinarily rare in a large undergraduate course. Most students never experience it.

The goal became a chatbot aligned with my specific course and with evidence-based teaching practice: something that could offer adaptive, personalized feedback at any hour, without requiring the student to wait for office hours or hope a TA happened to be online.

How long did it take to create?

A couple of days to build the prototype. Two years to iterate and improve.

The work at every stage depended on Liam Mueller, whose advice and contributions have been invaluable throughout, and on two undergraduates, Grace Constantian and Albert Nguyen, who built much of the documentation that made adoption possible at all. I also had a less expected collaborator: the AI itself.

Building tools with AI is harder than the ads suggest. AI literacy is a skill as hard to develop as any other, and it takes intentional practice. Getting a prototype running is the easy part — communicating precisely what you want, testing the output and iterating until the product actually behaves as intended takes an order of magnitude longer. And that's before you put it in front of students to find out whether it's achieving what you designed it to achieve.

The design process doesn't stop there either. Each new AI model generation changes what the tool can do and what needs to be recalibrated. What worked well with one version doesn't always hold for the next.

What makes Schema Study different than other AI-based teaching apps?

Most AI teaching tools make one of two compromises: they're flexible and customizable but require programming experience most instructors don't have, or they're accessible but locked behind proprietary systems you can't inspect or modify. Schema Study was built to avoid both.

The key is giving theAI system two things it doesn't have by default. The first is behavioral instructions: directives that redirect it away from the sycophantic information dump that general chatbots typically provide and toward Socratic dialogue grounded in evidence-based teaching practice. The second is course content: a two-column spreadsheet where instructors list the terms and concepts they want students to learn alongside the context that gives those terms meaning within the scope of my classroom. No general-purpose LLM is trained on that material, and it's what keeps the conversation anchored to what actually matters in the course.

The landscape of AI teaching apps is evolving by the week, but the pattern is clear. Open-source tools that offer real flexibility typically require coding skills. Polished proprietary platforms are accessible but expensive, opaque and largely unreviewed. Schema Study occupies the gap: free, open-source, no coding required, and published in a peer-reviewed journal so the evidence behind it can be examined and contested. For our institution, the entire cost came to $177.16 for roughly 300 students over 11 weeks, a fraction of what individual student subscriptions to commercial AI tools would have cost.

Why is it important to actually teach rather than just funnel information?

Learning is hard. It takes time and effort. Handing students an encyclopedia of information to memorize doesn't change that.

Pedagogical research is consistent on what does work: learning deepens when students are motivated, when they have multiple ways to engage with material and when they're equipped to apply concepts to problems they haven't seen before. The default behavior of most LLMs cuts against each of those. These systems are optimized to satisfy the user, which in practice means delivering as much information as possible, as confidently as possible, without pushing backthe

A default chatbot can help a student memorize a definition. That's useful, but it's the floor, not the ceiling. Real expertise develops through scrimmage: applying foundational concepts to novel situations, making mistakes, getting feedback, revising. That's what Socratic dialogue does. It withholds the answer, presses for reasoning and forces the student to do the cognitive work themselves. AI can do that, but only if the tool is built around that goal rather than the easier one of keeping the user satisfied.

If I’m an instructor, why should should I use it?

Not all instructors should. A tool is only as appropriate as the context it's used in.

Start by asking whether your students actually have the support they need. Do all of them have equitable access to your instructional team? Can they get the scaffolding required to learn the material, or are there gaps you know exist but can't fill? Is your team spending most of its time mentoring and providing substantive feedback, or is it consumed by logistics, grading and fielding the same questions repeatedly? Studying effectively is a developed skill, and a first-year undergraduate needs considerably more support than a PhD student. If your team doesn't have the bandwidth to provide it, something else needs to fill that gap. Peer learning, tutoring services and supplemental instruction are all options— so is thoughtfully implemented generative AI.

Schema Study shifts the financial burden from the student to the institution. Every student gets equal access to the same model, regardless of whether they can afford a commercial subscription. Conversations are deleted from OpenAI's servers within 30 days and never used to train future models.

One distinction matters: Schema Study is for formative practice. It's where students explain, fumble, get questioned and revise. Mastery is still assessed through secure, independent evaluation outside the app, such as writent exams and oral presentations. If an instructor is weighing whether this fits their course, I'm happy to talk it through.

Are you continuing to refine the app?

Every quarter I use it in the classroom, and every quarter the data shapes the next version. Pre- and post-course questionnaires track how students' attitudes toward AI change over time. Comments and analytics tell me where the tool is working and where it isn't.

The most visible recent change is a row of prompt template buttons in the chat window. Students who aren't sure where to begin can click "Create a Study Plan," "Two Truths and a Lie," or "Schema Map," and the chatbot scaffolds the conversation from there. The other change is structural: the chatbot now asks exactly one question per turn rather than stacking three or four. Multiple simultaneous questions overwhelm a student's cognitive load, fragmenting the dialogue and letting them sidestep the concepts they find hardest. One question focuses their attention and holds them to it. Students reported preferring this version, and it more faithfully serves the Socratic approach the tool is built around.

The next priorities are expanding model support beyond OpenAI to include Anthropic's Claude and Google's Gemini, adding the ability to upload images and files, and giving students a way to download their chat transcripts.

What has been the reaction to Schema Study?

The student data has been encouraging. In our class, 72 percent of students said they'd choose to use Schema Study again in another biology course, and each additional day per week of use more than doubled their likelihood of recommending it. The pattern is consistent: regular, low-stakes engagement builds more value than occasional study marathons.

Outside the classroom, I've led workshops for biology instructors at the Society for the Advancement of Biology Education Research, and colleagues at other institutions have used Schema Study with their own students. One has extended the open-source code to improve the upload functionality, which I hope to fold back into the main version. That's the open-source model working as intended.

The measure of success here isn't how many instructors adopt Schema Study specifically. It's how many take the lessons from this work and use them to improve how their students engage with AI in general. If this work prompts that kind of thinking, it's done its job whether or not anyone uses Schema Study itself.

What’s been most sastisfying in the development of Schema Study?

The most satisfying discovery has been how much the tool depends on what surrounds it. I promote Schema Study as something students can use any time, but dropping an AI tool into a course without building a genuine learning ecology supporting it doesn't work. Students need structured opportunities to develop AI literacy, to encounter the tool's limits, and to build comfort before they can use it well independently. AI tools can't stand alone.

The work I've found most rewarding is designing assignments where students do specific things with Schema Study and then bring the results back to each other. In the first week, students use the tool to explore a concept that confuses or intrigues them, then their classmates independently verify whether what the chatbot told them holds up against reliable sources. Later in the course, students with no prior coding experience use Schema Study to generate code for a figure that tests a real hypothesis, and other students build on that shared code to test additional ones. For many of them, that figure is the first thing they've ever made with code, and it looks professional.

What those assignments seem to do is give students the practice and the confidence to use an open-ended AI tutor independently, when no instructor is available. That's a harder thing to design than the tool itself.

What are the next stages in the evolution of the app?

The app itself will keep getting simpler. The goal was never a more complex AI tutor, just a more effective one, and lowering the barrier for instructors to copy and adapt their own versions remains important work.

But the more pressing issue isn't the app. It's foundational. Students and educators need practice with this technology embedded in real learning contexts, not just access to it. And the conversations about how to bring AI into education deliberately, rather than simply reacting to it as it advances, need to include all of them: students, educators, administrators, and policymakers alike. The field also needs sustained investment in research that gives those choices an empirical foundation.

That's what this work is ultimately about. The currency of science is evidence, and that currency is getting harder to separate from the chaff. A Socratic tutor earns its place not by giving students more answers, but by helping them build the knowledge and habits of mind to ask better questions and pursue the answers themselves.

Acknowledgments

In keeping with this interview's themes, Reuther engaged generative AI as an editorial tool, specifically Gemini Pro 3.1, Claude Sonnet 4.6 and Claude Opus 4.7, to critique and copyedit his drafts. These models served as instruments of revision, not authorship. The ideas, arguments and pedagogical philosophies are entirely his own, and he bears full responsibility for the accuracy of the final text. Schema Study was supported by University of California, San Diego intramural grants TG114333 and RG113974.