AI Tutoring Needs More Than an AI Model

"AI tutoring" has become one of those terms that means whatever the speaker wants it to mean.
A recent EdTech Insiders piece, drawing on new research from Stanford's SCALE Initiative, made this point clearly: the label now covers everything from homework-help apps that solve a problem on demand, to AI companions for language practice, to systems that support a human tutor behind the scenes, to frontier chatbots placed in a "study mode." These products share a category name and almost nothing else.
That imprecision has consequences. When researchers, school leaders, or parents ask whether AI tutoring works, they are often unknowingly asking about five or six different things at once. So it's worth asking a more specific question, one that gets past the label: if tutoring is fundamentally about helping a learner make progress over time, what role should AI actually play in making that happen?
We think about this question constantly at HiLink, because it sits at the center of what we build. Our answer is not the one the AI industry defaults to. It is not "let the model do more of the tutoring." It is closer to the opposite: build the infrastructure that lets AI make the human side of tutoring work better.
A Tutor Does More Than Answer a Question
The distinction between a homework helper and a tutor is not a matter of degree. It's a difference in kind. A homework helper is reactive. A student brings a question, and the system resolves it. That is a useful thing to build, but it depends entirely on the student already knowing what to ask.
A tutor operates differently. A good tutor holds a picture of where a learner is in a curriculum, what they already understand, where their misconceptions live, and what should come next even when the student can't articulate any of that themselves. Students, as one researcher interviewed by EdTech Insiders put it, don't always know what they don't know. Tutoring works by supplying that judgment, not just an answer.
That judgment depends on things a language model does not automatically have: a record of prior sessions, a sense of the curriculum's sequence, an understanding of this particular learner's pace and confidence, and a relationship built on some continuity of trust. None of that is a property of the model. It's a property of the system built around the interaction.
This is not an argument that AI is incapable of anything resembling these functions. It's an argument that the quality of a tutoring experience is determined by far more than which model sits underneath it. A very capable model, dropped into a system with no memory of the learner and no connection to what came before, will still behave like a homework helper. Capability without context does not add up to tutoring.
The Missing Layer: Infrastructure Around the Tutor
If you trace what actually needs to happen for a tutoring interaction to be effective, a pattern emerges: student, live tutor, session, context, AI, next session, and back around again. The model is one node in that loop. The rest of the loop is infrastructure, and most of it is invisible until it's missing.
Before a session, someone needs to know what happened last time, what the learner is working toward, and what materials are relevant. During a session, what matters is the live exchange itself: the questions asked, the decisions a tutor makes in the moment, the evidence of whether something has landed. After a session, that interaction needs to turn into something usable: a summary, a signal about what to follow up on, an update to the learner's ongoing profile.
The important idea here is a simple one, and it's easy to overlook because it sounds obvious: a session should not disappear when the call ends. Too many tutoring tools treat each session as a one-off event rather than one entry in a longer relationship. What happens in a session should become the context that shapes the next one. When that connective layer doesn't exist, tutors are left reconstructing context from memory or from a student's own recap, which is exactly the kind of overhead that erodes the time available for actual instruction.
AI Should Work Behind the Tutor
There is a meaningful difference between AI that replaces a tutor and AI that supports one. Both get called "AI tutoring." They are not the same product, and they are not aimed at the same problem.
AI working behind the tutor can help prepare for a session by surfacing what a learner struggled with last time. It can turn a live conversation into a structured summary instead of leaving that work to a tutor's memory or a rushed note afterward. It can flag follow-up items, organize what's been covered against a curriculum, and reduce the administrative load that pulls time away from teaching. None of this requires AI to sit between the learner and the tutor. It sits behind both of them, making the interaction they already have more informed.
We'd rather describe what this could look like than claim it as a settled result. The opportunity is to give tutors better preparation, better memory, and less overhead, not to promise a specific outcome without evidence behind it. What we can say with more confidence is where the leverage sits: in the parts of tutoring that are currently manual, disconnected, and dependent on one person's memory across dozens of learners.
Why Live Instruction Still Matters
None of this is an argument that humans are always necessary in every interaction a student has. It is a narrower and better-supported claim: the strongest evidence for tutoring's effectiveness is tied to live, human-led instruction, and that evidence gets thinner as direct human involvement decreases.
Stanford's SCALE Initiative recently framed this as a spectrum of relational intensity rather than a binary between human and AI tutoring. At one end, a human leads and AI assists with preparation or analysis. At the other, a student works with an AI tutor with no human oversight at all. The research summarized by EdTech Insiders suggests the evidence for durable learning gains is strongest at the human-led end of that spectrum, and that even modest human involvement, like a person simply encouraging a student to show up, can meaningfully change whether a student engages with a tool at all.
That doesn't mean every learner needs the same amount of human contact, or that AI-only tools have no place. It means the more useful question isn't whether to remove humans from tutoring. It's how to make whatever human involvement already exists more consistent, better informed, and less burdened by administrative overhead. That is a technology question, and it's one that has very little to do with how convincing a model's responses sound.
From AI Feature to AI Infrastructure
Most education companies are now adding some form of AI: a chat assistant, an automated summary, a lesson-planning tool, a recommendation engine. These features are proliferating faster than the systems that would make them useful together.
The more interesting question isn't whether a product has an AI feature. It's where that intelligence lives and whether it connects to anything else. A summary that's generated once and never referenced again isn't building anything. A recommendation that doesn't know what happened three sessions ago isn't really informed. AI features add the most value when they're wired into the underlying workflow and the learner's ongoing context, rather than functioning as standalone add-ons bolted onto an existing product.
This is the shift from AI as a feature to AI as infrastructure, and it's the layer HiLink is built around.
What the Tutoring Stack Could Look Like
A useful way to think about this is as a stack, with the model as only one layer inside it. Human interaction sits at the center: the tutor and learner working together live. Around that sits collaboration infrastructure that supports the interaction itself. Session intelligence turns that interaction into structured, usable information. Learner context persists that information across sessions rather than losing it when the call ends. AI assistance draws on that context to help tutors prepare, teach, and follow up. And a continuous feedback loop lets each session inform the next one, so the system gets more useful over time rather than starting from zero every session.
The model matters. But it's one layer among several, and it's not the layer that determines whether tutoring actually works.
The Future: AI-Enhanced Human Tutoring
The future of tutoring doesn't have to be a story about a human tutor being replaced by AI. It can be a different story: a human tutor supported by AI, persistent context, live collaboration, and session intelligence working together.
The next generation of AI tutoring may end up being defined less by how convincingly AI can imitate a tutor, and more by how effectively technology can help real tutors understand, support, and teach real learners. That's a less dramatic story than AI replacing teachers, and it's a more useful one.
AI doesn't need to sit between the learner and the tutor. It can work behind the tutor, making the human interaction more informed, more continuous, and more scalable. That's the infrastructure problem HiLink is trying to solve, and we think it's the one that matters most.