
Confidence is not competence: When employees are certain but wrong
Confidence can add useful diagnostic context, but it is not proof of capability. Learn how to combine certainty with evidence and improve calibration.
Read article →
AI can generate lessons, questions, and feedback. Whether people actually learn depends on how the system makes them think, practice, and perform.
By Alireza Ibrahimi
9 min read
Research Brief
AI can create a course in seconds.
It can summarize a policy, write an explanation, generate examples, produce quiz questions, simulate a conversation, and offer feedback.
None of those actions, by themselves, prove that learning happened.
Learning occurs when something changes in the learner: knowledge becomes more durable, a skill becomes executable, judgment improves, or a person becomes better able to handle a relevant situation later.
That distinction matters because generative AI is unusually good at producing the appearance of learning. It can make answers easier to obtain, explanations easier to consume, and tasks easier to complete.
Sometimes that helps people learn.
Sometimes it helps people avoid learning.
The difference is largely a design problem.
A document can contain everything an employee needs to know and still produce little learning.
The same is true of AI-generated content.
A polished explanation can be accurate but poorly remembered. A quiz can contain correct answers but measure recognition rather than understanding. A personalized summary can reduce reading time without improving the learner's ability to make decisions later.
Week 1 of the Modern Learning Systems study defines learning as a relatively durable change in knowledge, skill, judgment, strategy, or capability that develops through experience and can be demonstrated later. It deliberately separates exposure and completion from learning.
That gives us a useful boundary:
AI can create learning experiences.
It cannot simply manufacture learning outcomes.
Those outcomes still have to develop in the learner and be demonstrated through evidence.
This is where AI creates a measurement trap.
Imagine an employee learning how to handle a difficult customer escalation.
Without AI, the employee must interpret the case, remember the policy, choose an action, and explain the reasoning.
With an AI assistant, the employee can paste the case into a chat window and receive the recommended action immediately.
Work performance may improve.
But what happened to learning?
We cannot answer from the successful task alone.
The employee may have studied the explanation and developed better judgment.
Or the AI may have performed most of the reasoning.
Research on learning versus performance has warned about this distinction long before generative AI. Temporary conditions can improve performance during practice without producing equally durable learning.
Source: Soderstrom and Bjork, “Learning Versus Performance”. Generative AI makes that old problem much more important because support can now be extremely capable.
The emerging evidence does not support a simple conclusion that AI either helps or harms learning.
Tool design matters.
In a field experiment involving nearly 1,000 high school mathematics students, researchers compared a general GPT-4 interface with a version designed with learning safeguards. AI access substantially improved performance while students could use it. But students using the unrestricted version later performed worse than the control group when AI assistance was removed. The safeguarded tutoring design largely mitigated that negative effect.
Source: Bastani et al., “Generative AI Without Guardrails Can Harm Learning”. This study should not be generalized automatically to workplace learning or every AI system. It examined a particular population, subject, tool design, and assessment context.
But it demonstrates an important possibility:
AI can improve supported performance while weakening independent performance.
Other research shows the opposite outcome when AI is designed around good pedagogy.
Stanford researchers evaluated Tutor CoPilot, an AI system that supported human tutors during live mathematics tutoring. In a randomized controlled trial involving more than 700 tutors and roughly 1,000 students, students whose tutors had AI support were more likely to reach the study's topic-mastery outcome. The system also shifted tutor behavior toward practices such as asking students to explain their reasoning.
Source: Tutor CoPilot study. The lesson is not that one product proves AI tutoring works.
The more useful conclusion is:
The educational behavior produced by the AI matters as much as the intelligence of the AI itself.
Consider two AI interactions.
In the first:
Employee:
What should I do in this situation?
AI:
Choose option B. Here is the completed answer.
In the second:
AI:
Which part of the policy seems relevant?
Employee responds.
AI:
What evidence in the case supports that conclusion?
Employee responds.
AI:
There is one condition you may have missed. Review the exception section and try again.
Both systems are helpful.
But they distribute cognitive work differently.
The first primarily optimizes task completion.
The second preserves more of the interpreting, retrieving, comparing, and deciding for the learner.
This distinction becomes especially important when the person eventually needs to perform without AI.

Learning should not be unnecessarily difficult.
AI can remove plenty of friction:
Removing those barriers can be valuable.
But some effort serves the learning goal.
Learners often need to:
Retrieval practice, for example, can strengthen long-term retention rather than merely measure it.
Source: Roediger and Butler, “The Critical Role of Retrieval Practice”. An AI system designed for learning should therefore distinguish between friction worth removing and thinking worth preserving.
The question “Should AI teach this?” is too broad.
A better design process assigns AI a role.
AI as authoring assistant
AI helps a training creator convert a source into objectives, explanations, examples, questions, or scenarios.
The learner is not yet involved.
Success should be measured by instructional quality, source accuracy, review efficiency, and alignment—not learner outcomes.
AI as explainer
AI helps a learner make sense of difficult material.
Useful when the explanation responds to prior knowledge and confusion rather than simply adding more words.
AI as practice partner
AI creates cases, asks questions, role-plays situations, or adjusts task difficulty.
This can create repeated opportunities for active learning.
AI as feedback provider
AI analyzes a response and helps the learner identify what is correct, missing, or poorly reasoned.
The strongest design often encourages revision rather than replacing the learner's response.
AI as performance support
AI helps someone complete real work.
Here, independent memorization may not be the goal at all.
The system may intentionally carry information that would be inefficient or unsafe to memorize.
These roles require different success measures.
An authoring assistant should not be evaluated like a tutor. A performance-support system should not be criticized merely because the employee relies on it.
The objective determines the appropriate relationship between human and AI.
One of the cleanest design questions is:
If the AI were unavailable tomorrow, what should the person still be able to do?
For some tasks, the answer may be:
Everything.
A worker should recognize an immediate safety threat even if a chatbot is unavailable.
For others:
Very little needs to be memorized.
A specialist may appropriately use a reliable system to retrieve a rarely used technical specification.
And many jobs fall between those extremes.
The learner might need to remember:
The system can provide:
This creates a partnership between learned capability and external intelligence.
Before adding AI to a training workflow, define the outcome first.
Ask:
This approach changes AI from a content feature into an instructional design decision.
UNESCO's guidance on generative AI in education similarly emphasizes human-centered use and pedagogical design rather than adopting generative systems solely because they are technically available.
Source: UNESCO guidance on generative AI in education and research.
LoreGraph currently helps organizations turn SOPs, policies, manuals, handbooks, PDFs, Word documents, and training decks into structured course drafts with lessons and quizzes. Creators can review and edit AI-generated material before publishing, then assign training and track learner activity and quiz performance.
Source: LoreGraph. That solves an authoring problem.
It should not be confused with solving the entire learning problem.
The larger design challenge is moving from:
source → AI-generated content
toward:
source → intended capability → structured learning → meaningful practice → evidence → improvement
That distinction is important to how we think about the future of LoreGraph.
Learn more about LoreGraph. AI should make good learning design easier to create and maintain.
It should not allow us to pretend that generating a lesson means generating learning.
Research on generative AI and learning is developing quickly.
Current studies differ substantially in age group, subject, AI model, amount of human supervision, instructional design, outcome measure, and duration. Results from a mathematics classroom or tutoring platform should not automatically be transferred to compliance training, onboarding, clinical instruction, leadership development, or another workplace context.
AI systems also change rapidly.
For that reason, organizations should test the learning behavior of the complete system rather than assuming that a stronger language model automatically produces a stronger learning experience.

Choose one place where AI currently helps a learner.
Then separate the work into two columns:
Work AI should perform
and
Thinking the learner should perform
If AI is currently doing both sides, redesign one interaction.
Instead of immediately generating the answer, make the system ask for a prediction, explanation, decision, attempt, or reflection first.
Then test what the learner can still do when the support changes.
That is where the question becomes measurable.
AI can generate content.
Whether that content becomes learning depends on what happens next.
Alireza Ibrahimi
Founder, LoreGraph
Software engineer and Learning Engineering researcher building AI systems that transform workplace knowledge into measurable learning experiences.
Put it into practice
Upload a real document and create structured training with lessons, quizzes, and learner tracking.
Create free workspace
Confidence can add useful diagnostic context, but it is not proof of capability. Learn how to combine certainty with evidence and improve calibration.
Read article →

Learn why completion records participation rather than durable capability—and how to measure retention, transfer, and workplace application instead.
Read article →

A practical framework for designing learning systems around capability, practice, retention, transfer, evidence, workplace support, and responsible AI.
Read article →