
July 28, 2026
Why course completion does not prove learning
Learn why completion records participation rather than durable capability—and how to measure retention, transfer, and workplace application instead.
Read article →
Use a seven-level framework to match training claims with evidence—from participation and reaction to workplace application and results.
By Alireza Ibrahimi
Published August 3, 2026
12 min read
Framework
Most training dashboards answer the easiest question:
Did employees complete the course?
That question matters, but it is only the beginning.
Completion does not show whether employees understood the material, remembered it later, applied it to a new situation, used it at work, or contributed to a meaningful business result.
Each of those conclusions requires a different kind of evidence.
The Seven Levels of Learning Evidence is a practical LoreGraph framework for matching training claims with the evidence that can reasonably support them. It helps learning teams avoid two common problems: collecting convenient metrics that say very little and making claims that are stronger than their data.
The framework is an original LoreGraph synthesis, not a formally validated measurement instrument.
It adapts the familiar distinction among reaction, learning, behavior, and results in the Kirkpatrick Model. It expands the learning portion into immediate performance, delayed retention, and transfer, while also separating basic participation from learner reaction. The established Kirkpatrick framework uses four levels: reaction, learning, behavior, and results.
The expansion matters because performance during or immediately after instruction may not represent durable learning. Research on learning versus performance shows that temporary success during practice can differ from the longer-term changes needed for retention and transfer.
The framework contains seven levels:
These levels answer different questions. They should not be treated as a guaranteed sequence in which success at one level automatically causes success at the next.
Level 1: Participation
The question is:
Did the person access or complete the learning activity?
Possible evidence includes:
Participation data helps teams understand reach, access, operational compliance, and where learners leave a program.
It can support a statement such as:
Ninety percent of assigned employees completed the course.
It cannot independently support:
Ninety percent of employees learned the procedure.
A person may complete every screen without paying close attention. Completion confirms that the defined workflow finished—not that a durable capability developed.
Level 2: Reaction
The question is:
What did learners think or feel about the experience?
Possible evidence includes:
Reaction data can reveal confusing instructions, irrelevant examples, frustrating navigation, or low confidence. It is valuable for improving the learning experience.
However, learners may enjoy training that produces weak learning. They may also find demanding practice uncomfortable even when it helps them remember more later.
A positive rating supports a statement such as:
Learners considered the examples relevant to their roles.
It does not independently support:
The training improved employee performance.
Level 3: Immediate performance
The question is:
Can learners produce a correct response during or immediately after instruction?
Possible evidence includes:
This is the first level that directly examines what the learner can do.
But the conditions matter.
Was the answer still visible? Were hints available? Could learners retry the same item repeatedly? Did the assessment copy the wording of the lesson? Was the task completed with an instructor or independently?
Research distinguishes observable performance during practice from durable learning. Immediate success can be supported by recent exposure, prompts, familiar examples, or other temporary conditions.
A score at this level should therefore be labeled accurately:
Immediate quiz performance was 85 percent.
Avoid stronger labels such as mastery unless the system uses a documented model containing additional evidence.
Level 4: Delayed retention
The question is:
Can learners still demonstrate the capability after time has passed?
Possible evidence includes:
The delay helps separate short-term accessibility from more durable knowledge.
The appropriate interval depends on the outcome. A procedure used every day may receive natural workplace reinforcement. A rare emergency process may require deliberately scheduled practice because employees may not use it for months.
Delayed retention is stronger evidence than an immediate score when the training claim concerns lasting knowledge.
It can support a statement such as:
Two weeks after training, employees still identified the correct first action without reopening the SOP.
It does not yet prove that they can adapt the knowledge to a different situation.
Level 5: Transfer
The question is:
Can learners use the knowledge or skill in a new but relevant situation?
The National Academies defines transfer as extending what has been learned in one context to new contexts.
Possible evidence includes:
Repeating the same quiz item is not strong transfer evidence. A learner may remember the answer without understanding the principle.
Transfer becomes more convincing when the new task differs from practice but still represents the situations the employee is expected to handle.
It can support a statement such as:
Employees applied the escalation rule correctly to unfamiliar customer cases.
Transfer is especially important for judgment, troubleshooting, compliance decisions, leadership, and other work that cannot be reduced to one predictable script.
Level 6: Workplace application
The question is:
Did the learner use the capability in the real work environment?
Possible evidence includes:
This level moves beyond what the employee can do in a course or simulation and examines what they actually do at work.
However, workplace performance depends on more than learning.
An employee may know the correct action but lack time, permission, equipment, staffing, software access, managerial support, or psychological safety. The environment may support or block the behavior.
A workplace observation can support:
Employees used the approved reporting process during relevant incidents.
It should not automatically be interpreted as a pure measurement of learning. It reflects capability operating inside a larger system.
Level 7: Organizational result
The question is:
Did a meaningful business, safety, quality, workforce, or customer outcome change?
Possible evidence includes:
This level matters because organizations invest in learning to support meaningful outcomes—not merely to produce quiz scores.
But an improved result does not prove that training caused the change.
The CDC distinguishes outcome evaluation from impact evaluation. Outcome evaluation examines whether intended outcomes occurred, while impact evaluation attempts to determine whether the program caused those outcomes by comparing them with an estimate of what would have happened without the program.
A company may introduce training while also changing its software, staffing, supervision, incentives, policies, or reporting system.
The careful statement is:
Error rates declined after the training and process changes were introduced.
A stronger causal claim requires a stronger evaluation design.

The framework does not mean:
Participation causes reaction.
Reaction causes learning.
Learning automatically causes workplace behavior.
Workplace behavior automatically creates business results.
The levels represent different questions, not guaranteed links.
A learner may dislike a course but still learn from it. An employee may learn the correct procedure but be unable to use it because the required tool is unavailable. A business result may improve even when employee learning remains unchanged because a software control removed the opportunity for error.
A useful working chain is:
Learning experience → change in capability → application in context → contribution to a workplace result
Every arrow in that chain is a hypothesis.
Evaluation should examine the links that matter instead of assuming they are automatic.
Higher levels are not always necessary.
A low-risk announcement about an office closure may require evidence that employees received and understood the message.
A product-knowledge update may require an immediate check and a small delayed sample.
A customer-support procedure may require an independent task and a varied scenario.
A safety-critical process may require immediate performance, delayed retention, transfer, and workplace verification.
Leadership development may require repeated behavioral evidence collected over time because complex behavior depends heavily on organizational context.
The goal is not to collect every possible metric.
The goal is to collect enough evidence to support the claim responsibly.
Use these questions when selecting the minimum evidence level:
The CDC recommends creating an evaluation plan early by defining the purpose, questions, stakeholders, and data-collection methods rather than adding evaluation after the training is complete.
The level alone does not determine evidence quality.
An immediate task may be well aligned and carefully designed. A workplace observation may be inconsistent, biased, or unrelated to the training objective.
Before interpreting evidence, examine these conditions:
Alignment
Does the activity measure the stated objective?
A recall question is poorly aligned with an objective requiring independent equipment operation.
Independence
What help was available?
A task completed with step-by-step coaching means something different from the same task completed independently.
Delay
How much time passed after instruction?
A test given two minutes later and one given two weeks later answer different questions.
Variation
Is the task new, or is it an exact repetition of practice?
New but relevant cases offer stronger evidence of transfer.
Authenticity
Does the assessment resemble the real decision or task?
A definition question may not represent the pressure and ambiguity of a workplace situation.
Consistency
Would different evaluators or repeated measurements produce a similar conclusion?
One manager’s informal impression may be less dependable than a shared observation rubric.
Context
What helped or blocked performance?
Tools, staffing, incentives, workload, supervision, and team culture can affect workplace evidence.
A metric should always be reported with enough context to explain what was actually measured.

Imagine a customer-support team introducing a new refund-escalation procedure.
The objective is:
Given a customer case containing an exception, the employee will select the correct refund or escalation path and explain which policy condition applies.
The organization could collect evidence at all seven levels.
Participation
Did assigned employees complete the course?
Reaction
Did employees find the cases relevant and the procedure understandable?
Immediate performance
Could employees select the correct action immediately after training?
Were they using the policy document, hints, or feedback?
Delayed retention
Could they explain the main escalation rules one or two weeks later without reopening the lesson?
Transfer
Could they handle new customer cases containing details that were not shown during training?
Workplace application
Did quality reviews show employees using the correct escalation path during real interactions?
Organizational result
Did incorrect refunds, avoidable escalations, resolution time, or customer complaints change?
The organization should not collapse these measures into one “training effectiveness score.”
Each measure tells a different part of the story.
Suppose course completion is high, delayed performance is moderate, and workplace errors remain unchanged. The next action may involve more retrieval practice, better job aids, clearer system permissions, or process redesign.
The framework helps the team investigate rather than jump to one conclusion.

Calling completion learning
Completion supports an access or participation claim. Add evidence aligned with the intended capability.
Using satisfaction as proof of effectiveness
Use reaction data to improve relevance and usability—not as evidence of retention or performance.
Calling one immediate quiz mastery
A mastery claim should rely on a documented rule using multiple aligned pieces of evidence under defined conditions.
Using identical questions to claim transfer
Change the case, wording, context, or task while preserving the underlying objective.
Always aiming for organizational results
Some programs are too small, early, or low-risk to justify expensive impact evaluation. Choose proportionate evidence.
Ignoring available support
Record whether employees used hints, references, AI assistance, checklists, or coaching. Support changes what the result means.
Ignoring the work environment
Training cannot compensate for missing tools, conflicting incentives, unsafe workloads, or unclear authority.
Collecting unnecessary employee data
Measure only what is needed for the evaluation purpose. Consider privacy, access, retention, fairness, and the consequences of incorrect interpretation.
A document-to-course system should not treat every click and quiz score as equivalent evidence.
LoreGraph’s mission is to help organizations turn existing documents and workplace knowledge into structured learning, practice, assessment, and measurable progress.
The framework suggests several practical product principles:
This approach helps course creators ask a better question before publishing:
What evidence would justify the claim we intend to make?
Teams exploring document-based workplace learning can review LoreGraph for business, but the framework remains useful regardless of which platform or authoring process they use.
Choose one important objective from an existing onboarding, compliance, SOP, or skills course.
Write down:
Then change one dashboard label, assessment, or evaluation step so that the claim matches the evidence.
Do not begin by asking how many metrics you can collect.
Begin by asking what you need to know.
Alireza Ibrahimi
Founder, LoreGraph
Software engineer and Learning Engineering researcher building AI systems that transform workplace knowledge into measurable learning experiences.
Put it into practice
Give your team training with practice, assessment, and progress you can actually see.
Explore team training
July 28, 2026
Learn why completion records participation rather than durable capability—and how to measure retention, transfer, and workplace application instead.
Read article →

July 30, 2026
Learn why workplace training fades and how retrieval, spacing, relevant practice, and performance support can build more durable capability.
Read article →

July 28, 2026
Learn how to design training that employees remember later and can apply when the task, timing, or workplace context changes.
Read article →