Skip to article content

The seven levels of learning evidence

Use a seven-level framework to match training claims with evidence—from participation and reaction to workplace application and results.

By Alireza Ibrahimi

12 min read

Framework
A professional moves through seven connected checkpoints from participating in training to contributing to a workplace result.

Most training dashboards answer the easiest question:

Did employees complete the course?

That question matters, but it is only the beginning.

Completion does not show whether employees understood the material, remembered it later, applied it to a new situation, used it at work, or contributed to a meaningful business result.

Each of those conclusions requires a different kind of evidence.

The Seven Levels of Learning Evidence is a practical LoreGraph framework for matching training claims with the evidence that can reasonably support them. It helps learning teams avoid two common problems: collecting convenient metrics that say very little and making claims that are stronger than their data.

What this framework is

The framework is an original LoreGraph synthesis, not a formally validated measurement instrument.

It adapts the familiar distinction among reaction, learning, behavior, and results in the Kirkpatrick Model. It expands the learning portion into immediate performance, delayed retention, and transfer, while also separating basic participation from learner reaction. The established Kirkpatrick framework uses four levels: reaction, learning, behavior, and results.

The expansion matters because performance during or immediately after instruction may not represent durable learning. Research on learning versus performance shows that temporary success during practice can differ from the longer-term changes needed for retention and transfer.

The framework contains seven levels:

  1. Participation
  2. Reaction
  3. Immediate performance
  4. Delayed retention
  5. Transfer
  6. Workplace application
  7. Organizational result

These levels answer different questions. They should not be treated as a guaranteed sequence in which success at one level automatically causes success at the next.

Key takeaways

  • Completion is evidence of participation, not proof of learning.
  • Satisfaction and confidence describe the learner’s reaction, not necessarily their capability.
  • Immediate quiz performance should be reported separately from delayed retention.
  • Transfer requires learners to handle a new but relevant situation.
  • Workplace performance depends on learning and the surrounding work environment.
  • Organizational outcomes may have several causes, so improvement alone does not prove that training created it.
  • The required evidence should match the importance, risk, and intended outcome of the training.

The seven levels

Level 1: Participation

The question is:

Did the person access or complete the learning activity?

Possible evidence includes:

  • enrollment;
  • attendance;
  • pages or lessons accessed;
  • activities submitted;
  • course completion;
  • time recorded in the system;
  • certificate issuance.

Participation data helps teams understand reach, access, operational compliance, and where learners leave a program.

It can support a statement such as:

Ninety percent of assigned employees completed the course.

It cannot independently support:

Ninety percent of employees learned the procedure.

A person may complete every screen without paying close attention. Completion confirms that the defined workflow finished—not that a durable capability developed.

Level 2: Reaction

The question is:

What did learners think or feel about the experience?

Possible evidence includes:

  • perceived relevance;
  • satisfaction;
  • confidence;
  • difficulty;
  • usability;
  • clarity;
  • willingness to recommend the training.

Reaction data can reveal confusing instructions, irrelevant examples, frustrating navigation, or low confidence. It is valuable for improving the learning experience.

However, learners may enjoy training that produces weak learning. They may also find demanding practice uncomfortable even when it helps them remember more later.

A positive rating supports a statement such as:

Learners considered the examples relevant to their roles.

It does not independently support:

The training improved employee performance.

Level 3: Immediate performance

The question is:

Can learners produce a correct response during or immediately after instruction?

Possible evidence includes:

  • end-of-unit questions;
  • demonstrations;
  • short explanations;
  • practice tasks;
  • simulations;
  • scenario decisions;
  • worked examples completed by the learner.

This is the first level that directly examines what the learner can do.

But the conditions matter.

Was the answer still visible? Were hints available? Could learners retry the same item repeatedly? Did the assessment copy the wording of the lesson? Was the task completed with an instructor or independently?

Research distinguishes observable performance during practice from durable learning. Immediate success can be supported by recent exposure, prompts, familiar examples, or other temporary conditions.

A score at this level should therefore be labeled accurately:

Immediate quiz performance was 85 percent.

Avoid stronger labels such as mastery unless the system uses a documented model containing additional evidence.

Level 4: Delayed retention

The question is:

Can learners still demonstrate the capability after time has passed?

Possible evidence includes:

  • delayed recall;
  • an explanation given without reviewing the lesson;
  • a procedure performed after several days;
  • a new quiz administered later;
  • a repeated simulation;
  • a short follow-up scenario.

The delay helps separate short-term accessibility from more durable knowledge.

The appropriate interval depends on the outcome. A procedure used every day may receive natural workplace reinforcement. A rare emergency process may require deliberately scheduled practice because employees may not use it for months.

Delayed retention is stronger evidence than an immediate score when the training claim concerns lasting knowledge.

It can support a statement such as:

Two weeks after training, employees still identified the correct first action without reopening the SOP.

It does not yet prove that they can adapt the knowledge to a different situation.

Level 5: Transfer

The question is:

Can learners use the knowledge or skill in a new but relevant situation?

The National Academies defines transfer as extending what has been learned in one context to new contexts.

Possible evidence includes:

  • a scenario with different surface details;
  • a new customer case;
  • a simulation containing an unfamiliar exception;
  • a task performed with a changed sequence;
  • a problem requiring the same principle in another context;
  • an explanation of when a rule does and does not apply.

Repeating the same quiz item is not strong transfer evidence. A learner may remember the answer without understanding the principle.

Transfer becomes more convincing when the new task differs from practice but still represents the situations the employee is expected to handle.

It can support a statement such as:

Employees applied the escalation rule correctly to unfamiliar customer cases.

Transfer is especially important for judgment, troubleshooting, compliance decisions, leadership, and other work that cannot be reduced to one predictable script.

Level 6: Workplace application

The question is:

Did the learner use the capability in the real work environment?

Possible evidence includes:

  • direct observation;
  • reviewed work samples;
  • system logs;
  • quality-assurance records;
  • manager verification;
  • peer review;
  • completed operational tasks;
  • repeated behavior across relevant occasions.

This level moves beyond what the employee can do in a course or simulation and examines what they actually do at work.

However, workplace performance depends on more than learning.

An employee may know the correct action but lack time, permission, equipment, staffing, software access, managerial support, or psychological safety. The environment may support or block the behavior.

A workplace observation can support:

Employees used the approved reporting process during relevant incidents.

It should not automatically be interpreted as a pure measurement of learning. It reflects capability operating inside a larger system.

Level 7: Organizational result

The question is:

Did a meaningful business, safety, quality, workforce, or customer outcome change?

Possible evidence includes:

  • error rates;
  • incident rates;
  • customer-service quality;
  • time to proficiency;
  • rework;
  • compliance findings;
  • employee retention;
  • productivity;
  • operational cost;
  • service consistency.

This level matters because organizations invest in learning to support meaningful outcomes—not merely to produce quiz scores.

But an improved result does not prove that training caused the change.

The CDC distinguishes outcome evaluation from impact evaluation. Outcome evaluation examines whether intended outcomes occurred, while impact evaluation attempts to determine whether the program caused those outcomes by comparing them with an estimate of what would have happened without the program.

A company may introduce training while also changing its software, staffing, supervision, incentives, policies, or reporting system.

The careful statement is:

Error rates declined after the training and process changes were introduced.

A stronger causal claim requires a stronger evaluation design.

Seven connected evidence levels progress from participation to organizational results.

The levels are not a causal staircase

The framework does not mean:

Participation causes reaction.

Reaction causes learning.

Learning automatically causes workplace behavior.

Workplace behavior automatically creates business results.

The levels represent different questions, not guaranteed links.

A learner may dislike a course but still learn from it. An employee may learn the correct procedure but be unable to use it because the required tool is unavailable. A business result may improve even when employee learning remains unchanged because a software control removed the opportunity for error.

A useful working chain is:

Learning experience → change in capability → application in context → contribution to a workplace result

Every arrow in that chain is a hypothesis.

Evaluation should examine the links that matter instead of assuming they are automatic.

Choose the minimum sufficient evidence

Higher levels are not always necessary.

A low-risk announcement about an office closure may require evidence that employees received and understood the message.

A product-knowledge update may require an immediate check and a small delayed sample.

A customer-support procedure may require an independent task and a varied scenario.

A safety-critical process may require immediate performance, delayed retention, transfer, and workplace verification.

Leadership development may require repeated behavioral evidence collected over time because complex behavior depends heavily on organizational context.

The goal is not to collect every possible metric.

The goal is to collect enough evidence to support the claim responsibly.

Use these questions when selecting the minimum evidence level:

  • What capability is the training expected to produce?
  • How costly would false confidence be?
  • How long should the capability remain available?
  • Must the employee handle variation?
  • Will the work environment provide tools or job aids?
  • Is real workplace observation practical and ethical?
  • What claim will leaders make from the result?

The CDC recommends creating an evaluation plan early by defining the purpose, questions, stakeholders, and data-collection methods rather than adding evaluation after the training is complete.

Evidence quality changes what each level means

The level alone does not determine evidence quality.

An immediate task may be well aligned and carefully designed. A workplace observation may be inconsistent, biased, or unrelated to the training objective.

Before interpreting evidence, examine these conditions:

Alignment

Does the activity measure the stated objective?

A recall question is poorly aligned with an objective requiring independent equipment operation.

Independence

What help was available?

A task completed with step-by-step coaching means something different from the same task completed independently.

Delay

How much time passed after instruction?

A test given two minutes later and one given two weeks later answer different questions.

Variation

Is the task new, or is it an exact repetition of practice?

New but relevant cases offer stronger evidence of transfer.

Authenticity

Does the assessment resemble the real decision or task?

A definition question may not represent the pressure and ambiguity of a workplace situation.

Consistency

Would different evaluators or repeated measurements produce a similar conclusion?

One manager’s informal impression may be less dependable than a shared observation rubric.

Context

What helped or blocked performance?

Tools, staffing, incentives, workload, supervision, and team culture can affect workplace evidence.

A metric should always be reported with enough context to explain what was actually measured.

The same learner performs with guidance, independently, after a delay, and in a changed workplace situation.

A hypothetical workplace example

Imagine a customer-support team introducing a new refund-escalation procedure.

The objective is:

Given a customer case containing an exception, the employee will select the correct refund or escalation path and explain which policy condition applies.

The organization could collect evidence at all seven levels.

Participation

Did assigned employees complete the course?

Reaction

Did employees find the cases relevant and the procedure understandable?

Immediate performance

Could employees select the correct action immediately after training?

Were they using the policy document, hints, or feedback?

Delayed retention

Could they explain the main escalation rules one or two weeks later without reopening the lesson?

Transfer

Could they handle new customer cases containing details that were not shown during training?

Workplace application

Did quality reviews show employees using the correct escalation path during real interactions?

Organizational result

Did incorrect refunds, avoidable escalations, resolution time, or customer complaints change?

The organization should not collapse these measures into one “training effectiveness score.”

Each measure tells a different part of the story.

Suppose course completion is high, delayed performance is moderate, and workplace errors remain unchanged. The next action may involve more retrieval practice, better job aids, clearer system permissions, or process redesign.

The framework helps the team investigate rather than jump to one conclusion.

A customer-support employee progresses from training participation to applying a refund procedure during real work.

Common measurement mistakes

Calling completion learning

Completion supports an access or participation claim. Add evidence aligned with the intended capability.

Using satisfaction as proof of effectiveness

Use reaction data to improve relevance and usability—not as evidence of retention or performance.

Calling one immediate quiz mastery

A mastery claim should rely on a documented rule using multiple aligned pieces of evidence under defined conditions.

Using identical questions to claim transfer

Change the case, wording, context, or task while preserving the underlying objective.

Always aiming for organizational results

Some programs are too small, early, or low-risk to justify expensive impact evaluation. Choose proportionate evidence.

Ignoring available support

Record whether employees used hints, references, AI assistance, checklists, or coaching. Support changes what the result means.

Ignoring the work environment

Training cannot compensate for missing tools, conflicting incentives, unsafe workloads, or unclear authority.

Collecting unnecessary employee data

Measure only what is needed for the evaluation purpose. Consider privacy, access, retention, fairness, and the consequences of incorrect interpretation.

Applying the framework in LoreGraph

A document-to-course system should not treat every click and quiz score as equivalent evidence.

LoreGraph’s mission is to help organizations turn existing documents and workplace knowledge into structured learning, practice, assessment, and measurable progress.

The framework suggests several practical product principles:

  • Every important objective should map to an evidence type.
  • Analytics should distinguish access, completion, immediate performance, retention, transfer, and workplace evidence.
  • Assessment records should include conditions such as hints, retries, source access, and delay.
  • Delayed checks should use new items rather than repeating the same questions.
  • “Mastery” should be presented as an inference from a documented model.
  • High-risk training should require stronger evidence and human review.
  • Workplace evidence should be collected only when appropriate, proportionate, and interpretable.
  • Organizational result indicators should be labeled as outcomes with possible multiple causes.

This approach helps course creators ask a better question before publishing:

What evidence would justify the claim we intend to make?

Teams exploring document-based workplace learning can review LoreGraph for business, but the framework remains useful regardless of which platform or authoring process they use.

Next step

Choose one important objective from an existing onboarding, compliance, SOP, or skills course.

Write down:

  1. The exact capability the learner should develop.
  2. The current evidence being collected.
  3. The strongest claim that evidence can honestly support.
  4. The minimum additional level needed for the risk and purpose.
  5. The conditions that could weaken the interpretation.
  6. The workplace factor most likely to support or block performance.

Then change one dashboard label, assessment, or evaluation step so that the claim matches the evidence.

Do not begin by asking how many metrics you can collect.

Begin by asking what you need to know.

Sources and further reading


Alireza Ibrahimi

Founder, LoreGraph

Software engineer and Learning Engineering researcher building AI systems that transform workplace knowledge into measurable learning experiences.

Put it into practice

Create structured training with measurable progress

Give your team training with practice, assessment, and progress you can actually see.

Explore team training

Keep reading