Perspective

Kirkpatrick Fails When We Use It Too Late.

Evaluation fails when it becomes a post-launch reporting ritual. Better evidence starts while the problem, behavior, and performance system are still being defined.

Most organizations do not begin thinking seriously about evaluation until the learning experience is nearly finished.

The course has been built. The launch date is approaching. Stakeholders want to know what will be measured.

So the learning team reaches for the Kirkpatrick model.

Level 1: Did people find the experience useful?

Level 2: Did they learn something?

Level 3: Did their behavior change?

Level 4: Did the organization achieve better results?

The framework provides a familiar structure, but by this point, some of the most important evaluation decisions have already been made. They just were not recognized as evaluation decisions.

The business problem was defined without establishing a baseline. The desired behavior was never made observable. The learning solution was selected before the performance barriers were understood. The organization launched a course without deciding what evidence would demonstrate that it had actually worked.

Evaluation is expected to recover what the design process failed to define. It cannot.

The problem is not necessarily Kirkpatrick

Kirkpatrick is often criticized because organizations rarely move beyond Levels 1 and 2. Learning teams collect satisfaction surveys, report completion rates, and occasionally administer a knowledge check. Behavior change and business results remain much harder to prove.

But the problem is not simply that teams stop too early in the model.

The deeper problem is that they start too late.

You cannot wait until after launch to decide what successful performance looks like. You cannot measure behavior change if the target behavior was never clearly defined. You cannot prove business impact if nobody established the conditions that existed before the learning intervention.

By the time the evaluation plan is created, the team may discover that it does not have access to the necessary data, managers are not prepared to observe the behavior, or the business metric is influenced by several factors beyond learning.

Those are not reporting problems. They are design problems.

Evaluation should begin before the solution exists

Evaluation should start when the request first arrives.

When a stakeholder says, “We need a course,” the first question should not be about the content, length, or delivery date.

The first question should be:

What should people be doing differently as a result?

That question begins the evaluation process because it forces the team to define the performance change before building the learning product.

From there, we need to understand what people are doing now, what they should be doing instead, what is preventing that behavior today, what evidence would show that the behavior has changed, and what organizational result should improve if it does.

These questions are often treated as front-end analysis. They are also the foundation of evaluation.

Analysis, design, and evaluation should not operate as separate phases. They are different views of the same performance problem.

Start with evidence, then design backward

A stronger approach is to define the evidence of success before deciding what to build.

If employees need to make better judgment calls, what would better judgment look like in their work?

If managers need to provide more effective feedback, what observable behaviors would demonstrate improvement?

If a new process is supposed to reduce errors, what is the current error rate, and what change would be meaningful?

Once the evidence is clear, the learning experience can be designed to produce it.

Practice activities can mirror the decisions people must make. Feedback can address the reasoning behind those decisions. Performance support can be placed inside the workflow. Managers can be prepared to reinforce and observe the new behavior. Data collection can begin before launch instead of being improvised afterward.

Key principle

Evaluation stops being a report attached to the end of the project. It becomes part of the product architecture.

Completion is not evidence of performance

One reason evaluation is often delayed is that our systems make activity data easy to collect.

We can quickly report how many people enrolled, completed the course, passed the quiz, or rated the experience positively.

Those numbers may tell us whether people interacted with the learning product. They do not necessarily tell us whether the product improved performance.

A learner can complete a course and continue making the same mistakes. They can pass a knowledge check and still struggle to apply that knowledge under pressure. They can give the experience five stars because it was short, polished, and easy to complete.

None of that proves the original problem was solved.

This does not make completion, satisfaction, or knowledge data useless. It means we need to stop asking those measures to carry more meaning than they actually hold.

Learning does not operate by itself

There is another reason we need to begin evaluation earlier: learning is only one part of the performance system.

People may know what to do and still be blocked by confusing processes, inadequate tools, competing priorities, missing feedback, or incentives that reward the wrong behavior.

If those conditions are not identified during analysis, the learning team may be held responsible for a result that training alone could never produce.

Beginning with evidence helps reveal those dependencies.

This is the behavior learning can help develop. These are the conditions the organization must provide. This is the evidence we will examine together.

That is a much more credible evaluation strategy than launching a course and hoping a business metric improves.

Use Kirkpatrick as a design conversation

Kirkpatrick can still be useful, but we should stop treating it as a staircase we begin climbing after launch.

Its questions should influence the project from the beginning.

Before development starts, we should already understand the organizational result we hope to influence, the behavior required to support that result, the capabilities people need to demonstrate, and the experience necessary to help them build those capabilities.

In other words, we should think backward.

Not because every learning initiative can prove direct causation at Level 4. Many cannot. Business results are complicated, and learning rarely operates in isolation.

But we can create a defensible chain of evidence.

We can show that the experience developed a relevant capability. We can examine whether that capability appeared in the workplace. We can determine whether the behavior contributed to a meaningful operational result.

That is far stronger than presenting completion rates as proof of success.

Evaluation is not the final step

The most important shift is not adopting a new evaluation model.

It is changing when evaluation begins.

Evaluation begins when we define the problem. It becomes more precise when we identify the behavior. It becomes measurable when we establish the evidence. It becomes useful when that evidence shapes what we design.

If we wait until launch to ask how we will evaluate the work, we are not beginning evaluation.

We are trying to reconstruct it.

And by then, the evidence we needed may already be gone.

Application Pack · Inside The Lab

Evaluation & Evidence Planner

Turn the argument into a working evaluation strategy before development begins. Define the performance problem, target behavior, baseline, evidence chain, system dependencies, and measurement plan in one connected workflow.

My Library lives inside The Lab, where members can access their resources and application packs in one place.

Included
  • Performance Problem Framer
  • Behavior Evidence Map
  • Baseline Planning Worksheet
  • Performance System Check
  • Evaluation Conversation Guide

Source notes

  1. Donald L. Kirkpatrick and James D. Kirkpatrick, Evaluating Training Programs: The Four Levels, third edition, Berrett-Koehler Publishers, 2006.
  2. James D. Kirkpatrick and Wendy Kayser Kirkpatrick, Kirkpatrick’s Four Levels of Training Evaluation, ATD Press, 2016.
  3. Robert O. Brinkerhoff, The Success Case Method: Find Out Quickly What’s Working and What’s Not, Berrett-Koehler Publishers, 2003.
  4. Dana Gaines Robinson and James C. Robinson, Performance Consulting: A Practical Guide for HR and Learning Professionals, second edition, Berrett-Koehler Publishers, 2008.

Originally discussed on LinkedIn

Want to add your take?

Join the conversation about moving evaluation upstream and designing evidence before launch.

Follow the conversation →