Skip to content
Jay Moore

Design rationale

Most courses can only prove someone finished. This one can prove someone changed.

The simulation is the artifact. This page is the reasoning behind it: what the objective demanded, why the authoring tool could not carry it, and where the design puts its guardrails.

Evaluation designed · Level 3

Objective

Given a defensive employee, deliver below-expectations feedback that leads with observed behavior and closes on a dated commitment.

The Bloom's verb is demonstrate, not identify and not describe. That distinction is the whole design. A learner can identify good feedback on a multiple-choice item without ever being able to produce it while someone is pushing back at them.

The hard part of this skill is not knowing the rule. Nearly every manager can recite “be specific.” The hard part is staying specific in the ninety seconds after someone says “that is not fair.” That moment is what has to be practiced, and it cannot be practiced against static content.


Tool ceiling

Why Rise could not carry this one.

I build in Rise, and two of the case studies on this site are Rise courses. It handles responsive, text-and-media training well, and when the tool fits I use the tool.

What Rise cannot do is respond to a sentence nobody wrote in advance. Scenario blocks and branching are still scripted: a human authors every path, and the learner's job quietly becomes finding the path the author already picked. That is recognition practice wearing the costume of production practice. It teaches learners to select the good answer from a list, which is not the skill.

It also does not scale in the direction the content needs. Every additional realistic response is more authored branches, and the tree grows faster than anyone will maintain it.


Rubric

Each item is an observable behavior, not a quality.

A rubric item earns its place only if a manager's peer could watch the real conversation and agree on whether it happened. That rules out “showed empathy” and rules in the four below.

Led with specific observed behavior. On the job: the employee can picture the moment being described. Character judgments start an argument about identity; observations start a conversation about work.

Avoided generalizing language. On the job: “always,” “never,” and “everyone thinks” hand the employee a factual claim to disprove, and they usually can. One counterexample and the conversation is now about the exception.

Checked the employee's understanding. On the job: the manager stops and finds out what landed before moving on. Skipping it is how a manager leaves a meeting believing an agreement exists when it does not.

Agreed a concrete, dated next step. On the job: this is the item that makes the other three matter. Without a specific commitment with a when attached, the conversation produces feelings and no change.


Evaluation

What Level 3 measurement would look like in a real rollout.

Nothing on this page claims the simulation changes behavior. It has not been run with a cohort, and inventing a number would defeat the point of the artifact.

What it is built to support is the measurement. In a real rollout, Level 3 evidence is behavior on the job, not performance in the practice: rate of performance conversations that end with a documented dated commitment, manager self-report at thirty and ninety days paired with the direct report's read of the same conversation, and the gap between the two.

The simulation feeds that by emitting the things a Level 3 study needs a baseline for: which rubric items a manager meets unprompted, which ones they only reach after the character pushes back, and how those move on a second attempt weeks later. Per-item pass data with the transcript span attached is a far better pre-measure than a completion record, because it says which specific behavior to go look for at work.


Trust

What happens when the model is confidently wrong.

The scoring pass is a second model call, and the failure mode that matters is not a wrong score. It is a fabricated quote: the model asserting the learner said something they never said, in a format that looks exactly like evidence.

A learner who is shown a quote they do not recognize stops believing the entire debrief, and they are right to. So every returned quote is checked against the learner's own turns before anything renders. If the span is not literally in the transcript, it is dropped and the item is reported without evidence, labeled as such. The system is allowed to be uncertain. It is not allowed to invent the record.

This is the same trust pattern as the missed-call response system in my product work: surface the confidence, quote the source verbatim, and leave a human an override path. A system that says “I could not find support for this” is more usable than one that is confidently wrong, because the first can be trusted about everything else it says.


Limits

What this is not.

One scenario, no accounts, no saved history, and a hard stop at twelve turns. There is no LMS behind it and no analytics dashboard. Those were scoped out on purpose: they are build work, not design evidence, and adding them would not make the argument any stronger.

The rubric is also deliberately short. Four observable items that a peer could agree on beats twelve items that need a scoring guide, and a short rubric is what makes the evidence quotes readable.