Improving Assessment Quality using AI Feedback

5 minute read
summary

Picture a cockpit. High above the clouds, a pilot suddenly faces an engine vibration and a flickering dashboard. At that moment, they don’t reach for a mouse to “Select Option C.” They recall training, synthesize data, and act. A pilot who passes a 100-question multiple-choice test has shown they can read a manual. A pilot who survives a simulated bird strike has shown they can fly.

The multiple-choice question has been the comfort food of the L&D industry for decades. Predictable, easy to digest, and it grades itself. For simple, binary knowledge recall, it still has a place. But problems begin when we ask it to measure human judgment,  the kind of empathy and contextual reading needed to de-escalate a conflict with an exhausted customer. That, an MCQ simply cannot hold.

Research by Karpicke and Roediger (2008) on the Testing Effect points to a structural flaw in the standard approach: multiple-choice questions often measure recognition memory: the ability to spot a familiar answer, rather than genuine recall or application. Many learners pass courses because they’ve become “test-wise” mastering the skill of eliminating obviously wrong distractors. However, in reality, life lacks an “Option D: All of the above.”

Historically, evaluating nuanced, open-ended responses required a human mentor, a financial and logistical restriction at scale. Two shifts are changing that: Vibe Coding and AI-enabled assessment.

We made the wider case for retiring the default quiz in death of the multiple choice question, and for treating assessment as a design discipline in approaching eLearning assessment design like eLearning content design.

Vibe Coding

For years, assessment design was constrained by whatever question types came out-of-the-box in your authoring tool. Building a custom interactive sandbox meant hiring a developer and burning a large budget.

As we detailed in Vibe Coding: Support for Instructional Design Dreams, that friction has largely disappeared. Instructional designers can now use natural language to steer AI into writing functional code for custom interactions; no complex syntax required. And because instructional designers are already skilled at defining learning outcomes and breaking complex problems into steps, we are natural vibe coders. We provide the pedagogical context; the AI builds the interaction.

Consider a “Paint by Num-Birds” warbler identification tool we built directly into an Articulate Rise code block using AI-assisted code. Identifying warblers is notoriously difficult, even experienced birders struggle. Rather than fall back on standard multiple-choice image questions, we built a performance sandbox. The learner receives a blank bird diagram alongside field notes and anatomy files. They select body parts and assign colours from memory. Once they’ve coloured at least five parts, the ID engine cross-references their choices and returns possible species matches. Active interaction. Immediate feedback. We described the vision; AI wrote the code; and voila! The technical barrier to creativity has gone.

Warbler ID Buddy

AI-supported Assessment

When a learner submits a response to a complex scenario, they receive instant, multi-dimensional coaching. We call this the Artha 4-Point Return:

A Quantitative Score: How well did the response align with specific criteria.

Narrative Feedback: A conversational explanation of why the answer worked or didn’t.

Specific Strengths: Highlighting exact positive behaviours. For example, acknowledging a customer’s frustration immediately.

Identified Gaps: Pinpointing precise areas for improvement, like forgetting to mention the 24-hour turnaround policy.

This is feedback that previously required a skilled coach sitting across the table. Now, it scales to 2,000 learners as easily as it does to 20.

Judgment-free practice at scale

Take a leadership practice scenario as an example. A manager-in-training opens with: “I’ve noticed tension between you and Alex. Can we talk about it?” The AI, playing a defensive colleague, fires back: “Sure, but Alex is the one making things difficult, not me.” The learner now has to navigate, not pick from a list of four responses. No script, no safety net. They make a move, the AI reacts, and the conversation unfolds in real-time based on their choices. This is the kind of practice that changes behaviour. 

Leadership Practice Scenario

You can explore more of what we’ve built and what is possible at aiforlnd.com; the feedback engine, the roleplay simulations, and digital coach, all designed to embed directly into existing courses.

The research supports the instinct. In a study from Stanford and MIT tracking 5,179 customer support agents, researchers found that AI assistance boosted worker productivity by 14% on average, with the biggest gains among newer, lower-skilled employees, who effectively compressed six months of experience into two. The mechanism is the same: better feedback, in the moment, accelerates the experience curve.

Turning interactions into intelligence

The value of this approach doesn’t stop at the individual learner. Every open-text submission becomes a data point. Analyzing patterns across hundreds of responses gives learning teams a clear picture of where the workforce is strong and where it’s fragile. Imagine spotting a trend where fifty employees consistently fail to apply a critical policy in their scenario responses. That’s a substantiated, targeted signal, that moves from just tracking completion rates.

AI-driven assessments don’t just improve individual performance, they surface the organizational gaps that traditional metrics never captured. Traditional L&D metrics tell us who clicked through a module. This tells us what they actually understand.

Assessment across domains

We apply this approach well beyond compliance training. For a medical education client, we built a clinical case coach that puts healthcare professionals inside a real patient scenario; full history, medications, comorbidities, and a curveball. In one case, a 50-year-old female patient with Type 2 diabetes, hypertension, and a grandfather who died from diabetes complications. The catch: she resists medication and fears its side effects. The learner isn’t asked to pick a treatment option. They’re asked: “How would you use motivational interviewing to explore her concerns and engage her in treatment planning?”

The AI evaluates the response across multiple criteria; did the learner explore her values, affirm her preference for natural approaches, resist the righting reflex, and invite collaborative planning? The feedback is specific enough to coach: “You missed an opportunity to affirm the patient’s preference for natural approaches. Try acknowledging their desires before moving to recommendations.” This provides meaningful feedback, exactly as a clinical supervisor would do it.

The same logic applies across professional domains. In a communication course, learners draft a social media post following best-practice marketing principles and receive feedback on whether the call-to-action lands. In a business context, a market analysis becomes the prompt, and the learner recommends a product launch strategy. Theory doesn’t stay theory for long when you’ve had to apply it under realistic pressure.

Stealth Assessment

The most powerful shift here is temporal. Traditional assessments are terminal; they happen at the end and report a grade. AI-driven assessments are continuous. Competencies are measured as they emerge, woven into the workflow itself. Learners often don’t notice the assessment is happening because the learning and the practice have become the same thing.

Where to begin?

Continuing to rely on the multiple-choice question is a choice to measure what’s easy to measure but may not be what really matters. The argument for improving assessment quality isn’t new. What’s new is that AI feedback makes it achievable at scale, without a dedicated developer, without a coaching budget, and without waiting for the next platform cycle.

Find one assessment in your next module where an MCQ fails to capture what the job actually requires. Replace it with an open-text scenario. Define the criteria. Let the AI carry the evaluation. That’s one piece of feedback that now tells you something real.

Everything described here runs on AIReady, which embeds the feedback engine, roleplay simulations and digital coach directly into your existing courses.

Better feedback produces better learners. Better learners produce better outcomes. The tools are there. The only thing left is the decision to use them well.

References

  1. Karpicke, J. D., & Roediger, H. L. (2008). The critical importance of retrieval for learning. Science, 319(5865), 966–968. https://doi.org/10.1126/science.1152408
  2. Li, D., Brynjolfsson, E., & Raymond, L. (2023). Generative AI at Work. NBER Working Paper. As reported by CNBC, April 25, 2023. https://www.cnbc.com/2023/04/25/stanford-and-mit-study-ai-boosted-worker-productivity-by-14percent.html

Sign up for our LinkedIn newsletter to receive updates on new eBooks, exclusive content, and the latest trends in learning and development.

Recent Blog Posts

Rather inconveniently, custom eLearning doesn’t have a rate card. Real projects range from about $4,000 for a focused microlearning piece...

By Garima Gupta, Founder & CEO, CTDPThere is a point in many branching scenario projects where things start to get...

By Ken Wheadon, Creative Production ManagerMany instructional designers and eLearning developers are discovering a frustrating truth: Building an AI agent...