EnlightenWorks Free guide

Commissioning Learning & Development in the UK Public Sector

A 2026–27 guide for people who buy training.

By EnlightenWorks ~15 min read Reviewed August 2026

In short

Who this is for

You have been handed a budget and a problem: a team that needs to work differently, a new duty to land, a capability gap someone has named. A supplier will design the training. Your job is to commission it well: scope it, specify it, choose between bids, and judge whether what came back is any good.

You do not need to know how to design a course to do that well. You need to know what good looks like, what to ask for, and what to be wary of. That is all this document is. It assumes you know your own governance, and assumes no learning-design knowledge.

Where this guide stops: procurement law

This is a guide to judging the quality of what you buy, not to the law of buying it. Most procurements started in England, Wales and Northern Ireland since 24 February 2025 run under the Procurement Act 2023 and the Procurement Regulations 2024; Scotland has its own regime. Thresholds, frameworks and dynamic markets, transparency notices, conflicts, social value and your standing orders sit with your procurement team.

One rule from that world matters throughout this guide: award criteria and scoring must be clear, proportionate, related to the contract and disclosed in advance. Anything here you intend to score, including the ten markers and the seventeen questions, belongs in your published evaluation methodology, not in an undisclosed checklist.

Before you commission anything: is training the answer?

The most expensive commissioning mistake is buying a course for a problem a course cannot fix. Done well, training changes what people know, what they can do, and how they understand and approach the work. What it cannot change is the conditions around the work. If people already know what to do but do not do it, the cause is usually somewhere else: the process is unclear, the tools are poor, the incentives point the other way, or no one has the time. A course will not touch any of that, and you will have spent the budget proving it.

Before you write a line of specification, ask:

  1. What, specifically, would people be doing differently if this worked? Name a behaviour, not a topic.
  2. Do they not do it now because they cannot, or because something stops them? If something stops them, that is not a training problem.
  3. If I changed only the conditions around the work and trained no one, would the problem shrink? If yes, fix the conditions first, or alongside.

A good supplier will ask you these questions unprompted, and will sometimes tell you that you do not need them.

Rule of thumb

Training is most likely to help where people lack knowledge, skill, practice or confidence, and have the support to apply it. Where the blocker is elsewhere (unclear process, wrong tools, weak incentives, no time), commission the fix, not a course.

The 2026–27 context you are buying into

Working assumptions as of August 2026. Check each against your own organisation; several are contested and move quickly.

What good looks like, so you can recognise it in a bid

You do not need to design the training to judge a proposal. These are the ten markers of well-designed work; a strong bid shows most without being asked.

  1. It starts with behaviour, not topics. Objectives say what a learner will be able to do, with a condition and a standard. "Understand" is too vague on its own; look for an observable demonstration (explain, apply, compare, diagnose). Its overuse is a tell.
  2. It is designed backwards from the outcome. Decide the evidence first, design the experience to produce it. Slides come last.
  3. It shows a nuanced appreciation of the learner journey. The design works within your time constraints and organisational context, and thinks about what happens before, what follows after, and what the learner keeps using.
  4. Every live minute is justified. Synchronous time is spent on what needs people together: practice with feedback, real-time discussion, emotionally weighty content.
  5. It produces something people keep. A job aid or template that outlives the session.
  6. It is accessible by design. WCAG 2.2 AA, captions, plain English, inclusive examples, built in from the start.
  7. It is evaluated at the level of behaviour. How you will know, at 30–90 days, whether anything changed.
  8. It is honest about what it cannot do. A proposal with no risks listed has not been thought through.
  9. It is specific to your context. Generic content that could be sold to anyone probably was.
  10. A named human stands behind it. Someone puts their name to the design and the facts, alongside the organisation's own quality assurance.
Rule of thumb

Score a proposal 0–2 on each marker, as a working heuristic rather than a validated threshold. Below about 12 out of 20, ask the supplier to strengthen the weak areas first; if the score will affect the award, put it in your published evaluation methodology. It is cheaper to fix a bid than a delivered course.

Telling genuine expertise from AI slop

Many suppliers now use AI somewhere in production, and that is fine: used well, it makes good designers faster and sharper. The problem is not AI. It is AI slop: generic, unchecked, outcome-free content dressed up as expertise. As a commissioner you need to tell one from the other, because you are paying for judgement, not for text.

The tell is not whether AI was used. It is whether human judgement, verification and context are visible in the work.

Green flags versus slop, at a glance.

Expertise (AI-assisted or not)Slop
Objectives with verbs, conditions, standards, tied to your problem Objectives full of "understand", "be aware of"
Specific, current facts, visibly checked and sourced Confident statistics with no source
Named trade-offs and honest limitations No risks or limits listed
An evaluation plan that could embarrass them if it failed A happy-sheet, or no evaluation at all
Examples that could only be about your context Content that fits anyone, names swapped
A named person accountable Design myths as fact: learning styles, the 10%/90% pyramid

This is measurable. In July 2026 Dr Philippa Hardman tested four frontier AI models against a 36-point instructional design standard: execution quality varied widely, design judgement scored far lower in every model, and not one questioned whether a course was the right solution. Three of the four also introduced errors of omission (required duties and escalation routes silently dropped) that a fact-accuracy review would not catch (Hardman, AI Can Build Your Course, but Can it Design the Learning?, July 2026).

Signs a supplier is at the leading edge

The flags above are the baseline. A smaller group of suppliers is further ahead in how they work with AI, and it shows when they talk about their own practice. None of these is essential, but each is a good sign that you are dealing with an informed, adaptive user rather than a subscriber:

Rule of thumb

Ask one question: "What in this proposal did a human decide, verify, or take responsibility for?" If the honest answer is "very little", you are being sold slop at a specialist's price.

The paradox worth naming to your panel: good AI-assisted work and pure slop can look similar at a glance. The difference shows up in the specifics, the verification and the accountability. Commissioning well means looking there, not at the polish.

What to put in your specification

A tight spec gets you better bids and a fairer evaluation. Each line below can be lifted into a tender.

Purpose and outcomes

Design and delivery

Accessibility, inclusion and the law

AI use by the supplier

Many suppliers now use AI somewhere in production. The spec should not ban it; it should set the terms, and it should keep legal duties separate from the controls you choose to impose.

Evaluation, commercials and rights

Seventeen questions to ask a supplier

Sharp questions expose weak proposals faster than a scoring grid. If the answers will affect the award, build the questions and their scoring into your published evaluation methodology rather than applying them as undisclosed criteria.

  1. What must a learner be able to do differently after this, and how will we see it?
  2. Is training the right fix here? What part of the problem will it not touch?
  3. How does the design fit our learner journey, our time constraints and the context it will land in: what happens before and after the day itself?
  4. Which parts of the live time could not be done any other way, and why?
  5. What will a learner still be using in three months?
  6. How have you designed this for someone neurodivergent, using a screen reader, or joining from home?
  7. How will we know, at 60 to 90 days, whether behaviour changed?
  8. Which of your success measures could look good even if the course failed?
  9. How did you use AI in producing this, and how did you verify what it produced?
  10. What did you decide not to include, and why?
  11. Who, by name, stands behind the content and its facts?
  12. What are the three biggest risks to this working, and your plan for each?

And on AI:

  1. Where does AI sit in your production process, and what of ours goes into it?
  2. Is any of our data processed by an AI tool? Under what data processing agreement, with which sub-processors, and where is the data held?
  3. Will anything we give you be used to train a model?
  4. Who verifies AI-produced content, and who is accountable, by name, when it is wrong?
  5. Where do you expect AI to create a material productivity benefit, and how does that assumption affect your price, timetable and quality?
Rule of thumb

A credible supplier answers proportionately, and says where an answer depends on final design, security review or contract terms. Evasive, inconsistent or unsupported answers, particularly on evaluation (7, 8 and 9) and AI (13 to 17), warrant clarification and may signal risk.

Evaluation: what evidence to require

Most training is evaluated with a "happy-sheet" and nothing more. That tells you whether people enjoyed the day, which is a weak guide to whether they learned anything or changed how they work. Require more, and specify it up front.

Two cautions for the panel: these four are not a chain: a warm reaction does not predict learning, and learning does not guarantee changed behaviour. And ask for two leading indicators (drop-off, engagement, participation) that tell you it is working before the behaviour data lands.

Accessibility, inclusion and the law you are accountable for

For a public sector body, accessibility is a legal duty, and the accountability sits with you as much as with the supplier. Specify it, and check it on delivery.

Rule of thumb

If a proposal treats accessibility as a single reassuring sentence, treat that as a red flag. Compliance is detailed or it is absent.

Green flags and red flags at a glance

A one-look guide for reading a proposal.

Look forBe wary of
Objectives with verbs, conditions, standards Objectives full of "understand" / "be aware of"
Designed backwards from a named outcome Starts from content and topics
Fits the learner journey, with follow-through A single event, nothing after it
Live time justified minute by minute A day of talking at people
A lasting job aid or template Slides and a certificate
Accessibility built in, in detail One reassuring sentence about access
Behaviour-level evaluation in the bid Happy-sheet only, or no plan
Specific to your context Generic, sector-agnostic, name-swapped
Verified facts, cited sources Confident, unsourced statistics
A named human accountable No one's name on it

Sources checked

Time-sensitive claims in this guide were checked against the following sources on 5 August 2026.

A glossary for new commissioners

ADDIE
A systematic design approach (Analyse, Design, Develop, Implement, Evaluate), often run sequentially though it can be iterative. Suits settled, high-stakes content.
SAM
Successive Approximation Model. Iterative and prototype-driven; often used for fast-moving topics.
Backwards Design
Designing from the outcome and its evidence first, activities last. A good sign in a proposal.
Bloom's taxonomy
A ladder of cognitive verbs (remember, understand, apply, analyse, evaluate, create) used to write assessable objectives.
Kirkpatrick's four levels
Reaction, learning, behaviour, results: the standard, imperfect frame for evaluation. Four questions, not a guaranteed chain.
Synchronous / asynchronous
Live (a workshop, a call) versus self-paced (reading, video, e-learning).
Job aid
A checklist, template or one-pager used at the point of work. Often the most valuable output.
WCAG 2.2 AA
The accessibility standard your digital content must meet.

A one-page commissioning checklist

Need a hand?

EnlightenWorks designs and delivers professional training. If you would like help shaping a specification, talk to us.