AI at work

After an AI workshop, measure one real task with three tests

A small task-level trial can show whether workshop learning improves real work, respects information boundaries and holds up under human review.

An AI workshop can finish with full attendance, lively discussion and a folder of promising examples. A week later, managers still need to know whether anyone can use the learning on real work.

That question affects learning teams, operational leads, managers and the colleagues expected to change how they work. Attendance shows reach. It does not show whether a task became quicker, more accurate or safe enough to repeat.

OpenAI's Academy deployment guide suggests looking at five signals: completion and awareness, then application, adoption and progression. Anthropic says its enterprise guide covers adoption, efficiency, quality and satisfaction. Together, those categories give a useful frame. The practical unit of measurement should be a recognisable task with a clear human owner.

Decide What

Decide what success means before the trial

Choose one recurring task that appeared in the workshop. It should be common enough to matter and small enough for a reviewer to inspect from source to final output. For example, a team could test whether an approved AI tool can turn cleaned meeting notes into an action summary for internal circulation.

Write down the current method before anyone uses AI. Record the active time it takes, the checks a competent colleague performs and the mistakes that would make the result unusable. This is your baseline. It need not be an elaborate study; it needs to be consistent enough for a fair comparison.

Use a small scorecard that keeps different kinds of evidence separate:

  • Participation: who attended or completed the learning.
  • Application: who could complete the chosen task with the method taught.
  • Efficiency: active working and review time compared with the current method.
  • Quality: factual errors, omissions and corrections required before use.
  • Satisfaction: whether the user and reviewer found the method workable.
  • Safety: whether participants respected the agreed information boundary.

Do not compress these measures into one success percentage. A quick output with a wrong deadline has a different consequence from a slower output that needs a minor wording edit. Keep the categories visible so the team can decide which problems training can fix and which ones make the workflow unsuitable.

Smallest Useful

Run the smallest useful trial

Use a handful of workshop participants who already perform the task. Give each person the same cleaned source pack, the same instructions and the same approved tool. Ask them to complete the task once using the current method and once with AI support. Record active working time and review time separately.

For the meeting summary example, the source pack might contain an agenda and notes. It can also include attendees and an agreed output format. Use artificial or properly anonymised material for the first trial. The purpose is to test the method without introducing live client, employee or commercially restricted information.

Keep the trial narrow. You are trying to answer whether this task can be completed to the agreed standard. Wider usage data becomes more useful after the team has evidence that the underlying work is sound.

Give Tool

Give the tool enough context to do the task

The tool needs the working context that an experienced colleague would use. A short prompt alone rarely contains enough information for consistent work.

Provide:

  • The purpose of the output and who will read it.
  • The cleaned source material the output must rely on.
  • Definitions for fields such as owner, due date and decision.
  • The required format and acceptable length.
  • Relevant house style or process rules.
  • One approved example if the organisation owns and can share it.
  • An instruction to mark uncertainty and request missing information.

Tell the tool what it must not do. For this example, it should not infer an owner from seniority, convert a discussion into a decision or choose between conflicting dates. Those boundaries make review easier because the expected behaviour is explicit.

Keep Unnecessary

Keep unnecessary information outside the tool

The first trial does not need live sensitive data. Keep passwords, access tokens, private keys and authentication details out of prompts. Exclude unnecessary personal information and any client or commercial material that the organisation has not approved for use in that tool.

An approved business account does not remove the need for an input boundary. The team still needs to know which tool and account it may use, what categories of information are allowed and where outputs may be stored. Ask the people responsible for information security, data protection and the relevant business process to set those rules.

If the task only works when participants paste unrestricted records into an unapproved tool, the method has failed the trial. Redesign the input, change the approved environment or choose a different task.

Make Human

Make the human review observable

A statement such as “a person will check it” is too vague to measure. Give the reviewer the source material, the draft output and a short checklist. The reviewer should be someone who understands the task and has authority to reject the result.

For an action summary, the review can cover:

  • Trace every decision, action, owner and date to the source notes.
  • Mark each unsupported addition, omission and changed meaning.
  • Check that the output follows the agreed format and audience.
  • Confirm that restricted information has not entered the input or output.
  • Approve, return for correction or reject the result.

Record corrections by type and severity. A spelling change is not equivalent to an invented commitment. Also record how long the review took. If drafting time falls while checking time rises sharply, the workflow may have moved work rather than reduced it.

The human reviewer remains responsible for release. The AI output is a draft until that person has compared it with the source and applied the agreed standard.

Test Method

Test the method three ways

A polished example from the workshop only shows that one favourable case worked. Use the same instructions and review method on three cases:

  • Normal test: complete notes with one clear owner and deadline for each action. The output should capture the source accurately in the required format.
  • Missing information test: remove an owner or due date. The output should mark the field as unknown or ask for clarification. Assigning a plausible person or date is a failure.
  • Realistic edge case: include a deadline that changes during the meeting, such as Friday in the discussion and Tuesday in the final exchange. The output should flag the conflict or use the final decision only when the notes make that status explicit.

For every case, capture active time and review time. Record correction count, error severity and whether the result could be used after review. This turns a demonstration into a basic test of reliability.

The missing information case often teaches more than another perfect example. It shows whether the instructions let the tool admit uncertainty and whether participants recognise when to stop and ask a colleague.

Read Results

Read the results without rewarding activity

Course completion and workshop attendance help you understand reach. OpenAI also recommends collecting workflows, questions, survey responses and team examples to assess application. Its guide presents workspace usage as an adoption measure and repeatable workflows as a sign of progression. These are vendor recommendations, and the organisation should interpret them alongside its own task evidence.

Look for patterns across the scorecard. Strong attendance with weak task quality points to a gap between learning and application. Faster drafting with frequent factual corrections calls for better context, tighter instructions or a different task. High user satisfaction with low reviewer confidence means the method is not ready to spread.

Set a threshold that reflects the consequence of error. A low risk internal draft may tolerate small edits. A customer commitment, employment decision or regulated record needs a different process and may be unsuitable for this kind of trial. A business owner should make that call with the relevant specialists, rather than leaving it to the workshop facilitator or the tool.

OpenAI's guide suggests sharing a deployment readout after 8 to 12 weeks or monthly. That later view can show whether the practice lasts. Keep the first task trial close to the workshop so participants can use what they learned and surface missing support.

Turn Successful

Turn a successful trial into a team method

A repeatable workflow includes more than saved instructions. Package the parts that made the result reviewable:

  • The task definition, intended audience and named owner.
  • The approved tool, account and allowed information categories.
  • The input template and redaction steps.
  • The versioned instructions and an approved example.
  • The human review checklist and release authority.
  • The normal, missing information and edge case test pack.
  • The measures to collect and the person who reviews them.

Store this method where the team keeps normal process documentation. Give it a version and record why it changes. Run the three tests again when the instructions or source format changes. Repeat them after a tool, model or information policy change that could affect the result.

Managers can reinforce the practice by asking for examples in team meetings and directing difficult cases to an office hour or named support channel. This follows the reinforcement approach in OpenAI's guide while keeping the discussion attached to work people perform.

As more colleagues use the method, track whether quality holds across users and whether review effort remains acceptable. That is stronger evidence of adoption than a rising login count on its own.

What After

What to do after the next workshop

Before the session closes, ask each team to nominate one recurring task and the person who owns its quality. Choose one task for the first trial. Prepare a cleaned source pack and agree the information boundary with the relevant internal owners.

Then run the normal case, the missing information case and the realistic edge case. Compare the AI assisted route with the current method using the same scorecard. Keep the work human led until the reviewer can trace the output to its source, identify uncertainty and reject weak results.

The most useful post workshop measure is whether a team can repeat one defined task under clear boundaries and accountable review. If it can, document the method and widen access carefully. If it cannot, the failed test tells you what the next training or process change needs to address.

Checked sources

Read the original reporting.

  1. Champion deployment guideOpenAI Academy
  2. Enterprise AI Transformation GuideAnthropic learning resources