Start With
Start with a work problem the team already recognises
A leadership team often receives the information it needs in fragments. Finance sends a forecast, sales provides pipeline commentary and operations reports delivery risks. Someone then spends hours turning those updates into a briefing for a monthly decision meeting.
That preparation affects the analyst assembling the material, the departmental leads checking it and the executives relying on it. It also makes a useful candidate for a controlled AI pilot because the work has a defined audience and a finished output that a person can review.
The OpenAI Academy guide for executive sponsors recommends connecting a rollout to one or two organisational priorities. It also advises starting with substantial knowledge work that can use approved context to produce a reviewable deliverable.
Choose one recurring briefing rather than a broad ambition to use AI across leadership. A suitable first workflow should meet most of these conditions:
- The team already completes it often enough to compare attempts.
- The source material has clear owners and can enter the approved tool.
- A named person can review the whole output.
- Errors can be caught before the work affects a customer or employee.
- The team can describe the current time, effort and quality well enough to form a baseline.
A board briefing, options paper or synthesis of departmental updates may fit. An employment decision, legal conclusion or response to a live security incident creates a harder first test because the consequences and information boundaries require more support.
Smallest Useful
Run the smallest useful trial in shadow mode
Start with one existing briefing and one person using the AI tool. Give the result to a reviewer who understands the subject. Keep the output inside the pilot and prepare the real briefing through the normal process.
This shadow exercise answers a narrow question: can the tool produce a useful first draft from approved material without increasing the reviewer’s total workload or obscuring uncertainty?
Define the trial before anyone opens the tool:
- The output will be a two page decision briefing for the monthly leadership meeting.
- The pilot owner will use only the approved briefing pack.
- The tool may summarise, compare and identify gaps.
- A named reviewer will check every material claim against its source.
- An executive will decide whether the briefing is suitable for the meeting.
- The team will stop if the tool repeatedly invents facts, mishandles restrictions or requires more correction than the current method.
Anthropic’s Enterprise AI Transformation Guide describes targeted pilots intended to demonstrate value within 30 to 60 days. A leadership team can use that as an outer window while keeping the first exercise much smaller. The window should contain a few real instances of the same workflow where its normal cadence allows.
Give Tool
Give the tool a bounded briefing pack
The tool needs enough context to perform the task without receiving unrestricted access to company information. Assemble a briefing pack for the chosen meeting rather than connecting every available folder or system.
Include:
- The decision the leadership team needs to prepare for.
- The audience and the action they may take after reading.
- Approved source documents with dates, owners and reporting periods.
- Definitions for internal terms and important measures.
- The required structure, length and level of detail.
- Known constraints such as budget limits or an agreed planning horizon.
- An instruction to cite each source and label missing or conflicting information.
The quality bar should describe a usable deliverable. For example, the briefing must separate observed facts from departmental forecasts, show the source for each major claim and identify questions that leaders need to resolve.
Avoid asking the tool to find the answer from whatever it can access. That makes it harder for the reviewer to know which evidence shaped the draft.
Keep Unnecessary
Keep unnecessary and restricted information outside
The pilot owner should agree the information boundary with the people responsible for security, privacy and the relevant business records. The provider’s terms and the organisation’s own policy will determine what the approved environment may receive.
Exclude material that the task does not require. Common examples include:
- Passwords, access tokens and security configuration.
- Personal employee or customer records when aggregated information will do.
- Legally privileged advice unless the organisation has approved that use.
- Undisclosed transactions, investigations or personnel matters outside the pilot’s authority.
- Whole mailboxes or shared drives supplied for convenience.
If nobody can confirm whether a source may enter the tool, leave it out and record the gap. The pilot can then test whether the workflow remains useful with a smaller information set.
Pass Human
Use a two pass human review
The first pass checks source fidelity. The reviewer should open the source pack and trace every material statement in the draft. They should check names, dates, figures and reporting periods. They should also mark unsupported interpretations and note any relevant source content the draft omitted.
The second pass applies business judgement. The reviewer asks whether the briefing frames the decision fairly, gives uncertainty enough prominence and distinguishes evidence from recommendations. They also check whether the tool compressed two different departmental views into a false consensus.
Record corrections by type:
- Factual error or unsupported statement.
- Missing qualification or source.
- Important omission.
- Misleading emphasis.
- Style or formatting change.
- Decision or recommendation that required human judgement.
The reviewer then records how long the checks and corrections took. A quick draft has little practical value if an expert must rebuild it before use.
The executive responsible for the meeting keeps final approval. The AI tool can prepare and organise material. It should not decide which commercial risk to accept or whose disputed forecast to trust.
Test Normal
Test the normal case, a gap and a conflict
Run the same workflow against three cases. Use the same output requirements and review method so the results can be compared.
- Normal test: Provide a complete approved pack with consistent reporting periods. Check whether the tool produces a well sourced briefing that needs limited correction.
- Missing information test: Remove a source needed for one section, such as the latest cash forecast. The useful response identifies the missing evidence and leaves the conclusion open. A fabricated estimate or an old figure presented as current should fail the test.
- Realistic edge case: Include two departmental updates that use different definitions for the same measure or give conflicting dates for a delivery milestone. The output should surface the disagreement, cite both sources and send the issue to a person for resolution.
These tests reveal more than a polished demonstration. The missing information case shows whether the workflow handles uncertainty. The edge case shows whether it preserves disagreement that leaders need to see.
Review Evidence
Review evidence rather than usage
OpenAI advises sponsors to examine useful outputs, value, adoption, friction, support needs and readiness to scale. Its guide warns against treating activity or credit consumption as proof of success. Anthropic lists adoption, efficiency, quality and satisfaction among the areas organisations can measure.
For this pilot, compare each attempt with the normal process:
- Was the final briefing usable after review?
- How much preparation time and review time did each method require?
- How many material corrections did the reviewer make?
- Did the briefing help leaders reach the meeting prepared?
- Did the same intended users return to the workflow?
- Which access, source quality or policy problems slowed the work?
Keep observations separate from interpretation. Two faster drafts show what happened in those two instances. They do not establish an annual saving or prove that another team will get the same result.
At the review meeting, leadership should choose one of five actions: expand, improve, narrow, pause or stop. The evidence may support a smaller workflow even when the original version performs poorly.
Make Useful
Make the useful version repeatable
A team can repeat the pilot once the workflow no longer depends on one person remembering how it works. Write a short workflow card that records the owner, intended users and approved tool. Add the required inputs, excluded information and output format.
The card should also contain the review checklist, the three test cases and the route for policy or access questions. Save an approved example beside it so new users can see the expected standard without treating that example as a universal template.
Assign responsibility clearly. The workflow owner maintains the instructions and source requirements. Information owners approve their material. A subject expert reviews the output, and the relevant leader approves any consequential use.
When the team changes the instructions, sources or tool configuration, rerun the missing information and edge case tests. A workflow that worked with one source pack may behave differently after a new data connection or reporting format appears.
Bring Page
Bring a one page proposal to the next leadership meeting
The proposal only needs to settle the first controlled trial:
- Name the recurring work problem and the business priority it supports.
- Select one output, one owner and one reviewer.
- List the approved sources and explicit exclusions.
- Define the baseline and the evidence the team will collect.
- Schedule the normal, missing information and edge case tests.
- Set the date when leaders will expand, improve, narrow, pause or stop the workflow.
A controlled pilot earns wider use when the team can show a useful output, a manageable review burden and clear information boundaries. Leadership remains responsible for the scope, exceptions and final decisions.
Checked sources
