Method

Deciding what AI is allowed to do

Generic AI training tells you how to write a prompt. It does not tell you what your organization should actually approve. This is the process I use to answer that question, and the one I teach inside every engagement.

The idea, in one paragraph

Before a team uses AI on something, I ask two questions. First, how bad would it be if AI got this wrong? Second, how sensitive is the information involved? The answer to those two questions tells you whether a use case is fine to run today, fine with a person double-checking it, worth a small careful trial, or not worth attempting yet. That is the whole method. Everything below is how I apply it in practice, and how I help your organization’s own decision-makers apply it after I am gone.

Why a method rather than a policy document

Most organizations I meet have either no AI policy or a policy written once, in the abstract, that nobody consults. Both produce the same outcome: staff use these tools anyway, privately, with no guidance about what is safe. A policy tells people what happened to be decided last quarter. A method lets them decide the case in front of them, and lets whoever approves new tools at your organization, legal, privacy, risk, or a dedicated compliance function, apply the same standard consistently.

The steps below come from building and shipping AI in situations where being wrong is expensive, and from the evaluation work that any high-stakes AI deployment actually requires.

The rubric

How bad, versus how sensitive

Every proposed use case lands in one of four situations. Each one carries a standing answer, so the decision does not get relitigated every time someone has an idea.

Low stakes, low sensitivity

Go

Use it. Drafting internal summaries from public or non-sensitive material. This is not where your attention needs to go.

High stakes, low sensitivity

Go with a check

Allowed, with a person reviewing the output before it leaves the building. Most valuable professional work lands here.

Low stakes, high sensitivity

Try it carefully

Only in a tool cleared for that kind of data, and only after the retention and training terms are confirmed in writing.

High stakes, high sensitivity

Not yet

Sensitive data plus a costly mistake if it goes wrong. Revisit once the tool, the contract, and the evidence of accuracy all improve. Not before.

Teams leave a workshop with this rubric, branded and theirs to keep, along with a draft policy written against their own environment.

How I apply it, step by step

  1. 01

    Map what kind of data you are dealing with

    Every organization already has data categories with different rules attached. In a health system that is protected health information, internal-only material, and public material. In a firm it is privileged material, client-confidential material, and public material. I write down which sanctioned tool each category may enter, and which it may not. This becomes a one-page safe-use map specific to your environment, and it is what stops the most common failure: a capable person pasting the wrong thing into the wrong window.

  2. 02

    Score the use case on two questions

    How bad would it be if AI got this wrong, and how sensitive is the data involved? Every proposed use case gets scored on those two questions, which places it in one of four situations below, each with a standing answer. This takes about ninety seconds per use case once a team has the rubric, and it turns an anxious open-ended debate into a quick decision.

  3. 03

    Decide how the output gets checked

    The question is not whether the model is right. It is how a professional confirms it quickly. For each approved use case I name the checkpoint: ask the model to point to the source it relied on, ask it to flag the parts it is least sure about, and have a person spot-check the highest-stakes items by hand. The work still gets reviewed. It just arrives faster, and everyone knows in advance who is checking what.

  4. 04

    Ask vendors the questions that actually matter

    Before a tool is approved I want answers on data retention, whether your data trains their model, whether they will sign a business associate agreement where one is required, audit logging, access control, and what evidence they can show that the tool is accurate. A vendor who cannot answer that last one is asking you to take safety on faith.

  5. 05

    Roll it out with guardrails, then keep checking in

    Approved use cases go out with an owner, a guardrail level from the rubric below, a way to measure success, and a date to check back in. After that, someone has to keep owning the ongoing questions: what happens when the tool changes and starts behaving differently, how new use cases get reviewed, and how staff who are already using AI on their own get brought inside the policy instead of pushed further outside it.

Where this comes from

I spent two decades inside healthcare before AI was the job, at Kaiser Permanente, Huron, and McKinsey’s healthcare practice. I then shipped large language model products in a regulated environment at Verily, with legal, privacy, and security sign-off.

Today, in my work on AI safety, I hold systems to a bar where more than 99 percent of outputs have to pass human review, because when AI is wrong in a clinical or legal context, the cost is real. This method is that same discipline, human review, structured risk scoring, and clear ownership, translated into terms a leadership team can actually run.

Full background

Want this applied to your organization?

The workshop program walks a leadership team through this method on their own use cases, and they leave with the rubric, a safe-use map, and a 90-day plan.