Insights

AI Risk Assessment: A Template That Ends in Decisions

A six-column AI risk assessment for companies that use AI tools rather than build them, why crude scoring beats a likelihood grid, and the column almost every assessment leaves out.

By Alexej Pikovsky  ·  Updated

Most AI risk assessments are inventories wearing a risk costume. They list the tools, apply a likelihood and impact score, produce a colour-coded grid, and stop. Nothing is decided, nobody owns anything, and the document's real function is to exist when someone asks whether you have assessed your AI risk.

This is a template that ends in decisions instead. It is deliberately short, because a long assessment for a company that uses AI tools rather than building them is padding, and padding is how these documents become annual theatre.

Decide the scope before you assess anything

The first question is which of two assessments you are doing, because they share a name and almost nothing else.

If you build or deploy AI systems, meaning you train models, embed them in a product, or make automated decisions about people, the risk assessment is about the system: its data, its failure modes, who it affects, and what happens when it is wrong. That is the assessment ISO 42001 and regulatory regimes are asking for and it is a substantial exercise.

If you use AI tools, meaning your staff use somebody else's model through a browser or an app, the risk assessment is about exposure: what information reaches which third party under what terms, and what happens if it is retained or disclosed. This is the situation most companies are actually in, and it is what the template below covers.

Doing the second one and labelling it the first is a common and expensive mistake, usually discovered in a procurement questionnaire.

Where AI risk assessments go wrong ยท alexejpikovsky.com
lists AI tools in usemost do
scores likelihood and impactmost do
names an owner per risksome do
records accepted risksalmost none
Qualitative, from reading assessments rather than counting them. The marked row is the one that turns an assessment into a decision. Without it, an unrecorded choice to do nothing is indistinguishable from not having noticed.

The assessment

Six columns. One row per AI tool or use case in the business. If you cannot fill the first column, you have a discovery problem before you have a risk problem, and the detection methods are here.

AI risk assessment, one row per tool or use case
ColumnWhat goes in it
1. Tool and useWhat it is and what it is genuinely used for here, not what it could be used for
2. Data reaching itThe most sensitive category that realistically goes in. Judge the actual practice, not the intended policy
3. Account and termsCompany or personal account, and whether that tier's terms permit training on the content. This single column decides more than the scoring does
4. What goes wrongOne concrete sentence. "Client contract text retained by a third party and used for training" beats "data exposure"
5. OwnerA named person. Not a department
6. DecisionAccept, mitigate, or stop. If mitigate, the specific action and a date. If accept, say so explicitly and who accepted it

Column six is the whole point. An assessment that ends without a decision per row has not assessed anything, it has described. And an accepted risk that is written down as accepted, by a named person, is a governance artefact. The same risk accepted silently is an oversight waiting to be discovered.

On scoring, and why to keep it crude

Likelihood-times-impact grids give small companies false precision. You do not have the data to distinguish a 3 from a 4, and the arithmetic obscures the reasoning that actually matters.

Three levels is enough. High means stop or mitigate now. Medium means mitigate on a date. Low means accept and write down that you accepted it. If a row is hard to place, the difficulty is telling you something the number would have hidden.

Where scoring does earn its place is comparison across many systems, which is a large-organisation problem. At fifty people, the ranking that matters is which three rows you are going to do something about this quarter.

How often, and what triggers one

Quarterly for the review, alongside the policy review so it is one meeting rather than two. Plus a trigger-based reassessment on any of: a new tool approved, a vendor switching an AI feature on inside something you already use, an incident, or a customer or insurer asking a question you could not answer.

That last trigger is the most useful one, because it is the only one that comes from outside and it tells you what the market now expects you to have.

Where this sits

The assessment is one of three artefacts, and on its own it is the least useful of them. It tells you where you stand. The policy tells people what to do. The review makes both of them live documents rather than a filing exercise. If you are doing this because a standard asks for it, ISO 42001 covers what certification actually requires and whether you need it.

The governance checklist this sits inside

The assessment is one item on a shorter list. If you want the whole picture on one page, this is it, in the order that works. Each item is done when you can point at the artefact, not when you have discussed it.

AI governance checklist, in the order that works
ItemDone when
1A sanctioned tool exists on a business tierStaff have a compliant option they would actually choose
2A named owner, with contact detailsSomeone answers when a person asks "can I use this"
3A written policy, signedYou hold signatures, not just a circulated document
4Discovery of what is actually in useA list produced from evidence rather than assumption
5This risk assessment, with a decision per rowEvery row says accept, mitigate or stop, with an owner
6A tool request route with a response timeSomeone has used it and got an answer inside the window
7A quarterly review in the calendarThe second one has actually happened

Item seven is the one that separates governance from a project. Everything above it can be completed in a fortnight and be worthless by spring. The review is what keeps the rest alive, and the second occurrence is the honest test, because the first one always happens.

Item one is placed first deliberately. A policy issued before a sanctioned option exists has nothing compliant to point at, which is the sequencing mistake most companies make.

FAQ

What should an AI risk assessment include?

For a company using AI tools rather than building them, six columns per tool or use case: what it is and how it is actually used, the most sensitive data category that realistically reaches it, whether the account is a company or personal tier and what that tier's terms permit, one concrete sentence on what goes wrong, a named owner, and a decision of accept, mitigate or stop. The decision column is what separates an assessment from an inventory.

How do you score AI risk in a small company?

Crudely and deliberately so. Three levels, high meaning stop or mitigate now, medium meaning mitigate by a date, low meaning accept and record the acceptance. Likelihood-times-impact grids create false precision in organisations that lack the data to distinguish adjacent scores, and the arithmetic tends to obscure the reasoning that actually drives the decision.

How often should an AI risk assessment be reviewed?

Quarterly, scheduled alongside the AI policy review so it is one meeting. Additionally on four triggers: a new tool being approved, a vendor enabling an AI feature inside software already in use, an incident, or an external party asking a question that could not be answered. The last is the most informative, because it signals what the market now expects an organisation to have.

Is an AI risk assessment the same as an AI impact assessment?

They overlap but answer different questions and the distinction matters in procurement. An assessment of AI tool usage examines what information reaches which third party under what terms. An impact assessment, which is what standards and regulatory regimes generally mean, examines an AI system you build or deploy: its data, failure modes, who it affects and what happens when it is wrong. Presenting the first as the second is a common and expensive mistake.