Back to blog

What a human review actually checks

Human reviewed is the claim every AI disclosure leans on, and most teams cannot say what the reviewer looked at. A checklist turns a signature into evidence.

WhoWorked team6 min read
Modular terracotta blocks stacked to different heights on a workbench.

A client read a report we had sent and asked a reasonable question: who checked this? A senior consultant had, and saying so felt like a complete answer for about four seconds. Then came the follow-up, which was the real question. Checked for what?

Nobody could say. The consultant had read it, thought it was good, and sent it. That is a genuine review by any ordinary meaning of the word, and it is also unrepeatable, unauditable, and impossible to describe to the person paying for the work.

Every AI disclosure ends up resting on this. The policy says material AI use is reviewed by a human. The attribution statement says human reviewed. The invoice note says reviewed before delivery. All three are load-bearing claims, and all three are hollow until someone can name what the review examined.

A work record listing contributors, workstreams, evidence and human time, with a verifiable human total.
The review is only worth what the record of it can show. A named reviewer against a workstream, with the evidence attached.

Example record

One delivery, three distinct facts

Owner

Maya Chen

Responsible for the final delivery

AI contribution

Generated

Prepared the first test-suite draft

Review

Verified

Assumptions checked and brittle tests corrected

Outcome: test suite approved and mergedHuman accountable

Why a signature is not a review

Reviewing a colleague's draft and reviewing a model's output are different activities, and treating them the same is where most teams go wrong. With a colleague you are checking judgement, because you can assume basic competence and honest sourcing. The draft may be wrong, but it is wrong in ways a person would be wrong.

Model output fails differently. It is fluent where it is least reliable, confident about citations that do not exist, and plausible in exactly the register that makes a busy reader stop checking. The failure modes are not distributed like human failure modes, so a reading habit tuned to human drafts will pass over them.

This is why review has to become explicit at the moment AI enters the workflow. Not because people got worse at their jobs, but because the thing they are checking changed shape and their instincts have not caught up.

The six things a review checks

Six categories cover almost every deliverable an agency or consultancy ships. Not all six apply every time, and deciding which ones do not apply is itself part of the review.

CategoryThe questionWhat a failure looks like
FactsIs every specific claim true?A number, date, or name that nobody traced to a source
SourcesDoes every citation exist and say this?A real-looking reference to a paper that does not exist
PrivacyDid client or personal data leave where it should stay?Confidential material pasted into an unapproved tool
RightsCan we use and license what we are delivering?Generated material too close to an identifiable original
ConstraintsDoes this obey what the client actually asked for?Correct work that ignores a brief nobody re-read
AccountabilityIs a named person willing to own this?A deliverable everyone touched and nobody signed

Sources is the one that catches people. A citation check is slow, boring, and the single highest-yield thing on the list, because a fabricated reference in a client deliverable is the failure that turns into a phone call from their legal team rather than a correction in the next draft.

Accountability is last on purpose. If the first five pass and no one is willing to put their name on the result, the review has found something real, and the answer is not to add a name.

Risk sets the depth

A checklist that demands the same rigour from an internal meeting summary and a regulatory filing gets abandoned within a fortnight, and rightly so. The depth has to scale with what happens if the work is wrong. If you want a read on where your own record stands, the attribution readiness assessment scores review alongside the five other things it depends on.

Three tiers are enough. Low risk means internal or easily corrected, and a single competent read is proportionate. Medium risk means client facing and reputationally costly, so facts and sources get checked line by line and a second person sees it. High risk means legal, financial, medical, or safety consequences, and there the reviewer needs to be qualified in the domain, not merely senior in the firm.

The tiering is also what makes the checklist survive contact with a deadline. A team that can say this one is low risk has a defensible reason to move fast, which is much better than a team that quietly skips the process and hopes.

Where the checklist lives

A checklist in a wiki is a checklist nobody opens. It has to sit where the work is accepted, which usually means the review step in whatever system already holds the deliverable, with the outcome recorded against the workstream rather than in a separate compliance folder.

What gets recorded is small: which tiers applied, what was checked, who checked it, and what happened as a result. Four fields, filled in at the moment of review, which is the only moment anyone actually knows the answers.

The result field is the one teams leave off and the one that pays. Accepted, corrected, or rejected, recorded consistently over a quarter, is the difference between believing your review process works and being able to show it.

What a review is not

It is not a second opinion on whether the work is any good. Quality review already exists in most firms and does not need renaming. The AI review answers a narrower question about whether the specific failure modes of generated content made it through.

It is not a check on the person who used the tool. A review that reads as an accusation produces one reliable outcome, which is people recording less AI use, and the record is the only thing making any of this real.

  • Not a productivity measure: how long a review took says nothing useful about the reviewer.
  • Not a prompt audit: what someone typed is rarely relevant and frequently confidential.
  • Not a gate on every piece of work: low risk work that is easy to correct does not need ceremony.

The reviewer has to be able to say no

A review process where rejection is socially expensive is a rubber stamp with extra steps. If the reviewer is junior to the person who produced the work, or the deadline has already been promised to the client, the checklist will be completed honestly right up until the first time it matters.

Two things fix this, and neither is a document. The reviewer needs enough standing to return work, and the schedule needs enough slack that returning it is not a crisis. Firms that get this right tend to have built the review time into the estimate rather than hoping it fits in the gaps.

Start with one work type

Pick the deliverable your firm produces most often, decide its risk tier, and write the checklist for that one thing. A page of specifics about client research reports will be used. A general framework covering everything the firm might ever produce will not.

Run it for a month and look at what the result field says. If nothing was ever corrected, the checklist is too shallow or the reviews are not happening. If everything was corrected, the work is going to review too early. Either answer is worth more than the document you started with.

Related posts

AI did the work. Who gets the credit?

Human timesheets and AI usage logs each record only half the work. A useful attribution model keeps accountability human while making material AI contributions visible.

WhoWorked team8 min read

The AI attribution maturity model

Teams can progress from ad hoc AI disclosure to governed work records without collecting every prompt or deploying every integration at once.

WhoWorked team8 min read

Start counting all the work.

30-day free trial. Bring the time history you already have.