Part of a series
Decisive AI & Governance
“The machines have the answers. You have the call.”
Career pillar
"Human in the Loop" Is Not a Governance Model
"Human in the loop" sounds like governance—but five times the draft volume on zero extra review makes oversight notional.

Ask a program leader how they're governing AI and you'll get four words back: we keep humans in the loop.
It sounds like a safeguard. Operationally it's a workload — and in most programs it's the mechanism by which oversight quietly stops working.
Here's the arithmetic nobody runs before the rollout.
Your review capacity did not scale
Your team adopts AI assistance. Drafted output goes up three to five times. Status narratives, risk entries, meeting summaries, stakeholder updates, RAID additions.
Your review capacity goes up zero times. Same senior people, same calendars, five times the volume.
One of two things happens next, and neither gets announced.
Reviews become skims. And AI-generated work is specifically good at not containing obvious errors — it's fluent, organized, and reads like a competent person wrote it. The failure mode isn't a typo. It's a confident summary of a situation that isn't quite true.
Or reviews become signatures. The artifact gets approved because it came through the agreed process, and approving is faster than interrogating.
Both look identical on a governance chart. Both get labeled human in the loop.
This is judgment getting outsourced before governance catches up, and it happens without a single bad decision anyone can point to. For the wider pattern, see Decision Debt 2.0 and the Decision Debt pillar.
Can and should are different questions
The Judgment Line is the boundary between your judgment and the machine's output. Almost every AI failure in program delivery is a failure to draw it deliberately — the tool can do something, so it does, and nobody asked whether it should.
In program management the line is unusually easy to find, because most of our work sorts cleanly into two piles: work that produces material for a decision, and work that is the decision.
The machine belongs on the first pile. All of it.
Where teams get into trouble is that the two piles produce artifacts that look identical. A meeting summary and an escalation recommendation are both a tidy paragraph. One is retrieval. The other is a call. Fluency hides the difference, which is why the line has to be drawn in advance rather than judged output by output. The can vs. should labeling rule is the meeting-room version of the same discipline.
Start by not automating the status report
The status report is the first thing every program tries to automate, because it's the most hated recurring task. It's the wrong first target.
The document was never the point.
The value of a status report is that producing it forces a human to go reconcile reality — chase the dependency that's been amber for three weeks, ask whether "on track" means what it meant last week, notice that two workstreams are quietly assuming different launch dates.
The report is a byproduct of that reconciliation. Automate the byproduct and you keep the artifact and lose the forcing function. Status looks better. The program gets worse. It takes about two quarters before anyone connects the two.
If reporting is painful because the underlying information is scattered, AI will produce a fluent summary of scattered information. The symptom disappears. The debt compounds.
Architect: design the decision before the machine sees it
The first discipline of the A.R.C. Protocol is the one that pays off fastest in program delivery, and it's the least glamorous.
Architect means framing the question, setting the criteria, and naming the constraints before anything gets delegated. In practice that turns most of a program manager's AI usage into support work rather than decision work — which is exactly where the honest time savings live:
The pre-read. Getting up to speed before a meeting: reading back through threads, reconstructing what was decided six weeks ago, finding who owns the open item. Retrieval and synthesis, almost no judgment content. A five-minute brief before a stakeholder call is worth more than an automated status deck, and no one talks about it because it isn't impressive.
Format translation. The same information rendered for an engineering standup, a steering committee, and a board appendix. Three audiences, one underlying truth. Tedious and genuinely mechanical.
First-pass capture. Transcript to structured notes. Let the machine produce the raw material. Don't let it decide what mattered.
Retrieval against your own history. What did we decide about the vendor dependency in March, and who was in the room? The most underrated use in program management, and the one that most directly reduces rework.
Every item on that list produces material for a decision. None of them produce the decision. That's not a coincidence — it's what architecting first does to the work. The full A.R.C. stack — Architect, Reserve, Calibrate — is in Decisive AI, Vol. 5 of the Decisive Edge series.
Reserve: the escalation problem
Reserve means the tool advises inside the frame you built, and the named owner decides. In program management there's one category where this is worth more than all the others combined, and it's the one most teams hand over first.
Language models are trained toward agreeableness. Ask one to draft a risk escalation and you'll get something measured, diplomatic, and professionally worded — a slightly softer version of what you meant.
Each instance is fine. Nobody objects to a well-worded email.
Run it for a quarter across every risk memo, every amber narrative, every dependency warning, and your organization's risk signal attenuates. The escalations still go out. They land with less force. By the time something is genuinely red, the language has been smoothed so consistently that red doesn't carry the weight it used to.
You can't detect this from any single artifact. You detect it when leadership stops reacting.
Write your own escalations. It's the one category where the blunt human sentence is doing load-bearing work.
The rest of the sort is what the Decision Rights Charter is for. Three tiers, applied to program artifacts rather than enterprise decisions:
Delegate — the machine produces it, humans audit periodically rather than individually. Transcripts, pre-reads, format conversions, retrieval, assembly from approved source material.
Augment — the machine drafts, a named human decides and owns. Status narratives, risk register entries, stakeholder communications, resource forecasts. The name is the entire mechanism. Not "the PMO reviewed it" — a person who would be asked about it by name.
Reserve — human only. Escalation decisions. Resource tradeoffs between teams. Dependency arbitration. Anything that changes a commitment to another group. Anything touching a person's performance.
Reserve items exclude recommendations, not just decisions. Once a room has seen what the model suggested, every subsequent human thought gets measured against it, and disagreement starts to feel like arguing with arithmetic. For a narrow set of calls, the cleanest safeguard is not asking.
An unwritten boundary is a boundary the tool will renegotiate daily. Build your charter in the browser — or start from the three-tier essay if you want the logic before the tool.
Calibrate: the actual answer to review capacity
Here's where most governance advice stops short, and where the review-capacity problem actually gets solved.
You don't fix review overload with more review. You fix it by moving work out of Augment and into Delegate — legitimately, on evidence.
Calibrate means trust is earned per task, not granted per vendor. Track where the tool performs and where it slips. Expand delegation where it's been earned; claw it back where it hasn't. Meeting-note extraction might earn Delegate status in six weeks. Risk register drafting might never leave Augment. Both outcomes are fine. What isn't fine is assuming one and never checking.
The Trust Calibration Scorecard exists for exactly this — keeping the batting average per task so calibration is data rather than vibes.
Programs that skip Calibrate get stuck in permanent Augment, which is where the review-capacity arithmetic eventually crushes them. Everything requires a human, no human has time, and the review becomes a signature.
The debt underneath
There's a second-order effect that surfaces around month six.
The socially acceptable form of not deciding has always been asking for more analysis. Nobody gets criticized for wanting more information.
AI made more analysis nearly free.
You can now generate a comparative options matrix, three scenario models, and a risk-weighted recommendation in less time than it used to take to schedule the meeting where you'd have decided. The analysis is genuinely better. The decision still hasn't been made.
Watch for this in your own program. If your options analyses improved dramatically this year and your decision latency didn't, that isn't a productivity gain. That's Decision Debt accruing faster, in a nicer format.
Monday
Three moves, in order.
Name the owner. For every AI-assisted artifact type your program produces, write down the human who owns the output. If you can't name one, that artifact has no governance regardless of what the policy says.
Count the review load. How many AI-drafted artifacts reach your senior reviewers weekly, and how long would a real review take? If the second number exceeds their available hours, your oversight is already notional.
Move escalations to Reserve. Whatever else you change, write your own risk memos.
Then publish the charter. Not because the document is impressive — because the fight your team is currently having about tooling is usually a fight about ownership wearing a costume, and drawing the map is what ends it.
The tools will keep getting better. The Judgment Line doesn't move because the tool improved. It moves because someone decided to move it, on evidence, and wrote it down.
Leadership first. AI second.
Operational next steps
- Take the Leadership Assessment — seven questions and a read on where your Decision Debt is concentrated.
- Build your Decision Rights Charter
- Trust Calibration Scorecard
Go deeper: Decision Debt: the compounding cost of deferred decisions · Decisive AI — the Four Surrenders, the Judgment Line, and the A.R.C. Protocol in full.
Join The Bridge — 3 failures, 2 protocols, 1 strategy every Friday. Follow Decisive Leader on LinkedIn.
Was this piece useful?