While designing GuardianAI and FleetMind, I kept running into the same uncomfortable question: a model can look confident and still be making a decision I would never want it to make alone. That is why I started separating model confidence from permission to act.
THE SIGNAL
AI teams are quickly moving from “the model can answer” to “the system can act.” Microsoft’s 2025 Work Trend Index reported that 82% of leaders expected digital labor to expand workforce capacity within 12–18 months, while 46% said their organizations were already using agents to fully automate workstreams or business processes.
That makes one product question much more important:
When should an AI system be allowed to act without asking a human?
PwC’s guidance on responsible enterprise agents explicitly recommends considering both an agent’s degree of autonomy and the potential impact of its decisions when assigning risk tiers. NIST’s AI Risk Management Framework similarly treats risk as contextual: the right controls depend on intended use, impact, tolerance and lifecycle stage.
THE QUESTION
A lot of AI product conversations still collapse this into one number:
“If confidence is above 0.9, let the agent act.”
I do not think that is enough.
A model can be highly confident and still be operating in a domain where one wrong action is expensive, irreversible or personally consequential.
The product decision is not:
How confident is the model?
It is:
Given the evidence, consequence, reversibility and authority granted to this system, what should happen next?
THE OBVIOUS ANSWER
The obvious answer is to set a threshold.
Below 0.7 → human.
Above 0.7 → AI.
Thresholds are useful. They are also incomplete.
A 0.95-confidence recommendation to reorder office stationery and a 0.95-confidence recommendation to change a worker’s payout should not receive the same autonomy.
THE TENSION
I think autonomy needs at least four dimensions:
| Dimension | Product question |
|---|---|
| Evidence | Does the system have the facts needed to justify the action? |
| Consequence | What happens if it is wrong? |
| Reversibility | Can the outcome be easily undone? |
| Authority | Has the user or organization explicitly permitted this class of action? |
Confidence belongs inside evidence quality. It does not replace the other dimensions.
This is also why “human in the loop” should not be a generic checkbox. The human needs to appear at the point where authority changes.
A support AI can explain policy autonomously.
A travel agent might re-rank options autonomously.
A financial agent might prepare a recommendation autonomously.
But changing money, cancelling a booking, exposing sensitive information or executing a difficult-to-reverse decision may require a very different product contract.

MY PRODUCT TAKE
I would design autonomy as a ladder, not a switch:
Suggest → Recommend → Ask approval → Act inside a bounded policy → Verify → Escalate when the boundary breaks
The product becomes more autonomous only when evidence shows it deserves that autonomy.
That also means post-action verification matters. “The API returned success” is not always the same as “the user’s problem is solved.”
WHAT I WOULD TEST
Before expanding autonomy, I would build an eval set that deliberately includes:
- high confidence + low consequence;
- high confidence + high consequence;
- low confidence + reversible action;
- conflicting evidence;
- stale tool output;
- permission mismatch;
- adversarial instruction;
- post-action state that fails to match the expected outcome.
The most interesting failures are not always “the model answered incorrectly.” Sometimes the model is correct but the system gave it too much authority.
WHAT WOULD CHANGE MY MIND
If a class of actions proves consistently low-risk, reversible, well-observed and strongly preferred by users without manual approval, I would move that action further toward autonomy.
That is the point: autonomy should be earned by evidence, not granted by enthusiasm.
Secondary research
- NIST AI RMF / GenAI Profile — https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- PwC, Responsible adoption of AI agents — https://www.pwc.com/us/en/tech-effect/ai-analytics/responsible-ai-agents.html
- Microsoft Work Trend Index 2025 — https://blogs.microsoft.com/blog/2025/04/23/the-2025-annual-work-trend-index-the-frontier-firm-is-born/