Inside regulated institutions, the question isn’t how fast frontier labs should move but how much decision-making authority it must hold. It already sits on loan-approval or fraud-detection pipeline. While it is driven by the regulator and internal business priorities,more often, it happens by default rather than on purpose.
The Biggest Releases
In lending, fraud and risk systems, the conversation often begins with – can the model catch this pattern? It ends a few discussions later with a harder question: “who is allowed to override it and under what conditions?” The recent two biggest model releases made opposite bets on it. OpenAI’s GPT-6 Astra is built to take an open-ended goal and run largely unsupervised for days, choosing its own method and recovering when something breaks. On OSWorld 2.0, a benchmark for operating a real computer, it scored 72.6 per cent against 65.7 per cent for its predecessor. Anthropic’s Claude Fable 5.1 makes the opposite calls: queries flagged for cybersecurity or biology get rerouted to a less capable model instead of an answer from Fable itself. Deliberately lower autonomy is where the downside is severe.
Autonomy Needs Accountability
Autonomy should scale with reversibility, not with model confidence. A loan declined on a marginal profile, or a fraud hold that freezes a small business’s working capital, costs a lot. The same underlying model can be trusted to act alone on the first and shouldn’t be trusted to act alone on the second, no matter how good its accuracy numbers look in a validation deck. Explainability matters here for a specific reason, not a compliance-checklist reason. If a model cannot produce a reason, it should not get autonomy in that category. And in high- value financial decisions, this accountability matters.
Need for Precision
Everyone agrees a human should be in the loop. Fewer people ask what that human is actually doing once the review queue is too large. A reviewer skimming approvals at that volume is hardly exercising any judgment but just rubber-stamping with a audit trail that will carry a human sign- off either way.
Institutions that slow themselves down on autonomy, while competitors, or foreign vendors selling into India, will lose ground on cost and speed. But the answer isn’t holding every AI system back. It is being precise about which decisions are cheap to get wrong and which aren’t and automating only the first kind without a human gate. This is roughly the split Fable already ships by default.
Before adding another AI system to a lending, fraud, or customer-operations workflow, ask a narrower question: if this specific decision goes wrong, can it be undone before real harm reaches the customer? If yes, let the system run.


