Poornima_Gupta_aZ3h
Members
-
Joined
-
Last visited
Solutions
-
Poornima_Gupta_aZ3h's post in Should AI Decide Which Projects Deserve to Survive? was marked as the answerPosition: View B — Continue the project. The AI's prediction is a warning to investigate, not a verdict to execute.
I take View B without qualification. The project should continue. The AI recommendation should be logged, the underlying signals should be examined, and the sponsor should be required to respond to them — but the termination decision itself must stay with humans. To hand it to an AI in the conditions described is to commit a category error about what AI can and cannot know. Below is why, the eight cases that prove it, the strongest objections answered, where View A is genuinely right, and the framework I would put on the table on Monday morning.
The real issue the question is actually asking about
The dilemma is presented as "trust the data" versus "trust the sponsor." That framing flatters the AI. The genuine question is narrower and harder: on what reference class is the AI's failure prediction trained, and is the project in front of it a member of that class?
Every AI prediction model — whether logistic regression, gradient-boosted trees, or a transformer — works by pattern-matching the present situation against historical outcomes. Its accuracy is bounded by what statisticians call the stationarity assumption: that the future will resemble the past. For routine, high-volume, well-understood processes (call-centre staffing, fraud scoring, loan default prediction), that assumption holds and AI prediction is genuinely superior to human judgement. For transformational initiatives — the word is in the question itself — the assumption collapses. The whole point of a transformation is that it is not drawn from the historical reference class.
There is a deeper, less obvious flaw buried in the training data itself: survivorship bias. The AI learns "what failure looks like" from the projects in the organisation's history that ran long enough to produce a recorded outcome. But the boldest transformations are precisely the ones most likely to have been killed early in the past — so they never generated a "success" label for the model to learn from. The model is therefore structurally taught that ambitious, slow-burning, signal-noisy projects fail, because the counter-examples were terminated before they could prove otherwise. The AI is most confident about killing exactly the category of project on which it has the least valid evidence. It is reading the bullet holes on the planes that came back, and concluding the engines are safe.
This is reinforced by Clayton Christensen's argument in The Innovator's Dilemma (HBR, 1995, expanded 1997): the projects most likely to disrupt an organisation are precisely those that look like underperformers by conventional metrics in their early years, because they serve a market the existing measurement system was never designed to see. And Nassim Taleb's distinction in The Black Swan (Random House, 2007) gives it a name: AI prediction lives in Mediocristan (predictable, average-driven worlds where extremes are bounded), while transformational initiatives live in Extremistan (outlier-dominated worlds where a single result dominates everything else). Stopping a project in Extremistan because its early signals look like a typical Mediocristan failure is the textbook error. You are asking a forecaster trained on coin flips to rule on a lottery ticket. An AI trained on the existing measurement system will systematically flag disruptive initiatives for termination. It is not malfunctioning. It is doing exactly what it was built to do — and that is the problem.
Eight real-world cases where killing the project on the early signals would have destroyed the prize
#
Industry
Project
What the early "AI signals" would have shown
What patience actually produced
1
Aerospace
SpaceX Falcon 1, 2006–08
Three consecutive launch failures, $100m of personal capital exhausted, no commercial revenue, milestone slippage in every quarter
Fourth launch reached orbit on the last funded attempt (Sept 2008); NASA's $1.6bn CRS contract followed in Dec 2008; SpaceX now performs more orbital launches annually than any other launch provider on Earth
2
Pharmaceuticals / biotech
Katalin Karikó's mRNA research, 1989–2013
Continuous grant rejections, four formal demotions at the University of Pennsylvania, no commercial output for two decades — every conventional milestone signalled failure
Underpinned the BioNTech-Pfizer and Moderna COVID-19 vaccines; Nobel Prize in Medicine 2023; estimated millions of lives saved
3
Consumer products
Dyson Dual Cyclone vacuum, 1979–93
5,127 failed prototypes over 15 years; wife working as art teacher to fund household; no licensee in the UK industry — Bob Sutton noted this was a "textbook case" of what an AI would call escalation of commitment
Created a category that disrupted the entire global vacuum market; Dyson is now a privately held conglomerate worth over £20bn
4
Streaming / media
Netflix streaming pivot, 2007–11
Cannibalised the profitable DVD-by-mail business; the 2011 Qwikster split lost ~800,000 subscribers in a single quarter and the stock fell ~75%; Hastings publicly apologised
Foundation of the modern subscription economy; Netflix shares rose 6,744% from end-2009 to end-2020 vs. S&P 500's 237% over the same period
5
Banking — UK / global
HSBC Dynamic Risk Assessment (with Google Cloud), c.2018–2021
A long, data-intensive ML build to replace a legacy rules-based AML system; the hardest, slowest part was getting years of fragmented transaction and KYC data fit to train on — milestone slippage, sustained spend, no production output for an extended period
On completion, detected 2–4x more genuinely suspicious activity than the legacy system while cutting false-positive alert volumes by over 60% and compressing analysis from weeks to days; now monitors over 1bn transactions/month and won Celent Model Risk Manager of the Year 2023
6
Banking — UK
First Direct (Midland Bank), launched 1989
First two years showed customer-acquisition costs running far ahead of forecast and contribution margin deeply negative; internal scepticism that a "branchless" bank could work in the UK
First Direct became the highest-rated bank in the UK for customer satisfaction for over two decades; the template for every UK digital bank that followed, including Monzo, Starling and Revolut
7
Industrial / energy
Tesla Model 3 production ramp, 2017–18
Musk publicly called it "production hell"; multiple missed targets; cash burn so severe that analysts including Goldman Sachs predicted insolvency; an AI trained on automotive launches would have triggered termination
Model 3 became the best-selling electric vehicle in the world; Tesla's market capitalisation crossed $1 trillion in 2021
8
Pharma — gene therapy
Novartis CAR-T / Kymriah, 2012–17
Patient enrolment delays, FDA back-and-forth on manufacturing, treatment costs that seemed commercially unviable, multiple stoppages
First FDA-approved gene therapy in the US (2017); foundation of an entire treatment modality for paediatric leukaemia and lymphoma
The pattern is consistent. Every transformation that mattered looked, in its third quarter, exactly like a failing project. The signals the AI in the question is being asked to weigh — milestone delays, budget consumption, decision bottlenecks — are the very signals that a transformation, by definition, generates while it is being built. They are not symptoms. They are the work.
The asymmetric payoff that the AI cannot see
The case for View B is fundamentally a payoff-asymmetry argument, not a probability argument. Even if the AI is technically correct that the project has, say, a 70% probability of failure, the question is not "what is P(failure)?" — it is "what is the expected value, weighted by the asymmetry of outcomes?"
Decision
If project would have succeeded
If project would have failed
Continue (View B)
Captures the full upside — potentially transformative (Karikó, Dyson, SpaceX, HSBC's AML detection)
Loses incremental investment from today onward — finite, bounded, recoverable
Terminate (View A)
Loses the transformative outcome forever; loss is unmeasured because it never appears on any P&L
Saves incremental investment from today onward
In Mediocristan projects, the two losses are symmetric — and View A wins. In Extremistan projects, the upside loss is potentially infinite (a vaccine that saves millions, a launch capability that reshapes an industry, an AML system that catches multiples more financial crime) and the downside loss is finite (a few more quarters of burn). When the payoffs are this asymmetric, expected value mathematics inverts the apparent verdict of the probability model.
This points to the real fix, and it is not "put a human in the loop to overrule the machine" — that is babysitting a system whose objective is wrong. The fix is to redesign what the AI optimises for. A failure-predictor trained to maximise P(success) is answering the wrong question. The decision-relevant quantity is expected value under a convex payoff — the mathematical property that makes the upside disproportionately large relative to the bounded downside. Formally, the routing function should maximise:
E[V] = α · P(success) · V(success) − β · (burn rate × time remaining) + γ · O
where V(success) is the magnitude of the upside (the term that explodes in Extremistan and which a pure P(success) model discards entirely); the middle term is the downside, which is finite, bounded, and recoverable — you only ever lose the forward burn; and O is the option value of keeping the bet alive to learn more before committing further (Dixit & Pindyck's real-options logic, Investment Under Uncertainty, 1994). The coefficients α, β, γ are set by the board's risk appetite, not by the model.
A multi-objective routing function is the right instinct. But it has to be anchored in this — expected value under convexity — and not in proxy objectives like "capability gain" or "resource-utilisation depth." Those proxies are themselves Mediocristan metrics: easy to count, and therefore exactly the kind of measurable-but-incomplete variable the McNamara Fallacy warns against. Optimising a transformation decision on proxy metrics is a more sophisticated way of making the same mistake. The only objective that survives the Extremistan critique is one in which V(success) — the size of the prize — is a first-class term. A model that cannot represent the magnitude of what it might be killing has no business recommending the kill.
There is a name for the underlying error, and it is worth stating because it is exactly what the AI is doing. The McNamara Fallacy — named after Robert McNamara, the US Defense Secretary who measured success in the Vietnam War by enemy body count because it was the variable he could most easily quantify — describes the trap of treating what is measurable as the whole of what matters. The fallacy runs in four steps: measure what is easy to measure; disregard what cannot be measured; then assume what cannot be measured is unimportant; and finally conclude that what cannot be measured does not exist. The AI in the question measures milestone delays, burn, and bottlenecks because those are countable — and is structurally blind to strategic optionality, organisational learning, and the sheer size of the eventual prize, because those are not. It will report that the war is being won on the numbers, right up to the point the organisation loses it.
This is the same logic that underpins venture capital portfolio construction (Sequoia's published doctrine that a single 100x return justifies a fund full of zeros), and it is why the question framing — "high probability of failure" — is a red herring. Probability is not the deciding variable.
The four strongest objections to my position — and why none survives contact
Intellectual honesty requires meeting View A at its strongest, not its weakest. Here are the four most serious objections, each conceded on its own terms, then answered.
Objection 1: "You are rationalising the sunk-cost fallacy. Staw's escalation-of-commitment research is real, and View B is exactly the cover a biased sponsor uses to keep digging." Conceded — fully. Escalation of commitment is real, well-evidenced (Staw, 1976), and is precisely what kills organisations that confuse persistence with progress. But this objection defeats passive continuation, not my position. My Step 3 requires the sponsor to commit, in writing, to forward-looking termination triggers — specific, measurable, time-bound. That is the documented antidote to escalation: it removes the sponsor's discretion to move the goalposts. I am not defending the sponsor's right to persist; I am replacing their judgement with pre-committed kill criteria. The objection lands on View B's caricature, not on the protocol.
Objection 2: "Survivorship cuts both ways. For every Dyson and Karikó there is a graveyard of zealots who persisted into bankruptcy. You are showing me the winners and hiding the losers." Conceded — completely, and it is the strongest objection. Yes: most persistent bets fail, and a list of survivors proves nothing on its own. But this is why my argument is built on payoff asymmetry, not on success rates. In a convex-payoff portfolio you do not need most bets to win — you need the rare winner's magnitude to exceed the sum of the bounded losses. Venture capital is a standing proof: most investments return zero, and the model is still rational because one outcome can return the fund many times over. The graveyard is not evidence against the strategy; the graveyard is the strategy's accepted cost. The objection assumes we are counting wins. We are weighing magnitudes.
Objection 3: "Then just retrain the AI on transformation data. The problem is a bad model, not the principle of AI-led termination." Conceded in principle — defeated in practice. If a valid reference class of comparable transformations existed, the AI's outside view would be sound and View A would win. But genuine transformations are, by definition, low-frequency, non-stationary, and heterogeneous — there are too few, too dissimilar, and the world they operated in no longer exists. This is not a data-volume problem that more training fixes; it is a structural property of the phenomenon. No amount of more data or better architecture repairs it, because the failure is in the reference class, not the model. That is what makes this objection's remedy unreachable rather than merely difficult.
Objection 4: "Your protocol hands every sponsor a permanent excuse. 'It's Extremistan, the AI can't judge it' becomes the universal defence, and nothing ever gets killed." Conceded — this is the real danger, and View A is right to fear it. A framework that protects transformations must not become a framework that protects everything. That is exactly why Step 4 makes the AI's signal escalate — louder, more frequent, board-level — rather than disappear, and why "Extremistan" is not a label a sponsor may simply assert. It must be argued in Step 2 against explicit criteria (low-frequency, non-stationary, convex payoff), and the burden is on the sponsor to demonstrate membership, not merely claim it. The routine majority of initiatives are Mediocristan; they fail that test and remain squarely in the AI's domain. The protocol protects the rare transformation precisely by refusing to protect the routine project.
The conclusion is therefore earned, not assumed: View B survives its four strongest objections, and each objection, properly answered, turns into a feature of the framework rather than a hole in it.
Where View A is genuinely right — and why this case is not one of them
I will not pretend View A has no domain. It does, in a precise zone:
Routine IT migrations with rich historical reference classes — RBS 2012, where a corrupted update to the bank's overnight batch-processing system locked 6.5m customers out of their accounts and left 100m payments unprocessed (£125m in remediation, £56m in fines), and TSB 2018, where a "big bang" cutover to a new core banking platform went live with 2,000 known defects, locking out millions and even exposing some customers' accounts to strangers (£330m loss, 80,000 customers gone, CEO resigned). Both were flagged internally by engineers before go-live, and both fit AI's Mediocristan zone of competence. Stopping these — or at least delaying go-live — would have been the right call.
Compliance and operational-resilience projects where the universe of possible outcomes is bounded and well-characterised.
Cost-reduction programmes with linear, additive payoffs.
The distinguishing feature of the View A zone is that the project's success criteria are well-defined upfront, the reference class is rich, and the payoff distribution is roughly symmetric. In those conditions, AI prediction outperforms human judgement — particularly biased human judgement (Staw, "Knee-Deep in the Big Muddy," 1976) — and View A is correct. In practice this means an explicit allocation: AI-led termination for the routine majority of initiatives that are Mediocristan-class, human-led judgement for the minority that are genuine transformations.
The case in the question, however, is described as a transformation initiative with strong executive sponsorship (which signals strategic significance, not just political protection) and political importance (which signals that the organisation has staked its forward narrative on it). These are the markers of Extremistan, not Mediocristan. View A does not apply here.
The reframing: the AI does not decide. It triggers an investigation.
The question implies a binary — kill or continue. This is a false choice. The correct response to an AI failure prediction on a transformation initiative is a Mandatory Investigation Protocol — neither passive continuation nor automated termination. The AI leads on detection and on escalation cadence; the human leads on the termination decision itself.
Step 1 — Signal logged, sponsor informed within 48 hours. The AI's prediction and the underlying feature attribution (which signals drove the score) are sent to the sponsor and to an independent reviewer. No automatic action is triggered.
Step 2 — Diagnostic, not verdictive. The sponsor and a small independent panel ask three questions: (a) Are the AI's signals symptoms of real failure (e.g., team disengagement, vendor instability) or symptoms of normal transformation friction (milestone slippage during architectural change)? (b) Has the project's reference class been correctly identified, or is this an Extremistan initiative being judged on Mediocristan benchmarks? (c) What new information would update us in either direction, and when can we get it?
Step 3 — Pre-committed kill criteria, not predictive ones. If the project continues, the sponsor must commit in writing to forward-looking termination triggers — specific, measurable, time-bound events whose occurrence would close the project. Eric Ries calls these "innovation accounting" milestones (The Lean Startup, 2011). Andy Grove called them "strategic inflection points" (Only the Paranoid Survive, 1996). They turn an open-ended commitment into a series of bounded bets.
Step 4 — The AI's role is escalated, not authoritative. Each subsequent AI re-prediction is logged and forces a board-level review at fixed cadence. The AI gets louder over time. It never gets the final word.
This protocol gives you the best of both worlds: the AI's signal cannot be politically suppressed (Step 1 makes it visible to independent reviewers), and the AI's signal cannot prematurely kill a transformation (Step 4 keeps the decision human). It directly addresses the failure mode the question is worried about — sponsor capture — without falling into the opposite failure mode of algorithmic over-reach.
Why this answer matters specifically for banking
In banking, this argument has unusual weight — and the sharpest illustration comes from inside the AML function itself.
HSBC's Dynamic Risk Assessment is the case that should give every View A advocate pause. HSBC set out, with Google Cloud, to replace its legacy rules-based AML transaction-monitoring system — the kind of system across the industry that closes more than 95% of its alerts as false positives — with a machine-learning system. The build was long and data-intensive; the hardest and slowest part was not the model but getting years of fragmented transaction and KYC data into a state fit to train on. Through that period the project displayed exactly the signals an AI failure-predictor weighs most heavily: milestone slippage, sustained spend, and no production output. An AI trained on historical IT-migration reference classes — on RBS 2012 and TSB 2018 — would have recommended abandonment with high confidence.
It would have been catastrophically wrong. HSBC piloted the system in 2021 and is now finding two to four times more financial crime than it did previously, with much greater accuracy. It was first implemented in the UK in 2021 and has since been deployed across six markets, covering 80% of the bank's customers. The system reduced alert volumes by more than 60% while detecting 2-4x more suspicious activity, and cut the time needed to analyse billions of transactions across millions of accounts from several weeks to a few days.
Note the reflexive twist that makes this case unique: the transformation being built was itself an AI — and an AI failure-predictor, judging that build against legacy reference classes, would have killed it. The cost of that termination would not have been a write-off on a P&L. It would have been the choice to keep detecting a fraction of the money laundering the bank can now see — a Type I error (a false-positive "this will fail" verdict) whose consequence is societal, not merely financial.
The wider banking record reinforces the point. First Direct (Midland, 1989) survived two years of negative contribution to become the highest-rated UK bank for customer satisfaction for over twenty years. By contrast, HSBC's own Connected Money app, JPMorgan's Finn and NatWest's Bó were all killed early on exactly the kind of signals an AI would flag — and ceded digital-deposit territory to neobanks and to rivals who persisted.
For a UK bank operating under SS1/21 (PRA operational resilience) and SMCR, the right framing is therefore not "should the AI be allowed to stop a project" but "should the accountable senior manager be required to engage with AI signals on the record before approving continuation?" The answer to that is unambiguously yes — and is materially different from the question being asked. The first preserves accountability with the human. The second outsources it to the model.
Crucially, the banking failures most often cited in support of View A — RBS 2012 and TSB 2018 — were not killed by AI; they failed because the humans ignored signals their own engineers had raised. The lesson there is not "let the AI decide." It is "force the humans to act on the evidence." That is what View B's protocol does. View A solves the wrong problem.
Why normalising AI-led termination is the deeper institutional danger
There is a cost that sits above any single project. If an organisation normalises AI-led kill decisions, it teaches its most capable people that ambition is futile — that any initiative bold enough to matter will be flagged and stopped before it can prove itself. Over time the best sponsors stop proposing transformations at all, because they learn the model will end them at the first noisy quarter. That is an irreversible ratchet: the organisation quietly loses the muscle to attempt hard things, and capability of that kind cannot be switched back on when it is finally needed — it has to be rebuilt over years, by which point the disruptor has already arrived. The danger of View A is not that it kills one good project. It is that, normalised, it trains an entire institution out of the capacity for transformation while every dashboard still shows green.
Conclusion
Continue the project. Not because executive sponsors are infallible — they are not, and Staw's escalation-of-commitment literature is real. Continue it because the AI is operating outside its zone of competence, its training data is survivorship-biased against exactly this kind of project, the payoff distribution makes probability the wrong variable, and the historical pattern is clear: every transformation that mattered looked, at this stage, exactly like this one. The right response to the AI's signal is a structured investigation that forces the sponsor to defend the project on forward-looking criteria — not an automated termination that confuses a forecaster trained on the past with an oracle of the future.
The AI is telling you something is unusual. That is useful information. It is not telling you the project will fail. It cannot.
-
Poornima_Gupta_aZ3h's post in Rare but Critical — Should AI Remove the Safeguard? was marked as the answerPosition: Retain the Approval Step — View B
The AI in this scenario has done its job perfectly. It found the delay. It found the 1% intervention rate. It found the catastrophic consequence in those rare cases. It reported everything.
The dilemma is not an AI failure. It is a human decision-making failure waiting to happen.
The organisation is now looking at accurate data and considering the wrong conclusion. They are reading a 99% confirmation rate and seeing an unnecessary step. They should be reading a less than 1% catastrophic prevention rate and seeing an irreplaceable safeguard. Same data. Completely different categorisation. Completely different outcome.
This is the real root cause. Not a measurement error. Not a design failure. A risk categorisation failure — made by humans, not the AI.
To understand why that categorisation failure matters — and why it changes everything — you need to understand the difference between two fundamentally different types of risk.
The Risk Categorisation Framework
High-frequency low-consequence risks — credit card fraud, customer service errors, data entry mistakes — should be managed for speed and volume. Getting it wrong occasionally is acceptable and recoverable. View A works here.
Low-frequency high-consequence risks — severe misdiagnosis, drug approval failures, nuclear safety, aircraft structural integrity — must never be managed for frequency. Getting it wrong once can be catastrophic and irreversible. View B is non-negotiable here.
The approval step in this scenario exists entirely to manage a low-frequency high-consequence risk. That categorisation should have been defined by humans before any conclusion was drawn from the AI findings. It was not. The AI therefore presented accurate data that was misread through entirely the wrong lens.
"The AI reported everything correctly. The organisation is about to conclude everything wrongly. That gap — between accurate data and sound judgement — is where the risk categorisation failure lives. And it is entirely a human problem."
Banking Learned This the Hard Way — Barings Bank
In February 1995, Barings Bank — Britain's oldest merchant bank, founded in 1762 — was sold for £1 and ceased to exist overnight. Nick Leeson had been given the dual role of managing both the trading floor and the settlements division — a clear violation of standard banking procedure. This concentration of power allowed him to bypass checks and balances entirely, creating fictitious trades and hiding losses from management.
The approval step — segregation of duties — was effectively removed for a star performer. The step would have changed nothing in the vast majority of trades. No one in management accepted responsibility for Leeson's activities between October 1993 and January 1995. Then the 1% arrived. 233 years of history gone in weeks.
Nobody categorised Leeson's trading oversight as a low-frequency high-consequence safeguard before removing it. They read the data — a step that rarely changed outcomes for a consistently profitable trader — and drew the wrong conclusion from accurate information. The categorisation failure cost a 233-year-old institution its existence.
Every High-Consequence Industry has This Story — NASA Challenger
Barings is not an isolated case. Every high-consequence industry has its version — the moment a rarely-triggered safeguard was bypassed in the name of speed and the rare event arrived.
On 28 January 1986, Challenger broke apart 73 seconds after launch. Seven crew members were killed. Engineers at Morton Thiokol had formally flagged the O-ring risk the night before and recommended delay. The risk had never caused a catastrophic failure before. The data was accurate — the O-ring had performed without incident in the vast majority of launches. NASA managers read that data and drew the wrong conclusion.
The Rogers Commission identified the failure as normalisation of deviance — the gradual acceptance that because the rare catastrophic event has not happened yet, it probably will not. The O-ring risk had never been formally categorised as low-frequency high-consequence before the launch decision was made. It was treated as a manageable operational concern by people who had accurate data and reached a catastrophically wrong conclusion.
That single categorisation failure cost seven lives.
Three industries. Three warnings. Three times humans read accurate data and drew the wrong conclusion. Three times the rare event arrived. Zero times the damage could be undone.
The Healthcare Warning Is Already Playing Out — UnitedHealth nH Predict
This is not hypothetical. It is in federal court right now. And it is the closest direct parallel to the scenario in this question.
UnitedHealth deployed an AI model called nH Predict to evaluate patient care claims. A 2023 lawsuit alleged the company knowingly used this model to deny elderly Medicare Advantage patients care that their own physicians had determined was medically necessary — and that the AI model had a 90% error rate. Nine out of ten denials that were challenged were ultimately reversed. Yet the system continued to override physician judgement at scale. UnitedHealthcare's post-acute care denial rate more than doubled — from 8.7% to 22.7% — between 2019 and 2022, coinciding directly with the rollout of their algorithmic tool.
Elderly patients discharged prematurely. Families depleted savings. Patients worsened and died. UnitedHealth gave an AI system authority over clinical decisions without categorising those decisions as low-frequency high-consequence risks requiring human expert oversight. The outcome is a Senate investigation, a federal lawsuit, and irreversible patient harm.
This is what happens when accurate data meets uncategorised risk. The AI reported what it found. The humans drew the wrong conclusion. The patients paid the price.
When One Specialist Got the Risk Categorisation Right — Thalidomide and Frances Kelseys
History also shows what happen when a single specialist holds the line. This is the most powerful healthcare example available.
In the 1950s and 60s, Thalidomide was approved across Europe and prescribed to pregnant women without adequate specialist review of rare but catastrophic side effects. Over 10,000 children were born with severe birth defects across 46 countries. In the United States, a single FDA specialist reviewer named Frances Kelsey refused to approve it. She was seen as causing unnecessary delay for a drug that appeared safe in the vast majority of cases. She was pressured repeatedly to remove her objection and speed up the process. She refused.
The United States was largely spared.
Frances Kelsey did not have better data than the European regulators. She had better categorisation. She recognised that drug approval for pregnant women was not a high-frequency low-consequence process. It was a low-frequency high-consequence decision where the rare catastrophic outcome was irreversible. Same data available to everyone. One person categorised the risk correctly. An entire country protected from an irreversible catastrophe.
This is not an argument about bureaucracy slowing progress. This is an argument about one person with specialist expertise standing between a population and an irreversible outcome. That is exactly what the approval step in this scenario represents.
When Risk Is Categorised Correctly — Design Follows Automatically
The Four Eyes Principle in banking is the most powerful proof that correct risk categorisation leads directly to correct design. Every major bank — NatWest, HSBC, Barclays, Deutsche Bank — correctly categorised large transactions and critical approvals as low-frequency high-consequence risks decades ago. The design response to that categorisation was immediate and permanent — no single person can initiate and approve a critical transaction. Two independent pairs of eyes on every decision that carries catastrophic potential.
The true value lies not just in catching errors, but in creating an environment where accuracy becomes embedded in organisational culture.
Nobody has questioned this design since. Not because it catches problems frequently. But because the categorisation that created it has never changed. Large financial transactions remain low-frequency high-consequence risks. The design therefore remains permanently in place.
This is the sequence the healthcare organisation in this scenario has reversed. They looked at the design — the approval step — and questioned whether it was necessary. They should have looked at the risk category first. Had they correctly categorised the specialist approval step as a low-frequency high-consequence control — as every bank does with the Four Eyes Principle — the design conclusion would have been automatic. You do not remove low-frequency high-consequence controls. You protect them. And when they are slow you redesign them to be faster. You never remove them.
The specialist approval step in this scenario is the medical equivalent of the Four Eyes Principle. Not bureaucracy. A structural design response to a correctly categorised risk — and the last line of defence for a category that demands it.
To Be Fair — When Does View A Actually Work?
View A is not always wrong. Banking proves it on both sides — and the distinction is exactly what makes the healthcare case so clear.
When you swipe your card at a grocery store, approval takes 200 milliseconds. The manual referral step was removed entirely. View A is correct there — because the risk has been correctly categorised as high-frequency low-consequence. The delay causes measurable harm to commerce. The consequences are fully recoverable — a fraudulent charge reversed with one click under Zero Liability policies. And alternative post-transaction safeguards catch catastrophic fraud after the fact.
View A's one valid point in this scenario is the 8 to 10 hour delay. That is genuinely harmful to patients. It deserves a direct response.
The answer is not removal. The answer is redesign.
Go back to the scenario for a moment. A senior specialist approval step adds 8 to 10 hours to a treatment decision. The AI has correctly identified that delay as harmful. But look at what is actually causing those 8 to 10 hours. It is not the specialist. It is everything that happens before the specialist sees the case. The case notes gathered manually. The patient history retrieved separately. The frontline doctor's findings written up and passed across. The specialist starting from scratch on context that AI could have assembled in seconds.
The specialist is not the problem. The information gap before the specialist sees the case is the problem.
Use AI to triage which cases genuinely need specialist review based on complexity and risk markers. Use AI to pre-summarise the patient case, surface relevant history, and flag historical misdiagnosis patterns before the specialist opens the file. The 8 to 10 hour delay becomes a targeted 90-minute review for the cases that warrant it. The safeguard is retained. The speed problem is solved. Both at the same time.
DBS Bank validated this principle in a different context. Rather than removing human oversight from 250,000 monthly customer interactions they built AI to make the human faster and better informed. The human stayed in control. Speed improved dramatically. The safeguard was not removed. It was redesigned. Healthcare can and should apply exactly the same logic.
This is why banking uses View A for coffee and groceries but View B for global wire transfers, corporate lending, and Mergers and Acquisitions. The categorisation determines the approach. Always.
In healthcare there is no Zero Liability policy. There is no reverse button for a severe misdiagnosis. We are dealing with biological systems, not digital ledgers. You cannot call the patient the next day and tell them the error has been credited back to their account.
View A works for the £50 transaction because you can fix it later. View B is for the 1% event where later is too late.
Final Verdict
The approval step must be retained. Not streamlined. Not reviewed. Not reduced. Retained — because it exists precisely for the moment when everything else has already passed the case and got it wrong.
Barings Bank. Accurate data on a star performer's trading. Wrong conclusion drawn. A 233-year-old bank destroyed overnight.
NASA. Accurate data on O-ring performance history. Wrong conclusion drawn. Seven lives lost.
UnitedHealth. Accurate AI analysis of claims data. Wrong conclusion drawn. Senate investigation. Federal lawsuit. Irreversible patient harm.
Frances Kelsey. The same data as every European regulator. Correct conclusion drawn. An entire country spared.
Four cases. Same quality of data. One variable. Whether the humans reading it correctly categorised the risk.
The approval step in this scenario is not a bottleneck. It is the O-ring. And we already know what happens when you decide the O-ring is not worth the delay.
One question — and only one — before removing any critical control:
"What category of risk does this safeguard exist to manage — and what is the consequence if it fires and nothing is there?"
In banking — a 233-year-old institution destroyed overnight.
In space — seven lives lost because a cold morning felt manageable.
In healthcare — a patient receives the wrong treatment and cannot be made whole again.
In drug approval — 10,000 children harmed across 46 countries because one country categorised the risk correctly and 45 others did not.
The AI gave you the data. The categorisation is yours to make.
Categorise it wrong and the rare event will arrive. It always does.
And in healthcare — unlike banking, unlike digital ledgers, unlike a fraudulent charge reversed with one click — there is no undo.
-
Poornima_Gupta_aZ3h's post in Efficiency Up, Experience Down — Should AI Win? was marked as the answerPosition: Reject or Rethink the Change — View B
This is a classic case of the AI efficiency versus customer experience trade-off — and it is one of the most common mistakes organisations make when deploying AI for the first time. This organisation measured three things — handling time, cost per interaction, and cases per day. None of those three metrics measure whether the customer actually got what they came for. That is the entire problem.
Why the Efficiency Gains Here Are an Illusion, Not a Win
Imagine you run a bakery. You find a way to serve customers 30% faster. But the bread tastes worse, customers feel rushed, and half of them stop coming back. Did you win? No. You just moved the cost from the counter to the empty shop.
That is exactly what is happening here.
The core error in View A is treating cost-per-interaction as a profit driver when it is actually a cost-deferral mechanism. When first-contact resolution drops and satisfaction falls 8–10%, the organisation is not saving money. It is moving costs downstream. Customers who feel rushed and misunderstood do not disappear. They call back, escalate to supervisors, churn, and tell others. Every single one of those outcomes costs more than the handling time saved.
This is called the efficiency-satisfaction trap — internal numbers look better while the outcomes customers actually care about quietly get worse, until the real cost shows up in churn, complaints, and lost revenue.
The most damning number in this scenario is not the 8–10% satisfaction drop. It is the first-contact resolution decline. That one number tells you everything. Customers are not getting their problems solved. They are calling back. Every repeat call costs more than the time saved on the first one. The efficiency gain is already negative — the organisation just has not counted it yet.
Real World Evidence — From Banking, Where I Work
1. NatWest Cora+ — The Closest Mirror to This Scenario
NatWest launched a virtual assistant called Cora back in 2017. Sound familiar? It handled routine queries quickly but customers felt like they were talking to a wall. Interactions felt cold, transactional, and unhelpful. The system pointed people to existing content instead of actually solving their problem. Efficiency was there. The customer experience was not.
Rather than shrugging and moving on, NatWest stopped and rethought the entire design. In 2024 they rebuilt Cora into something called Cora+ — powered by generative AI and IBM watsonx technology. The new version actually understood what customers were asking and answered them directly in plain language. The guiding principle was beautifully simple — when someone is worried about their money, they need to feel understood, not processed.
The results spoke for themselves. Customer satisfaction improved by 150%. Human intervention dropped. Efficiency went up. Experience went up. Both at the same time. NatWest also saw Customer Lifetime Value double and Net Promoter Score triple — not soft feel-good numbers, but hard revenue outcomes. By H1 2025 they had deployed 24 more AI models, all built around the same idea — better experience first, efficiency as the reward.
The lesson: You do not have to choose between efficiency and experience. Fix the experience and efficiency follows.
2. DBS Bank — The Most Direct Comparison to This Scenario
DBS Bank in Singapore faced the exact same challenge described in this question — how to handle over 250,000 customer service calls every single month without losing quality. Their answer was not to make conversations shorter. It was to make agents smarter.
They built a Gen AI tool called CSO Assistant — a live co-pilot that sits alongside the agent during every call. It listens, transcribes in real time, searches the knowledge base instantly, and surfaces the right answer while the conversation is still happening. The agent stays in control. The customer still feels heard. And the problem gets solved faster because the agent is not scrambling to find information.
Pilots showed transcription and solutioning accuracy of nearly 100%, call handling time reduced by up to 20%, and close to 90% of customer service officers said it had a positive impact on their work. Across all its AI initiatives in 2024, DBS delivered SGD 750 million in economic value — more than double the previous year.
The design difference is everything. DBS reduced time without removing the human. The scenario we are discussing removed the human without solving the problem. That single decision is what separates success from failure.
3. Bank of America Erica — The Gold Standard
Bank of America had a simple rule when building their AI assistant Erica — start with what the customer needs, not what is easiest to automate. Erica has now handled over 2.5 billion customer interactions with a 98% success rate. Customers either get their answer from Erica or are passed smoothly to a human. No dead ends. No frustration. Bank of America was ranked the most satisfying mobile banking app of any national bank.
The sequencing lesson here is critical. Experience and revenue came first. Efficiency came second. That is the right order. That is why it worked.
The Real Issue — The Scorecard Was Wrong From the Start
Here is a simple truth. The metric you choose to measure determines the result you will get. If you measure how fast you close a call, you will get fast call closures. If you measure whether the customer actually got what they needed, you will get satisfied customers. The organisation in this scenario chose speed. They got speed. And they lost trust.
HSBC proves what good metric design looks like — twice.
1.When HSBC used AI to transform their KYC onboarding process, they did not just measure how fast documents were processed. They measured accuracy alongside speed. The result — processing time dropped from 12 days to under 24 hours while accuracy jumped from 87% to 99%. Both improved because both were measured from day one.
2.When HSBC partnered with Google Cloud on their AML anti-money laundering system, most banks are still drowning in false positive alerts — flagging innocent customers and wasting thousands of investigator hours chasing nothing. HSBC measured what actually mattered — how accurately real criminals were being caught, how much time investigators spent on genuine cases, and how many innocent customers were being unnecessarily disrupted. The system identified two to four times more real suspicious activity while cutting false alerts by 60%. Investigators focused on actual crime. Innocent customers faced fewer unnecessary checks. Everyone won.
HSBC built the scorecard before they built the system. The organisation in this scenario did it the other way around. They measured what was easy to count — speed and cost — and got exactly those things. The problem was never the AI. The problem was the scorecard.
To Be Fair — When View A Can Actually Work
I do not broadly support View A. But a strong argument acknowledges the other side — because showing when something works makes it clearer why it fails here.
Amazon is the textbook example. When Kiva robots rolled into fulfillment centers in 2012, things got worse before they got better. Delivery updates became impersonal. Handling exceptions became harder. But picking efficiency jumped by over 50% and costs per order fell sharply. Here is the crucial part — Amazon did not keep those savings. They reinvested every penny into building next-day and same-day delivery. Today, fast reliable delivery is the number one reason customers love Amazon. The short-term experience dip funded a permanent experience upgrade.
Lloyds Banking Group shows the same thinking applied in banking. Lloyds automated data entry, transaction processing, and basic back-office inquiries. Tasks customers never see or feel. Nobody notices whether a human or a machine processed their data in the background. What customers noticed was faster, more accurate service. The efficiency metric and the satisfaction metric pointed in exactly the same direction because Lloyds chose the right processes to automate.
But this only works when all four conditions are true:
Efficiency gains are reinvested into customer experience — not kept as profit
The experience decline is temporary and fixable, not permanent
Customers have little reason to switch during the difficult period
AI is applied to routine back-office tasks — not emotionally sensitive conversations where trust matters
In the scenario presented, every single one of these four conditions fails.
The savings were not reinvested. The satisfaction decline is not a temporary blip — it is a structural signal that customers consistently feel unheard. Banking customers who lose trust do switch, and winning them back costs far more than any handling time saving ever delivered. And most critically, the AI was placed in exactly the wrong conversations — the moments when someone is worried about their money and needs a human being who actually listens.
That is not one mistake. That is four.
Final Verdict
The change should not be accepted. Not because efficiency does not matter — it absolutely does. But because this organisation has not yet earned the right to claim it.
The banks that get AI right — NatWest, DBS, Bank of America, HSBC — all made the same decision before they wrote a single line of code. They decided what success actually looked like from the customer's point of view. They built the scorecard first. Then they built the system.
NatWest asked — does the customer feel understood?
DBS asked — does the agent have everything they need to help?
Bank of America asked — does the customer get what they came for?
HSBC asked — are we catching criminals or just closing alerts?
The organisation in this scenario asked — how fast can we close the call?
That one question, measured alone, produced exactly the outcome described. Faster calls. Lower costs. Unhappy customers. Declining trust.
Fix the question you are measuring and you fix the outcome. That is the lesson. That is the only lesson.