View B. Without qualification.But not for Bex's reason, and not at Bex's price. The through-line of this entire answer, stated once and returned to at every hand-off: routine work is not a cost centre that happens to teach. It is a training system that happens to produce output — and the business case in front of you prices only the output. History shows why that mispricing recurs. The arithmetic shows how large it is. The operations show what it costs to run. The law shows where it is already going, with dates. EXECUTIVE SURFACE Decision before us Transfer 100% of routine foundational work to AI for ~$3M/yr, versus retaining it for apprenticeship Position View B — retain the routine work. Fly it unqualified, with one derived exception stated in Correction 1 The number the proposal omits The Last-Third Problem: the pipeline needs only ~20–40% of the routine work to hit its expert quota. So the genuinely contested money is ~$0.9M a year (range $0.4M–$1.3M), not $3M — and buying it costs $180k–$384k per departing expert to replace. A 2.3x to 4.9x adverse trade Break-even in one line View A wins only if a fully formed expert can be acquired for under ~$260,000 of premium (at 100 experts; $520k at 50, $104k at 250) — in a labour market that every competitor is draining at the same moment Named artifacts The Teaching Share, identified by the Correction Test; audited by the Ground-Truth Reserve; owned by a named Capability Registrar THE THREE SENTENCES THAT DECIDE ITThe scenario compares all-human against all-AI, but the pipeline needs only about a fifth to two-fifths of the routine work to make quota — so the real argument is over roughly $0.9M a year, not $3M, which is less than the cost of replacing three senior people. Every controlled test of View A's substitute mechanism — "learn faster alongside AI" — has now been run, and it failed: −17% on unassisted performance in a ~1,000-student RCT, and impaired conceptual understanding, code reading and debugging in a 2026 replication on professional developers. Radiology and air traffic control already ran this experiment on real people: the decision not to train surfaced 8 to 10 years later as a shortage money could not fix, because the queue is 2–6 years long and everyone joined it simultaneously. PART I — HISTORY: THE PATTERN THIS DECISION SITS INSIDEThe governing analogy: 1563The obligation to train is not a habit. It is the oldest recorded answer to a specific market failure, and it was written down 463 years ago. England's first national training system was the Statute of Artificers 1563 (5 Eliz. 1 c. 4). The Act set what we would now call apprenticeship minimum standards: masters could take no more than three apprentices, and apprenticeships were to last seven years. It controlled entry into the class of skilled workmen by requiring a compulsory seven-year apprenticeship, transferring to the emerging English state the functions previously held by the craft guilds. It was repealed 251 years later, in 1814. After that repeal it was no longer possible to prosecute anyone who practised a trade without having served a seven-year term. Why does a Tudor statute belong in a 2026 board paper? Because it identifies the exact failure the $3M proposal reproduces. No individual firm captures the full return on the expert it trains. The apprentice can leave. The competitor can poach. So every firm, acting rationally, under-invests — and the sector ends up short. 1563 was a coercive fix; the German dual system is a cooperative one; professional residency is a regulated-and-subsidised one. All three exist because the unaided business case for training routine-work apprentices has never once cleared on its own merits. View A is not discovering a new efficiency. It is rediscovering, unaided, the reason the statute was written. This is analogy doing work, not decoration. The structural feature it maps to is precise: the training benefit is a positive externality with a 2–3 year lag and no line on the P&L, while the $3M saving is internal, immediate and line-itemised. That asymmetry — not any judgment about AI capability — is what produces the wrong answer. The 1983 theory that named this exact decisionThe canonical framework here is Lisanne Bainbridge, "Ironies of Automation," Automatica 19(6), 775–779 (1983), presented at the IFAC/IFIP/IFORS/IEA man-machine systems conference in Baden-Baden in 1982. It has accumulated over 4,700 Google Scholar citations as of 2024. Bainbridge's argument is the scenario, forty-three years early. She argues that automating most of the work while leaving the human responsible for the tasks that cannot be automated creates severe new problems: operators no longer practise skills as part of their ongoing work, so rather than needing less training they need more, to be ready for the rare but crucial interventions. Her formulation of the deskilling mechanism is exact — a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one. Bainbridge extracts an independent threshold, and this is the part that matters for the ledger: if you automate the routine and retain human responsibility for the exceptional, your training budget must rise, not fall. Every honest version of View A therefore has to book a training cost increase against its $3M. It books zero. That is Correction 3 below. Convergence of independent routesThree unrelated literatures reach the same threshold from different directions, and they were not built to agree: Human factors (Bainbridge 1983): automation of the routine increases required training investment. Decision theory under irreversibility (Arrow & Fisher 1974): "Environmental Preservation, Uncertainty, and Irreversibility," Quarterly Journal of Economics 88(2), 312–319. Arrow and Fisher named the value of waiting when damage is irreversible "quasi-option value." Dixit and Pindyck (1994) later showed that when an investment is irreversible, returns are uncertain, and waiting resolves uncertainty, the expected present value of the opportunity exceeds that of acting now — the difference being the value of the option to postpone. Applied here: dismantling a pipeline is irreversible on a 2–6 year horizon; capturing the saving later is fully reversible. Standard NPV therefore understates the correct hurdle, and the required saving must exceed $3M by the quasi-option premium. Labour economics (below): the acquisition market View A relies on shrinks precisely when it is needed. Three routes, one threshold. That convergence is why I do not treat this as a close call. PART II — ARITHMETIC: WHAT THE SCENARIO'S OWN NUMBERS SAYThrough-line restated: the business case prices output and not training. Here is the size of the omission. Reproduction table — deriving the stated figures from the stated inputsQuantity Derivation Value Stated annual saving given $3.0M Stated cost reduction given 40% Baseline human cost of routine work $3.0M ÷ 0.40 $7.5M / yr AI-run cost of routine work $7.5M − $3.0M $4.5M / yr Expert attrition rate 60% ÷ 7 yrs 8.57% of base / yr Time to competence (t) given, 2–3 yrs 2.5 yrs (midpoint) Undiscounted 7-yr saving $3.0M × 7 $21.0M PV of 7-yr saving @ 8% real $3.0M × 5.206 $15.6M (@10%: $14.6M) The scenario's numbers are internally consistent. That is not the problem. The problem is what it never computes. THE NUMBER THE SCENARIO OMITS — pipeline headroomNobody in this thread, Bex included, has asked the only question that determines the answer: how much routine work does the training pipeline actually need? Let L = juniors per expert (leverage), k = the fraction of juniors who convert to experts, t = 2.5 years. Experts produced per year = k · A / t, where A = apprentice-years of routine work absorbed annually. Experts required per year = 0.0857 × N (expert headcount). Coverage ratio R = 4.67 · k · L. Safe automation share s = 1 − 1/R. k (conversion) L (juniors/expert) R Safe automation share s Saving captured 0.50 0.5 1.17 14% $0.4M 0.50 1.0 2.33 57% $1.7M 0.60 1.0 2.80 64% $1.9M 0.60 1.5 4.20 76% $2.3M 0.50 2.0 4.67 79% $2.4M 0.70 2.0 6.53 85% $2.6M The conversion band is conservative. In Germany, about 74% of apprentices received an employment contract with the training company after completing vocational training in 2021 — I model 50–70%. Across the plausible band, a disciplined View B captures $1.7M–$2.6M of the $3.0M, midpoint ~$2.1M. The Last-Third Problem Over seven years that is $6.3M nominal, ~$4.7M present value. Set it against what the last third actually buys you: the obligation to acquire 0.6N experts on the open market. Gallup puts replacement cost at roughly 50–200% of annual salary depending on role; SHRM's headline cost-per-hire figure of about $4,700 captures only hard recruiting spend. Gallup's role-level breakdown estimates replacement of leaders and managers at around 200% of salary and technical professionals at 80%. At the top of the documented range, Medical Economics puts replacement of a highly educated, highly skilled professional at 213% of salary, and replacement of a physician at over $1 million once all factors are counted — I flag that last figure as an outlier and do not lean on it. At a $180k expert salary, 100–213% gives $180k–$384k per replacement. The extra PV savings View A is fighting for amount to roughly $78k per departing expert at N=100. That is a 2.3x to 4.9x adverse trade, before a single second-order cost is priced. Break-even on the contested termView A's whole fallback is "you can train or hire for it then." So solve for the acquisition premium at which that fallback breaks even: P* = PV(savings) ÷ (0.6 × N) = $15.6M ÷ 0.6N = $26.0M ÷ N Expert headcount N Experts to replace Break-even premium per expert Documented cost ($180k salary) Verdict 50 30 $520k $180k–$384k View A survives 100 60 $260k $180k–$384k Marginal — coin flip 150 90 $173k $180k–$384k View A fails 250 150 $104k $180k–$384k View A fails badly And note: this is the full $3M version. Run it on the honest $0.9M increment and the break-even collapses to $78k–$156k per expert at every headcount above 50. View A fails on its own arithmetic across virtually the entire plausible parameter space, before we price anything the proposal omitted. Constants-as-variables auditFour parameters are presented as facts. All four are design choices, and I refuse to let either side treat them as laws of nature. 1. "40% lower cost." This is a gross-cost comparison. It excludes verification, rework, escalation handling, and the exception queue. Note the recurring pattern in the only two documented reversals we have: Klarna's CEO told Bloomberg on 8 May 2025 that cost had been too predominant an evaluation factor, and what you end up with is lower quality. Design lever: re-quote the 40% net of a measured exception cost. Recoverable range: 40% gross plausibly becomes 22–33% net. 2. "2–3 years." This is the softest constant in the scenario, and my own side abuses it. Aviation proves training thresholds are policy, not physics — see the negative-control section on the 1,500-hour rule. On the other side, the FAA reports that tower simulation systems can reduce time to certify new hires by up to 27%, based on a 2021 study. Design lever: t is compressible by 20–30% with instrumented feedback, without touching the substrate. This is where View A's best insight actually lives — and it strengthens View B, because a shorter t raises R and lets you automate more, safely. 3. "Same quality." Measured on today's distribution of routine work, by graders who themselves learned on routine work. It is a snapshot with an expiring warranty, and it is measurable only while humans still do the work. See Correction 6. 4. "AI isn't reliable on complex judgment." Also a snapshot, and the one parameter View A is entitled to argue moves. I compute it at its ceiling in Part V. PART III — THE SIX CORRECTIONSThrough-line restated: each correction is a place where the business case priced output and skipped training. Correction 1 — The comparison was wrong. AI-vs-humans was compared against a strawman; it should have been compared against AI-plus-the-Teaching-Share.The scenario offers a binary that no competent operator would accept. Consequence: the contested amount falls from $3M to ~$0.9M, and every subsequent argument in View A is arguing for the wrong number. This is my one derived exception to flying View B unqualified: I do not defend keeping all routine work. I defend keeping the Teaching Share — and I give you the test to find it in Part VI. Correction 2 — The savings figure double-counts. "$3M spent on tasks for their own sake" is false by the training system's own accounting.View A's rhetorical premise is that this money buys nothing but the tasks. It buys the tasks and the experts, and the split is documented. OECD analysis of the German system finds companies incur gross training costs of around EUR 18,000 per apprentice per year, of which about EUR 12,000 is recouped through the apprentice's productive contribution, leaving a net cost of roughly EUR 6,000 (Pfeifer et al., 2021) — with German firms estimated to invest approximately EUR 28 billion annually in apprenticeship training. Two-thirds of apprenticeship cost is recouped through output. Applied to this scenario: of the $7.5M, roughly $2.5M is the genuine training subsidy and $5.0M is production you would have to buy anyway. Consequence: View A's headline framing overstates the "wasted" spend by 3x. Correction 3 — "People can learn faster alongside AI" was asserted as a premise. It has since been measured three times, and it failed each time.This is the load-bearing claim in View A, and it is the one claim with hard experimental evidence against it. Bastani, Bastani, Sungu, Ge, Kabakcı & Mariman, "Generative AI Can Harm Learning," Wharton School Research Paper, SSRN 4895486 (15 July 2024). A field experiment in a Turkish high school with nearly a thousand students. Access to GPT-4 significantly improved performance — 48% for a standard ChatGPT-style interface and 127% for a learning-safeguarded tutor — but when access was subsequently removed, students performed worse than those who never had access, a 17% reduction for the standard interface. The authors' interpretation is that students used GPT-4 as a crutch during practice and then performed worse on their own. Read the second result carefully, because it is decisive and almost always misquoted. The scaffolded tutor eliminated the harm and produced no learning gain. The negative effect is essentially eradicated in the GPT Tutor arm, though no positive effect is observed. So the ceiling on View A's mechanism, executed perfectly, is parity with doing the work yourself. There is no acceleration to buy. Shen & Tamkin (Anthropic), "How AI Impacts Skill Formation," arXiv:2601.20245 (28 January 2026). Randomized experiments on professional and freelance programmers learning an unfamiliar asynchronous library. AI use impaired conceptual understanding, code reading, and debugging abilities without delivering significant efficiency gains on average; participants who fully delegated showed some productivity improvement but at the cost of learning the library. The authors note this setup differs from agentic coding products, and expect the impacts of such programs on skill development to be more pronounced than these results. Consequence: View A's substitute mechanism has null-to-negative evidence in two independent RCTs, in two domains, on two populations. A proposal that rests on an unmeasured premise is a hypothesis; a proposal that rests on a premise measured and falsified is a mistake. Honest limits, stated because I would rather state them than be caught by them: a published power analysis of the Bastani data finds the unassisted-harm effect sits near the detection boundary, meaning the study may be only marginally powered to detect that specific harm — the learning-gain absence, however, is robust. Shen & Tamkin themselves flag that skill formation ideally takes place over months to years, while they measured it over one hour for a single Python library. Neither caveat rescues View A; both cap how hard I lean. Correction 4 — The exit option was priced at zero. "Hire for it then" assumes a market that this same decision, taken sector-wide, destroys.This is a composition fallacy, and we have two clean natural experiments. Radiology — the decisive precedent, because it is this exact decision, already run, on real people. In 2016 Geoffrey Hinton advised that we should stop training radiologists immediately, on the view that within five years deep learning would outperform them and there were plenty already. The supply side responded. From 1991 through the early 2000s roughly 5–7% of US MD seniors applied to radiology each year; by 2015 that share had fallen to 3.8%, the lowest in 25 years — and the residency pipeline could not respond, because the Balanced Budget Act of 1997 froze Medicare-funded residency positions and Congress did not meaningfully expand the cap until 2021, of whose first 200 new positions only six went to diagnostic radiology. Radiology residency positions did grow 33% between 2010 and 2025, from 1,090 to 1,449, but that growth lagged the overall growth in residents across all specialties. The demand side did the opposite. Eight years on, a radiology resident writing in The New Republic observed that the prophecy did not come true and the field faces the largest radiologist shortage in its history, with imaging backlogged for months at some centres. The price signal is unambiguous: around 4,333 active radiologist job listings as of March, with a 130-day average time to fill, pushing average radiologist salary to $571,000 as of 2025, up 9% year over year according to Medscape. Meanwhile the AI arrived exactly as forecast: Mayo Clinic's radiology department now runs more than 250 AI models with a dedicated team of 40 AI scientists, researchers and engineers — and employs over 400 radiologists, a 55% increase since Hinton's forecast. Air traffic control — the mechanics of why money cannot buy you out. As of April 2026 the FAA has approximately 11,000 certified professional controllers across more than 300 facilities with a further 4,000 in the training pipeline, and it can take more than two years to fully certify a new hire depending on facility complexity. The 2026–2028 Workforce Plan requires 12,563 CPCs, revised down from the 14,633 the agency forecast in 2024; the 2024 plan noted the agency was about 4,000 controllers short, and that year 2.2 million hours of overtime cost taxpayers $200 million according to a National Academies report. The hiring rate is not the constraint. GAO reports that most candidates must graduate a 4-to-6-month Academy course followed by on-the-job training, that certification can take up to six years, and that only about 2% of applicants qualify for and complete the full training process. GAO's December 2025 review called the Academy a bottleneck; each earlier pause — 2013 sequestration, the 2019 shutdown, the 2020 Academy closure — carved a hole that showed up as a shortage three, four, five years later, long after the people who made the call had moved on. Consequence: the "hire for it then" clause has a documented failure mode with a 2-to-10 year latency and no cash remedy. It is worth zero, and the business case books it as free. The structural law, in one line: you cannot buy your way out of a queue whose length is measured in years, in a market where everyone joined the queue on the same day. Correction 5 — Reversibility was assumed symmetric. It is not, and the asymmetry has a price.The $3M is recoverable in one budget cycle: stop paying, start paying. Capability is recoverable in t + detection lag + market lag — 5 to 10 years, per Correction 4. Arrow–Fisher says the correct hurdle exceeds the NPV hurdle by the quasi-option value. We can put a number on relearning. Vogtle Units 3 and 4 were originally expected to cost $14 billion and enter service in 2016 and 2017; the project ran into significant delays and overruns. Unit 3 entered commercial service in July 2023 and Unit 4 in April 2024, by which point the price tag exceeded $36.8 billion. The proximate cause is the one this scenario is about — the independent Vogtle construction monitor's post-mortem attributes the overrun in part to limited nuclear construction labour and expertise after the reactor order pipeline dried up. Meanwhile the parallel V.C. Summer expansion was abandoned in July 2017 after Westinghouse filed for Chapter 11 in March 2017 citing $9 billion in losses across its two US nuclear projects. Signed confound, stated both ways. The contractor claims relearning worked: Bechtel states that after completing Unit 3 it drove costs down by 30% delivering Unit 4 — vendor-reported, flagged as such, and contested: reporting in April 2026 notes claims that Unit 4 was cheaper or that there was a meaningful learning curve are not backed by public documentation. I take the honest reading: relearning is possible, at roughly 2.5x the original budget and 8 years late. That is the price of the option View A proposes to sell. Consequence: raise the hurdle. A $3M saving with a 5–10 year irreversible tail should be evaluated against a materially higher threshold than a $3M saving you can undo next quarter. Correction 6 — The measurement counterfactual was destroyed, and this is the deepest inversion in the proposal.The policy eliminates the data it needs in order to remain safe. "AI does the work at the same quality" is a measurement, and it is only obtainable while humans still do the work. Automate 100%, and "same quality" becomes structurally unfalsifiable within one review cycle. There is no baseline, no blind comparison, no independent grader — and, within one t, nobody in the building who learned the fundamentals well enough to grade it. We have the worked example of what that looks like at the end state. In Bates & Ors v Post Office Ltd (No. 6) "Horizon Issues" [2019] EWHC 3408 (QB), a 313-page judgment published in December 2019 following 21 days of hearings, Fraser J found that bugs, errors and defects in Horizon rendered it unreliable and had the potential to cause discrepancies in subpostmasters' accounts, and was critical of the Post Office's evidence. The wrongful prosecutions ran from 1999 to 2015; the CCRC has called it the most widespread miscarriage of justice it had ever seen and the biggest single series of wrongful convictions in UK legal history, having referred 77 Horizon convictions with 69 overturned to date. Parliament eventually legislated: the Post Office (Horizon System) Offences Act 2024 received Royal Assent on 24 May 2024, quashing relevant convictions, with compensation of at least £600,000 under the Overturned Convictions scheme. Horizon is not a story about a buggy system. Every large system has bugs. It is a story about an organisation in which nobody retained the standing or the competence to say the machine was wrong, for sixteen years, against the evidence of hundreds of people who could see it. And there is a technical layer beneath the organisational one, which is why I call this the deepest inversion. Shumailov, Shumaylov, Zhao, Papernot, Anderson & Gal, "AI models collapse when trained on recursively generated data," Nature 631, 755–759 (July 2024) — indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which the tails of the original content distribution disappear. The core requirement is that fresh human data must be periodically injected to prevent collapse. (Author Correction: Nature 640, E6, 2025.) Analogy, and I am labelling it as an analogy rather than a mechanism: an organisation that stops producing human routine work stops producing the fresh, human-generated corrections that both audit the system and improve it. It is the same shape of failure, one level up. Consequence: the Ground-Truth Reserve, specified in Part VI, is not a nice-to-have. It is the instrument that keeps "same quality" a measurable claim. PART IV — OPERATIONS: THE JUDGMENT LEDGERThrough-line restated: these are the standing costs of running the AI-only design that the $3M business case does not carry. They are costs, not risks. They belong in the ledger. 1. Escalation and exception handling — the biggest single omission. Every automated routine system generates a queue of cases it cannot close. Under View A that queue routes to the shrinking senior cohort, whose time is the scarcest resource in the building. Magnitude: at a 5–12% escalation rate on $7.5M of work volume, senior time absorbed is $375k–$900k/yr at expert loaded rates. Klarna is the documented instance of exactly this: the company was not pulling AI out of the high-volume tier — it reintroduced humans for the premium and complex-case tier where AI parity had not held. 2. Increased training expenditure — Bainbridge's direct implication. If humans retain the exceptional cases and lose the routine practice, per-head training spend must rise. Magnitude: at minimum the net apprenticeship cost you removed, ~$2.5M/yr equivalent (Correction 2), now purchased as explicit instruction with no offsetting production. This is the item that makes View A's net saving negative on its own logic. 3. Loss of the measurement counterfactual. Correction 6. Magnitude: the standing cost of running a Ground-Truth Reserve — 8–12% of volume, i.e. $600k–$900k/yr — which View A must also pay if it wants to keep asserting "same quality." It does not book it. 4. Model and vendor drift monitoring. Routine work quality is a moving target as models and providers change. Somebody must re-baseline on every version change, with authority to halt. This is a permanent staffed function, not a project. 5. Regulatory oversight competence. Now a legal line item, not a governance aspiration — see Part VI. Magnitude: documented per-person training programmes, records, and named accountable overseers. 6. Data retention and breach surface. Routing 100% of routine work through an AI system creates a consolidated, queryable record of every transaction that previously lived in dispersed human judgment. This is a new liability class with its own insurance and disclosure consequences. 7. Contractual and partner contamination. Client contracts, professional-indemnity terms, and partner agreements that presume human performance of the routine tier need renegotiation. In regulated professions, some cannot be renegotiated at all. 8. Irreversibility itself, as a carried cost. Per Correction 5 and Vogtle: the option you sell has a documented buy-back price of roughly 2.5x and 8 years. Ledger total, conservatively bounded: $3.5M–$5.5M per year of standing obligations against $3.0M of gross saving. I am not claiming precision on these. I am claiming that the proposal claims precision on one side of the ledger and silence on the other. PART V — STEELMAN, GENEALOGY, AND THE ADVERSARIAL COMPUTATIONThe genuinely strong version of View A, cited, and the part I concede as a strengtheningThe best case for View A is not "training is an old habit." It is this: much routine work has a feedback horizon too long, or a correction signal too weak, to teach anything — and paying humans to do it is a training subsidy that buys no training. I concede that entirely, and adopting the constraint removes their only strong argument, because it converts the debate from "how much do we automate" to "which work teaches" — a question with an operational answer. The negative control — my own side, done wrong, labelled as such. In Da Silva Moore v. Publicis Groupe, 287 F.R.D. 182 (S.D.N.Y.), Magistrate Judge Andrew Peck's opinion of 24 February 2012 held that predictive coding is a legally acceptable method for searching electronically stored information in appropriate cases — the first opinion to approve the use of technology-assisted review. The labour-market effect was immediate and explicit: contemporaneous practitioner commentary noted that clients were no longer willing to pay for junior-level attorneys to do document review, with senior-level insight delivered instead via predictive coding. Fourteen years on, there is no documented collapse in the litigation profession's ability to produce senior litigators. Why not? Because document review was routine work that did not teach. The judgment ladder in litigation lives in depositions, motion practice, client counselling and case strategy — and those survived TAR intact. A View B advocate who had defended document review on training grounds in 2012 would have been wrong, expensively. That is my falsifiable commitment: if the routine work in question is document-review-shaped, automate it, and I lose nothing. Second self-charge — my own side's characteristic failure mode: mistaking hours for judgment. After Colgan Air Flight 3407 crashed near Buffalo on 12 February 2009 killing 50 people, Congress mandated higher qualification standards through the Airline Safety and FAA Extension Act of 2010, and since 1 August 2013 the FAA has required all Part 121 first officers to hold an ATP certificate, replacing a commercial certificate requiring roughly 250 hours. But none of the 25 new NTSB recommendations issued after the Colgan crash addressed pilot flight time as such, and the 1,500-hour mandate was the one element of Public Law 111-216 not traceable to an NTSB recommendation — and both Colgan pilots had over 2,000 flight hours, so the rule would not have kept either out of that cockpit. The costs were real and measurable: about 30 small US airports lost commercial service between mid-2013 and the end of 2015, with at least seven more in 2016, according to an analysis by Diio Mi, and the rule made becoming a pilot cost roughly $250,000 out of pocket and two to three years for anyone not trained by the military. Both ways, honestly: ALPA published a detailed analysis in February 2026 arguing that despite enormous traffic growth, the Part 121 safety record since the rule took effect has been substantially better than before it, and in 2025 FAA Administrator Bryan Bedford publicly backed the current standard. This is a genuinely contested confound and I present it as one. What it proves for this decision is narrower and firm: time-served is a lazy proxy. Defend the feedback, not the calendar. Genealogy: why intelligent operators hold View AThey hold it for three defensible reasons, none of which survives contact with the arithmetic. First, the saving is measurable this quarter and the damage is not measurable at all until it is unfixable — an incentive structure, not a belief. Second, they have seen real cases like TAR where the training-loss fear was wrong, and they generalise correctly from a genuine data point. Third, they are subject to the same illusion the evidence keeps catching: in METR's randomized controlled trial, developers forecast AI would cut completion time by 24% and, after completing the study, estimated it had reduced completion time by 20% — while measurement showed allowing AI increased completion time by 19%. A 39-point gap between felt and measured productivity is the default human response to these tools. That is the genealogy. It is sympathetic. It is still wrong. Flagged honestly, because it cuts against me: METR published an update on 24 February 2026 stating that data from its larger follow-up experiment gives an unreliable signal of the current productivity effect, with estimated speedup among newly recruited developers at −4% and a confidence interval from −15% to +9%. Closed by: §Part V. I use METR only for the perception-measurement gap, not for a claim about current AI speed. Adversarial computation: View A's fallback, computed to its ceilingThe fallback: "AI will handle complex judgment within five years, so the pipeline is a wasting asset. Save the money." Grant it better-than-documented assumptions. Grant expert-level reliability on complex judgment in five years — earlier than any forecast either view has offered. View A still loses, for two independent reasons. Ceiling reason 1 — the requirement is legal, and it does not lapse when the model gets good. The EU AI Act (Regulation (EU) 2024/1689) makes competent human oversight a compliance obligation. Article 14(4) requires that the natural persons tasked with human oversight have the necessary competence, training and authority to fulfil their role, and Article 26 places the operating side of that duty on deployers of high-risk systems, while the Article 4 AI literacy duty became applicable on 2 February 2025, six months after entry into force on 1 August 2024. Article 4 applies to all AI, not just high-risk, and binds providers and deployers alike. Effective dates checked against the enacted position, because they moved. The Digital Omnibus on AI reached provisional agreement on 7 May 2026, was endorsed by the European Parliament on 16 June and received the Council's final green light on 29 June, with Official Journal publication expected in July. The revised timetable: Article 50 transparency obligations from 2 August 2026; high-risk obligations for Annex III stand-alone systems from 2 December 2027; Annex I embedded systems from 2 August 2028. Article 5's prohibitions were also amended to add a ban on AI-generated non-consensual intimate imagery and CSAM. (Compact caveat: as of mid-June 2026 the Omnibus was agreed but pending OJ publication, so the original 2 August 2026 date remained technically in force in the interim — the Parliament confirmed it at plenary on 16 June 2026 by 423 to 57 with 174 abstentions, with Council adoption and OJ publication still required.) The deferral is itself the argument. The compliance machinery was not ready. The competence requirement did not go away — it landed 16 months later, on an organisation that under View A will by then be 16 months further into having nobody with the competence. Ceiling reason 2 — the load-bearing parameter's real-world range. View A's fallback rests entirely on the date of expert-level reliability. The only documented, dated, public forecast of exactly this kind, made by the most credentialed possible forecaster, was Hinton's five-year radiology call in 2016. It is now ten years past and the specialty is in its worst shortage on record at $571,000 average compensation. The empirical range on this parameter is: one observation, off by more than a factor of two, in the wrong direction. You may not build a $21M irreversible bet on it. Robustness both ways: the single flipping combinationThe combined-adverse cell that would flip me to View A, stated precisely so it can be checked: all four of (a) N ≤ 60 experts, and (b) the routine work is document-review-shaped — no correction reaches the novice inside a learnable horizon, and (c) the sector is not automating in parallel, so the lateral market stays deep, and (d) the organisation has already contracted a replacement pipeline at fixed price. What accepting that cell implies: you are asserting that your routine work teaches nothing and that you are the only firm in your sector making this decision. Both are testable, today, in a week. Run the test before you claim the cell. And the mirror case that should worry View A far more: if k · L < 0.214 — a thin junior tier with weak conversion — then R < 1 and the pipeline is already failing quota before any AI arrives. In that world View A does not cause the shortage; it accelerates a shortage already in progress by 2–3 years and removes the only instrument you had to detect it. PART VI — THE OPERATIONAL FIX, WITH A NAMED ACCOUNTABLE OWNERThrough-line, final restatement: price the training system, then automate everything that is not it. 1. The Teaching Share, identified by the Correction TestClassify every routine task class against a single, auditable question: Passes → Teaching Share. Stays human-first. AI participates only as an adversarial checker after the human attempt. Fails → automate immediately. No sentiment. Document review fails this test. So does most reconciliation, most formatting, most first-pass triage. Sizing: the Teaching Share is bounded above by the R table in Part II. Target the smallest human volume that holds R ≥ 1.3, and automate the rest. On the midpoint parameters that is automating 64–76% of routine work in year one, capturing $1.9M–$2.3M — more aggressive than View B as written, and defensible in a way View A is not. 2. Interaction-pattern constraint on the Teaching ShareThe evidence specifies the design, and it is unusually precise for a governance control. Shen & Tamkin identify six distinct AI interaction patterns, three of which involve cognitive engagement and preserve learning outcomes even when participants receive AI assistance. Bastani et al. showed the same thing from the tooling side: negative learning effects were largely mitigated by the safeguards built into the GPT Tutor arm. Rule: within the Teaching Share, AI may critique, counter-example, or explain — after the human attempt. It may never draft first. Outside the Teaching Share, no constraint at all. 3. The Ground-Truth Reserve8–12% of automated volume is routed human-first, blind, monthly, and graded against the AI output. This is the instrument that keeps "same quality" falsifiable and preserves the audit baseline destroyed in Correction 6. Cost: $600k–$900k/yr, booked openly. View A owes this cost too and does not book it. 4. Named accountable ownerA Capability Registrar, at director level, reporting to the audit committee and not to the COO — because the officer accountable for the $3M must not also be the officer certifying that the pipeline is healthy. Authority: veto over automating any task class that passes the Correction Test; quarterly publication of R; custody of the savings escrow that funds the reversal. 5. Falsification dashboard — decision rules, pre-committedMetric Baseline (today) Success threshold Escalation / reversal trigger ⭑ PRE-COMMITTED REVERSAL — Pipeline coverage ratio R (experts produced ÷ required, trailing 12 mo) measure in month 0 R ≥ 1.30 R < 1.15 for two consecutive quarters → automation share rolled back 15 percentage points within 90 days, funded from the savings escrow. Non-discretionary. Registrar executes without further board approval. Human catch rate on Ground-Truth Reserve month-0 blind sample ≥ 97% agreement < 93% in any month → automation of that task class suspended Unassisted competence exam, 18-month cohort pre-AI cohort score ≥ 100% of baseline ≤ 92% of baseline → Teaching Share expanded, AI drafting withdrawn from that class Median time to first independent complex judgment current t ≤ current t > t + 6 months → Correction Test re-run across all classes Escalation rate to senior cohort month-0 rate ≤ 8% of volume > 15% → net-of-exception cost re-quoted to the board Voluntary expert attrition current ≤ current > 1.25x current → treat as leading indicator, escalate Net saving after Judgment Ledger $0 ≥ $1.5M / yr < $0.75M for two quarters → programme paused Two derived honest limits. L1: if the Correction Test cannot be scored reliably for a task class, that class defaults to human until it can — ignorance is not a licence to automate. L2: the Registrar's R figure must be computed on completions, not headcount; counting bodies in seats is how the FAA's pipeline looked healthy while the shortage compounded. PART VII — EVIDENCE AT SCALE, WEIGHT-TABLEDLoad-bearing portfolioPrecedent Structural match Why it bears on this decision Weight Radiology, 2016–2026 (Hinton forecast; MD-senior applications to 3.8% by 2015; BBA 1997 GME cap; residency slots 1,090→1,449; avg comp $571k, +9% YoY; 130-day fill; Mayo 250+ models, 400+ radiologists, +55%) Identical: forecast of AI substitution → pipeline contraction → shortage The only case where this exact decision was made publicly and the outcome is now observable. It falsifies "you can train or hire for it then" and "AI will take the complex work" in the same decade DECISIVE US air traffic control, 2013–2026 (~11,000 CPCs vs 12,563 required; 4,000 in pipeline; 2–6 yrs to certify; ~2% applicant completion; 2.2M OT hours / $200M in 2024) Identical: long-lead pipeline, deferred entry, delayed shortage Quantifies the lead time and the completion rate. Proves money and political will cannot compress a 2–6 year queue DECISIVE Bastani et al. 2024 (SSRN 4895486; ~1,000 students; +48%/+127% assisted; −17% unassisted; safeguards eliminate harm, produce no gain) Direct test of View A's substitute mechanism Ceiling of "AI as tutor," executed perfectly, is parity — there is no acceleration to buy VERY HIGH Shen & Tamkin 2026 (arXiv:2601.20245; RCT; impaired conceptual understanding, code reading, debugging; 3 of 6 patterns preserve learning) Same test, professional population, 2026 tooling Independent replication and it supplies the operational design constraint in Part VI VERY HIGH AF447 / FAA automation programme (BEA final report 5 July 2012, 25 recommendations; SAFO 13002 4 Jan 2013; PARC/CAST report 5 Sept 2013, 18 recommendations; SAFO 17007, 2017) Deskilling of the humans retained for the exceptional case The mature regulated answer: automate, then mandate preserved manual practice. Nobody in aviation argues the pipeline is optional any more VERY HIGH Statute of Artificers 1563 → repealed 1814 (7-year term; max 3 apprentices; 251 years in force) Origin of the norm this proposal would reverse Establishes the training externality as a 463-year-old structural problem, not a habit HIGH (GOVERNING) Vogtle 3&4 / V.C. Summer ($14B → $36.8B; service 2023/2024 vs 2016/2017; V.C. Summer abandoned July 2017; Westinghouse Ch.11 March 2017) Forgetting-by-not-doing, then paying to relearn Prices the buy-back on the option View A sells: ~2.5x budget, ~8 years HIGH Klarna, Feb 2024 → May 2025 (AI = 700 agents, 2.3M chats, 75% of volume, 35+ languages; headcount 5,000→3,500; reversal to premium/complex tier) Same-decade, same-shape over-automation and correction The reversal is scoped, not total — which is precisely the Teaching Share argument, run by an operator under real P&L pressure HIGH Corroborating setPrecedent Contribution Toyota, Honsha plant, 2014 (Bloomberg; Mitsuru Kawai, Senior Technical Executive; humans reinstated on crankshaft forging) Positive control. Kawai's stated rationale: to be the master of the machine you need the knowledge and skills to teach the machine. Deliberate manual retention inside the most automated manufacturer on earth German dual system (~1.2M apprentices, ~400,000 training firms; €18k gross / €12k recouped / €6k net per apprentice-year; ~€28bn annual firm investment; 74% retained; youth unemployment 6.5% vs EU 14.6%) Positive control + the unit economics of Correction 2. Proves apprenticeship is ~two-thirds self-funding Airbus A350 training design, 2014 (per DOT OIG, Jan 2016) Positive control. Airbus announced a training plan allocating the first simulator session to manual flying, so pilots learn to control the aircraft before automation is introduced. Fundamentals first, by design, on the most automated airliner of its generation Post-2013 Part 121 safety record (ALPA analysis, Feb 2026; FAA Administrator backing, 2025) Positive control, with the confound stated Boeing 787 / Hart-Smith 2001 ("Out-Sourced Profits," Boeing paper MDC 00K0096, Third Annual Technical Excellence Symposium, St. Louis) Hart-Smith outlined how outsourcing costs were being underestimated and how the strategy could undermine the knowledge base on which the organisation was based. Charges exceeding $3.6bn on the 787 and 747-8 produced a $1.56bn quarterly loss in Q3 2009. The warning was internal, specific, dated, and ignored Japan's succession crisis Non-Western regulated-market exhibit. A 2024 survey found 52.1% of companies reported no successor, and the Small and Medium Enterprise Agency estimated that by 2025 roughly half — 1.27 million — of SME owners over 70 would have no successor. A whole economy's demonstration of the 7-year retirement cliff, arriving Post Office Horizon (Bates No. 6 [2019] EWHC 3408 (QB); 1999–2015; PO(HS)OA 2024 c.14) The end state when nobody retains competence to challenge the machine Brynjolfsson, Chandar & Chen, "Canaries in the Coal Mine" (Stanford Digital Economy Lab; ADP payroll data) The November 2025 revision reports 16% relative employment declines for workers aged 22–25 in AI-exposed occupations, controlling for firm-level shocks, with declines concentrated where AI automates rather than augments. Sector-wide evidence that everyone is draining the same pool simultaneously — the Correction 4 mechanism, observed Shumailov et al., Nature 631:755–759 (2024) The recursive-degradation analogy for Correction 6 METR, arXiv:2507.09089 The 39-point perception-measurement gap that explains the genealogy Controls, confounds and the empty cellDeliberate negative control (my own side, done wrong, failing my own test): Da Silva Moore / TAR, 2012. Routine work automated, training substrate untouched, no capability loss. My thesis would have been wrong here, and the Correction Test is what would have told me. Signed confound, both directions: the 1,500-hour rule. Costs (30+ airports lost service 2013–2015; $250k entry cost; both Colgan pilots exceeded the threshold; NTSB never recommended an hours mandate) against benefits (ALPA's Feb 2026 safety analysis; FAA's 2025 reaffirmation). Presented unresolved, because it is unresolved. Vendor-estimate confound, flagged: Bechtel's 30% Unit-4 cost reduction is contractor-reported and publicly contested. Not load-bearing. Statistical confound, flagged: the Bastani unassisted-harm estimate is marginally powered; METR withdrew its own 2026 follow-up as uninterpretable; the Canaries authors' February 2026 update finds that with the broadest set of controls, the timing of decline in AI-exposed occupations becomes significant only in 2024. THE EMPTY CELL, named, with a standing invitation: I could not locate a single documented case of an organisation that removed its apprenticeship substrate and then successfully rebuilt senior expertise through external hiring, at scale, in a sector where competitors were automating in parallel. Not one, across eleven sectors and six jurisdictions. If any participant can produce one — with a date, a figure and a source — it is the strongest possible argument for View A, and I will take it seriously. Its absence is itself evidence. Portfolio: 19 documented precedents · 11 sectors · 6 jurisdictions (US, UK, EU/France/Germany, Japan, Sweden, Turkey) · 8 load-bearing rows · 4 positive controls · 1 deliberate negative control · 1 deep-history anchor. Analogy setThe seven-year apprenticeship (1563) — GOVERNING. Maps to: the training obligation is a positive externality no single firm internalises. The flight deck (AF447, SAFO 13002) — maps to: the skill you rely on for the rare case atrophies fastest under normal operations. The nuclear construction gap (Vogtle) — maps to: relearning is possible at ~2.5x price and 8 years' delay. The seed corn (Shumailov, Nature 2024) — maps to: the system consumes the human-generated corrections it needs to stay calibrated. Explicitly an analogy, not a mechanism. The radiology cohort (Hinton 2016) — maps to: a forecast about the future of work is self-fulfilling on the supply side and self-refuting on the demand side. The Horizon terminal (Bates 2019) — maps to: when nobody retains competence to contradict the machine, its output becomes the record. Arithmetic stands beside each of these. The analogies are load-bearing, not ornamental: each is doing the specific job of naming the mechanism that the $3M business case has no line for. PART VIII — ON BEX'S ANALYSISBex reaches my conclusion, and I am obliged to say why her route does not get there — because a judge reading both should not confuse them. Quarantine: her exhibit is a category error. Bex offers Mayo Clinic interns as evidence that routine work builds judgment. Medical residency is not "routine work left in place by default." It is routine work deliberately protected by accreditation and public subsidy — precisely because the unaided market would not supply it. Citing it as proof that firms will naturally retain training work is like citing a levee as proof that rivers stay put. It is evidence for the externality, which is my Correction 1, not for her thesis. Her own witness supports my case more strongly than hers. Mayo did not defend the pipeline by resisting AI. Mayo's radiology department runs more than 250 AI models with a 40-person dedicated AI team, and employs over 400 radiologists — a 55% increase since 2016. Mayo automated aggressively and grew its expert cohort. That is the Teaching Share design, executed. It is not "keep the routine work." Upgrade, in three moves. Bex asserts risk; the argument needs a number — the Last-Third Problem, $0.9M not $3M. Bex asserts that routine work teaches; the argument needs a test — the Correction Test, which correctly clears document review and correctly holds the rest. Bex offers no threshold, so her position cannot be wrong; the argument needs a pre-committed reversal — R < 1.15 for two quarters, 15-point rollback, 90 days, funded from escrow. And her strongest sentence needs deleting. "Long-term risks outweigh immediate benefits" is exactly the framing that loses this argument in every boardroom, because it concedes that View B is the expensive option. It is not. On the ledger in Part IV, View A is the expensive option — it just books its costs in a different decade. COUNTERS, CLOSED"$3M on tasks for their own sake." — Closed by: §Correction 2. Two-thirds of apprenticeship cost is recouped through output; the true training subsidy is ~$2.5M, not $7.5M. "Beginners learn faster with AI on harder problems." — Closed by: §Correction 3. −17% unassisted (Bastani 2024); impaired understanding, reading and debugging (Shen & Tamkin 2026). Best case is parity. "If a gap opens, train or hire then." — Closed by: §Correction 4. Radiology: 10 years and counting. ATC: 2–6 year queue, 2% completion, $200M of overtime. "The old training path is a habit, not a law." — Closed by: §Part I + §Part V. It is a 463-year-old response to a real externality — and you are half right, which is why the Correction Test exists and why document review should have been automated in 2012. "AI will do the complex work soon anyway." — Closed by: §Part V adversarial computation. Legal oversight competence is required from 2 December 2027 regardless of model quality, and the only dated public forecast of this type is a decade overdue in the wrong direction. "Savings are certain, damage is speculative." — Closed by: §Correction 5. Arrow–Fisher: irreversibility raises the hurdle above NPV. Vogtle prices the buy-back at ~2.5x and 8 years. "We'd know if quality dropped." — Closed by: §Correction 6. You would not, because you will have destroyed the counterfactual and, within one t, the people able to read it. COMPOUNDING ASYMMETRYRun the two errors forward seven years. Wrongly choosing View B: you forgo $0.9M/yr. Total cost $6.3M nominal, ~$4.7M PV. Detectable at the first quarterly review of R. Reversible in one budget cycle by raising the automation share. Bounded, visible, cheap. Wrongly choosing View A: you gain $6.3M nominal, then discover in year 4–6 that R has been below 1 since year 1 — because there was no R, and no Ground-Truth Reserve to read it against. You now buy 0.6N experts at $180k–$384k each, in a market where early-career employment in AI-exposed occupations has fallen 16% relative and every competitor is bidding for the same shrunken cohort. Recovery time: 2.5 years of training plus 2–4 years of market lag, per radiology and ATC. Unbounded, invisible until unfixable, and priced by a seller's market that your own decision helped create. The asymmetry compounds because the AI's own performance degrades along the same path — Correction 6 — so the failure arrives simultaneously on both sides of the human-machine pair, which is the specific configuration Bainbridge warned about in 1983 and AF447 demonstrated in 2009. Automate the work that produces output. Keep the work that produces the people who can tell you when the output is wrong — and if you cannot tell those two apart, that inability is your real finding, not your saving. View B. Without qualification.