Chapter 46: Machine Intelligence
Era span: 1958 perceptron → present · Difficulty: extreme
Requires: Ch 12, Ch 35, Ch 42
Unlocks: Ch 47
Data snapshot: volatile figures in this chapter (prices, capacities, deployment counts, regulation, and capability claims) reflect published sources through 2024 unless dated otherwise; check current data before planning.
Machine learning is statistics grown muscular: fitting flexible functions to data at scale. Demystified, its power remains enormous — and its limits (hallucination, bias amplification, brittleness) are engineering facts to design around, not mysteries.
46.1 The Learning Arc
- Perceptron (1958): a tunable weight layer learns classifications; Minsky/Papert prove single layers can't do XOR → funding winter (lesson: capability claims outrunning math invite corrections).
- Backpropagation popularized (1986): gradient descent through layered networks trains multi-layer models — the algorithmic unlock.
- Convolutional nets read images (LeNet handwriting); then AlexNet (2012): GPUs + ImageNet data + depth crushed the benchmark — hardware-software-data CO-EVOLUTION became the pattern (Ch 45's sensors, Ch 43's power economics all feeding one curve).
- Transformers (2017, “Attention Is All You Need”): attention supports parallel sequence processing and scalable training. Scaling produced large capability gains in many settings, but “bigger is better” is incomplete: data quality, compute, optimisation, post-training, evaluation, safety, and deployment cost all matter. One architecture supports many tasks, not guaranteed general reasoning.
| Wave | Unlock | Limit found | Lesson |
|---|---|---|---|
| Perceptron 1958 | Learned weights | No XOR (single layer) | Math audits hype |
| Expert systems 80s | Encoded rules (XCON $25M+/yr) | Brittle, unmaintainable | ROI mechanical ≠ general |
| Backprop 1986 | Deep training | Compute-starved | Wait for GPUs |
| AlexNet 2012 | GPU + data + depth | Hungry for labels/power | Co-evolution wins |
| Transformers 2017+ | Parallel scale | Cost, hallucination | Per-domain adoption |
46.2 What Learning Actually Is
Keep the demystified frame:
- A model is a function with millions-to-trillions of adjustable parameters.
- Training adjusts them to minimize error on examples (gradient descent finds slopes downhill).
- Generalisation—performance on unseen, representative data—is necessary. It does not establish safety, calibration, fairness, robustness, usefulness, or value; those require additional metrics and testing.
- RLHF-style alignment training shapes model behavior toward human preferences after base training — capability and behavior are separately tuned dials.
Evaluation doctrine: held-out sets never touched in training, leakage audits (near-duplicates count as cheating), per-domain scorecards (aggregate means hide the failing specialty), human-verified samples for consequential outputs. A benchmark that becomes a target stops measuring (Ch 47 Goodhart) — rotate and re-blind regularly.
46.3 Proven Wins
- AlphaGo/AlphaZero: superhuman game play via self-play reinforcement — search + learned intuition.
- AlphaFold: protein structure prediction solved-ish (50-year grand challenge) — structural biology accelerated by years-decades; drug targets arrive pre-computed.
- Weather models beating physics-based ensembles on speed and sometimes accuracy — forecasting democratized.
- Chip-design floorplanning, code assistants (measurable productivity gains in controlled studies), industrial inspection (Ch 45), drug-candidate screening (Ch 44).
| Win | Method | Why it transferred |
|---|---|---|
| AlphaGo → Zero | Self-play RL, no human games | Rules closed, simulator perfect |
| AlphaFold2 (92 GDT) | Attention + geometry | 50-yr labeled data (PDB) |
| Weather ML | Train on reanalysis | Physics ensembles as teachers |
| Code assistants | Next-token at scale | Repositories = labeled corpus |
46.4 Failure Modes (Engineering Facts)
- Hallucination/confabulation: fluent falsehoods generated with confidence — retrieval grounding, citation requirements, human verification for consequential outputs.
- Bias inheritance: training data's prejudices amplify silently — audits, representative data, outcome monitoring.
- Distribution shift: performance collapses off-distribution (the world keeps moving); monitor drift continuously.
- Adversarial fragility: tiny perturbations flip classifications — security mindset applies (Ch 42).
- Evaluation gaming: Goodhart's law (Ch 47) — when a benchmark becomes the target, it stops measuring capability.
Doctrine: AI augments verification-capable humans best; autonomous deployment belongs where errors are cheap and checkable.
| Failure | Guard | Cost of skipping |
|---|---|---|
| Hallucination | Retrieval + citations + human sign-off | Confident falsehoods in records |
| Bias | Audits + representative data + outcome stats | Silent discrimination at scale |
| Drift | Monitoring + retrain triggers | Decaying accuracy, nobody notices |
| Adversarial | Threat model + input validation | Stickers fooling classifiers |
| Gaming | Rotated blind evals | Benchmarks green, users red |
46.5 Compute Economics and Governance
- Frontier training runs cost $M→$B; inference optimization (quantization, distillation, caching) determines deployment economics — the fab race (Ch 35) now doubles as strategic policy (export controls treat chips like oil once was).
- Governance snapshots: risk-tiered regulation (EU AI Act pattern), safety institutes evaluating frontier models, liability allocation debates unresolved. Engineering stance: measurable evaluations, incident reporting norms, red-teaming as standard practice — aviation's post-mortem culture (Ch 33) ported to software intelligence.
Deployment arithmetic: training cost ÷ lifetime inferences + serving watts (Ch 43) = per-answer price. Distill/quantize/cache till the price clears the domain's wage comparison (Ch 45 $/hr logic, now per decision). Chips as strategy: fabs, power, and talent are the oil fields of this curve (Ch 51 pattern recurring).
46.6 The Honest Uncertainty Section
Trajectory debates are genuinely open: ceiling heights, timeline distributions, alignment difficulty all contested among experts. Planning stance for this book:
- Treat AI as a powerful amplifier whose exact ceiling is unknown — build institutions that benefit from strong-but-bounded systems while remaining robust to stronger ones.
- Possible feedback loops among model design, chips, data centres, funding, and scientific use could accelerate progress, but the size and timing of such loops are uncertain. Physical fabs, energy, data, institutions, and regulation remain independent constraints.
Capability gate: adopt AI domain by domain. Define a domain-specific task, baseline, error cost, human escalation, monitoring, and reliability target. AI can accelerate adoption when those conditions are met; aggregate forecasts about “AGI” do not substitute for measured performance.
46.7 The Machine-Intelligence Papers
- McCulloch & Pitts modeled neurons as logic gates (1943); McCarthy coined "artificial intelligence" at Dartmouth (summer 1956). The New York Times reported the perceptron (1958) as an embryo machine expected to "walk, talk, see, write, reproduce itself" — hype cycles have primary sources.
- Minsky & Papert's Perceptrons (1969) proved single-layer limits (XOR) and chilled neural-network funding; the Lighthill Report (1973) and DARPA cuts then brought the first broad AI winter of the mid-1970s. Expert systems boomed in the 1980s, and their late-1980s collapse brought the second: MYCIN matched expert physicians on bacterial meningitis studies yet never entered clinics (liability/regulatory friction — capability ≠ adoption); DEC's XCON configurer demonstrably saved $25M+/year, proving deployment where ROI was mechanical.
- Backpropagation: Werbos' 1974 thesis preceded the Rumelhart–Hinton–Williams Nature papers (1985–86) that popularized it; LeNet's convolutional nets read a significant share of US bank checks by the late 1990s — neural networks paid rent decades before their fame.
- Deep Blue beat Kasparov May 11, 1997 (match 3.5–2.5) using brute-force search plus handcrafted evaluation — a different paradigm than what followed. AlexNet (2012): top-5 error 15.3% versus runner-up 26.2%, trained on two consumer GPUs in under a week — hardware availability, not new theory, flipped the field (Ch 35's compute curve meeting Ch 42's datasets).
- AlphaGo's move 37 (game 2 vs Lee Sedol, March 2016): policy network assigned a human probability of ~1/10,000 — a documented moment where machine valuation visibly diverged from centuries of professional intuition. AlphaZero generalized via pure self-play; AlphaFold2 (CASP14, 2020) reached median backbone accuracy ~92 GDT — atom-level structure prediction for most targets, converting a 50-year problem into a web service.
- Transformers ("Attention Is All You Need," 2017) scaled predictably enough to become infrastructure; GPT-3's 175B parameters (2020) demonstrated few-shot behavior; ChatGPT (November 30, 2022) reportedly crossed ~100M users in two months (third-party estimate) — adoption speed itself became a policy input. The EU's AI Act (adopted 2024) tiered regulation by risk class; export controls turned advanced chips into strategic goods (Ch 51's resource-politics pattern recurring around computation).
46.8 Domain Scorecard (Run Quarterly)
Per domain (code, triage, dispatch, inspection, drafting): reliability vs verified baseline → error cost × volume → wage comparison → adopt / assist / hold. Aggregate AGI forecasts excluded from the meeting — domains decide, prophecies observe from the hallway.