Appearance
History and Philosophy of Machine Learning
机器学习的历史与哲学
Introduction
引言
EN: Machine learning (ML) is more than a branch of computer science; it is a profound intellectual project at the intersection of mathematics, statistics, neuroscience, and philosophy. Its history is a story of bold ideas, dramatic winters, and explosive springs. Its philosophy grapples with questions that have occupied thinkers for millennia: What does it mean to know? How does intelligence arise from experience? Can a machine truly understand?
ZH: 机器学习(ML)不只是计算机科学的一个分支;它是一项深刻的思想工程,位于数学、统计学、神经科学与哲学的交汇处。它的历史充满了大胆的构想、戏剧性的寒冬与爆发式的春天。它的哲学则直面困扰思想家数千年的问题:知道意味着什么?智能如何从经验中产生?机器能否真正理解?
EN: This article traces the major milestones in ML’s evolution and explores the philosophical currents that have shaped—and continue to shape—its development.
ZH: 本文追溯机器学习演进中的主要里程碑,并探讨那些塑造了、且仍在塑造其发展的哲学思潮。
Part I: A History of Machine Learning
第一部分:机器学习史
Antecedents (Pre-1950): The Mathematical Foundations
前身(1950年前):数学基础
EN: Long before computers existed, the mathematical tools that would power machine learning were already being developed. In 1763, Thomas Bayes’s work on probability was published posthumously, laying the groundwork for what would become Bayes’ theorem. In 1805, Adrien-Marie Legendre described the method of least squares for data fitting. Pierre-Simon Laplace formalized Bayes’ theorem in 1812. In 1843, Ada Lovelace envisioned Charles Babbage’s Analytical Engine as capable of processing not just numbers but any form of symbolic data—music, text, or logic—planting an early seed for the idea of a general-purpose thinking machine. In 1847, Augustin-Louis Cauchy first described gradient descent. Andrey Markov introduced Markov chains in 1913.
ZH: 早在计算机出现之前,最终将驱动机器学习的数学工具就已经在发展之中。1763年,托马斯·贝叶斯关于概率的著作在其身后出版,为后来所谓的贝叶斯定理奠定了基础。1805年,阿德里安-马里·勒让德描述了用于数据拟合的最小二乘法。皮埃尔-西蒙·拉普拉斯于1812年将贝叶斯定理形式化。1843年,阿达·洛芙莱斯设想查尔斯·巴贝奇的分析机不仅能处理数字,还能处理任何形式的符号数据——音乐、文本或逻辑——为通用思维机器的观念埋下了早期种子。1847年,奥古斯丁-路易·柯西首次描述了梯度下降。安德雷·马尔可夫于1913年引入了马尔可夫链。
1940s–1950s: The Birth of Neural Networks and the Coining of “Machine Learning”
20世纪40—50年代:神经网络的诞生与“机器学习”一词的提出
EN: The modern era began in 1943, when neuroscientist Warren McCulloch and logician Walter Pitts proposed the first mathematical model of an artificial neuron—the Threshold Logic Unit. This was followed in 1949 by Donald Hebb’s learning principle, which explained how neural connections could be strengthened through repeated activation.
ZH: 现代纪元始于1943年,当时神经科学家沃伦·麦卡洛克和逻辑学家沃尔特·皮茨提出了第一个人工神经元的数学模型——阈值逻辑单元。1949年,唐纳德·赫布提出了学习原理,解释了神经连接如何通过反复激活而得到强化。
EN: In 1950, Alan Turing proposed the Turing Test as a criterion for machine intelligence. In 1952, Arthur Samuel wrote the first computer learning program—a checkers-playing program that improved with experience. Samuel also coined the term “machine learning” in 1959. In 1956, John McCarthy coined the term “Artificial Intelligence” at the Dartmouth Workshop, formally establishing AI as a research field.
ZH: 1950年,阿兰·图灵提出了图灵测试,作为机器智能的一项标准。1952年,阿瑟·塞缪尔编写了第一个计算机学习程序——一个能通过经验不断改进的跳棋程序。塞缪尔还在1959年创造了“机器学习”这一术语。1956年,约翰·麦卡锡在达特茅斯研讨会上创造了“人工智能”一词,正式将AI确立为一个研究领域。
EN: In 1957, Frank Rosenblatt introduced the Perceptron, the first artificial neural network capable of learning from data. In 1958, Rosenblatt published his work on the Mark I Perceptron, a neural network computer.
ZH: 1957年,弗兰克·罗森布拉特提出了感知机,这是第一个能够从数据中学习的人工神经网络。1958年,罗森布拉特发表了他关于Mark I感知机——一台神经网络计算机——的研究。
1960s–1970s: Progress, Limits, and the First AI Winter
20世纪60—70年代:进展、局限与第一次AI寒冬
EN: The 1960s saw the introduction of Bayesian methods for probabilistic inference in ML. Donald Michie implemented a machine that could play Tic-Tac-Toe via reinforcement learning in 1963. The nearest neighbor algorithm for pattern recognition was developed in 1967. In 1969, Bryson and Ho introduced multistage backpropagation.
ZH: 20世纪60年代,贝叶斯方法被引入机器学习中的概率推断。1963年,唐纳德·米基实现了一台能够通过强化学习玩井字棋的机器。1967年,用于模式识别的最近邻算法被提出。1969年,布赖森和何毓琦引入了多阶段反向传播。
EN: However, in 1969, Marvin Minsky and Seymour Papert published Perceptrons, a book that mathematically demonstrated the limitations of single-layer neural networks. This triggered the first “AI Winter”—a period of reduced funding and pessimism about ML’s potential.
ZH: 然而,1969年,马文·明斯基和西摩·帕普特出版了《感知机》一书,从数学上证明了单层神经网络的局限。这引发了第一次“AI寒冬”——一个资金减少、对机器学习潜力普遍悲观的时期。
1980s: The Backpropagation Revival
20世纪80年代:反向传播的复兴
EN: The 1980s brought a dramatic resurgence. In 1980, Kunihiko Fukushima developed the Neocognitron, which laid the groundwork for Convolutional Neural Networks (CNNs). In 1982, John Hopfield introduced Recurrent Neural Networks (RNNs). More importantly, backpropagation—the algorithm that allows multi-layer neural networks to learn—was rediscovered and popularized. In 1986, David Rumelhart, Geoffrey Hinton, and Ronald Williams described backpropagation in its modern form; Yann LeCun had independently developed a similar approach in 1985. This decade also saw the rise of expert systems—rule-based systems that encoded human knowledge.
ZH: 20世纪80年代出现了戏剧性的复兴。1980年,福岛邦彦开发了神经认知机,为卷积神经网络(CNN)奠定了基础。1982年,约翰·霍普菲尔德引入了循环神经网络(RNN)。更重要的是,反向传播——使多层神经网络得以学习的算法——被重新发现并推广。1986年,大卫·鲁梅尔哈特、杰弗里·辛顿和罗纳德·威廉姆斯描述了现代形式的反向传播;杨立昆则在1985年独立开发了类似方法。这十年还见证了专家系统的兴起——一种将人类知识编码为规则的基于规则的系统。
1990s: The Shift to Data-Driven Learning
20世纪90年代:转向数据驱动学习
EN: The 1990s marked a fundamental shift from a knowledge-driven to a data-driven approach. Toward the end of the 1980s, in 1989, Chris Watkins developed Q-learning, which improved reinforcement learning methods. In 1995, Corinna Cortes and Vladimir Vapnik introduced Support Vector Machines (SVMs), which became widely used for classification tasks. The random forest algorithm was also introduced in 1995. In 1997, Sepp Hochreiter and Jürgen Schmidhuber introduced the Long Short-Term Memory (LSTM) network, a type of RNN capable of learning long-term dependencies. Also in 1997, Freund and Schapire proposed AdaBoost, an effective ensemble learning method.
ZH: 20世纪90年代标志着从知识驱动方法向数据驱动方法的根本转变。在20世纪80年代末的1989年,克里斯·沃特金斯开发了Q学习,改进了强化学习方法。1995年,科琳娜·科尔特斯和弗拉基米尔·万普尼克提出了支持向量机(SVM),它被广泛用于分类任务。随机森林算法也在1995年被提出。1997年,塞普·霍赫赖特和于尔根·施米德胡贝提出了长短期记忆(LSTM)网络,这是一种能够学习长期依赖关系的RNN。同样在1997年,弗罗因德和沙皮尔提出了AdaBoost,一种有效的集成学习方法。
2000s: Kernel Methods and the Birth of “Deep Learning”
21世纪初:核方法与“深度学习”的诞生
EN: The 2000s saw the widespread adoption of kernel methods and unsupervised learning techniques. In 2006, Geoffrey Hinton coined the term “deep learning” to describe new algorithms that allowed computers to “see” and distinguish objects in images. Crucially, the internet provided the vast amounts of data needed to train large models, while advances in GPUs and Moore’s Law provided the computational power.
ZH: 21世纪初,核方法和无监督学习技术得到广泛采用。2006年,杰弗里·辛顿创造了“深度学习”一词,用以描述使计算机能够“看见”并区分图像中物体的新算法。至关重要的是,互联网提供了训练大型模型所需的海量数据,而GPU的进步和摩尔定律则提供了算力。
2010s: The Deep Learning Revolution
2010年代:深度学习革命
EN: The 2010s were defined by the triumph of deep learning. In 2012, AlexNet—a GPU-accelerated Convolutional Neural Network developed by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton—won the ImageNet competition by a dramatic margin. This was a watershed moment that proved the power of deep neural networks at scale. In 2012, deep learning also surpassed traditional models in speech recognition. The Variational Autoencoder (VAE) was proposed in 2013, and Ian Goodfellow introduced Generative Adversarial Networks (GANs) in 2014. In 2016, AlphaGo—developed by DeepMind—defeated Lee Sedol, one of the world’s top Go players. That same year, Google Translate switched to deep neural networks. In 2017, the Transformer architecture was introduced in the paper “Attention Is All You Need,” which would become the foundation for large language models.
ZH: 2010年代以深度学习的胜利为标志。2012年,由亚历克斯·克里泽夫斯基、伊利亚·苏茨克维和杰弗里·辛顿开发的GPU加速卷积神经网络AlexNet,以巨大优势赢得ImageNet竞赛。这是一个分水岭时刻,证明了大规模深度神经网络的力量。2012年,深度学习在语音识别上也超越了传统模型。变分自编码器(VAE)于2013年被提出,伊恩·古德费洛于2014年引入了生成对抗网络(GAN)。2016年,DeepMind开发的AlphaGo击败了世界顶尖围棋棋手之一李世石。同年,谷歌翻译转向深度神经网络。2017年,Transformer架构在论文《Attention Is All You Need》中被提出,它后来成为大型语言模型的基础。
2020s: Generative AI and Foundation Models
2020年代:生成式AI与基础模型
EN: The 2020s have been defined by the rise of generative AI. GPT-3 (175 billion parameters) was released in 2020. Scaling Laws for neural language models were formalized. AlphaFold 2 solved the protein folding problem in 2021. Generative AI has led to revolutionary models, including advanced chatbots and text-to-image systems. Machine learning has entered the wider public consciousness, and the commercial potential of AI has driven massive increases in company valuations.
ZH: 2020年代以生成式AI的崛起为标志。GPT-3(1750亿参数)于2020年发布。神经语言模型的缩放定律被形式化。AlphaFold 2于2021年解决了蛋白质折叠问题。生成式AI催生了革命性模型,包括先进聊天机器人和文本到图像系统。机器学习已进入更广泛的公众意识,而AI的商业潜力推动了公司估值的巨大增长。
Part II: The Philosophy of Machine Learning
第二部分:机器学习的哲学
The Empiricist Tradition
经验主义传统
EN: One of the most profound philosophical connections is between machine learning and empiricism—the philosophical tradition, associated with thinkers like John Locke and David Hume, that holds that knowledge comes primarily from sensory experience.
ZH: 机器学习与经验主义之间存在着最深刻的哲学联系之一。经验主义是一种与约翰·洛克、大卫·休谟等思想家相关的哲学传统,认为知识主要来自感官经验。
EN: Cameron J. Buckner’s book From Deep Learning to Rational Machines argues that recent breakthroughs in deep learning can be understood as a realization of classical empiricist philosophy of mind. Empiricists argued that general psychological faculties—perception, memory, imagination, attention, and empathy—enable rational agents to extract abstract knowledge from sensory experience. Buckner shows how deep neural networks can be seen as modeling these very faculties.
ZH: 卡梅伦·J·巴克纳的著作《从深度学习到理性机器》认为,深度学习的近期突破可以被理解为古典经验主义心灵哲学的一种实现。经验主义者主张,一般心理官能——知觉、记忆、想象、注意和共情——使理性行动者能够从感官经验中提取抽象知识。巴克纳展示了深度神经网络如何可被视为对这些官能本身的建模。
EN: This connection is not merely historical. Philosophers such as Aristotle, Ibn Sina (Avicenna), William James, and Sophie de Grouchy developed faculty psychologies that are now being computationally instantiated in deep learning systems. As Buckner puts it, computer scientists can “mine the history of philosophy for ideas and aspirational targets,” while philosophers can see how “historical empiricists’ most ambitious speculations can be realized in specific computational systems.”
ZH: 这种联系不仅仅是历史性的。亚里士多德、伊本·西那(阿维森纳)、威廉·詹姆斯和索菲·德·格鲁希等哲学家所发展的官能心理学,如今正在深度学习系统中被计算化地实例化。正如巴克纳所说,计算机科学家可以“从哲学史中挖掘思想和理想目标”,而哲学家则可以看到“历史经验主义者最雄心勃勃的思辨如何在具体的计算系统中得到实现”。
The Nativism vs. Empiricism Debate
先天论与经验主义之争
EN: A perennial philosophical debate concerns the origins of abstract knowledge: is it innate (nativism) or learned from experience (empiricism)? Prominent scientists evaluating deep learning’s potential have explicitly cited this debate. Deep learning, with its emphasis on learning from data, represents a powerful instantiation of the empiricist position—though the debate continues over whether certain architectural priors constitute a form of “innateness.”
ZH: 一个长期存在的哲学争论关乎抽象知识的起源:它是先天的(先天论),还是从经验中习得的(经验主义)?评估深度学习潜力的著名科学家曾明确援引这一争论。深度学习强调从数据中学习,是经验主义立场的一种有力实例——尽管关于某些架构先验是否构成某种“先天性”,争论仍在继续。
Epistemological Challenges: Correlation vs. Causation
认识论挑战:相关与因果
EN: A central epistemological concern in ML is the distinction between correlation and causation. ML algorithms excel at identifying correlations in data, but they lack an inherent mechanism for discerning causal relationships. Misinterpreting correlations as causations can lead to erroneous conclusions, particularly in high-stakes fields like healthcare and finance.
ZH: 机器学习中一个核心的认识论关切是相关与因果的区分。机器学习算法擅长识别数据中的相关性,但缺乏辨别因果关系的内在机制。将相关性误认为因果性可能导致错误结论,尤其是在医疗和金融等高风险领域。
EN: This is not a minor technical issue—it is a fundamental epistemological limitation. As one analysis notes, “ML operates purely on statistical inference, relying on patterns rather than structured reasoning.” While classical AI aspires to build systems capable of conceptual abstraction and logical inference, “ML remains tied to empirical data, making its epistemological foundation markedly different.”
ZH: 这不是一个次要的技术问题——而是一种根本性的认识论局限。正如一项分析所指出的,“机器学习纯粹依靠统计推断运作,依赖模式而非结构化推理。”古典AI渴望构建能够进行概念抽象和逻辑推理的系统,而“机器学习仍然受缚于经验数据,使其认识论基础显著不同”。
Induction and the Problem of Generalization
归纳与泛化问题
EN: Machine learning is fundamentally an exercise in inductive reasoning—generalizing from specific training data to make predictions on new, unseen data. This raises the classic philosophical problem of induction, famously articulated by David Hume: How can we justify generalizing from past observations to future cases?
ZH: 机器学习从根本上是一种归纳推理实践——从特定训练数据中泛化,以对新的、未见过的数据作出预测。这引出了经典哲学中的归纳问题,大卫·休谟对此有著名表述:我们如何能够证明从过去观察推广到未来案例是正当的?
EN: In ML, this manifests as the challenge of inductive bias. Models may not account for new or unseen situations that differ from the training data, leading to inaccurate predictions and limited adaptability. The reliance on induction poses “unique epistemological challenges” that are central to understanding ML’s capabilities and limitations.
ZH: 在机器学习中,这表现为归纳偏置的挑战。模型可能无法考虑与训练数据不同的新情况或未见情况,从而导致预测不准确、适应性有限。对归纳的依赖带来了“独特的认识论挑战”,这些挑战对于理解机器学习的能力与局限至关重要。
The “Black Box” Problem: Epistemic Opacity
“黑箱”问题:认识论不透明性
EN: Deep learning models are often described as “black boxes”—their internal workings are opaque even to their creators. This raises profound epistemological questions:
ZH: 深度学习模型常被描述为“黑箱”——其内部运作即使对创造者而言也是不透明的。这引出了深刻的认识论问题:
EN: - Model-model understanding: How do ML models function internally?
- Model-world understanding: How does ML contribute to knowledge about the world?
ZH: - 模型—模型理解:机器学习模型内部如何运作?
- 模型—世界理解:机器学习如何贡献关于世界的知识?
EN: These questions touch on the nature of scientific representation. Some philosophers argue that ML models function as “highly idealized toy models” that can provide epistemic success despite lacking similarity to their targets. Others argue that ML models are “instruments that we use to facilitate our epistemic activities in science”—they do so “without scientific representation.”
ZH: 这些问题触及科学表征的本质。一些哲学家认为,机器学习模型作为“高度理想化的玩具模型”发挥作用,尽管与其目标缺乏相似性,仍能提供认识论上的成功。另一些哲学家则认为,机器学习模型是“我们用来促进科学认识活动的工具”——它们“无需科学表征”也能做到这一点。
The Theory-Free Ideal
无理论理想
EN: A provocative philosophical claim is that ML enables a form of “theory-free inductive inference.” This is the idea that ML can discover patterns directly from data without needing pre-existing scientific theories. Critics argue this is an illusion—that all learning involves prior assumptions and biases—but the debate continues over whether ML represents a genuinely new kind of scientific methodology.
ZH: 一个颇具挑衅性的哲学主张是,机器学习使某种“无理论的归纳推断”成为可能。这种观点认为,机器学习可以直接从数据中发现模式,而无需预先存在的科学理论。批评者认为这是一种幻觉——所有学习都涉及先验假设和偏置——但关于机器学习是否代表一种真正新的科学方法论,争论仍在继续。
Unsupervised Learning and Ontology
无监督学习与本体论
EN: Unsupervised learning—where models find patterns in data without labeled examples—raises unique philosophical questions. These methods “raise unique epistemological and ontological questions” about how and whether we can identify natural kinds, infer essential and contingent properties, and imagine unrealized possibilities. Some philosophers argue that unsupervised learning is “ontologically fundamental” compared to supervised or reinforcement learning.
ZH: 无监督学习——模型在无标注样本的数据中发现模式——提出了独特的哲学问题。这些方法“提出了独特的认识论和本体论问题”,涉及我们如何以及能否识别自然种类、推断本质与偶然属性,并想象未实现的可能。一些哲学家认为,与监督学习或强化学习相比,无监督学习在“本体论上更为根本”。
Philosophy-Informed Machine Learning (PhIML)
哲学引导的机器学习(PhIML)
EN: A recent development is Philosophy-Informed Machine Learning (PhIML), which “directly infuses core ideas from analytic philosophy into ML model architectures, objectives, and evaluation protocols.” PhIML promises “new capabilities through models that respect philosophical concepts and values by design.” This represents a growing recognition that philosophy is not merely an abstract exercise but can actively shape the design of ML systems.
ZH: 近期的一项发展是哲学引导的机器学习(PhIML),它“将分析哲学的核心思想直接注入机器学习模型架构、目标和评估协议中”。PhIML承诺“通过在设计上尊重哲学概念与价值的模型,带来新的能力”。这代表了一种日益增长的认识:哲学不只是抽象练习,还能主动塑造机器学习系统的设计。
The Dynamic Relationship Between History and Philosophy
历史与哲学之间的动态关系
EN: The history and philosophy of machine learning are deeply intertwined. Each major advance in ML has raised new philosophical questions, and each philosophical insight has opened new avenues for research.
ZH: 机器学习的历史与哲学深度交织。机器学习中的每一次重大进展都提出了新的哲学问题,而每一种哲学洞见也都开辟了新的研究路径。
EN: The empiricist tradition, centuries old, has found its most powerful computational expression in deep learning. The problem of induction, debated since Hume, is now a practical engineering challenge in generalization. The nature of representation, discussed by philosophers from Plato to the present, is now being tested in neural network architectures.
ZH: 历经数百年的经验主义传统,在深度学习中找到了其最强大的计算表达。自休谟以来争论不休的归纳问题,如今成为泛化中的实际工程挑战。从柏拉图到当代哲学家所讨论的表征本质,如今正在神经网络架构中接受检验。
EN: As we move forward, this interdisciplinary dialogue will only intensify. The history of philosophy offers a rich resource for thinking about the future of AI—and the future of AI will, in turn, reshape our philosophical understanding of mind, knowledge, and intelligence.
ZH: 随着我们向前推进,这种跨学科对话只会愈加深入。哲学史为思考AI的未来提供了丰富资源——而AI的未来也将反过来重塑我们对心灵、知识与智能的哲学理解。