Learn With Nathan Sign in

Can a Machine Think?

Alan Turing replaced an argument over an undefined word with an observable game—a move that made progress measurable without making philosophy disappear.

Chapter 3

Can a Machine Think?

replaced an argument over an undefined word with an observable game—a move that made progress measurable without making philosophy disappear.

Picture two closed doors. Behind one is a person. Behind the other is a machine. You cannot see either participant or hear a natural voice. You exchange written messages. Both answer questions, tell stories, make mistakes, challenge your assumptions, and respond to jokes. After several minutes, you must decide which is which.

This is the popular shape of the Turing test, although Alan Turing's 1950 paper presented a more careful and historically specific argument. The enduring power of the setup comes from its discipline: instead of inspecting a mysterious inner essence called thought, the evaluator judges observable behavior under stated conditions.

Businesses make versions of this move every day. “Is this candidate a good leader?” becomes a structured interview, references, and observed decisions. “Is this process secure?” becomes a threat model and a set of tests. The measurement is never identical to the underlying quality, but it allows evidence to replace some intuition.

Turing's move does not prove that intelligence is nothing more than imitation. It shows how a difficult question can become experimentally approachable. That distinction—between an operational test and a complete definition—is the heart of this chapter.

3.1

When the definition blocks the investigation

Turing opened “Computing Machinery and Intelligence” with the question “Can machines think?” He immediately saw trouble. What counts as a machine? What counts as thinking? If the answer depends on how ordinary people use two elastic words, a survey of language may replace an investigation of capability.

He therefore proposed changing the question. This was not evasion. Science and management frequently replace an inaccessible idea with an operational definition: a stated procedure that determines how the idea will be observed or measured. Temperature becomes a reading produced under defined conditions. Customer satisfaction becomes a specified survey or behavior. Model quality becomes performance on a chosen evaluation set.

A funnel transforms the vague question can a machine think into defined participants, communication channel, task, time limit, and decision rule, producing observable evidence rather than a final metaphysical answer.
Figure 3.1 · From a vague question to an operational test. Defining participants, conditions, evidence, and a decision rule makes inquiry possible—but the result answers the operational question, not every meaning of “think.”
Audiobook description

A cloud contains the vague question “Can a machine think?” A funnel narrows it through five rings: participants, channel, task, time, and decision rule. The output is a box labeled observable evidence. A note underneath says evidence under these conditions, not a complete theory of mind.

An operational definition creates leverage and risk. It creates leverage because different systems can be compared and claims can be challenged. It creates risk because people may forget that the measure is a designed proxy. Once a score becomes a target, builders optimize for the score. If the test captures only part of the real objective, impressive results can coexist with failure outside the test.

Plain English

Operational definition

An operational definition specifies how an idea will be observed or measured. It turns “good,” “safe,” or “intelligent” into a procedure that can produce evidence. The procedure is useful only to the extent that it represents the underlying quality we care about.

3.2

The imitation game Turing actually described

Turing began from a conversational party game with three participants: a man, a woman, and an interrogator separated from them. The interrogator tried to identify which participant was which through typed questions, while the participants had different aims in shaping the judgment. Turing then asked what would happen if a machine took one participant's place.

The text-only channel was essential. It removed appearance, mechanical voice, and physical performance from the judgment. A machine would not fail merely because its casing did not look human or because it could not reproduce human speech. The experiment concentrated attention on linguistic interaction.

An interrogator exchanges text with two hidden respondents labeled A and B, one human and one machine, then makes an identity judgment without access to appearance or voice.
Figure 3.2 · The behavioral core of the imitation game. The evaluator has access to a controlled transcript, not to bodies, voices, design documents, or inner experience.
Audiobook description

An interrogator at the bottom sends typed questions through two lines to closed rooms A and B. One contains a human and one a machine, but the labels are hidden from the interrogator. A curtain across the diagram blocks sight and voice. The interrogator's only evidence is the text conversation.

In the paper, Turing forecast that by around the year 2000, a machine with a specified storage capacity could have a meaningful chance of misleading an average interrogator after five minutes of questioning. The exact forecast matters less here than his willingness to make the claim testable. He also examined a wide range of objections—from mathematics and consciousness to originality and human inconsistency—and discussed the possibility of building learning machines rather than programming an adult mind fact by fact.

The imitation game is therefore smaller and larger than its popular caricature. It is smaller because it does not test every form of intelligence. It is larger because Turing used it to open a method of inquiry: judge a machine through what it can do, confront objections, and improve the experiment.

Where the shorthand breaks: “If a chatbot fools one person, it passes the Turing test” leaves out the protocol. Results depend on the interrogators, prompts, duration, comparison group, incentives, and success threshold. Without those conditions, “passed” is more publicity phrase than scientific result.
3.3

What behavioral testing gets right

A behavioral test can remain useful even when people disagree about the mechanism. A procurement team does not need a complete theory of language to test whether a system summarizes contracts accurately. A customer does not need access to model weights to discover that a support assistant repeatedly invents return policies. Observable outcomes are part of the evidence that deployment decisions require.

Behavioral evaluation also limits favoritism toward a particular technique. If the task is to route claims correctly, a rule system and a learned model can be tested on the same cases. The organization can compare error patterns, cost, speed, maintainability, and robustness rather than assuming that the newer method is superior.

Two different systems, one rule-based and one learned, receive the same test cases and are compared through outcomes, errors, cost, speed, and robustness without requiring identical internal mechanisms.
Figure 3.3 · Compare behavior without assuming one mechanism. A shared evaluation can reveal practical differences between systems built in different ways.
Audiobook description

A common stack of test cases splits into two paths. One enters a box labeled rule-based system; the other enters a box labeled learned model. Their results meet in a comparison table with five rows: outcomes, errors, cost, speed, and robustness. The internal boxes look different, but the external questions are shared.

Good behavioral evaluation has at least four properties. It uses cases that represent the intended environment. It defines success before seeing results. It examines different kinds of failure, not only one average score. And it repeats over time because the environment, data, model, or users may change.

AI in the Wild

A blind comparison for business writing

Suppose a company is evaluating an AI drafting assistant. Reviewers receive anonymized drafts: some from the current process, some from the proposed system, and some from the system plus human editing. They score factual fidelity, clarity, policy compliance, editing time, and serious errors. The blind comparison reduces brand excitement and focuses attention on the workflow outcome.

3.4

Fluency changes the judge

The imitation game is not only a test of the machine. It is also a test involving a human evaluator. That matters because people bring expectations, biases, fatigue, knowledge, and social instincts into the interaction.

A system can appear human by making mistakes, using informal language, or avoiding difficult questions. An evaluator may mistake confidence for competence, verbosity for depth, or agreement for helpfulness. The system may perform better with a novice interrogator than with a specialist. A short conversation can hide contradictions that become obvious over time.

The same fluent answer passes through different evaluator lenses—novice, expert, rushed, and skeptical—producing different judgments and showing that evaluation depends on the observer and protocol.
Figure 3.4 · The evaluator is part of the measurement system. A judgment reflects the system's behavior, the evaluator's knowledge and expectations, and the test conditions.
Audiobook description

One polished answer appears in the center. Four magnifying lenses surround it: novice, domain expert, rushed evaluator, and skeptical evaluator. Each lens points to a different judgment ranging from convincing to unsupported. The figure emphasizes that apparent intelligence is relational, not a property revealed independently of the judge.

Deception also complicates the goal. In Turing's setup, successful imitation involved making an identification difficult. In many business systems, pretending to be human is unnecessary or undesirable. A reliable assistant should identify itself appropriately, disclose uncertainty when useful, and make escalation easy. The best system may be the one that helps a person complete a task—not the one that most successfully hides its nature.

Hype Check

Humanlike is not the universal target

A calculator is valuable because it exceeds ordinary human arithmetic, not because it imitates our errors. A risk dashboard should communicate evidence, not stage a personality. Evaluate whether humanlike behavior improves the task and preserves user agency; do not treat imitation as an automatic product virtue.

3.5

Benchmarks are descendants of the same move

Modern AI evaluation rarely consists of one open conversation with a hidden machine. Researchers and companies use collections of problems, labeled datasets, simulated environments, human preference studies, red-team exercises, and task-specific acceptance tests. Yet the family resemblance remains: a broad claim is converted into observable performance under defined conditions.

A is a standardized set of tasks and scoring rules. It enables comparison and makes progress visible. But every benchmark draws a boundary around the world. A question-answering test may reward choosing the correct option without revealing whether the model used a sound method. A coding benchmark may use clean, self-contained problems unlike a large company's changing codebase. A safety test may miss a new attack pattern.

A small rectangular benchmark sample sits inside a much larger irregular real-world environment, with gaps labeled changing context, rare cases, human behavior, and consequences.
Figure 3.5 · Every benchmark is a sample of a larger reality. Standardization enables comparison; the gap between the test and the deployment environment determines how far the result can travel.
Audiobook description

A neat rectangle labeled benchmark contains a limited set of test cases. It sits inside a much larger uneven shape labeled deployment reality. Areas outside the rectangle are labeled rare cases, changing context, human behavior, and real consequences. An arrow from benchmark score to deployment decision crosses a gap marked validate locally.

Scores can also become less informative as developers repeatedly tune systems against a public test or as benchmark material appears in training data. The problem is similar to studying the answer key. Performance may rise without a matching increase in the ability we hoped to measure.

For a business leader, the remedy is not to reject benchmarks. It is to build an evidence ladder: start with published evaluations, then test representative internal cases, then run a controlled pilot, then monitor the production workflow. Confidence should increase only as evidence moves closer to the intended use.

Decision Lens

Read a score like a contract

  • What exact task and population does the score represent?
  • What baseline or alternative was used for comparison?
  • Which failure types disappear inside the average?
  • Could the test cases have influenced training or system tuning?
  • What local evidence is required before the score supports deployment?
3.6

What a convincing performance cannot prove

Suppose a machine sustains a conversation so convincing that expert judges cannot reliably distinguish it from a person. What follows?

We would have strong evidence of performance in that conversational setting. We might have evidence of broad linguistic skill, depending on the questions and duration. We would not automatically have evidence that every answer was true, that the system could act reliably in the physical world, that it used the same process as a human, or that it possessed subjective experience.

A solid circle labeled supported by conversational behavior is surrounded by separate dashed circles for factual truth, broad real-world competence, humanlike process, and consciousness, showing they require additional evidence.
Figure 3.6 · Respect the boundary of the evidence. A result supports the claim actually tested. Broader conclusions need broader and different evidence.
Audiobook description

A solid green center circle says “convincing conversation under test conditions.” Around it are four separate dashed circles: factual truth, real-world competence, humanlike internal process, and consciousness. No arrow automatically connects the center to the outer claims. Each outer circle is labeled additional evidence required.

This restraint is sometimes described as moving the goalposts. But refusing an unsupported conclusion is not the same as denying an achievement. A machine that produces useful language across many domains is economically and scientifically important even if we remain uncertain about the best theory of its cognition. Capability deserves to be measured accurately, not inflated into a claim the evidence cannot carry.

The same discipline protects people. If a screening model matches past hiring decisions, that does not prove those decisions were fair. If a medical assistant writes like a clinician, that does not prove its advice is safe. If an agent appears confident, that does not prove it has checked the latest policy. Behavioral resemblance is a starting point for evaluation, not a transfer of authority.

Where the analogy breaks: Human conversation is unusually rich evidence because people share biology, development, social life, and a world. A machine may generate similar language through a very different history and mechanism. The same output can therefore carry different implications about the speaker behind it.
Chapter close

Remember this

Turing changed the form of the question. He replaced a dispute over definitions with an observable game.

An operational test is a tool, not a complete theory. It produces evidence under specified conditions.

The evaluator and protocol matter. Knowledge, expectations, duration, channel, and scoring can change the judgment.

Benchmarks sample reality. Published scores should begin—not end—a business evaluation.

Evidence has a boundary. Conversational success does not automatically prove truth, general competence, humanlike process, or consciousness.

Five-rung evidence ladder rising from vendor claim to public benchmark, representative internal test, controlled pilot, and monitored production use.
Figure 3.7 · The evidence ladder for business adoption. Move from general claims toward evidence collected in the actual workflow, with explicit success criteria and continued monitoring.
Audiobook description

A ladder has five rising rungs. From bottom to top they read vendor claim, public benchmark, representative internal test, controlled pilot, and monitored production use. Confidence rises with relevance, while a side arrow says consequences and controls must be evaluated at every rung.

Five key terms

  • Imitation game
  • Operational definition
  • Behavioral test
  • Benchmark
  • Evaluation set

Reflection questions

  1. Which broad quality in your organization is measured through a convenient proxy, and where can the proxy mislead?
  2. Would an AI product you use become more or less valuable if it stopped trying to sound human?
  3. What is the next rung of evidence needed before a promising AI demonstration becomes a responsible deployment?

Register to read all chapters for free.

Use your name and email to unlock the complete and growing Academy library. No fee and no credit card required. You can close this window and keep reading this sample chapter.

By continuing, you agree to receive the account email needed to secure access. Already registered? Sign in.

More information