Can the Turing Test Really Evaluate Artificial Intelligence?

In this blog post, we’ll examine whether the Turing Test is a suitable standard for evaluating the intelligence of artificial intelligence, and explore its limitations and potential areas for improvement.

 

In 2014, news that “Eugene Goostman,” a chatbot created by Russian developers, had passed the Turing Test came as a shock and caused confusion for many people. This was a clear demonstration that artificial intelligence is no longer merely a product of an imaginary future society seen only in science fiction movies. AI has already become deeply embedded in our lives and is advancing at a rapid pace, in line with Ray Kurzweil’s “Law of Accelerating Returns.” However, we cannot simply stand by and watch as AI, which is evolving at the blink of an eye, becomes firmly established in society. We need to determine whether AI can indeed think as humans do, and we must also consider the direction in which it should develop in the future.
These issues are linked to the Turing Test. The Turing Test is an evaluation method that originated from the question: “Is it possible to create a computer that we can say possesses intelligence and a mind?” It is commonly known that Alan Turing proposed it in 1950 to answer the question, “Can machines think?” However, Turing actually rephrased the question itself as “Can machines behave in the same way humans do?” when designing the experiment. The test is conducted by having a judge pose questions to both a human and an AI, then determining whose responses are more human-like. In the Turing Test, various methods can be used to deceive the judge, and it is permissible to slow down response times or continue the conversation in a chat-like format. Furthermore, the test is considered passed if the AI can deceive approximately 30% of the judges within the limited scope and topic of the discussion.
I questioned whether a Turing Test with these conditions and rules is truly an appropriate measure for assessing AI intelligence. In other words, I believe the Turing Test is not suitable for evaluating AI. To support this view, I will present three arguments and propose the “Reverse Turing Test” as a solution. Before doing so, however, it is necessary to first examine other notable examples of AI that have passed the Turing Test besides “Eugene Goostman.”
“ELIZA,” developed by Joseph Weizenbaum in 1966, is the most widely known example among early natural language dialogue programs. This AI was designed to identify key keywords in input sentences and return corresponding responses; if no appropriate keywords were found, it would repeat the user’s words or provide a generic reply.
There is also “PARRY,” developed by Kenneth Colby in 1972. This AI was modeled after the conversational style of a paranoid patient, and experiments were even conducted in which several psychiatrists attempted to distinguish between actual patients and PARRY.
From here on, I will examine why the Turing Test is not a suitable standard for evaluating AI.
First, the Turing Test does not accurately assess intelligence itself. In other words, rather than evaluating how intelligently an AI thinks, this test determines through conversation whether it behaves like a human. There are two problems with this.
First, not all humans always behave intelligently. Here, “intelligent” refers to intellectual activity that involves understanding a phenomenon or object and processing it rationally. Furthermore, according to Alan Turing’s paper, some intellectual abilities of AI may manifest in ways different from those of humans. Simply put, if an AI solves a problem that humans cannot, the evaluator is actually more likely to realize that the other party is a computer. Ultimately, the Turing Test fails to properly evaluate forms of intelligence that surpass human capabilities.
Second, the evaluation criteria of the Turing Test are subjective, and the experimental conditions are overly restrictive. In other words, the test results can be significantly influenced by the evaluator’s attitude, experience, skills, and knowledge rather than the AI’s actual level of intelligence, making it difficult to ensure objectivity.

Ultimately, the entity determining whether something is an AI or a human is a person, so it is difficult to completely eliminate the evaluator’s subjectivity. Furthermore, the questions asked within the limited timeframe of about five minutes are also very restricted. As long as the evaluator is human, there is always the possibility of being influenced by various external factors.
Additionally, most AIs participating in the Turing Test are modeled after humans with specific situations or characteristics. Therefore, even if an AI passes the test, it is difficult to conclude that it behaves in the same way as humans in general. For example, “Eugene Goostman,” introduced earlier, was developed with the premise that it was a 13-year-old boy living in Ukraine. This was a deliberate attempt to create a situation where even somewhat clumsy English would be accepted as natural. “PARRY” was modeled after a paranoid patient, and “ELIZA” was designed to mimic a professional psychological counselor.
In this way, when an AI is set up to represent a human with specific characteristics and then engaged in conversation, the judges evaluate it with those characteristics in mind. For example, even if a somewhat strange response comes up while conversing with an AI programmed to represent a person with a mental illness, the judges are likely to accept it as a natural reaction consistent with that characterization. In other words, while judges are supposed to evaluate without bias, they are actually placed in an environment where doing so is difficult.
Of course, people with mental illnesses are capable of thinking. However, just because an AI modeled after a person with a mental illness passes the Turing Test does not mean it thinks more “human-like” than a human. To remain faithful to the question posed by Alan Turing—“Can a machine behave in the way a human does?”—the test must be designed based on general situations that the majority of people can accept.
Third, AI can pass the Turing Test simply by manipulating symbols rather than actually understanding the content. A prime example is “ELIZA.” For a person to think about a topic, they must understand that topic and make judgments based on relevant experience. For this reason, philosopher John Searle also argued that the Turing Test alone cannot prove whether a machine is actually thinking.
Furthermore, since artificial intelligence operates according to predetermined algorithms, it can easily reveal errors when faced with unexpected situations. The following is an example illustrating this.
Interlocutor: Do you still support Bill Clinton?
Judge: I’ve never met Clinton.
Interlocutor: You really dislike Bill Clinton. At least, it’s clear that you dislike him. In my opinion, it seems like most people have simply been uninterested or haven’t given it much thought so far.
Judge: Have you seen ‘The Martian’?
Interlocutor: Let’s take it easy. I’m not here to answer stupid questions. That’s not it. What were we talking about? Right. Does what you just said mean that the underlying issues—like Whitewater and the special counsel—are serious?
This conversation is between a judge and “Catherine,” an AI designed to continue a dialogue based on pre-programmed information about events that had become news the day before the Turing Test was held. If the judge had continued the political conversation related to Bill Clinton, the dialogue could have remained quite natural. However, when the judge suddenly shifted the topic to something else, the AI was unable to properly handle the unprogrammed subject and revealed the error shown above.
Thus, we cannot conclude that an AI behaves like a human simply because it is capable of a certain level of communication. While it may appear to engage in human-like conversation on the surface, it is not actually understanding the content but merely responding in a way that makes it seem human. Therefore, it falls short of the concept of “intelligence” as defined earlier. The current Turing Test, which passes AI systems that merely mimic human behavior without actually understanding the meaning, has clear limitations.
In that it evaluates AI by comparing it to humans, the Turing Test is a meaningful experiment that explores the boundary between humans and AI. Considering that humans themselves cannot clearly define the essence of thought or the mind, this attempt in itself holds significant value. On the other hand, since AI operates based on relatively simple structures and principles, it is possible to measure and analyze it to a certain extent, even if not perfectly.
However, the Turing Test fails to accurately assess intelligence, and its evaluation criteria are also overly subjective. A closer look at the actual testing process reveals that it is difficult to view it as an environment capable of comprehensively evaluating an AI’s intelligence. Even if multiple judges participate and an institution oversees the evaluation process, the number of evaluators is limited, and it is not easy to reach an objective conclusion that takes all variables into account. Therefore, there are fundamental limitations to deriving a perfect result that everyone can accept.
The Turing Test holds significant importance in that it demonstrated that AI can, to a certain extent, replicate the cognitive abilities once considered unique to humans. Therefore, rather than completely dismissing this test, there is a need to establish evaluation criteria and procedures that are more accurate and acceptable to a wider audience than those currently in use.
In this regard, I would like to propose the “Reverse Turing Test” as an alternative. In the Reverse Turing Test, the evaluator is a computer rather than a human. In other words, it is a method in which a computer evaluates a human or another computer while interacting with it. Introducing this approach into existing tests would reduce the problem of human subjectivity and allow evaluations to be conducted in a more objective environment. Of course, the fact that the evaluator is an AI also raises the possibility of new issues.
Regardless of the subject, tests designed for evaluation and judgment are unlikely to satisfy everyone, and it is not easy to establish perfect criteria that everyone can accept. Therefore, we must continually explore ways to improve them. In particular, since the Turing Test is a crucial evaluation method that could significantly influence the future of industry and the direction of AI technology development, it must evolve into a more efficient and reliable form.
Today, AI is advancing at a pace incomparable to the past and is already being utilized in every aspect of our daily lives. When evaluating such AI, we should not judge it solely on whether it can deceive humans, but rather assess its ability to understand meaning, grasp context, and communicate naturally across a variety of situations. Of course, since this is not a field with clearly defined “correct answers” like a math problem, establishing objective criteria is not easy. Nevertheless, to evaluate AI that will coexist with humans in the future, the Turing Test must evolve to comprehensively consider ethical factors and the ability to communicate effectively.

 

About the author