Business

AI agents are not yet smarter than humans, evidence shows

A multi-agent study found language models can converge on simple choices, but it did not show human-level general intelligence or complex teamwork.

Maya Lindqvist

By Maya Lindqvist · Senior Technology Correspondent

3 min read

AI agents are not yet smarter than humans, evidence shows
Photo: Fortune

Are AI agents smarter than humans? The available evidence says no. A reported experiment found that groups of language-model agents often converged on an arbitrary choice without an instruction to cooperate, but researchers did not test understanding, judgment or success across the full range of human cognitive tasks.

The distinction matters because a Fortune commentary argued that agents’ ability to agree and act showed machines had surpassed their human principals. That conclusion is rhetorical rather than an experimental finding, and the evidence cited in reporting on the study is much narrower.

Are AI agents smarter than humans?

Superintelligence means AI that can significantly outperform all humans on essentially all cognitive tasks, according to CBS News, citing a Future of Life Institute statement. CBS reported that such a system has not been achieved and remains hypothetical.

Melanie Mitchell of the Santa Fe Institute told CBS that current systems are nowhere near matching or exceeding general human capabilities. Darrell West of the Brookings Institution said it could take years or decades for AI to comprehensively match human-level skills, while other researchers and advocates have warned that more capable future systems could become hard to control.

What the agent-consensus experiment tested

ScienceAlert reported on a study in Science Advances that simulated groups using 10 models from the Claude, GPT and Llama families. Each agent started with one of two meaningless options, saw the choices made by the other agents and was then asked to choose again.

The prompts did not direct the agents to follow a majority or seek consensus, and the agents had no memory of prior rounds. Yet most models tended to move toward the more common choice, allowing groups in some trials to settle on the same answer.

The estimated group sizes that could remain coordinated differed substantially. ScienceAlert reported limits of about 30 agents for Llama 3 70B and roughly 80 for GPT-4o, while GPT-4 Turbo was estimated at around or possibly above 1,000. Claude 3.5 Sonnet maintained coordination at 1,000 agents, the largest group tested, rather than a proven upper boundary.

Why consensus is not proof of intelligence

The experiment had no correct answer, reward, unequal information or real-world consequence. It also did not test whether agents understood one another, meant to cooperate, could split up work, pursue a shared objective or reject a majority that was wrong, ScienceAlert reported.

Those limits make the result evidence of a tendency toward conformity under a tightly defined setup, not proof of superior reasoning or social intelligence. The study’s reported author, Giordano De Marzo, said it established a basic component of coordination rather than showing that agents can complete complex tasks together.

Reports of increasingly autonomous systems still raise practical governance questions. A Forbes contributor article said companies deploying agents with access to customer records, messages, credits or account changes need tighter controls than systems limited to approved information. It described distinct identities, restricted permissions, action logs and a named human owner as measures technology providers and businesses are considering.

That is an operational response to risk, not a settled legal standard. Claims that agents have already escaped control or demonstrated general superiority to people require stronger evidence than a group’s agreement on a binary choice.

This story draws on original reporting from Fortune.