Sekėjai

Ieškoti šiame dienoraštyje

2025 m. gegužės 5 d., pirmadienis

A.I. Is Getting More Powerful, but Its Hallucinations Are Getting Worse


"Last month, an A.I. bot that handles tech support for Cursor, an up-and-coming tool for computer programmers, alerted several customers about a change in company policy. It said they were no longer allowed to use Cursor on more than just one computer.

In angry posts to internet message boards, the customers complained. Some canceled their Cursor accounts. And some got even angrier when they realized what had happened: The A.I. bot had announced a policy change that did not exist.

“We have no such policy. You’re of course free to use Cursor on multiple machines,” the company’s chief executive and co-founder, Michael Truell, wrote in a Reddit post. “Unfortunately, this is an incorrect response from a front-line A.I. support bot.”

More than two years after the arrival of ChatGPT, tech companies, office workers and everyday consumers are using A.I. bots for an increasingly wide array of tasks. But there is still no way of ensuring that these systems produce accurate information.

The newest and most powerful technologies — so-called reasoning systems from companies like OpenAI, Google and the Chinese start-up DeepSeek — are generating more errors, not fewer. As their math skills have notably improved, their handle on facts has gotten shakier. It is not entirely clear why.

Today’s A.I. bots are based on complex mathematical systems that learn their skills by analyzing enormous amounts of digital data. They do not — and cannot — decide what is true and what is false. Sometimes, they just make stuff up, a phenomenon some A.I. researchers call hallucinations. On one test, the hallucination rates of newer A.I. systems were as high as 79 percent.

These systems use mathematical probabilities to guess the best response, not a strict set of rules defined by human engineers. So they make a certain number of mistakes. “Despite our best efforts, they will always hallucinate,” said Amr Awadallah, the chief executive of Vectara, a start-up that builds A.I. tools for businesses, and a former Google executive. “That will never go away.”

For several years, this phenomenon has raised concerns about the reliability of these systems. Though they are useful in some situations — like writing term papers, summarizing office documents and generating computer code — their mistakes can cause problems.

The A.I. bots tied to search engines like Google and Bing sometimes generate search results that are laughably wrong. If you ask them for a good marathon on the West Coast, they might suggest a race in Philadelphia. If they tell you the number of households in Illinois, they might cite a source that does not include that information.

Those hallucinations may not be a big problem for many people, but it is a serious issue for anyone using the technology with court documents, medical information or sensitive business data.

“You spend a lot of time trying to figure out which responses are factual and which aren’t,” said Pratik Verma, co-founder and chief executive of Okahu, a company that helps businesses navigate the hallucination problem. “Not dealing with these errors properly basically eliminates the value of A.I. systems, which are supposed to automate tasks for you.”

Cursor and Mr. Truell did not respond to requests for comment.

For more than two years, companies like OpenAI and Google steadily improved their A.I. systems and reduced the frequency of these errors. But with the use of new reasoning systems, errors are rising. The latest OpenAI systems hallucinate at a higher rate than the company’s previous system, according to the company’s own tests.

The company found that o3 — its most powerful system — hallucinated 33 percent of the time when running its PersonQA benchmark test, which involves answering questions about public figures. That is more than twice the hallucination rate of OpenAI’s previous reasoning system, called o1. The new o4-mini hallucinated at an even higher rate: 48 percent.

When running another test called SimpleQA, which asks more general questions, the hallucination rates for o3 and o4-mini were 51 percent and 79 percent. The previous system, o1, hallucinated 44 percent of the time.

In a paper detailing the tests, OpenAI said more research was needed to understand the cause of these results. Because A.I. systems learn from more data than people can wrap their heads around, technologists struggle to determine why they behave in the ways they do.

“Hallucinations are not inherently more prevalent in reasoning models, though we are actively working to reduce the higher rates of hallucination we saw in o3 and o4-mini,” a company spokeswoman, Gaby Raila, said. “We’ll continue our research on hallucinations across all models to improve accuracy and reliability.”

Hannaneh Hajishirzi, a professor at the University of Washington and a researcher with the Allen Institute for Artificial Intelligence, is part of a team that recently devised a way of tracing a system’s behavior back to the individual pieces of data it was trained on. But because systems learn from so much data — and because they can generate almost anything — this new tool can’t explain everything. “We still don’t know how these models work exactly,” she said.

Tests by independent companies and researchers indicate that hallucination rates are also rising for reasoning models from companies such as Google and DeepSeek.

Since late 2023, Mr. Awadallah’s company, Vectara, has tracked how often chatbots veer from the truth. The company asks these systems to perform a straightforward task that is readily verified: Summarize specific news articles. Even then, chatbots persistently invent information.

Vectara’s original research estimated that in this situation chatbots made up information at least 3 percent of the time and sometimes as much as 27 percent.

In the year and a half since, companies such as OpenAI and Google pushed those numbers down into the 1 or 2 percent range. Others, such as the San Francisco start-up Anthropic, hovered around 4 percent. But hallucination rates on this test have risen with reasoning systems.

DeepSeek’s reasoning system, R1, hallucinated 14.3 percent of the time. OpenAI’s o3 climbed to 6.8.

(The New York Times has sued OpenAI and its partner, Microsoft, accusing them of copyright infringement regarding news content related to A.I. systems. OpenAI and Microsoft have denied those claims.)

For years, companies like OpenAI relied on a simple concept: The more internet data they fed into their A.I. systems, the better those systems would perform. But they used up just about all the English text on the internet, which meant they needed a new way of improving their chatbots.

So these companies are leaning more heavily on a technique that scientists call reinforcement learning. With this process, a system can learn behavior through trial and error. It is working well in certain areas, like math and computer programming. But it is falling short in other areas.

“The way these systems are trained, they will start focusing on one task — and start forgetting about others,” said Laura Perez-Beltrachini, a researcher at the University of Edinburgh who is among a team closely examining the hallucination problem.

Another issue is that reasoning models are designed to spend time “thinking” through complex problems before settling on an answer. As they try to tackle a problem step by step, they run the risk of hallucinating at each step. The errors can compound as they spend more time thinking.

The latest bots reveal each step to users, which means the users may see each error, too. Researchers have also found that in many cases, the steps displayed by a bot are unrelated to the answer it eventually delivers.

“What the system says it is thinking is not necessarily what it is thinking,” said Aryo Pradipta Gema, an A.I. researcher at the University of Edinburgh and a fellow at Anthropic.” [1]

1.  A.I. Is Getting More Powerful, but Its Hallucinations Are Getting Worse. Metz, Cade; Weise, Karen.  New York Times (Online) New York Times Company. May 5, 2025.

Dirbtinio intelekto agentai mokosi bendradarbiauti --- Atsižvelgiant į inovacijų tempą, įmonės turėtų ruoštis pokyčiams


„Įmonės turėtų pradėti planuoti kitą dirbtinio intelekto (AI) etapą: kelių agentų koordinavimą visoje savo įmonėje.

 

Dauguma įmonių vis dar aiškinasi, kaip įdiegti bent vieną dirbtinio intelekto valdomą agentą, kuris galėtų atlikti užduotį savarankiškai arba koordinuodamas darbą su žmonėmis.

 

Tačiau kūrėjai kuria protokolus, kaip šiuos agentus įtraukti į komandas, kurios tvarko viską – nuo ​​klientų aptarnavimo ir kodavimo iki tiekimo grandinės, logistikos, finansų, rinkodaros ir verslo strategijos.

 

Atsižvelgiant į inovacijų tempą ir laiką, kurio reikia organizacijoms prisitaikyti, įmonės padarys sau paslaugą, jau dabar ruošdamosi daugiaagentėms sistemoms, kurios vis labiau bus prieinamos vėliau šiais metais.

 

„Accenture“ vyriausioji dirbtinio intelekto pareigūnė Lan Guan teigia, kad šiuo metu tik 10–15 % jos klientų naudoja daugiaagentines sistemas, tačiau ji tikisi, kad per 18–24 mėnesius šis procentas viršys 30 %.

 

Pavyzdžiui, profesionalių paslaugų įmonė sukūrė 15 agentų sistemą, naudojamą rinkodarai, kurią sudaro trys „superagentai“, atsakingi už 12 agentų koordinavimą“, agentų, apmokytų atlikti konkrečias užduotis.

 

Pasak Guano, ji gali planuoti rinkodaros kampaniją tokia tema, kaip „2025 m. tendencijos“, atlikti tyrimus, nustatyti panašias ankstesnes kampanijas ir atsakyti į klausimus, kaip žmogus.

 

Iš viso „Accenture“ šiandien turi daugiau, nei 50, daugiaagentinių sistemų įvairioms pramonės šakoms ir rinkoms ir tikisi, kad iki metų pabaigos šis skaičius pasieks daugiau, nei 100.

 

Įmonė teigė, kad tokie klientai, kaip automobilių gamintoja BMW, vartojimo prekių ženklų bendrovė „Unilever“ ir sporto milžinė ESPN šiuo metu diegia šias sistemas.

 

Praėjusį mėnesį „Accenture“ pristatė „Trusted Agent Huddle“, kuri, anot jos, leidžia agentams sąveikauti su tokiais partneriais, kaip technologijų bendrovės „Amazon Web Services“, „Google Cloud“, „Meta“, „Microsoft“, „Nvidia“, „Oracle“, „Salesforce“, SAP ir „ServiceNow“.

 

Daugiaagentinės galimybės netrukus taps plačiau prieinamos. „Salesforce“ ir „Google“ balandžio mėnesį vykusioje „Google Next“ konferencijoje paskelbė, kad dirba su protokolu, vadinamu A2A arba agentas-agentui. Protokolas, leidžiantis „Salesforce“ „Agentforce“ ekosistemos agentams sąveikauti tarpusavyje ir su išoriniais agentais, daugiausia dėmesio skiria tokioms sritims, kaip autentifikavimas, identifikavimas ir pranešimų perdavimas, teigia Gary Lerhauptas, „Agentforce“ produktų architektūros viceprezidentas. Su partneriais vyksta darbas, kuriant daugiaagentių prototipus, naudojant A2A, sakė jis.

 

„Keyway“, Niujorke įsikūrusi komercinio nekilnojamojo turto technologijų startuolis, teigia vienas iš įkūrėjų ir generalinis direktorius Matias Recchia, leidžia pažvelgti į tai, kaip ši koncepcija veikia praktiškai.

 

Ji siūlo turto valdytojams ir nekilnojamojo turto valdytojams daugiaagentę platformą, kuri naudoja koordinuotą sąveiką, kad išspręstų tokius klausimus, kaip nuomojamo turto kainos nustatymas arba patogumų ir paskatų taikymas.

 

Bendrovė pritraukė 45 mln. dolerių iš investuotojų, įskaitant „Canvas Ventures“, „Camber Creek“ ir „Thomvest“.

 

Nors „Keyway“ agentai yra specializuoti ir gali aktyvuoti vienas kitą per struktūrizuotą darbo eigą, bendrovė teigė, kad jie vis tiek veikia kontroliuojamai, iš anksto nustatyta seka, kuriai reikalinga žmogaus priežiūra, pavyzdžiui, raginimų nustatymas, rezultatų peržiūra ir sprendimų priežiūra.

 

Pasak Recchia, tikroje daugiaagentėje sistemoje agentai dinamiškai samprotauja, derasi arba bendradarbiauja realiuoju laiku, nereikalaudami žmogaus apibrėžtų darbo eigų, aiškių raginimų ar rankinio koordinavimo. Kitaip tariant, agentai imasi iniciatyvos, prisitaiko prie naujos informacijos ir sklandžiai sąveikauja su kitais agentais, nelaukdami žmogaus nurodymų.

 

Įmonės gali pradėti ruoštis daugiaagentėms sistemoms, kurdamos standartinius, savarankiškus agentus. Kai tinkami protokolai bus parengti, įmonės galės suorganizuoti šiuos agentus, kad jie spręstų sudėtingas, bendradarbiaujančias, sistemas.

 

„Principal Financial Group“ įdiegė atskirus dirbtinio intelekto agentus įvairiose srityse, įskaitant programinės įrangos inžinerijos kopilotus, pretenzijų santraukas ir analizę po skambučių, teigia vyriausioji informacijos pareigūnė Kathy Kay. Jie daugiausia veikia apibrėžtose srityse, tačiau investicijų valdymo ir draudimo bendrovė aktyviai kuria techninį pagrindą agentų bendradarbiavimui palaikyti, sakė Kay.

 

Tai reiškia duomenų srautų ir valdymo modelių kūrimą. Darbo eigos taip pat turės vystytis, kad būtų galima pritaikyti realaus laiko bendradarbiavimą tarp žmonių ir intelektualių, prisitaikančių, dirbtinio intelekto sistemų, sakė ji.

 

Kay mato didelį daugiaagentinių sistemų naudojimo potencialą pensijų paslaugose, tokiose, kaip pervedimų optimizavimas. Turto valdymo srityje ji tikisi, kad daugiaagentai analizuos nestruktūrizuotus rinkos duomenis, generuos investicinius naratyvus ir derins išvadas skirtinguose portfeliuose.“ [1]

 

 

1.  Business News: AI Agents Learn to Collaborate --- Given the pace of innovation, companies should prepare for change. Rosenbush, Steven.  Wall Street Journal, Eastern edition; New York, N.Y.. 05 May 2025: B3.

AI Agents Learn to Collaborate --- Given the pace of innovation, companies should prepare for change


"Companies should start planning for the next stage of artificial intelligence: the orchestration of multiple agents across their businesses.

Most companies are still figuring out how to deploy even one AI-powered agent that can perform a task autonomously or in coordination with humans.

But developers are creating protocols to harness these agents into teams that handle everything from customer service and coding to supply chain, logistics, finance, marketing and business strategy.

Given the pace of innovation and the time it takes for organizations to adapt, companies will do themselves a favor by getting ready now for multiagent systems increasingly available later this year.

Accenture's chief AI officer, Lan Guan, says only 10% to 15% of her clients currently use multiagent systems, but she expects that percentage to exceed 30% within 18 to 24 months.

The professional-services company has created a 15-agent system used for marketing, for example, comprising three "super agents" that are responsible for coordinating 12 agents trained for specific tasks.

It can plan a marketing campaign around a topic such as "2025 trends," conducting research, identifying similar past campaigns and addressing questions like a human would, according to Guan.

All told, Accenture has more than 50 multiagent systems today for a range of industries and markets and expects that number to hit more than 100 by the end of the year.

The firm said customers such as carmaker BMW, consumer-brands company Unilever and sports giant ESPN are currently adopting these systems.

Accenture last month introduced Trusted Agent Huddle, which it said allows agent-to-agent interoperability with partners such as technology companies Amazon Web Services, Google Cloud, Meta, Microsoft, Nvidia, Oracle, Salesforce, SAP and ServiceNow.

Multiagent capabilities are about to become more widely available. Salesforce and Google announced at the Google Next conference in April that they were working on a protocol called A2A, or Agent-to-Agent. The protocol, which allows agents within Salesforce's Agentforce ecosystem to interact with each other as well as external agents, focuses on areas such as authentication, identification and message passing, according to Gary Lerhaupt, vice president of product architecture for Agentforce. Work is under way with partners to develop prototype multiagents using A2A, he said.

Keyway, a commercial real-estate tech startup based in New York, provides a glimpse into how the concept works in practice, according to co-founder and Chief Executive Matias Recchia. It offers asset managers and property managers a multiagent platform that uses coordinated interactions to address questions such as how to price a rental property or target amenities and incentives.

The company has raised $45 million from investors including Canvas Ventures, Camber Creek and Thomvest.

While Keyway's agents are specialized and can trigger each other through structured workflows, the company said, they still operate in a controlled, predefined sequence that requires human oversight such as setting prompts, reviewing outputs and supervising decisions.

A true multiagent system, Recchia said, involves agents that dynamically reason, negotiate or collaborate in real time without requiring human-defined workflows, explicit prompts or manual coordination. In other words, the agents take initiative, adapt to new information and interact fluidly with other agents without waiting for human instruction.

Companies can start to prepare for multiagents systems by creating standard, stand-alone agents. Once the proper protocols are ready, companies can orchestrate these agents into tackling complex, collaborative systems.

Principal Financial Group has embedded individual AI agents across domains including software engineering co-pilots, claims summarization and postcall analytics, according to Chief Information Officer Kathy Kay. They largely operate within defined scopes, but the investment management and insurance company is actively building the technical foundation to support agent-to-agent collaboration, Kay said.

That means developing data pipelines and governance models. Workflows will also have to evolve to accommodate real-time collaboration between humans and intelligent, adaptive AI systems, she said.

Kay sees strong potential for the use of multiagent systems in retirement services such as rollover optimization. In asset management, she expects multiagents to analyze unstructured market data, generate investment narratives and align findings across portfolios.” [1]

1.  Business News: AI Agents Learn to Collaborate --- Given the pace of innovation, companies should prepare for change. Rosenbush, Steven.  Wall Street Journal, Eastern edition; New York, N.Y.. 05 May 2025: B3.