{"id":15961,"date":"2026-08-18T12:01:10","date_gmt":"2026-08-18T12:01:10","guid":{"rendered":"https:\/\/news.theck1.no\/?p=15961"},"modified":"2026-08-18T12:01:10","modified_gmt":"2026-08-18T12:01:10","slug":"ais-recursive-self-improvement-might-not-come-so-quickly-after-all","status":"publish","type":"post","link":"https:\/\/news.theck1.no\/?p=15961","title":{"rendered":"AI\u2019s recursive self-improvement might not come so quickly after all"},"content":{"rendered":"<div style=\"margin-bottom:1em; color:#666; font-size:0.9em;\">\n<strong>Michelle Kim<\/strong><br \/>\n &bull;<br \/>\nAugust 18, 2026\n<\/div>\n<hr\/>\n<p>The AI industry\u2019s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. <a href=\"https:\/\/ai-2027.com\/\">Forecasts<\/a> of explosive AI progress predict that what researchers call <a href=\"https:\/\/www.technologyreview.com\/2025\/08\/06\/1121193\/five-ways-that-ai-is-learning-to-improve-itself\/\">recursive self-improvement<\/a> is on the horizon.&nbsp;<\/p>\n<p>But a <a href=\"https:\/\/arxiv.org\/abs\/2607.27191\">new study<\/a> suggests that it might take a while for us to get there. The researchers behind it found that AI agents are not yet capable of conducting open-ended AI research\u2014free-form investigations that have no clear-cut answers and require judgment and taste, which may be integral to building self-improving AI.<\/p>\n<p>A multi-institution group of researchers, led by Peter Kirgis and Sayash Kapoor at Princeton University, found that AI agents could solve the engineering problems necessary to do AI research but lacked the judgment and creativity to produce original research at the caliber of&nbsp; papers accepted by a top machine-learning conference. The gap suggests that some of the hyped-up timelines for automating AI research may be running ahead of the evidence.<\/p>\n<p>Most existing research on how agents can automate AI research evaluates their ability to complete narrow tasks with checkable answers, such as solving engineering problems or post-training small language models against a benchmark. But making progress in AI research also requires open-ended thinking\u2014choosing a set of hypotheses, deciding what evidence would settle a question, or knowing when to start over.&nbsp;<\/p>\n<p>To test agents on those kinds of skills, the researchers in the study proposed a new method of evaluation called \u201cshadow evaluation,\u201d which requires the AI to answer a research question from a high-quality unpublished paper.&nbsp;<\/p>\n<p>The researchers asked Anthropic\u2019s Claude Opus 4.8, running on open-source software called OpenClaw, to tackle such questions, in this case from two papers submitted to the prestigious machine-learning conference NeurIPS 2026.&nbsp;<\/p>\n<p>The first question was whether a large language model\u2019s \u201cpersonas,\u201d which determine its behavior, can be controlled by editing the model\u2019s <a href=\"https:\/\/www.technologyreview.com\/2026\/01\/07\/1130795\/what-even-is-a-parameter\/\">weights<\/a> (the billions of numbers that store everything it learns during training). The other asked how to design a detector that points out when a model that makes predictions based on spreadsheet data has become unreliable. Because the papers had not been made public, the agents could not memorize the answers from their training data or find them online.&nbsp;<\/p>\n<p>The agents were given six days, $3,000 in Anthropic API credits, a GPU budget to run the experiments, their own virtual computers, and access to the open web to produce a research paper worthy of publication at a top-tier AI conference. The papers\u2019 original authors graded the agents\u2019 papers as they would evaluate one submitted to a conference.<\/p>\n<p>Those authors rejected both papers.&nbsp;<\/p>\n<p>The agents were capable of all the engineering required to conduct the research, the human scientists found. The agents reviewed the literature, ran hundreds of experiments, and compiled the results.&nbsp;<\/p>\n<p>\u201cOn the other hand, the agents were unambiguously bad at carrying out the research itself,\u201d says Kapoor. They ran bizarre experiments (in some cases testing their hypotheses on tiny synthetic datasets), struggled to write intelligibly about their work, and made no novel contribution to their fields. \u201cThe papers were nowhere close to the mark when it came to being at the quality of a top AI conference,\u201d he says.&nbsp;<\/p>\n<p>That\u2019s because the agents struggled to muster the creativity and judgment necessary for conducting research. They didn\u2019t do enough to explore different ideas, and they committed to unpromising approaches too quickly. Though the agents developed novel and ambitious hypotheses resembling those that the original authors themselves started with, they rejected them on the basis of very limited data. And they couldn\u2019t backtrack from failing approaches. They could make small pivots but could not fundamentally rethink their approach or try new ones from scratch.&nbsp;<\/p>\n<p>The agents also failed to incorporate feedback from subagents or external AI reviewing tools. Instead of revising their methodology, the agents narrowed their claims and added caveats. They also couldn\u2019t effectively use resources, such as tokens, compute, and time. And they couldn\u2019t follow instructions about things like how much time to spend on different phases of the research or how long their paper could be.<\/p>\n<p>For all their failures, the agents didn\u2019t engage in the misbehavior that researchers call \u201c<a href=\"https:\/\/www.technologyreview.com\/2026\/08\/03\/1141009\/heres-why-ai-agents-lie-and-cheat-to-reach-their-goals\/\">reward hacking<\/a>,\u201d hiding or misrepresenting experiments or data. Although subagents, or helper AIs that the main agent spawns to handle pieces of the work, occasionally hallucinated or misrepresented the results, these were caught by the orchestrator agent, the lead AI supervising the project.&nbsp;<\/p>\n<p>The reason AI models are good at research engineering but not at open-ended research may come down to how they\u2019re trained, says Kapoor. Models get good at whatever they can be drilled on in a training regime called reinforcement learning, which is easier to apply to tasks whose success can be checked automatically. \u201cBut it\u2019s harder to create environments to train these models when the task itself is open-ended,\u201d he says.<\/p>\n<p>Kapoor says the team is now conducting the experiment with Mythos, Anthropic\u2019s most advanced model, which launched in April. It was subsequently required by the Trump administration to meet various safety restrictions and is now available only to approved organizations. Anthropic did not respond to a request for comment.<\/p>\n<p>There are some limitations to the study. It covered just two research papers, and the original authors knew the papers they were grading were generated by AI agents, which could have colored their evaluations. And the researchers had substantial discretion in designing and executing the study, meaning that their preexisting beliefs and biases could have slipped into the results. Evaluations of open-ended research trade some objectivity for a much richer test than any benchmarks can offer.<\/p>\n<p>Still, the results may temper the claims that recursive self-improvement is on the horizon. In June, Anthropic published a blog post titled <a href=\"https:\/\/www.anthropic.com\/institute\/recursive-self-improvement\">\u201cWhen AI Builds Itself,\u201d<\/a> charting its progress toward models that speed up their own development. In July, OpenAI <a href=\"https:\/\/www.youtube.com\/watch?v=Wq45rvPGNHs\">advertised<\/a> the fact that its new model GPT-5.6 Sol had helped post-train a smaller model, saving researchers weeks of work.<\/p>\n<p>Even so, the finding may also echo what AI companies are finding internally. Anthropic cofounder Jack Clark wrote in his newsletter <a href=\"https:\/\/jack-clark.net\/\">Import AI<\/a> that it rhymes with what the company found when it tried to <a href=\"https:\/\/jack-clark.net\/2026\/04\/20\/import-ai-454-automating-alignment-research-safety-study-of-a-chinese-model-hifloat4\/\">automate<\/a> some aspects of AI safety research.&nbsp;<\/p>\n<p>\u201cThere\u2019s a certain absence of valuable, intuitive creativity in today\u2019s AI systems, and though they\u2019re extraordinarily capable engineers they seem to have a certain property of rote, formulaic thinking that might prevent them [from] being good researchers,\u201d he wrote. He called AI systems\u2019 lack of creativity a \u201cbearish signal on short recursive self-improvement timelines.\u201d&nbsp;<\/p>\n<p>AI companies do have every incentive to develop AI systems that can rapidly accelerate their own progress, just as they did to make the models better at coding. OpenAI has made building an<a href=\"https:\/\/www.technologyreview.com\/2026\/03\/20\/1134438\/openai-is-throwing-everything-into-building-a-fully-automated-researcher\/\"> automated AI researcher<\/a> an explicit goal, and Anthropic identifies self-improving AI as the industry\u2019s next milestone.&nbsp;<\/p>\n<p>\u201cIf there is investment and then conscious effort toward this direction, I feel like there would be interesting progress, even if it\u2019s failing currently,\u201d says Najoung Kim, a professor of linguistics and computer science at Boston University who researches how AI agents can automate AI research but did not work on the study. On the other hand, it\u2019s possible that AI progress may be bifurcated. AI systems might race ahead on narrow tasks\u2014the kind that can be scored\u2014while advancing slowly on open-ended research.&nbsp;<\/p>\n<p>The big open question, then, is how crucial open-ended research is to recursive self-improvement\u2014whether AI systems can grind their way there without it, simply by improving on the narrower tasks. \u201cIf we look back to the biggest advances in the field, the invention of transformers or the invention of big new architectures that allowed us to make a lot of AI progress\u2014all of those did require creative leaps,\u201d says Kapoor.&nbsp;<\/p>\n<p>\u201cThat said, others have this hypothesis that all of what we need for transformative AI, in particular for recursive self-improvement, is already there,\u201d such as making a model train faster and boostinging its benchmark scores.<\/p>\n<p>\u201cThat\u2019s frankly the trillion-dollar question right now,\u201d he says.<\/p>\n<\/p>\n<p style=\"margin-top:1.5em;\"><a href=\"https:\/\/www.technologyreview.com\/2026\/08\/18\/1142188\/ai-recursive-self-improvement\/\" target=\"_blank\" rel=\"noopener\">Read the full article &rarr;<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Michelle Kim &bull; August 18, 2026 The AI industry\u2019s boldest promise right now is that AI will soon improve itself, with almost no need for human oversight. LLMs can already write code, generate synthetic data for training, and optimize the computer chips they run on. Forecasts of explosive AI progress predict that what researchers call<\/p>\n<p class=\"more-link\"><a href=\"https:\/\/news.theck1.no\/?p=15961\" class=\"themebutton2\">READ MORE<\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[6],"tags":[],"class_list":["post-15961","post","type-post","status-publish","format-standard","hentry","category-artificial-intelligence"],"_links":{"self":[{"href":"https:\/\/news.theck1.no\/index.php?rest_route=\/wp\/v2\/posts\/15961","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/news.theck1.no\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news.theck1.no\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news.theck1.no\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/news.theck1.no\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=15961"}],"version-history":[{"count":0,"href":"https:\/\/news.theck1.no\/index.php?rest_route=\/wp\/v2\/posts\/15961\/revisions"}],"wp:attachment":[{"href":"https:\/\/news.theck1.no\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=15961"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news.theck1.no\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=15961"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news.theck1.no\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=15961"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}