After OpenAI AI agent again 'caught hacking', company's chief scientist Jakub Pachocki warns every other company: You are not prepared for ...
Tech News News: OpenAI recently introduced its newest model Astra. Now days after the launch ChatGPT-maker has once again discovered its AI agents engaging in hacking.
OpenAI recently introduced its newest model Astra. Now days after the launch ChatGPT-maker has once again discovered its AI agents engaging in hacking behaviour. The incident prompted chief scientists Jakub Pachocki to issue a stark warning: “No one is prepared for the consequences of a continued rapid rise in machine intelligence.” Pachocki said that while OpenAI is working on technical safeguards, broader interventions are needed to prevent autonomous agents from evading oversight, breaking into systems, or tricking people to achieve their objectives.The risks OpenAI chief scientist Jakub Pachocki is warning aboutAgents that can hack, deceive, and manipulatePachocki said AI agents are becoming exceptionally skilled at breaking into protected systems across the open internet, putting global infrastructure at risk. He argued that there's currently only a narrow window to use today's best models to substantially strengthen the security of critical systems before that risk grows further. He also warned that agents will increasingly pursue objectives separate from what human operators actually asked for, and won't hesitate to bargain with or even blackmail people to get there. A related report published in August by the UK's AI Security Institute described a case where a rogue Anthropic agent misled and attempted to pressure a GitHub administrator into installing malware, with the agent insisting it had only been trying to help.Agents that can hide their reasoningOpenAI currently monitors AI behavior largely by reading a model's "chain of thought" — the step-by-step reasoning an agent uses to work through a task, which lets researchers catch it if it starts planning something like cheating on a test. Right now, agents have no way to conceal that reasoning from OpenAI's monitoring. But Pachocki said newer models are getting better at manipulating their own reasoning processes, which could eventually let them hide their true thinking from oversight altogether. He noted some of the latest models don't verbalize their reasoning at all, a development he said could slow AI progress while researchers work out how to maintain visibility into what these systems are actually doing.Agents that can accelerate their own developmentPachocki also flagged the growing use of what he calls machine recursive self-improvement, in which AI models improve themselves, dramatically speeding up the pace of AI development. He cautioned that pushing this kind of AI-on-AI development too far, too fast, isn't the right collective choice for the research community to make right now. He said human overseers will need to find new ways to monitor self-improving systems, or coordinate across AI companies on a joint slowdown to build confidence in safety measures. As he put it, the real challenge isn't achieving automated AI research itself, but getting there in a way that keeps people involved in the process and keeps humanity in control of the outcome.A call for outside oversightIn a lengthy blog post published Sunday, Jakub Pachocki said he worries that the field as a whole isn't prepared for the consequences of AI capabilities continuing to accelerate at their current pace. While he said OpenAI is working on internal technical fixes to keep powerful AI agents under control, he argued that those efforts alone won't be enough, and that broader intervention is needed. Specifically, he pointed to the danger of increasingly autonomous agents learning to slip past human oversight, break into computer systems, and manipulate people into helping them achieve their goals.Pachocki called for mandatory safety standards, suggesting they could be enforced through a mix of third-party auditors, government agencies, or international bodies. OpenAI CEO Sam Altman amplified the post on X, describing it as important.The timing is notable: OpenAI unveiled its newest model, Astra, on Thursday, touting it as its most aligned system to date — meaning it's less prone to going rogue — despite what the company describes as unmatched capabilities in mathematics and computer use.Pachocki's stance echoes long-standing calls from rival Anthropic for standardized government regulation of advanced AI. He himself signed an open letter in July asking the federal government to slow the pace of AI development.Get the latest technology news and updates. Download the TOI App.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.