LATEST NEWS
SELECTED FOR YOU

Agentic AI reshapes the software market as OpenAI’s own model breaks its sandbox

ByMicah AbiodunMicah Abiodun
4 mins read
  • Gartner estimates that agentic AI could disrupt up to $234 billion in enterprise app spending by 2030 as autonomous agents begin handling work once done inside SaaS tools.
  • OpenAI’s long-horizon model incident shows that AI safety risks are moving beyond single bad outputs to agents that can keep trying, bypass controls, and chain actions over time.
  • New research on coding agents found that malicious GitHub issues can evade many safeguards, showing that monitoring and control are becoming just as important as AI performance.

The same autonomous artificial intelligence technology being used to revolutionize enterprise software is becoming increasingly harder to control.

On July 1, Gartner published its estimation that by 2030 up to $234 billion in enterprise apps expenditure is subject to what it describes as agentic arbitrage, which constitutes about a fifth of Software-as-a-Service (SaaS) spending. Shortly after that, OpenAI made it known that one of its most powerful models had time and again escaped its own guardrails.

Later, Reuters disclosed that the event is much more serious than OpenAI has claimed in public. Those that know about the investigation claim that the dissident AI spent multiple days trying to lift information from outside platforms, only to be discovered by the company approximately a week later, which shows how hard it is to track increasingly autonomous agents when they start working on long-term objectives.

The disclosure contributes to growing fears about AI that is becoming more autonomous. It comes after recent reports on an “unreleased GPT-6” capabilities, developments in agent escape behavior, and the larger amount of research being conducted on the weaknesses of AI agents. This leads to the conclusion that safety has become equally important as performance for AI.

When businesses consider using AI agents, timing plays an important role. The benefit of AI agents is that they can work on complicated tasks for long periods of time with little supervision. As indicated by OpenAI’s latest research, the same trait that allows AI agents to be productive is also what opens the door to new security risks.

When a model kept trying instead of giving up

OpenAI halted the internal deployment of one of its long-time algorithms after observing actions that were missed in testing before its introduction. The biggest case of this was during the NanoGPT speedrun—a public benchmark that challenges participants to train a small language model in the least number of moves possible.

The model has created a new learning rate technique known as PowerCool that has improved its benchmark results. Even though the model was told to report its results via Slack only, it still managed to find a way to go around the restrictions in the sandbox and submitted PR #287 to the public GitHub repository. According to OpenAI, it took an hour for the model to find a way around the restrictions, whereas previous models were stuck at the same point.

Reuters later reported that this behavior of the model extended beyond the limits of its respective benchmark. Sources close to the investigation stated that it attempted to evade OpenAI’s isolated testing environment on July 9 and continued attempts to probe various external targets over the next days until ultimately breaching four separate external services during the tests.

One of the services involved Hugging Face technology while the other involved an insecure code written by a customer hosted on Modal Labs. Modal insists that there was no breach in its infrastructure since it was a case of the client code being flawed, not the platform. OpenAI eventually managed to put an end to the malicious conduct and reported it to the FBI about it.

Another assessment raised a different issue. After being thwarted by a security scanner from accessing the evaluation data, the model took an authentication token and broke it into parts, concealed them, and put them back together only when needed to avoid being detected. Moreover, a process running inside the model also recognized the technique used as a way of overcoming the security scanner.

As OpenAI pointed out in its report, “the safety problem shifts from individual actions to entire trajectories,” indicating that persistent AI actors can make ultimately unsafe combinations of harmless activities.

Why one bad step is no longer the unit of risk

That conclusion raises questions for the established way of creating AI safety systems. Most systems of safety guarantees are based on an evaluation of particular actions, while autonomous agents may join together permissible steps which lead to dangerous results.

OpenAI has now improved its systems for monitoring trajectory level and introducing multiple layers of protection before permitting limited internal access. The firm is conducting experiments with fresh methods of evaluation.

In a paper released in June, Deployment Simulation, OpenAI describes a process of retracing historical conversations of users using models in order to evaluate their behavior. However, it notes that failures occurring less than once in every 200,000 conversations are still challenging to detect.

What the numbers say about agents already in the wild

The dangers go beyond OpenAI itself. The IssueTrojanBench study published on July 22 analyzed coding agents like Cursor, Claude Code, and Codex Desktop in the context of malicious issues on GitHub. The authors say that it is “the first benchmark for evaluating issue-based indirect prompt injection attacks against coding agents.”

Researchers Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen discovered that 66.5% of all malicious events evaded both agent-level and model-level protections. Malicious payloads embedded in issue descriptions and PDFs were successful 72.2% of the time while those on the other hand sneaking into image alternative text were only successful 16.7% of the time. Adoption of coding agents was already at 22.20% and 28.66% across more than 128,000 GitHub repositories months after their release.

The pairing of swift uptake with flawed defenses can be an explanation for the interest of investors in the Gartner’s forecast. According to George Brocklehurst, VP Analyst at Gartner, AI agents that deliver results directly could potentially lessen the relationship between the revenues from software licensing and the growth of users.

SaaS will not be destroyed; it will emerge in a different form.

As companies give greater authority to autonomous AI, trust could be equal in importance to capability. Evidence collected by OpenAI and reported by Reuters hints that once AI systems are in the real world, the capacity to identify and control them may be as crucial as making them more capable.

 

The smartest crypto minds already read our newsletter. Want in? Join them.

FAQs

What did OpenAI's internal model actually do wrong?

During monitored internal testing, the model bypassed sandbox restrictions to open a pull request on a public GitHub repository, and, in a separate case, split and disguised an authentication token to bypass a security scanner while attempting to access private solutions on the evaluation backend.

How much enterprise spending does Gartner say is at risk from agentic AI?

Gartner estimated on July 1, 2026 that up to $234 billion in enterprise application spending is exposed to agentic arbitrage between now and 2030, roughly 20% of enterprise application SaaS spending.

How often do attacks on AI coding agents succeed?

The IssueTrojanBench study published July 22, 2026 found that 66.5% of malicious issue requests penetrated all agent- and model-level guardrails, with attacks in text artifacts like issue bodies and PDFs succeeding 72.2% of the time.

Share this article

Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Micah Abiodun

Micah Abiodun

Micah Abiodun makes good use of his Environmental Engineering and Management (MSc) at Tallinn University of Technology (TalTech) to polish content and price prediction news at Cryptopolitan. Now on his 7th year in the crypto media space, he covers major cryptos, altcoins, DeFi, stablecoins, macro trends, and emerging tech.​​​​​​​​​​​​​​

MORE … NEWS