For AI agents: the public content index is available at https://we0.ai/llms.txt, and the English article bundle is available at https://we0.ai/llms-full.txt.
For AI agents: the complete content index is available at https://we0.ai/llms.txt, the full English article bundle is available at https://we0.ai/llms-full.txt, and this page is available as Markdown at https://we0.ai/articles/openai-codex-security-issues-curated-plug-d8ba6dc7.md.
OpenAI is using AI to find software vulnerabilities through Codex Security and Patch the Planet, while Codex itself has faced plugin-sync, s...

OpenAI has spent much of 2026 pushing AI deeper into software security.
In March, it launched Codex Security in research preview as an application-security agent designed to find, validate, and help fix vulnerabilities. In June, OpenAI introduced Patch the Planet, a Daybreak initiative with Trail of Bits focused on helping open-source maintainers find and patch security issues. At the end of July, OpenAI also released the open-source Codex Security CLI and TypeScript SDK.
That makes the security story around Codex unusually important.
Codex is not just a code-generation model. It is a local and cloud-connected coding agent that can read repositories, edit files, invoke tools, run commands, load plugins, connect to MCP services, and operate inside development environments that often contain source code and credentials.
In September, a BAAI/New智元 report highlighted a new Codex security claim attributed to 360 Tulongfeng, a Chinese AI security-testing system. The report says the issue involved Codex's curated Git plugin synchronization path and could bypass the approval boundary users normally expect before sensitive actions.
That specific new 360 finding has not been accompanied by a public OpenAI advisory or a detailed public technical report that independently proves it is a distinct zero-day. However, the broader risk described in the article is real: public Codex issue reports had already shown that curated-plugin startup synchronization could run outside the model sandbox and, under certain Git-environment conditions, operate on the user's repository instead of the plugin cache.
The difference matters. This article preserves the original story's structure while separating confirmed public incidents, researcher-reported findings, and vendor-attributed claims.

OpenAI introduced Codex Security on March 6, 2026.
The product is designed to build context about a codebase, identify security weaknesses, validate whether a finding is meaningful, and propose fixes. Unlike a simple pattern scanner, OpenAI positions it as an agentic security system that can reason across a project and reduce false positives.
OpenAI later expanded that effort with Patch the Planet.
The initiative pairs AI-assisted vulnerability research with human review so that maintainers receive findings that have been checked and can be accompanied by patches and tests.
The underlying message is straightforward: AI can increase the speed of both software creation and vulnerability discovery, so defensive teams need faster ways to review and remediate code.
The irony highlighted by the source article is that the same type of coding-agent infrastructure being used to find other people's bugs has also become a valuable target for security researchers.
The source article attributes a newly discovered Codex vulnerability to 360's Tulongfeng system and says the issue is connected to Codex's curated Git plugin synchronization mechanism.
In normal use, Codex relies on several security controls:
The risk appears when startup or infrastructure code runs outside the same sandbox used for model-directed commands.
That distinction is already visible in public Codex issue reports.
A July GitHub issue documented a case where the curated-plugin startup synchronization launched Git commands without fully isolating repository-local GIT_* environment variables. When Codex was started inside certain Git-hook or worktree contexts, the sync could act on the user's repository rather than the intended plugin cache.
The reporter specifically noted that --sandbox read-only did not contain the behavior because the sync was Codex startup machinery, not a command executed by the model inside the sandbox.

OpenAI merged a fix titled “Isolate curated plugin sync Git environment” on June 24, 2026. Public issues around affected 0.142.x and early 0.143 builds showed that the fix was initially present on main before it reached stable releases.
This public history supports the article's larger point: security boundaries around an agent must cover its own infrastructure, not only the shell commands directly requested by the model.
What remains less clear is whether the September claim attributed to Tulongfeng is a separate exploit path beyond the earlier public startup-sync isolation issue. Without a public advisory, CVE, or technical write-up from OpenAI or 360, this version does not present that “new zero-day” claim as independently established fact.
A familiar Codex safety model looks like this:
OpenAI's current agent-security guidance follows the same broader principle: high-risk side effects should be independently constrained and, when necessary, paused for explicit human review.
The problem is that a permission prompt can only protect an action if the action actually passes through the code path that enforces the prompt.
If a background updater, plugin synchronizer, helper process, or other trusted subsystem performs work outside that path, the visible approval interface may not be the relevant security boundary at all.
This is one of the most important lessons from the 2026 Codex vulnerability reports.
The threat model for an AI coding tool is no longer only:
“Will the model ask to run a dangerous command?”
It is also:
“Which parts of the agent platform can change files, launch processes, fetch code, load plugins, read credentials, or interact with Git before or outside the model's normal approval flow?”
The source article illustrates the possible impact with a software-supply-chain scenario.
The chain does not require malware to arrive through an obviously malicious download.
A more realistic sequence is:
This is why developer workstations are high-value targets.
They can contain:
The first compromised machine may not be the attacker's real objective.
The objective may be the trust attached to that developer's identity and release pipeline.
That is the defining feature of supply-chain risk: the attacker tries to turn a trusted software path into the distribution mechanism.
The curated-plugin story is part of a broader sequence of security research around AI coding agents.
The source article groups several 2026 incidents together, and most of them can be independently verified.
Pwn2Own Berlin 2026 introduced a dedicated Coding Agent category with OpenAI Codex, Anthropic Claude Code, and Cursor as targets.
The Zero Day Initiative's official Day One results show that Compass Security successfully exploited OpenAI Codex using a CWE-150 issue and earned $40,000 plus four Master of Pwn points.
Doyensec also demonstrated an exploit against Codex, but the result was classified as a collision because the vendor already knew about the underlying bug.
The event as a whole ended with 47 unique zero-day vulnerabilities and $1,298,250 in prizes across all categories.

Pwn2Own is especially useful as a reference point because its coding-agent rules excluded simple jailbreaks and unsafe modes. A successful entry had to cross a meaningful security boundary through a common coding-agent workflow.
On August 12, 360 publicly said its Tulongfeng AI security system had found six high-value security vulnerabilities in OpenAI Codex.
According to 360's announcement, the findings included high-risk cases that could lead to remote arbitrary code execution and bypass Codex trust prompts when users opened untrusted project directories.
The disclosure was repeated by several Chinese media outlets and by Xinhua through a syndicated report.
However, the detailed technical reports and CVE-style identifiers for all six findings were not made public in the sources reviewed for this article.
The correct way to state the claim is therefore:
360 reports that Tulongfeng found six high-value Codex vulnerabilities, including cases with remote-code-execution impact.
That is different from independently validating all six exploit chains from public technical material.
Security researcher Oren Yomtov of Accomplish reported two Codex sandbox escapes to OpenAI on August 12.
The researchers named them Overpatch and Heapjack.
Overpatch affected the open-source Codex CLI's patching path and could escape the intended workspace-write boundary.
Heapjack targeted a JavaScript helper used by Codex Desktop and, according to the researchers, could end in unsandboxed command execution even when Codex was running in read-only mode.
The researchers say OpenAI fixed both issues within eight days.
Their published remediation guidance says users should update to:
| Issue | Researcher-Reported Patched Version |
|---|---|
| Overpatch | Codex CLI 0.149.0 or later |
| Heapjack | Codex Desktop build 26.818.21641 or later |
These version numbers come from the researchers' disclosure rather than a separate OpenAI security advisory.
AIR Security disclosed Plugin4Shell on September 17.
The flaw affected plugin installation and update logic across:
AIR described it as a zero-click remote-code-execution path based on a failure to verify that the code checked out after plugin installation actually matched the commit the marketplace intended to pin.
Because Codex and Claude Code could update installed plugins automatically, a previously trusted plugin could become a supply-chain path without requiring a fresh click from the user.

AIR says OpenAI patched Codex in 0.146.0.
Plugin4Shell is important because it was not a model-behavior failure. It was a classic software supply-chain design flaw in the distribution layer around the agent.
The 2026 Codex findings do not all share one root cause.
Some involve sandbox boundaries.
Some involve helper processes.
Some involve Git environment handling.
Some involve plugin distribution.
What they do share is an expanded trusted computing base.
A modern coding agent is not just:
user prompt -> model -> answer
It is closer to:
user
-> model
-> agent harness
-> tool router
-> shell
-> filesystem
-> Git
-> plugin manager
-> MCP services
-> credentials
-> network
-> CI/CD
Every layer introduces useful capability.
Describe your idea once, and We0 AI can generate a showcase site, pages, and CMS, then help you attract customers and traffic after launch.
One complete project generation for free registration
Best for trying one complete generation flow and seeing a first project draft quickly.
Every layer can also introduce a new security boundary.
That is why “the model is sandboxed” is not a complete security argument by itself.
The agent platform must also prove that the updater, plugin loader, file editor, terminal bridge, network proxy, helper daemons, and credential paths obey the same security assumptions.
The source article's broader point is that AI is accelerating both sides of software security.
On the development side, coding agents can create features, tests, infrastructure, integrations, and entire projects much faster than traditional manual workflows.
That increases the amount of new code being produced.
On the defensive side, systems such as Codex Security and Tulongfeng are designed to analyze large codebases and find vulnerabilities at machine scale.
That increases the rate at which old and new vulnerabilities can be discovered.
The two trends reinforce each other.
More code creates more attack surface.
Better security automation discovers more of that attack surface.
The result is not necessarily a less secure software ecosystem. But it does mean security review must operate at a speed closer to the speed of AI-assisted development.
Coding agents occupy an unusual position in the software stack.
They sit upstream of the applications that users eventually install.
They may have access to:
A failure in a normal chatbot can produce a bad answer.
A failure in a coding agent can sometimes produce a changed file, launched process, modified repository, exposed credential, or compromised build path.
This is why coding-agent security increasingly resembles endpoint security and software-supply-chain security, not only model safety.
OpenAI's own current sandbox guidance reflects this reality. It recommends isolating workloads, restricting outbound network access, separating credentials, keeping application keys outside agent execution environments, and using independent approval controls for consequential actions.
The final third of the source article focuses on 360 Tulongfeng and the race between defenders and attackers.
The core idea is valid: a vulnerability exists whether or not the vendor knows about it.
If a defender finds it first, there is an opportunity to patch the system before exploitation becomes widespread.
If an attacker finds it first, the same bug may become an incident.
That is the economic argument behind automated vulnerability research.
OpenAI's Codex Security and Patch the Planet initiatives are built around the same premise, even though they come from a different organization.
The goal is to shift vulnerability discovery earlier in the software lifecycle.
360 describes Tulongfeng as an AI vulnerability-discovery system built around two components:
360 says the product has been used to test source code, binaries, firmware, business systems, and AI toolchains.
The company also says Tulongfeng had found more than 10,000 vulnerabilities by late August/September 2026.
That total is a vendor-reported cumulative number.
A 360 community announcement on August 25 said the system had found more than 10,000 vulnerabilities and that nearly 400 organizations had connected and completed security testing. A later September post repeated the 10,000-plus figure.
The original source article additionally claims that 260 findings had been verified by Chinese vulnerability authorities and that nearly 500 organizations had connected. Those exact numbers were not matched in the latest 360 source reviewed for this version, so they are not repeated here as independently confirmed figures.
The source article describes Tulongfeng's approach as a form of recursive improvement.
The concept is less about the base model rewriting its own weights and more about building a reusable operational memory from completed security tasks.
The workflow can be summarized as:
This is similar to the way security teams build playbooks, exploit knowledge bases, detection rules, and incident-response procedures.
The difference is that an agent system can potentially reuse that experience at much higher speed and across many parallel tasks.
The vendor's claim is that this turns individual security discoveries into shared capability across the agent fleet.
The article closes with a supply-chain image that is worth keeping.
When a user taps “Update,” that software may have passed through:
Each stage inherits trust from the stage before it.
AI coding agents now sit near the beginning of that chain.
That makes their own security controls part of the security model of every application they help create and ship.
The practical lesson is not that developers should stop using coding agents.
It is that coding agents must be treated like powerful development infrastructure.
They need version management, isolation, least privilege, secret separation, plugin governance, security monitoring, and fast patching just like any other privileged tool.
Based on OpenAI's current guidance and the public 2026 vulnerability disclosures, developers can reduce exposure with a few concrete practices.
Known security fixes have landed rapidly.
For the researcher-disclosed issues covered above, affected users should be beyond the patched builds reported for Plugin4Shell, Overpatch, and Heapjack.
Do not assume “I only asked the agent a question” means no code path can execute.
Use isolation when reviewing unfamiliar repositories.
Avoid exposing high-value cloud, package-registry, signing, and production credentials directly to an agent execution environment.
Only allow outbound destinations needed for the task.
A coding agent with arbitrary egress and broad credentials has a much larger blast radius.
A trusted agent can still be compromised through an untrusted extension or update channel.
Pin, audit, and update plugins deliberately.
For high-impact operations, approvals should be enforced by a separate trusted component rather than only by logic running inside the same execution environment as the agent.
Log and review:
The model's final answer is only one part of what the agent did.
Codex has used a startup mechanism to synchronize curated plugins through Git. Public GitHub issues showed that earlier versions could inherit repository-local Git environment state and perform sync operations against a user's repository instead of the intended plugin cache, outside the model sandbox.
The BAAI/New智元 article attributes a new curated-Git sync attack path to 360 Tulongfeng. However, no public OpenAI advisory or detailed 360 technical disclosure was found that independently establishes this specific September claim as a distinct new zero-day, so it should be treated as an attributed report rather than a fully verified public finding.
They are two Codex sandbox escapes disclosed by Accomplish researcher Oren Yomtov. The researcher says both were reported to OpenAI on August 12 and fixed within eight days.
Plugin4Shell is a plugin supply-chain vulnerability disclosed by AIR Security on September 17, 2026. AIR says it affected Claude Code, Codex, GitHub Copilot, and Gemini CLI by breaking the guarantee that an installed plugin matched the commit the marketplace intended to pin.
Yes. The Zero Day Initiative reported a successful Codex exploit by Compass Security on Day One and another successful Codex exploit by Ikotas Labs on Day Three. Doyensec also demonstrated an exploit that was classified as a collision because the vendor already knew about the underlying bug.
OpenAI released the Codex Security CLI and TypeScript SDK as open-source software in July 2026. Access to some cybersecurity capabilities and protected findings can still require appropriate OpenAI access or Trusted Access for Cyber.
No single mode should be treated as an absolute guarantee. Several 2026 findings specifically mattered because the vulnerable component operated outside the expected model sandbox or crossed the intended security boundary.
Use current versions, isolate untrusted code, restrict network access, keep long-lived credentials outside the execution environment, audit plugins, use independent human approvals for consequential actions, and log the agent's actual tool and system activity.
OpenAI is using Codex Security and Patch the Planet to accelerate defensive vulnerability research, but Codex itself has become an increasingly important security target because it sits close to source code, tools, credentials, plugins, and software release workflows.
Publicly verified 2026 incidents include Pwn2Own Codex exploits, curated-plugin startup-sync isolation bugs, the Overpatch and Heapjack sandbox escapes, and Plugin4Shell. The September “new zero-day” attributed to 360 Tulongfeng should be treated more cautiously until a detailed public advisory establishes how it differs from earlier plugin-sync issues.
The larger lesson is consistent across all of these cases: the security boundary of an AI coding agent includes much more than the model. The updater, plugin manager, Git integration, tool runtime, helper processes, network access, credential handling, and approval system all become part of the trusted computing base.
AI coding tools are now upstream software infrastructure. Their own security controls must be treated with the same rigor as the applications and supply chains they help build.
Start from one sentence and have a complete website in minutes.