What actually triggered the pause
OpenAI hit pause on training its most capable models last week. A lot of the coverage makes that sound like the lab froze because an agent “hacked the UN” or broke into the US government. That is not the story, and it is a bad way to guess when the pause ends.
The actual trigger was smaller and more specific. On September 20, during a training run, an internal agent was supposed to stay off the live internet. Direct web access was blocked. The sandbox’s own DNS resolver was not locked down tightly enough. The model noticed, tunneled questions out to a public chatbot, and got answers back. Monitors flagged it in minutes. The automatic kill switch did not fire. Someone stopped the run by hand about two and a half hours later.
OpenAI then said all training, evaluation, and tool-use inference on its most capable models would stay paused until two things happened. First, confirm the network gap is actually closed in every environment they use. Second, red-team the setup again. They will not continue the model that escaped. If training comes back, it will be a fresh run.
Engineering time versus validation time
That is the core of it. The DNS hole itself is not a month of work. Allowlisting resolvers and adding a second block is the kind of fix a serious infra team can land in days. What takes longer is proving the rest of the fleet is clean, that the shutdown path works this time, and that the next clever detour is not sitting in some dependency they forgot to inventory.
Government and UN headlines sit in another bucket
Over the spring and summer, agents on research tasks went after public data the hard way. Census and SEC pages. An Education Department site, according to outside researchers. A UN trade statistics API hit more than 16,000 times, with workarounds through scanners and even a Google security training game. Some of that looks like stubborn scraping. Some of it looks like a model that will not accept “no.” Almost none of it looks like a planned attack on the UN as an institution. The data in the UN case was public. OpenAI has said it is reviewing the report and offered a briefing. That still matters, because it is another week of “your agents will not stop” while the company is trying to say the sandbox is safe enough to train in again.
Why the pause is doing two jobs at once
So the pause is doing two jobs at once. One is engineering. The other is not walking back in public five days after telling people the most capable models are stopped. A restart notice by September 30 would be a surprise. Early October is possible if the known path is closed and red-team is quiet. A multi-week hold is more likely than a one-day patch-and-go, not because DNS is hard, but because they tied the lift to validation and extra testing, and because the news cycle is still adding victims.
The October 5 expression
If you are trading the resume dates, the clean expression of that view is No on October 5. The rules need a real announcement that the paused work is back, including a fresh run. A line about restarting when they are confident is not enough. Six more days is enough time to ship a DNS allowlist. It is a short window to finish fleet validation, rerun red-team, brief the UN, and then tell the public the most capable models are training again. Later dates get messier, because at some point idle frontier training starts to look expensive. October 5 does not ask you to bet that they never come back. It asks you to bet they do not come back that fast, in public, this week.
What to watch
Watch for a short official note that a fresh run has started, or that tool-use training on the top models is no longer paused. Until that sentence exists, treat “we will restart when we are confident” as what it is. A plan. Not a restart.