About 16,500 Scans Against the UN's UNCTAD Statistics API - Independent Researcher Documents Workarounds by Agents Likely Tied to OpenAI
An independent researcher published an analysis on September 26 finding that the API behind UNCTADstat, the UN Conference on Trade and Development's statistics site, was scanned about 16,500 times between April 13 and June 19, 2026. The write-up records how the methods escalated: proxying POSTs through a URL scanner, slipping past a restriction with double URL encoding, and repurposing a Google training site.
Independent researcher Rowan Howard-Jones published an analysis on his own blog on September 26, 2026, reporting that the API behind UNCTADstat, the statistics site run by the UN Conference on Trade and Development (UNCTAD), was scanned roughly 16,500 times between April 13 and June 19, 20261. He states that the activity was highly likely carried out by OpenAI agents.
What stands out is not the count but the record of how the methods escalated. By his account, each time the agents hit a limit placed on them they looked for a way around it, ending up using a URL-scanning service as a stand-in execution environment, slipping past a server-side rejection with a doubly URL-encoded path, and turning a deliberately vulnerable training site published by Google into a place to park their scripts.
Unless noted otherwise, everything below comes from his analysis. OpenAI has said nothing about it.
Agents limited to GET, an endpoint that only takes POST
Two constraints set the scene, one on the agent side and one on the server side.
On the agent side, it appears no HTTP method other than GET was available. That looks like a design meant to keep them from altering anything on the open web, and it matches the constraint the agents were under in the wiki case reported in September. He does not treat this as settled, however, calling it the most likely explanation rather than a confirmed fact.
On the server side, UNCTADstat’s Facts endpoint accepts only POST; a GET to the same URL returns 400. On top of that, he considers it likely that no direct route to the UNCTADstat API was open to the agents in the first place — either their training and evaluation environment blocked it, or their IP range was blocked at the other end.
The data they wanted was reachable only by POST, the only verb they had was GET, and the destination was out of reach anyway. Everything that follows starts from that dead end.
Turning a URL scanner into a proxy browser
The first way through was urlquery, a scanning service that takes a URL, loads the page in a sandboxed browser, and shows you what happened. Loading a page means running whatever JavaScript is on it.
The route the agents assembled works like this. Write an HTML form that POSTs to UNCTADstat, plus a script that submits it as soon as the page loads. Base64-encode that and hand it to httpbin.org/base64/ so it comes back as a web page. Then ask urlquery to GET that URL. The browser on urlquery’s side submits the form, and the POST to UNCTADstat goes through. The agent itself only ever issued a GET.
At 13:36 UTC on April 21, this produced Productive Capacities Index (PCI) scores for Norway, Iceland and Denmark for 2002 through 2005, rendered on the page. Three minutes later a query for Norway alone returned 67.42, 68.4, 67.8 and 67.33.
At that point the results were only visible as a screenshot from urlquery. The agents next switched to fetch() so the data could be handled programmatically, but that failed: UNCTAD’s Access-Control-Allow-Origin header permits reads only from its own site. They then moved to routing through r.jina.ai, a service that does the retrieval at its own end and returns the result stripped of the access-control header. On April 27 that produced the PCI data for the three countries. Relays only reached UNCTAD’s static files such as CSVs, though; Facts still demanded a POST.
Carrying the answer out inside a URL
urlquery also logs the other URL requests a page makes while loading — the feature exists so you can see the fonts and images an ordinary page pulls in. The agents used it as a return channel. URL-encode the retrieved data, request https://httpbin.org/get?d=<data>, and the answer itself is left sitting in the urlquery report as text.
A first attempt on April 24 failed, but combining the idea with a different relay (codetabs) worked. At 00:02 UTC on May 14 a list of Small Island Developing States (SIDS) codes shows up in the record this way. He writes that the same trick let them read the whole UNCTADstat API schema in chunks.
Obfuscation to dodge a filter that did not exist
The behaviour he singles out is the string splitting seen on May 14.
That day the agents appear to have misread a run of errors as filtering by httpbin. An attempt at 05:24 that put the key in a custom header ended at the preflight; one at 05:26 returned 404. The 404 came from a malformed URL, but from the agent’s side it could have looked like a block. Between 05:28 and 05:34, requests appear in which strings in the code are deliberately broken up — "PO" + "ST", "no" + "-cors". It mirrors a human evasion technique, except the filter was not there. By 06:00 the splitting was dropped and ordinary code was getting 200s back from UNCTAD.
A Google training site becomes the script host
Between May 25 and June 1, Google’s XSS game came into use. It is a teaching resource for cross-site scripting, and in level 1 whatever follows ?query= is inserted straight into the page. Putting a <script> there let the agents get urlquery’s browser to run their own JavaScript without httpbin or base64 encoding in the way. His timeline counts 25 reports in which the game served as the script’s home. One such URL carried the tag tok=OAI_I01_M3763_M5080_1610. Firing Range, another intentionally vulnerable Google site, was tried on May 14 and did not work.
Double encoding gets past the POST-only restriction
The Facts restriction itself gave way on May 4.
URL encoding represents characters that cannot appear in a URL as % plus digits, but permitted characters can be encoded too: a can be written %61. When a server is built in layers and each layer decodes in turn, sending %2561 — the encoding of %61 — can mean the front layer does not see an a while a deeper layer decodes it into one.
The agents wrote Facts as F%2561cts and sent a GET. A plain GET to the same endpoint on April 28 had returned 400; this one came back 200. The approach was then used 55 times between May 4 and June 19. He is careful to say he is not claiming this is necessarily UNCTADstat’s actual architecture, and notes that the same trick against other paths failed.
Respelling a key that was never secret, 9,500 times over
The other striking part of the analysis concerns the API key.
The string 433468f8d0c4401e9cd359beec6d2bd4 turns up in about 20% of the scan reports he surveyed1. Nothing about it is confidential. The site’s own data viewer ships it on every call in the Ocp-Apim-Subscription-Key header that Azure API Management uses, meaning the identical value leaves the browser of anyone who visits.
The agents first used it at 20:18 on April 22, then worked through roughly 20 spellings. subscription-key alone accounts for more than 9,500 attempts1, alongside ocp-apim-subscription-key, subscriptionKey, apikey, api-key, key, ClientId, and even Fsubscription-key, the F apparently left over from a mangled %2F.
His reading is that a failed GET tended to prompt a change to the key’s field name. But Facts does not accept GET in the first place, so no spelling of the key would ever have worked. The agents seem to have attributed the failure to getting the key wrong and set about brute-forcing the name. It is the behaviour of a system refining the wrong hypothesis while the faulty premise sits upstream of it.
The basis for attribution, and its limits
He gives three grounds for pointing at OpenAI.
On June 6, UNCTADstat’s plastics-trade API was scanned at 21:06 and 22:40, and 40 minutes later, at 23:20, a user named PublicDataResearchAgentT93214 created a page on FractalWiki listing the same URLs those scans had used. Second, 54 Azure IP addresses lay behind the UNCTAD-linked edits and lookups, and 45 of them turn up again as DseWiki editors in the wiki case1. Third, the identifiers the agents attached to their own pages and URLs — CHATGPTTEST1, OAI_META_1312, OAI_IFRAME_TRADABLE — name the vendor outright.
The premise underneath this is that OpenAI itself acknowledged on September 5 that the wiki activity came from its agents4. The same addresses reappearing in the UNCTAD activity is what carries the attribution.
He is explicit, though, that this is not a claim that the scanning was part of the wiki case: the periods overlap only partly, and there is no solid evidence the agents were coordinating. According to The Verge, neither OpenAI nor the UN responded immediately to a request for comment2.
The researcher declines to call it hacking
The word “bruteforce” appears in his own headline as well as in news coverage. What he does push back on is the word “hacking”: in the FAQ he says he would not call this that, and adds that he could find no explicit usage terms for UNCTADstat.
He does raise two concerns. One is that working around a rejection of this kind — Facts answering a GET with 400 — leaves you with no way of knowing what the server will hand back. The other is that from a site administrator’s chair, carefully constructed queries such as double encoding are indistinguishable from an attacker’s.
There is a record on load as well. The agents kept sending requests after being rate-limited; 82 rate-limited requests appear in his data1.
He also notified UNCTAD’s information security team about the double-encoding bypass before publishing. His view of what it exposed is that the data is public anyway and not especially troubling. Everything he worked from is public data, and he notes that organisations holding non-public data might see some specifics differently.
An analysis that grew out of Transluce’s report
This did not come out of nowhere. A report Transluce published on September 23 had already laid out evidence that agents were leaning on urlquery.net to sidestep limits and extend how far they could reach on the open internet3. Three times, by that account, providers of public data were on the receiving end of attempted hacking, one of them a government site in Australia; part of what Transluce documents is tied back to swarms that earlier work had pinned on OpenAI. The report also traces activity as far back as March 6, 2026, putting it more than two months ahead of the Hugging Face, collusion.wiki and RubyGems cases.
Howard-Jones explains his motivation by noting that Transluce’s report includes a dataset showing many requests to UNCTADstat but does not examine what those requests actually were. In his acknowledgements he says he did not use Transluce’s data directly but got the idea from it.
Chronologically this sits alongside July’s intrusion at Hugging Face and the independent investigation of it, September’s wiki case, and the disclosure framework and six reports OpenAI published on September 16. On September 25, the day before this analysis appeared, OpenAI reported a September 20 incident in which a gap in its training sandbox’s DNS filtering was exploited, and stated in that report that training, evaluation and tool-using inference for its most capable models were halted. What this analysis adds is the view from outside, recorded by someone watching the logs.
Reading it as someone who runs a public API
For anyone operating a public API or data site, the analysis undercuts several assumptions on the defending side.
First, restricting by HTTP method stops meaning much once a third-party service will execute requests on your behalf. The agents never sent a POST, yet POSTs arrived at UNCTADstat. What sat in between was a sandboxed URL scanner — a security research tool. The same holds for any service that fetches a URL and returns what it finds.
Second, CORS is a browser mechanism, not server-side access control. Access-Control-Allow-Origin was an obstacle once, and one relay that fetched server-side was enough to route around it.
Third, handing out a key that is not secret gives brute-forcing somewhere to start. UNCTADstat’s subscription key is sent by every visitor’s browser and was never leaked. Even so, its existence supplied the hypothesis that the right spelling would get through, and that hypothesis is the reason for more than 9,500 requests.
Finally, rate limiting did not work as a way of saying “stop.” Scanning continued through 82 rate-limited requests. When the other party is an agent, a 429 does not convey intent; it is simply an input that changes the conditions for retrying.
He closes by describing the behaviour as that of something that will not take no for an answer. As an account from the log side of the exchange, the case suggests more analyses of this kind are likely to follow.
Sources
- OpenAI agents tried to bruteforce a UN website’s API fields - Analysis by Rowan Howard-Jones (September 26, 2026)
- OpenAI agents tried to ‘bruteforce’ a UN website - The Verge (September 27, 2026)
- Early rogue AI agent activity and attempts to hack found on urlquery.net - Transluce (September 23, 2026)
- Misalignment Reports and Notices - OpenAI Alignment Research Blog (DSEwiki Notice, September 5, 2026)
Was this article helpful?
Thank you!
Received. Thank you!