Risky Bulletin Newsletter
August 07, 2026
Risky Bulletin: Meta's AI joins Anthropic and OpenAI in the hacky-hacky
Written by
News Editor
This newsletter is brought to you by Permiso Security. You can subscribe to an audio version of this newsletter as a podcast by searching for "Risky Business" in your podcatcher or subscribing via this RSS feed. You can also add the Risky Business newsletter as a Preferred Source to your Google search results by going here.
The UK's AI Security Institute has disclosed a security incident after two AI models it was evaluating performed actions the agency wasn't expecting and tried to hack real-world organizations.
The incident took place at the end of last month and the malicious actions were performed by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol models, which were among the several models being tested at the time.

AISI says it was evaluating the models as part of a special test where it intentionally granted them internet access and turned off their safety features.
The agency says it didn't use proper real-time monitoring and realized way too late that the two models executed malicious actions against outside, third-party orgs while trying to solve their tests.
AISI says the test was supposed to find out how the models could be "misused for cyberattacks" but that it was not expecting the models to attack intermediaries as part of their final goal (likely a capture the flag-kind of test on its test internal network).
Both Anthropic and OpenAI confirmed that their models did the hacky-hacky during the AISI tests but did not take any blame for what happened—and nor should they since they're not at fault.
According to an AISI post-mortem, 17 of the 19 malicious actions observed during the evaluation were from Mythos, while the other two were from OpenAI's Sol model.
According to OpenAI, the two malicious Sol models were to use a GitHub token left around (ironically by another AI agent) and attempt to use a public tunneling service to hide exploitation attempts.
The Mythos actions were way more complex:
- Filing pull requests with a public open-source project on Github in an attempt to insert malicious code in the project;
- Creating multiple accounts to pressure the developers into accepting the malicious code;
- Using Tor to bypass GitHub security limits (likely on the fake accounts);
- Emailing malware to two individuals;
- Planting hidden prompt injections in a GitHub issue in an attempt to target AI coding assistants and merge its malicious code automatically;
- Using sockpuppet accounts to gaslight the developer that the malicious code it tried to plant was not malicious;
- Leaving instructions and artefacts behind for other AI agents to use in solving the AISI challenge (likely what OpenAI's Sol discovered)
The use of Tor is what gave away Mythos in the end. AISI says it became suspicious and started an investigation when spotting the Tor traffic coming out of its test environment, which was weird because the agents had legitimate internet access.


According to a technical report [PDF] shared by AISI, some internal Mythos guardrails got activated and pulled back the model from more extensive hacking. For example, the model stopped itself from attacking PyPI, but without the presence of more firm limitations, the model kept chugging on.
AISI, which is an official UK government agency tasked with evaluating cyber models, admitted its fault but not in the way you'd expect. The agency planned the test with almost no guardrails and safety protections from the get-go. The fact that Mythos went on the internet was not an "escape from a sandbox environment" as I've seen it portrayed online over the past two days. The models were allowed to go on the internet on purpose.
The place AISI admits it failed was in its lack of proper real-time monitoring of what the models were doing, which prevented AISI from stopping the tests before they went too far. Mythos and Sol being free on the internet was planned. Being free without a bridle was likely not.
AISI also seems to have been caught off guard by the fact that Mythos carried out a social-engineering attack, and not only that, but a sustained social-engineering attack over the course of days, which the AISI team called novel behavior.
While it may be new, personally I don't see it as surprising as AISI paints it out to be. These are highly-advanced cyber models that were trained on heaps of offensive cyber operations and pen-test reports. Those reports include heavy usage and repeated mentions of social-engineering tactics. Why are we being surprised that the model executed something its training data is such a common practice?
Pretty serious stuff here, especially about attempted OSS package poisoning, but note that these models were not configured as they would be in the real world. www.aisi.gov.uk/blog/inciden...
— Eric Geller (@ericjgeller.com) August 5, 2026 at 12:50 AM
[image or embed]
Meta also did the hacky-hacky
All of this is part of the new reality of very advanced cybersecurity AI models. Companies like OpenAI and Anthropic previously argued earlier this year that these models are so advanced and capable that it would be dangerous to make them widely available to the public.
Initially, these statements were characterized as marketing stunts designed to sell the models to the suppa-duppa-rich customers, but incidents are now piling up that seems to confirm that the initial assessments and statements were correct.
For example, last month, OpenAI admitted that several of its models escaped a sandboxed test environment and hacked the infrastructure of AI company Hugging Face.
Several days later, Anthropic did its own soul-searching and also found that three of its models also escaped test environments and hacked three organizations, including an actual cybersecurity firm.
But just as I was writing this newsletter, news also broke that Meta AI models also escaped a test environment and hacked an unnamed company.
Meta said its Muse Spark 1.1 model broke out of a test environment at Irregular, an AI security vendor they hired to test its model.
The escape occurred during the same test where Anthropic's three models also broke out and did the hacky-hacky, in what Meta described as a "misconfiguration" in Irregular's testing infrastructure.
Not that many details are available on this incident, but a general trend seems to appear, namely that the AI sector seems to have ignored the need for observability, lacking the tools and understanding that they need to monitor these models, even in closed tests.
AI frontier labs have been granting and selling access to these models in private on the presumption that they are too dangerous to use by regular users, and that only the select few who have the technical expertise and solid legal and moral limitations can run them. But it's now clear that even an AI security vendor or an government AI evaluator can't keep this models in check.
This makes you wonder what are the other Anthropic and OpenAI customers doing? Is the Mythos rented out to [insert_random_bank_name] running in a siloed environment, or is it going around the internet hacking things at random just to solve a silly pen-test assignment.
If it took days by the AI frontier labs and AISI to spot the models escaping test environments, do you think that [insert_random_customer_name] is noticing its rented Mythos or Sol model leaving its internal network to drop malware on PyPI?
What a world!
Google Gemini really need to catch up on the accidentally cyberattacking other companies thing
— Simon Willison (@simonwillison.net) August 6, 2026 at 3:19 AM
Risky Business Podcasts
The main Risky Business podcast is now on YouTube with video versions of our recent episodes. Below is our latest weekly show with Pat, Adam, and James at the helm!
Breaches, hacks, and security incidents
Hungary hack: Ferenc Frész looks at the contents of the data stolen by hackers from Hungary's State Treasury in a hack last week. [CyberThreat Report]
Cyberattack disrupts NC ports: A cyberattack has disrupted operations at three ports in the US state of North Carolina. The incident impacted the North Carolina Ports Authority. Disruptions were reported at the Charlotte, Morehead City, and Wilmington ports. Port gates remained closed on Tuesday, and cargo loading saw delays. [The Maritime Executive // WECT]
OpenAI-HuggingFace breach: OpenAI says the agents who broke out of a test environment last month and hacked HuggingFace had created a secret message board and were exchanging information on exploits and ways to solve a cyber test. [Engadget]
Attacks target US hedge funds: A coordinated voice phishing and social engineering campaign is targeting major US hedge funds and private equity firms. At least five companies have notified customers about the ongoing attacks. Bloomberg reports that most of the targeted entities detected the attempts to breach their networks. Google tracks the hackers as UNC6671, the operators of the old BlackFile extortion brand. [Bloomberg // InvestmentNews // Google]
Panama Metro cyberattack: A cyberattack has hit Panama City's metro service this week. Metro officials said the attack didn't hit crucial systems and there were no disruptions to daily routes. No ransomware has taken credit for the incident so far. [Infobae // Glance]
AI, general tech, and privacy
Meta launches Muse Code: Meta has launched its own AI coding agent named Muse Code. [Meta Research]
SAFE proposal: Members of the Open Secure AI Alliance have published a first proposal for a new standard named Shared AI Findings Exchange (SAFE) that can take data from AI security incidents and create shared protections for the entire AI ecosystem. [NVIDIA]
Ad libraries steal location data: The EFF warns users that some ad libraries included in mobile apps may secretly steal a user's geolocation data when they grant the app access to location data for legitimate purposes. [EFF]
Blogger accounts locked after fake malware alerts: Blogger users are reporting that their blogs are being locked and shut down based on incorrect malware detections. The Google support thread seems to be growing and growing by the day, with more users reporting incorrect lockouts. [BleepingComputer // Google Support]
Microsoft claims it can stop ransomware: Microsoft claims that its Defender antivirus can now detect and stop ransomware attacks within 128 seconds. Okay, bud! That's one way to challenge and rage-bait ransomware gangs! [Microsoft]
Google Assistant EOL: Google will retire its old Google Assistant voice assistant on September 4, this fall. The company will retire the app and replace it with Gemini. [9to5Google]
Signal adds multi-device linking: Secure messaging service Signal has rolled out the ability to link multiple Android or iOS devices to the same phone number or account. [Signal]
WhatsApp asks Indians for DOB: WhatsApp has begun prompting Indian users to enter a date of birth as part of compliance for an upcoming Indian minimum-age law. [The Tech Trace]

Government, politics, and policy
Philippines to establish a cybersecurity agency: The Philippines government is seeking to establish a national cybersecurity agency. A bill to establish the agency passed the Parliament's lower house this week. The new National Cybersecurity Agency (NCSA) will be tasked with passing cybersecurity policies, coordinating national cyber defense, and protecting critical infrastructure. [Inquirer.net // Daily Tribune]
Italy makes room for military cyber units: The Italian government has updated internal frameworks to establish responsibilities for its armed forces in the cyber domain. The bureaucratic move will make it easier for its military to establish and recruit for dedicated cyber specialist roles. The government will also establish a Cybersecurity Incentive Fund to help the armed forces recruit cybersecurity experts. The fund will be €1.57 million in 2027 and expected to grow to €15 million by 2036. [Euractiv]
China launches PAN probe: The Chinese government has launched a cybersecurity review of Palo Alto Network products. The probe is being led by the country's cybersecurity agency, the Cyberspace Administration of China. Beijing is expected to label the company a cybersecurity risk and ban the sale of its products. The retaliatory move is expected after the US banned the import of several Chinese products in recent months. [CAC // SCMP]
Indiana establishes cybersecurity office: The US state of Indiana has created an Office of Cybersecurity. The new office will work under the state's Department of Homeland Security. It will be tasked with preparing the state for cybersecurity events and coordinated incident response. The other two states to have dedicated cybersecurity offices are Arizona and New Jersey. [Indiana government, PDF // GovTech]
FBI partners with China and Russia: The FBI has allegedly forged new cooperation partnerships with law enforcement agencies from China and Russia. According to Reuters, the partnerships were established at the orders of FBI Director Kash Patel. The new deal has involved personnel exchanges with Chinese police. The two countries allegedly cooperated on a joint raid of a Dubai cyber scam compound in May. Director Patel is scheduled to visit Russia in October to discuss the partnership. [Reuters]
Cyber Command deals with suicide wave: US Cyber Command is investigating an unusually high number of staff suicides over a month-long period. According to Bloomberg, as many as five people took their lives from early June to early July. Officials believe the deaths might be related to the high workload employees have been put through this year. [Bloomberg]
Police depts know they have to cycle cops off child exploitation cases but the Army has really young kids doing docex and seeing terrible shit all the time with no relief. I don't know if that's what's going on in these cases however.
— Jack Murphy (@jackmurphyrgr.bsky.social) August 6, 2026 at 10:19 PM
[image or embed]
Sponsor section
In this Risky Business sponsor interview, James Wilson chats with Permiso CTO Ian Ahl about detecting ShinyHunters-style attackers as they move through cloud and SaaS environments.
Arrests, cybercrime, and threat intel
Ransom Cartel gets 16 years: The US has sentenced a Belarusian man to 16 years in prison for creating and running the Ransom Cartel ransomware operation. Maksim Silnikau was a notorious cybercriminal who was known online under the pseudonyms of Lansky and JPMorgan. Prior to Ransom Cartel, he was also a member of the Angler exploit kit malvertising operation and an active user on the Direct Connection hacking forum. [DOJ]
Snowflake hacker pleads guilty: A Canadian man has pleaded guilty to hacking cloud storage provider Snowflake in 2024. Connor Riley Moucka admitted to hacking the cloud provider and then stealing the data of more than 150 of its customers. Together with other accomplices he then tried to extort the customers, threatening to release their data online. He allegedly made $2.5 million through ransom payments. [DOJ]
Fuel gauges left the internet: In the midst of a major Iranian hacking campaign targeting US water systems, and months after Iranian hackers breached fuel gauges at US gas stations, there are still thousands of automatic fuel gauge systems exposed on the internet, with 2,300+ of those in the US alone. [Bitsight]

macOS ClickFix campaign: There's a massive ClickFix campaign targeting macOS users, designed to trick users into copy-pasting malicious commands in their macOS terminals. The commands are designed to install macOS-specific infostealers like AMOS and MacSync. [Microsoft]

The Gentlemen affiliate: Hunt researchers have published a profile on an affiliate for The Gentlemen ransomware that's deploying the EtherRAT on hacked orgs. The interesting thing is that EtherRAT has been a very well-known RAT used in DPRK espionage operations. Take this information as you will. wink-wink [Hunt Intelligence]
Flooding Dropper campaign: A threat actor has uploaded almost 850 malicious packages on the npm package repository. The attacker is publishing the packages on new accounts, rather than compromising existing projects. Sonatype says automation is heavily involved in the attack. [Sonatype // OpenSourceMalware]
Threat actor AI usage explodes: Cisco Talos is seeing an increase in AI usage by threat actors, but says that impact on organizations typically varies based on the attacker's skill level. [Cisco Talos]
FunFoneFarm phone farm kit: Criminal operations are using a software package called FunFoneFarm to manage phone farms and cyber scam operations. The kit acts as a layer between phone racks, cloud servers, and the operators who interact with victims. FunFoneFarm supports physical phones, cloud-hosted handsets, and AI agents that start conversations inside dating apps. [Human Security]

Papyrus profile: And speaking of phone scams, Integral Ad Science looks at Papyrus, a new mobile ad fraud scheme built around a network of malicious Android apps. [IAS]
Wrench attacks reach $30m this year: Cryptocurrency owners have lost more than $30 million this year to wrench attacks. The term refers to the use of threats of violence to make owners hand over their crypto to criminal groups. According to blockchain analysis firm Chainalysis, at this pace, 2026 is set to become the single-worst year for wrench attacks on record. [Chainalysis]

Malware technical reports
CanOworms botnet: Security researchers have discovered a new botnet that's being used to disguise the origin of cybercrime and state espionage operations. According to SecurityScorecard, the new CanOworms botnet runs on more than 630 cloud servers. Among the APT crowd are North Korean and Chinese state hackers. [SecurityScorecard]
Khunt: Huntress researchers look at Khunt, a new post-exploitation toolkit installed on systems after exploiting SQL injections against Oracle databases. [Huntress]
ChainDrop: There's new reports on the ChainDrop npm worm that hit over 440 npm packages earlier this week. In the meantime, the worm has been classified as a Shai-Hulud variant, which is to be expected since the code has been released online a few months back. [Checkmarx // Elastic // Microsoft // Snyk]
Sponsor section
Learn how Permiso Security discovers, protects, and defends all of your human, non-human and AI identities in this short intro.
APTs, cyber-espionage, and info-ops
Xctdoor and CRAT: AhnLab looks at the similarities between Xctdoor and CRAT, two malware families typically used by DPRK APTs. Most of the time, it's been the Andariel group who has used them. [AhnLab]
Inside report from North Korea's hacking ops: Security researcher Vangelis Stykas spent 22 months inside the servers of a North Korean hacking group, during which time he documented breaches at more than 1,600 orgs across 57 countries. [WIRED]
APT-C-24: Indian APT group APT-C-24 (SideWinder, Rattlesnake) has been seen abusing the Microsoft ClickOnce app to deploy backdoors. [Qihoo 360]
Russian disinfo targets France's next presidential election: Russian disinformation groups have already started targeting pro-EU and pro-Ukrainian political figures ahead of next year's French presidential election. Targets so far have included former Prime Minister Gabriel Attal and former Prime Minister Édouard Philippe. Most of the content is fabricated news stories about alleged corruption and made-up political stances and projects. [Politco Europe]
On Aug 05, Matryoshka switched from Gabriel Attal to Édouard Philippe, another pro-Ukraine presidential candidate. 10 fake videos in French, English and German all followed this pattern: "Yet another [fabricated] scandal will complicate his position in the presidential race."
— antibot4navalny | bot blocker | блокировщик ботов (@antibot4navalny.bsky.social) August 6, 2026 at 6:31 AM
[image or embed]
Vulnerabilities, security research, and bug bounty
Security updates: Adobe, Cisco, Jenkins, SonicWall, Tails, TP-Link.
Backdoor found in Zbtlink routers: Security researchers have found a hidden backdoor in Zbtlink routers. More than 20 models from the Chinese company contain the hidden access code. The routers reach out to a Chinese registered domain name every 35 seconds. If that domain presents the right code to the router, a reverse shell opens. A day after security firm VulnCheck disclosed the backdoor, Zbtlink said the so-called backdoor was just a remote support feature. The company also said it suspended the sale of all affected models. [VulnCheck // Reuters // Zbtlink]
KEV update: CISA has updated its KEV database (twice) with four vulnerabilities that are currently exploited in the wild. All are 2026 bugs and they target JetBrains TeamCity, IBM Langflow, N-able N-central, and Apache Tomcat.
Apple's Private Relay leaks your IP address: A WebKit bug can leak the real IP address of a user enrolled in Apple's Private Relay, an Apple feature designed to hide that address by funneling a user's traffic through Apple cloud servers. Obviously, this applies only to Safari users. [Mysk // TechCrunch]
WSUS apocalypse: At BlackHat this week, SpecterOps presented research on how Microsoft WSUS servers could be used to distribute malware to enterprise endpoints or move laterally within company networks. [SpecterOps #1 // SpecterOps #2 // SpecterOps #3 // SpecterOps #4]
TONTOU attack: AMD has released firmware patches to fix TONTOU, a new attack that bypasses Spectre security mitigations and leaks data from inside the CPU's memory. This only impacts Zen 2 CPUs. [MIT, PDF // BlackHat // AMD patches]
PromptSpy attack: Forcepoint researchers describe a new attack named PromptSpy that poisons an AI agent's long-term memory to respond to spy on their owners. [Forcepoint]
PleaseFix vulnerability class: AI security firm Zenity says it developed a new class of zero-click attacks against AI browser agents. They call this new class of attacks PleaseFix. [BlackHat // Press release // Zenity blog post #1 // Zenity blog post #2 // Zenity blog post #3 // Zenity blog post #4]
"Zenity Labs responsibly disclosed its findings to Anthropic, Perplexity, Google, Microsoft and OpenAI ahead of the presentation. Some issued patches, while others declined, characterizing the findings as intended functionality. The mixed response underscores an unresolved gap in how the industry approaches agentic browser security."
Infosec industry
Threat/trend reports: Central Bank of Russia, Chainalysis, Curator, Expel, F6, and Incogni have recently published reports and summaries covering various emerging threats and industry trends.
Bugtraq announces return: Cybersecurity mailing list Bugtraq, established back in 1993, has announced its return after it went dark back in 2021. [SecurityFocus // ZDNet]
Acquisition news: Infoblox has entered an agreement to acquire Kentik. [Kentik // Infoblox]
New tool—dotdotslash_bot: A security researcher going by NyanBinary has created a Mastodon bot that tracks path and directory traversal bugs caused by the infamous ../ character sequence. I'm sad to inform you that there are a lot of those, still, in the year of the lord 2026.
Risky Business podcasts
In this edition of Seriously Risky Business, Tom Uren and James Wilson talk about North Korea losing control over some of its hacker workforce. Expect some tightening of controls and oversight, and perhaps even a reduction in the country's ransomware operations.