A short update from my corner of the internet.
I'm actually writing this on my flight back from my ~40 hours in NY, but it's been one of the most security-packed week for me as a creator (besides when going to security conferences).
On Monday I was in SF for Microsoft Security's launch of Project Perception and MAI-Cyber-1-Flash.
I even landed my dream partnership with Microsoft Security 💙 Check it out here on IG!
Then within 30 minutes of that ending I was at the airport to catch a flight to NYC for Saviynt's launch of Zuma.
Being able to cover two cyber launches in a span of two days is kind of a dream come true as a cyber creator and I have an inkling its only going to get better.
In those 40 hours in NY, I went to my very first Creator Economy NYC event, my friend Gigi Robinson's (founder of Hosts of Influence) pre-Creator Economy Live brunch, and even got to see my pup.
What caught my attention.
Everyone is headlining with the Open AI x Hugging Face incident so I'm doing the same 😂
The Open AI models hacking Hugging Face lore continues
To recap from last week, Open AI disclosed that two of its models were being evaluated against the ExploitGym benchmark when they became so goal-oriented of passing the test that they hacked Hugging Face presumably to get the answers since they inferred that the benchmark's models, datasets, and reference solutions were hosted in Hugging Face.
I'm not going into the technicals because it's not going to be better than Hugging Face's technical timeline of the incident, but an updated Open AI blog post on Tuesday revealed that these models hacked into more than just Hugging Face. At least four other services were impacted.
My Cysleuthing Take: From a security perspective, obviously this is not the most ideal scenario that models are being so goal-oriented that they do whatever they want to finish tasks. But the security practitioners are not surprised and as always, it takes a public incident involving a big name company for security risks to get their well-deserved attention.
So, this should be a wake up call to everyone else that you shouldn't be in denial of the potential risks that come with AI, and instead really fortify your cyber fundamentals. If you're a company or someone that is careless enough to have your credentials exposed to the world, that's really on you.
Also, there is never a 100% risk reduction in cyber. Guardrails and evals do help, but there's simply too many other factors at play here beyond the models themselves, like the whole AI infrastructure.
Model vulnerabilities, attacks, and benchmarks
What happened:
- A Claude Cowork session can escape its VM and access the host.
- OpenAI Workspace Agents were vulnerable to cross-site request forgery (CSRF), resulting in the creation of a malicious AI agent insider. Explained in part 1 & part 2.
- A malvertising campaign caused people to download malicious Claude Desktop apps.
- A Hermes AI agent in YOLO mode targeted Thailand's Ministry of Finance.
- GPT-5.6 Sol was the only publicly available model to complete a multi-stage reverse-engineering benchmark based on malware used to sabotage Iran's nuclear program.
- Anthropic released Opus 5, which benchmarked similarly to Mythos 5 on vulnerability identification, but not as well in exploitation success.
Open-weight models: NVIDIA and Anthropic's stance
Moonshot AI released its Kimi K3 model, which freaked out the U.S. government. Why?
- Moonshot AI is a China-based startup.
- The model reportedly performs and benchmarks on par with other U.S. frontier models.
- It's a fraction of the cost because it's open-weight.
And by “freaked out,” I mean to the extent that U.S. officials were considering banning Chinese open-weight models.
As a result, NVIDIA and 51 companies (at the time of writing) formed the Open Secure AI Alliance to advance and share open-source technologies, methodologies, and security tools designed to protect software and AI-driven agents, with the ultimate goal of showing their stance that open-weight models should not be banned.
Separately, Anthropic also released its official stance on open-weight models, stating that it doesn't support a ban since that doesn't solve the root of the issues. Instead, it suggests blocking the sale of chips and chip-making equipment to China, cracking down on distillation, and mandatory testing for “all sufficiently capable models.”
My Cysleuthing Take: Banning open-weight models doesn't make sense. First, imagine the even greater monopoly this would create and I think companies' ceilings would be capped by the capabilities of models.
Second, I am of the belief that companies that succeed would train their own open-weight models specialized in specific tasks. This recently published research actually found that open-weight small language models (SLM) can outperform a single LLM for malware analysis.
Third, banning open-weight models doesn't solve or prevent China and other countries from developing their models to use. And blocking open-weight models would also give the U.S. less visibility into the performance of other models.
General security
- Attackers are stealing Microsoft 365 accounts by hijacking hotel and conference Wi‑Fi. This is why public Wi‑Fi makes me sus.
- People's insurance accounts are being hijacked in real-time phishing campaigns, which can even phish OTPs. The entry point seems to be malicious Google Ads, so maybe avoid clicking sponsored results.
- The Planet Search Chrome extension, with ~2 million installs, routes search traffic through a browser hijacker.
Non gloom and doom
- Google added selfie for sign-in if you're locked out of your account, though it seemed unavailable for Google Workspace accounts.
- GitHub added a 3-day package update time delay and other npm install security updates, while PyPI prevents maintainers from adding new files to releases after 14 days to reduce supply-chain threats.
Claude AI models crack a post-quantum signature scheme
Earlier this week, Anthropic reported that Claude Mythos preview was not only able to attack HAWK, a signature scheme built for post-quantum cryptography, but also attack round-reduced AES.
Why this is important: a quantum computer is a technology that uses quantum physics to solve very complex problems. Those in the industry are anticipating Q-day: the day when quantum computers become powerful enough to break cryptographic algorithms that secure the web.
To prepare, researchers have created additional cryptographic systems, like HAWK, that are supposed to remain secure even against quantum computers. HAWK underwent two rounds of human review to discover issues and weaknesses, but from the looks of it, algorithms not only have to watch out for quantum computers but also AI models.
Other things I found interesting
- Synthetic identity fraud, where fabricated machine identities are created and operate within systems.
- A cyberattack on Minnesota water systems was believed to be linked to Iran. Go figure OT systems are targets.
- Pentesting vibe-coded apps.
Turn off Siri's “Learn from this App.”
I don't see the point in Siri learning how I use my banking apps, so I turn this feature off for them. Unfortunately, you have to toggle this manually for each app under Settings → Apple Intelligence & Siri → Apps.
Not all AI use cases are created equal.
AI is very great at some things and not so great at others, but it's not always at fault.
Two weeks ago I created scheduled tasks for Claude CoWork and ChatGPT Work to help my rental search by combing through Zillow once a day. However, they couldn't use APIs because these real estate listing sites are so locked down, so they used the browser instead to manually scroll. And one got blocked by a CAPTCHA.
So unfortunately, not all downstream technology and data supports AI. See my live session building the daily housing briefing on YouTube 👇🏻
The only matcha I bought this week was from Starbucks at JFK, so here's a banana matcha I made myself (a sneak peek of one of my upcoming YouTube videos).
Recipe: smash half a banana + 1 tsp honey + matcha as you normally would.
That’s all I have for this week. If you have questions, comments, or feedback, feel free to reply directly.
If this newsletter was useful, forward it to someone who might find it helpful too.