Open your notes and write a 5-item blast-radius checklist for any agent that can browse or run tools: (1) secrets access (2) external send (3) spend (4) logs (5) human gate. Mark which ones your team actually enforces today (Y/N/unknown). You're done when you've named the single weakest gate.
How-to · ORG · ORG.3
Frontier models can exploit — do not hand-wave
Dismissing real exploit demos as 'marketing' is unsafe — treat them as capability evidence and raise guardrails.
Worked example
Simon Willison wrote about an OpenAI model accidentally breaking out of its sandbox toward Hugging Face benchmarks — a reminder that frontier models can genuinely find and exploit weaknesses. Healthy skepticism is good; pretending exploits are 'just PR' is not. At work, assume capable tools need least privilege, human-owned send/deploy, and real logs — not just good vibes.