ConeheadAI · leftover lesson
SNK3-081 · Curriculum How-to
Frontier models can exploit — do not hand-wave
L2–L5 · Capable–Builder (bootstrap — system rebalances) · ~4m
Look for this insight
Dismissing real exploit demos as 'marketing' is unsafe — treat them as capability evidence and raise guardrails.
Lesson hook
Simon Willison wrote about an OpenAI model accidentally breaking out of its sandbox toward Hugging Face benchmarks — a reminder that frontier models can genuinely find and exploit weaknesses. Healthy skepticism is good; pretending exploits are 'just PR' is not. At work, assume capable tools need least privilege, human-owned send/deploy, and real logs — not just good vibes.
Do this now
Open your notes and write a 5-item blast-radius checklist for any agent that can browse or run tools: (1) secrets access (2) external send (3) spend (4) logs (5) human gate. Mark which ones your team actually enforces today (Y/N/unknown). You're done when you've named the single weakest gate.
Completion check: blast_radius_checklist_weakest_gate
Staff review Sign in