ConeheadAI · leftover lesson
SNK3-088 · Curriculum How-to
Benchmarks lag real tool work — test the skill you need
L2–L5 · Capable–Builder (bootstrap — system rebalances) · ~4m
Look for this insight
Cute single-shot demos can detach from real workplace AI use: multi-turn tool calling on your actual tasks. Measure what your job needs, not the viral benchmark.
Lesson hook
Simon Willison on Kimi: there's still something to learn from the pelican benchmark, even as it matters less than the thing that really counts — longer, tool-using conversations. Our translation: don't pick a chat or agent product based on a pretty demo alone. Re-run one real work task after any model or mode change. You still own publish and send.
Do this now
Open your notes and write two columns: Demo I have seen | Skill my job needs. Fill in 3 rows (for example, a pretty image versus multi-tool research). You're done when you've named the single job-skill test you'll re-run after any tool update, in 12 words or less.
Completion check: names_job_skill_to_test
Staff review Sign in