Open your notes and write two columns: Demo I have seen | Skill my job needs. Fill in 3 rows (for example, a pretty image versus multi-tool research). You're done when you've named the single job-skill test you'll re-run after any tool update, in 12 words or less.
How-to · SEL · SEL.1
Benchmarks lag real tool work — test the skill you need
Cute single-shot demos can detach from real workplace AI use: multi-turn tool calling on your actual tasks. Measure what your job needs, not the viral benchmark.
Worked example
Simon Willison on Kimi: there's still something to learn from the pelican benchmark, even as it matters less than the thing that really counts — longer, tool-using conversations. Our translation: don't pick a chat or agent product based on a pretty demo alone. Re-run one real work task after any model or mode change. You still own publish and send.