ConeheadAI · leftover lesson

SNK3-088 · Curriculum How-to

Benchmarks lag real tool work — test the skill you need

L2–L5 · Capable–Builder (bootstrap — system rebalances) · ~4m

generalevalsagentsbenchmarksSEL · J6 · S3C6.1C3.2

Look for this insight

Cute single-shot demos can detach from real workplace AI use: multi-turn tool calling on your actual tasks. Measure what your job needs, not the viral benchmark.

Lesson hook

Simon Willison on Kimi: there's still something to learn from the pelican benchmark, even as it matters less than the thing that really counts — longer, tool-using conversations. Our translation: don't pick a chat or agent product based on a pretty demo alone. Re-run one real work task after any model or mode change. You still own publish and send.

Do this now

Open your notes and write two columns: Demo I have seen | Skill my job needs. Fill in 3 rows (for example, a pretty image versus multi-tool research). You're done when you've named the single job-skill test you'll re-run after any tool update, in 12 words or less.

Completion check: names_job_skill_to_test

Open snack link →

Staff review Sign in

Related by competency

All snacks · Take Question 1 · Team training