Open your notes and list 5 checkable work tasks (not vibes). Pick two models or tools you have access to, and score each task Pass/Fail for model A and B — 10 minutes total, a quick smoke test. You're done when you can name which model wins on at least 3 tasks, or call it a tie that needs a deeper look.
How-to · SEL · SEL.4
Multipolar frontier — judge models by your eval, not the logo
When many labs ship near the frontier at once, loyalty to one brand is weak strategy — pick by checkable evals on your tasks (tools, multi-step, cost).
Worked example
swyx points to a multipolar frontier: lots of strong models all competing at once. That's good news if you're choosing AI tools at work — you can switch, route, or A/B test on your real tasks. Conehead rule: write down 5 real tasks from your job and score two models on them before you reorganize the whole company around one vendor. You still own send/deploy and spend caps.