The fall glut makes personal benchmarking urgent
Opus 5.5, Astra 6, Sol 6.1, faster tiers like Sonnet 5.5, and a wave of open-weight models have all landed in a few weeks — each with benchmark tables that say nothing about how a model fits your personal AI stack. The fix is a system for testing new models against your own work.