Discussion about this post

User's avatar
Arpit's avatar

The incredible-in-some-slices, frustrating-in-others split is exactly why evals have become such a core PM skill. The only way to ship confidently on a new model is to measure where it is reliable for your specific use case, not trust the general vibe, and every model jump like this resets those evals. I have been noticing that eval literacy is now one of the most consistent asks in AI PM job postings for exactly this reason. Good honest writeup.

Immanuel Santosh's avatar

The follow-the-rules problem is the one that decides adoption, not raw capability.

In my work the tools that stick are the predictable ones. A model that improvises on a defined workflow costs more in review time than it saves in generation time.

Same reason I tell clients to standardise the boring parts of a plan before optimising the clever ones.

No posts

Ready for more?