Illustration of two agents facing the same task with very different tool-use habits

Whether a Model Saves Tokens Is a Habit—Not How Hard the Task Looks

On the same OpenHands + SWE-bench Verified setup, token efficiency gaps between frontier models stay large even when every model succeeds—or every model fails.