Discussion about this post

User's avatar
Morgan Rhinehart's avatar

This was a nice read! Thanks for sharing. I am glad you all did not use the Claude Plan Mode to make it fair for Gemini and ChatGPT. I do wonder if that was turned on if Claude would do better at completing the task with less mistakes .

Peter Ashby Smith's avatar

What's interesting isn't so much that the tools produced different outputs, but the different philosophies of uncertainty.

Some preferred to guess. Others preferred to preserve the ambiguity. Others explained more than executed.

That suggests we're no longer just choosing AI models, but choosing how much uncertainty we're comfortable delegating to it.

1 more comment...

No posts

Ready for more?