GPT, Claude, Gemini: My Honest Take After Months With All of Them
September 2025
I did not plan to become someone with strong opinions about AI models. It happened the way these things usually happen: gradually, through repeated daily use, until I noticed I was automatically reaching for a specific model for a specific task without consciously deciding to. At that point I figured I might as well make the opinions explicit.
I want to be clear about what this is and is not. This is not a benchmark. I have not run structured tests or compared outputs on standardised prompts. This is what I actually notice when I use these models for real work: writing, research, strategic analysis, multi-part briefs with conflicting constraints, and the kind of messy, context-heavy prompts that real work tends to produce.
Claude is my default. I know that sounds like I am picking a favourite, and I should acknowledge that I am writing this using Claude Code, which is a conflict of interest I will flag and then proceed past anyway. The specific reason I keep returning to Claude is that the output requires less editing. For writing and reasoning work, Claude gives me something closer to what I actually wanted the first time. When I give it a complex, multi-part prompt with conflicting constraints, it handles the harder parts instead of quietly dropping them. That happens enough with other models to be noticeable.
GPT-4o is the model I reach for when I need multimodal capabilities quickly. The voice mode is genuinely impressive, better than anything else I have tried. Image understanding is reliable. Where I get tripped up is a subtle overconfidence: GPT-4o gives you a well-structured, fluent answer that is sometimes subtly wrong in a way that takes a moment to catch. For anything where accuracy matters, I verify. I have learned to do this the hard way.
Gemini 2.0 is better than it gets credit for in most conversations I am part of. The long-context capability is real and actually works at scale, which is not always true of claimed context windows. The Google Workspace integration is genuinely useful if you live in that ecosystem, and I do not fully, so that benefit is partially lost on me. My honest feeling is that the output lacks a certain distinctiveness at the top of the range. Competent, always. Memorable, less often.
The honest answer to which model you should use is that it depends on what you are doing, and the answer will be different in six months. The capability gap between the top models has been narrowing all year. The model I was dismissing in March is one I am taking seriously again in September. What I would suggest: pick one as your primary for the kind of thinking you do most, keep one as a backup for the tasks where it genuinely excels, and revisit the comparison every quarter. Model loyalty is not a virtue.
Sources
- LMSYS Chatbot Arena: lmarena.ai
- Artificial Analysis: AI model benchmarks and comparison, 2025
- Anthropic, OpenAI, Google DeepMind: official model release documentation