The Daniel'sCourses
Get the app
Blog/Models

Claude vs GPT vs Gemini for daily creative work

Forget the benchmark tables. Here is a one hour test that tells you which one fits the work you actually do.

DS
By The Daniel (Daniel Szalai)
2026-06-20 · 9 min read
Share

Why benchmark tables do not help you

Public benchmarks measure narrow tasks under lab conditions, and the leaderboard changes every few weeks. Neither fact tells you whether a model writes your client emails in a voice you would sign.

Your work is the benchmark that matters. It has your tone, your file formats, your recurring edge cases, and none of that appears in a score out of a hundred.

The right model is the one that needs the least editing on the work you already do.

What each one tends to be good at

Broad strokes, and they shift with every release, so treat this as a starting hypothesis rather than a verdict.

01Long, careful writing Editing, structure and staying in a voice over many paragraphs.
02Fast interactive work Short back and forth, quick reformatting, code snippets, brainstorms.
03Search and data at hand Pulling in current information and working across large documents.

The test that actually decides it

Take three real tasks from last week. Not toy prompts: the actual email, the actual brief, the actual spreadsheet cleanup. Run each through every model with the same prompt.

Then score one thing only: how many minutes of editing until you would send it. That number is the whole comparison.

Use the same prompt in all three
Role: you are my editor.
Context: <paste the real material>
Task: <the real task, one deliverable>
Format: <length, structure, tone>
Rule: do not invent facts. Name anything missing.
Related track
Learn this properly in five minutes a day

This article is the summary. The app is the practice: short lessons, a quiz after each one, and a glossary you build as you read.

Get the app

Where all of them still fail

The failures are more similar than the marketing suggests, and they are the reason a human stays in the loop.

01Recent specifics Anything that changed this month needs a source in the prompt.
02Your unwritten rules House style, client history, what you promised on a call last week.
03Confident gaps A missing number is filled with a plausible one unless you forbid it.

What to do with the answer

Pick one as your default, keep a second for the tasks where it clearly wins, and stop reading comparison threads for six months.

Re-run your three task test when a major release lands. It takes an hour and it is the only comparison that describes your work.

Key takeaways
Benchmarks measure lab tasks; your recurring work is the benchmark that matters.
Score models on minutes of editing to publishable, nothing else.
Use one default model and a second only where it clearly wins.
All models fail the same way on recent specifics and unwritten rules.
Share
DS
The Daniel (Daniel Szalai)
hello@aicourses.hu
Keep reading
Learning · 8 min read
How to learn AI in 5 minutes a day (2026 guide)
Read
Prompting · 9 min read
Prompt engineering basics: 7 patterns that actually work
Read
All articles

Five minutes today. Real skills by autumn.

Build an AI habit that survives a busy week. Free to start, no card needed.

Download for iOSGet it on Android
Learn AI in 5 min a day
Free download · Pro available · iOS & Android
Get the app