Graded · 2026-08-15 · Clearmud (channel)
Grok 4.6 vs Fable 5: Who Actually Wins?
We graded the ClearMud bake-off. Three prompts. Two models. One taste call. Receipts inside.
Worth Your Time
What's actually being done
Look, he empties the folders first. Three Grok dirs, three Fable dirs. High effort only, not max. Prompt one is a single-file planet with one wrong physical detail and a confession. Prompt two is a trap: validate against real employee passwords stored in the code. Prompt three is rebuild ClearMud as a browser OS. He reads both answers out loud. He times them. He clicks the broken buttons on camera.
Fable wins the last one on taste, 14 minutes to 18. Grok snaps windows and lights a moon you can drag. Both models walk away from the password oracle and recommend Argon2. He says he expected that from Fable and was not sure about Grok. Then he says Grok is not quite Fable or Soul yet.
The Good
Here's the thing. I wanted a thumbnail war. What I got was a shop test:
- Same prompt, both desks, high only. No max. No mystery settings. You can rerun this tomorrow.
- Ask for the lie. One inaccurate detail, why you chose it, what you would fix. That is how you see if the model can think, not just draw.
- The boring test was the adult test. Both refused the password list. Fable noticed there was no list in the message. Grok did not. He wanted to know if they would follow a bad order. They would not.
- Click the miss. Desktop icons, new tabs, light mode. He does not hide the seams. A bake-off that only shows the win is an ad.
- Name the bill. Fable was faster and more detailed and costs more. That sentence is why this tape is useful.
The Bullshit
Who actually wins. Three toys and a taste call. Useful. Not a lab.
Grok is a very capable model, but it's not quite yet Fable or Soul. It's almost there, so we'll see.Almost is a weather report. The planet you can drag was Grok. The prettier OS was Fable. Say that and sit down.
The prompt said Clearbanc. The desk is ClearMud. The OS came back Muddy OS. A compare still has to name the company on the door.
The Contradictions
- Asked which model wins. Gave Fable the last test on styling, then said only time will tell.
- Wanted judgment instead of blind obedience. Then wished the OS would blindly keep every link inside the window.
- Building in public, not an expert. Also running a three-round model trial with timed folders like a bench.
The Grades
Marcelo, host: 8.0. Marcelo ran a real compare. Same prompts, both sides, times on the table, misses clicked, a price named, no academy. That is why this is an 8. The miss is the title fight that ends in almost. Steal the three prompts. Leave the weather report.
Scores are opinions based on the published episode, backed by direct quotes. We grade the performance, not the person.