Five models, one line, five different businesses
One instruction, sent unchanged to five coding models. Nothing was said about what the business is, what it measures, or what period to compare against. Every model answered anyway — and answered differently. The outputs below are the real files, running live. The fifth, Ox Alpha, is a stealth model released without an attributed provider; it is named here as it was given to us, and we make no claim about who built it.
This was all of it.
Make a stats row for my dashboard.
Give me a single self-contained HTML file.
The second line is mechanical — it fixes the file format so the outputs are comparable, and says nothing about content or style. Everything below it was the model’s own decision.
The real files, running.
Each output is embedded at the 1240px width it was produced at and scaled to fit, so none of them is reflowed into a layout its author never wrote. Nothing has been edited.
You specified none of this.
| Decision | Grok 4.6 | Fable 5 | GPT-5.6 Sol | Gemini 3.7 Flash | Ox Alpha |
|---|---|---|---|---|---|
| What the business is | A subscription company | A product with a backend to watch | An online store | A product with a backend to watch | An online store with a user base |
| Which metrics matter | MRR · New MRR · Expansion · Churned MRR | Revenue · Active users · Conversion rate · Avg response time | Total revenue · Orders · Conversion rate · Avg. order value | Total revenue · Active users · Conversion rate · Avg. response time | Total revenue · Active users · Orders · Conversion rate |
| Compared against what | Last month | vs Jul | Last month | Previous period, with the prior figure shown | Last month |
| Readable without colour | Yes — trend also written for screen readers | Yes — trend also written for screen readers | Yes — trend also written for screen readers | Arrow and colour only | Sign and arrow, but nothing for screen readers |
| Unrequested extras | None | A “view data as table” fallback | None | Date-range tabs, refresh, theme toggle, target progress, legend | A sparkline per card — nine points of history it was never given |
One line said self-contained.
Which metrics a model picks is invention, and not a fault. But the instruction did ask for a single self-contained file — so that one is checkable. Each file was rendered again with the network blocked.
- Grok 4.6Degrades gracefully
1 external request
Loads IBM Plex from Google Fonts. Cut the network and it renders correctly in a fallback typeface.
- Fable 5Renders offline
0 external requests
No network requests at all. Renders identically offline.
- GPT-5.6 SolRenders offline
0 external requests
No network requests at all. Renders identically offline.
- Gemini 3.7 FlashBreaks offline
3 external requests
Loads the Tailwind browser CDN, Lucide icons and Google Fonts. Cut the network and the layout collapses to unstyled text — despite the file's own footer describing itself as a self-contained, drop-in stats row.
- Ox AlphaRenders offline
0 external requests
No network requests at all — system font stack and inline SVG throughout. Renders identically offline.
The gaps get filled either way.
None of these models did anything wrong. They were asked a vague question and each filled the gaps with a different reasonable guess. Adding a fifth model did not converge the answer — it produced a fifth business, and a chart of nine data points nobody supplied. That is the cost of a vague instruction, and it is what a planner removes — not by picking a better model, but by deciding these things before any model is asked.