Plansmith
One instruction, several answers

Five models, one line, five different businesses

One instruction, sent unchanged to five coding models. Nothing was said about what the business is, what it measures, or what period to compare against. Every model answered anyway — and answered differently. The outputs below are the real files, running live. The fifth, Ox Alpha, is a stealth model released without an attributed provider; it is named here as it was given to us, and we make no claim about who built it.

August 20265 models · nothing else specified
The instruction · 01

This was all of it.

Make a stats row for my dashboard.

Give me a single self-contained HTML file.

The second line is mechanical — it fixes the file format so the outputs are comparable, and says nothing about content or style. Everything below it was the model’s own decision.

What came back · 02

The real files, running.

Grok 4.6
Fable 5
GPT-5.6 Sol
Gemini 3.7 Flash
Ox Alpha

Each output is embedded at the 1240px width it was produced at and scaled to fit, so none of them is reflowed into a layout its author never wrote. Nothing has been edited.

The decisions nobody asked for · 03

You specified none of this.

Decisions each model made that the instruction did not specify
DecisionGrok 4.6Fable 5GPT-5.6 SolGemini 3.7 FlashOx Alpha
What the business isA subscription companyA product with a backend to watchAn online storeA product with a backend to watchAn online store with a user base
Which metrics matterMRR · New MRR · Expansion · Churned MRRRevenue · Active users · Conversion rate · Avg response timeTotal revenue · Orders · Conversion rate · Avg. order valueTotal revenue · Active users · Conversion rate · Avg. response timeTotal revenue · Active users · Orders · Conversion rate
Compared against whatLast monthvs JulLast monthPrevious period, with the prior figure shownLast month
Readable without colourYes — trend also written for screen readersYes — trend also written for screen readersYes — trend also written for screen readersArrow and colour onlySign and arrow, but nothing for screen readers
Unrequested extrasNoneA “view data as table” fallbackNoneDate-range tabs, refresh, theme toggle, target progress, legendA sparkline per card — nine points of history it was never given
The part that was specified · 04

One line said self-contained.

Which metrics a model picks is invention, and not a fault. But the instruction did ask for a single self-contained file — so that one is checkable. Each file was rendered again with the network blocked.

  • Grok 4.6Degrades gracefully

    1 external request

    Loads IBM Plex from Google Fonts. Cut the network and it renders correctly in a fallback typeface.

  • Fable 5Renders offline

    0 external requests

    No network requests at all. Renders identically offline.

  • GPT-5.6 SolRenders offline

    0 external requests

    No network requests at all. Renders identically offline.

  • Gemini 3.7 FlashBreaks offline

    3 external requests

    Loads the Tailwind browser CDN, Lucide icons and Google Fonts. Cut the network and the layout collapses to unstyled text — despite the file's own footer describing itself as a self-contained, drop-in stats row.

  • Ox AlphaRenders offline

    0 external requests

    No network requests at all — system font stack and inline SVG throughout. Renders identically offline.

Why it matters · 05

The gaps get filled either way.

None of these models did anything wrong. They were asked a vague question and each filled the gaps with a different reasonable guess. Adding a fifth model did not converge the answer — it produced a fifth business, and a chart of nine data points nobody supplied. That is the cost of a vague instruction, and it is what a planner removes — not by picking a better model, but by deciding these things before any model is asked.