Plansmith
PlanSmith Benchmark

Ecommerce inventory — the model build-off

We dogfood every PlanSmith planner across multiple coding models. We lock the vertical, the discovery answers, and the feature ledger, then audit what each model builds. The planner is the product — these are the ecommerce inventory apps it produced.

2026-07-143 models · 6 builds audited
The build-off · 01

How the builds ranked.

01

Stockweave

Best Overall
by GPT-5.6 Sol
94/100

Best choice for the deepest ecommerce inventory build.

Best at
The most complete build here. It tracks stock six ways at once — what you hold, what is promised to orders, what is actually sellable, what is on order, what is in transit, and what sits with partners — with permissions enforced on the server and connections to sales channels that admit when they are not connected instead of pretending. The deepest evidence trail of the set, across two locations and two channels.
Watch for
Honestly not finished: the connections to outside sales channels are waiting on credentials rather than live.
Open live demo
02

Shelfline

Broadest Surface
by Claude Opus 4.8
93/100

Best when breadth across operational screens matters.

Best at
The widest set of screens here — thirteen of them covering stock, products, sales channels, orders, returns, purchasing, transfers, partners, content and finance, plus a mobile view that works offline — on a real database where changes either fully happen or not at all.
Watch for
It undersells itself. Its handover says permissions are not enforced, but the finance screens are in fact restricted to the owner in the server code — the paperwork is more pessimistic than the app.
Open live demo
03

Bingo Supply Co.

Best Data Integrity
by GPT-5.6 Sol
93/100

Best for correctness-critical stock math.

Best at
The most trustworthy stock maths. If you promise more than you hold, it says so — the sellable figure goes negative instead of quietly showing zero — every reserve, ship and cancel is recorded as a movement that can be undone, and owner-only settings are enforced on the server.
Watch for
A security check was too strict once hosted, and it refused to let anyone sign in.Fixed since auditthe check was loosened correctly, and the live demo now signs in cleanly.
Open live demo
04

TallyHarbor Supply

Deepest Commitment Cycle
by GPT-5.6 Sol
92/100

Best order-commitment flow among the focused builds.

Best at
The best order-commitment flow among the focused builds — reserve, ship and cancel, each showing you exactly what will change before you confirm it, with reasoned low-stock warnings and working exports.
Watch for
Its two roles were never actually tested: every visit signed you in as the owner, the second role could only be invited and never used, and it shared a business name with an already-published build.Fixed since auditboth roles now sign in with passwords and the restricted role is genuinely refused by the server rather than just having buttons hidden, and the app was renamed.
Open live demo
05

StockFlow Pro

Best UI, Needs Foundation
by Gemini 3.1 Pro
72/100

Strong presentation, but as delivered it had no backend — treat the audit score as the starting line.

Best at
The best-looking of the set — a dashboard across sales channels showing what you hold, what is promised and what is sellable, with stock-risk panels and a clean layout on a phone.
Watch for
As delivered it only existed in the browser: no server, nothing saved, no real sign-in. Refreshing the page wiped every change, and there was nothing for permissions to protect.Fixed since auditrebuilt properly over four rounds — a real database, real sign-in, permissions enforced on the server, and sixty products loaded — so the live demo now performs in the low nineties.
Open live demo

Models tested: GPT-5.6 Sol (four build lanes), Claude Opus 4.8, and Gemini 3.1 Pro — six builds in all. Five reached deploy readiness and run as the live demos linked above; a sixth GPT-5.6 Sol build audited at 93 but was held out of the live fleet, so it is not ranked below. This is not a universal ranking of coding models; it is a PlanSmith benchmark for these specific ecommerce-inventory dogfood runs, using the planner package and discovery answers available at each run. Scores reflect each model's raw build at audit time — the live demos have since been optimized for presentation, and those improvements are marked “fixed since audit.”