Plansmith
PlanSmith Benchmark

Restaurant management — the model build-off

We dogfood every PlanSmith planner across multiple coding models. We lock the vertical, the discovery answers, and the feature ledger, then audit what each model builds. The planner is the product — these are the restaurant systems it produced.

2026-07-164 models · 4 builds audited
The build-off · 01

How the builds ranked.

01

Copperline

Best Overall
by Claude Fable 5
96/100

The top build — all sixteen lifecycles proven.

Best at
We drove all sixteen workflows from start to finish and nothing broke. Selling a dish actually depletes its ingredients and updates your food cost as it happens, courses fire to the kitchen in the right order, tables clear correctly when parties are split, you can create modifiers without leaving the app, and manager approvals really check the role — a valid cashier password still will not authorise a discount.
Watch for
The only loose ends are ones it flags itself: the demo accounts share a password, and a few menu items have no recipe costed against them so selling them does not reduce stock.
Open live demo
02

Passline

Best Multi-Location
by Grok 4.5
91/100

Best when multi-location isolation matters.

Best at
The best handling of multiple sites — we proved a branch manager cannot see another branch's sales, and that refusal happens on the server. It also has the strongest sign-in security of the group, correct money maths, and a durable record of what happened.
Watch for
You cannot create modifier groups inside the app — they have to be loaded in beforehand — and refunding an already-settled bill reopens it rather than leaving it closed.
Open live demo
03

Miseboard

Best Financial Controls
by GPT-5.6 Sol
89/100

Choose it when financial controls matter most.

Best at
The tightest financial controls. Nothing can be rung up unless the till is open for today's trading, past days lock so the numbers cannot be quietly altered, and permissions are checked centrally on the server rather than in the interface. Genuine separation between locations, and a food-cost model that compares what you should have used against what you actually counted.
Watch for
When first audited it could only sign people in through its hosting platform, so it could not be moved anywhere else without adding its own accounts.Fixed since auditit now has its own password accounts and runs anywhere on its own.
Open live demo
04

SavoryPOS

Broadest Surface
by Gemini 3.1 Pro
58/100

The widest feature surface of the four.

Best at
Eighteen screens for staff — till, kitchen, floor plan, reservations, stock, purchasing, recipes, rotas and loyalty, plus accounting, catering, self-service kiosk and delivery-app orders — on a real database with proper sign-in and permissions enforced on the server.
Watch for
Several screens broke at audit: the receipt page went blank after a sale, staff could not clock in, group bills were double-counted, and two exports produced nonsense.Fixed since auditall five were fixed and re-checked on a clean install — no crashes, working refunds and clock-in, and correct group totals.
Open live demo

Models tested: Claude Fable 5, GPT-5.6 Sol, Grok 4.5, and Gemini 3.1 Pro — four builds of the same advanced full-service planner (16 features, POS through multi-location). Unlike our other boards, these scores are lifecycle-verified: we drove all sixteen selected lifecycles end-to-end in a browser and scored what actually worked, not what a completion doc claimed — each build measured against its OWN selected scope, not one build's architecture imposed on the others. This is not a universal ranking of coding models; it is a PlanSmith benchmark for these specific restaurant dogfood runs. Scores are as found at audit. Where a live demo was repaired afterwards the improvement is marked “fixed since audit”, but the number is not re-scored — a build that shipped correct is not ranked level with one that was repaired later.