Appointment booking — the model build-off
We dogfood every PlanSmith planner across multiple coding models. We lock the vertical, the discovery answers, and the feature ledger, then audit what each one builds. The planner is the product — these are the apps it produced.
How the builds ranked.
salon & spa
Best OverallBest choice for deep appointment-booking builds.
- Best at
- Depth without gaps. A working calendar, bookings, payments, exports, staff accounts and an activity log — every screen the planner asked for, and it passed every automated check we ran.
- Watch for
- Not yet ready to take real money or send real messages: payments and notifications run in test mode and need live accounts connected before launch.
clinic & therapy
Best BalancedExcellent product depth with honest readiness.
- Best at
- A convincing clinic booking flow — patient intake forms, deposits, a cancellation policy, and separate tidy screens for the owner and each practitioner. It also went hunting for its own bugs and wrote down what it fixed.
- Watch for
- Covers less ground than the salon and fitness builds.
fitness studio
Feature-RichBest when breadth matters.
- Best at
- The most features of the set — classes with capacity limits, passes, waitlists, repeating bookings, payments and reports. It also proved that two people grabbing the last spot at once cannot both get it.
- Watch for
- A bigger, heavier app, and it hands over less tidily than the clinic build.
solo practitioner
Clean FocusedBest for smaller solo appointment apps.
- Best at
- Focused and actually finished. Booking, availability, payments, packages, waitlist, notifications and reports — each one carried through to the end rather than left half-done.
- Watch for
- Deliberately narrow, so there is less to show off than in the bigger builds.
care desk
Best Repair CandidateUsable spine; needs cleanup before handover.
- Best at
- A real public booking page, deposits and no-show fees, patient intake, and working spreadsheet and PDF exports — with permissions enforced on the server rather than merely hidden in the interface.
- Watch for
- It failed its own automated checks, the inbox overflowed its container on a phone, leftover test data appeared on customer-facing screens, and its handover notes contradicted each other about what was ready.Fixed since audit — it now builds cleanly, passes its checks, and has one consistent go-live checklist.
family clinic
Clean Narrow ScopeSafer for limited demos.
- Best at
- A tidy, coherent handover with a good-looking dashboard — achieved by staying inside a safe, limited scope.
- Watch for
- It sidestepped the hard parts: payments, intake, checkout and telehealth were all left out, so it proves less about how far the planner can push a model.
clinic front desk
Needs ReviewUse only with a strong audit and fix loop.
- Best at
- Quick to stand something up, with a working public booking page and a go-live checklist screen.
- Watch for
- Weak and inconsistent proof that logins, permissions and exports actually work, layout overflowing on a phone, and internal test labels left visible on customer-facing screens.Fixed since audit — login works and the demo surface has been cleaned up.
Models tested: Claude Opus 4.8, GLM 5.2, Composer 2.5, and Kimi 2.7. The four Claude Opus 4.8 entries are the same model across different planner runs and verticals — not different model versions. This is not a universal ranking of coding models; it is a PlanSmith benchmark for specific appointment-booking dogfood runs, using the planner package and discovery answers available at each run (which were not fully scope-identical). Scores reflect each model's raw build at audit time. The live demos have since been optimized for presentation; improvements are marked “fixed since audit.”