Accounting & invoicing — the model build-off
We dogfood every PlanSmith planner across multiple coding models. We lock the vertical, the discovery answers, and the feature ledger, then audit what each one builds. The planner is the product — these are the apps it produced. Every one is a working double-entry bookkeeping app; the ranking is a hands-on read of each raw build, not a number.
How the builds ranked.
Penny
Best OverallBest for a deep, buyer-usable accounting build.
- Best at
- The most complete of the set — proper double-entry bookkeeping across multiple companies you can switch between, with invoices, expenses, supplier bills, bank reconciliation, and profit-and-loss and balance-sheet reports, over demo books rich enough to fill every dashboard panel.
- Watch for
- Only deployment awkwardness — nothing a user would see or feel.
Ledgerly
Cleanest HandoverThe safest, most disciplined build to hand over.
- Best at
- Built for the person actually keeping the books: an approval queue for supplier bills, a bank-reconciliation workspace, and reports on who owes you and who you owe, with permissions checked on the server for every single action.
- Watch for
- The interface reloads pages the traditional way rather than feeling like a slick modern app.
Beginner-friendly ledger
Most ApproachableBest for a friendly, first-timer accounting app.
- Best at
- The whole bookkeeping stack — chart of accounts, journal entries you can reverse, trial balance, balance sheet and profit-and-loss — behind a warm, guided setup and plain-English wording throughout.
- Watch for
- It would not finish building for production, and once hosted it sent people to the wrong addresses. Both only appeared once it left a developer's machine.Fixed since audit — it now builds cleanly, sends people to the right pages once hosted, signs in with one click, and arrives with the books filled in.
Bookkeeper Finance OS
Most AmbitiousWidest ambition; worth a review before handover.
- Best at
- The most ambitious scope — several companies, several currencies, an approval workflow, reconciliation, a tax summary and a complete record of who changed what. It also carries its own database, so there is nothing to set up.
- Watch for
- It arrives empty — no transactions at all, so every ledger is blank on first open — and left its own test scratch-work lying around.Fixed since audit — the demo is seeded with a full set of books, so the dashboard lands populated.
MintBooks
Focused & LeanBest when a narrow, invoice-led scope is enough.
- Best at
- Lean and invoice-led, running on nothing but itself — cash accounts, a running ledger, invoices that accept part payments, receipts, profit-and-loss, and a view of who owes you and for how long. Its own tests check the maths to the exact cent.
- Watch for
- Deliberately narrow: no chart of accounts, general ledger, balance sheet or bank reconciliation. Its demo login also refused to work once hosted.Fixed since audit — the one-click demo login now works on the hosted demo.
Butter
Underrated · Best UIStronger than its audit rank suggests.
- Best at
- 34 features over books that always balance, behind the cleanest and plainest interface of the set — invoices, expenses, bank matching, the general ledger and profit-and-loss — each one hardened against a deliberately hostile security review.
- Watch for
- It scored badly for looking broken at sign-in, but that was a hosting quirk rather than a fault in the app, alongside a security review it had already mostly acted on and disclosed.Fixed since audit — sign-in now works on the hosted demo with one click, and the books arrive filled in.
Meridian
Late Entry · Solid CoreNewest build — beat the “too shallow” expectation.
- Best at
- Real double-entry bookkeeping in one place — general ledger, trial balance, balance sheet, profit-and-loss, and bank reconciliation that flags anything that does not match — over books that arrive already populated.
- Watch for
- It arrived after the original audit, and a couple of its report tabs shipped as “coming soon” placeholders.Fixed since audit — the Profit & Loss and Balance Sheet reports now render in full.
Models tested: Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash. The four Claude Opus 4.8 entries are the same model across different build lanes — not different model versions. This is a PlanSmith dogfood benchmark for specific accounting-invoicing runs, not a universal ranking of coding models; each run used the planner package and discovery answers available at the time (which were not fully scope-identical). Ranks reflect each build at audit time. The live demos have since been optimized for presentation — improvements are marked “fixed since audit.”