
Every ERP passes a feature comparison. That is what makes feature comparisons useless.
Ask 4 vendors whether they support multi-entity consolidation, multi-currency, approval workflows and audit trails, and you will get 4 yeses and 4 demos that prove it. The word covers a system where consolidation is a property of the ledger and a system where it is a monthly job with a nicer interface, and the spreadsheet cannot tell them apart.
The criteria that separate systems at this size are mostly about what happens after the demo, and they can be asked directly.
people in finance is the realistic size of the team running most of these projects
systems at Oper Credits, which is the shape of stack most evaluations are trying to fix
countries supported out of the box, from US GAAP to UK MTD VAT
Does it scale 10 times
The best framing of this decision came from a finance lead in the middle of it. All Gravy's Sebastian Sandorff Jacobsen describes the standing test as whether the foundation scales 10 times, and the failure mode as adding a tool per problem.
"We always have to look at the foundation and ask whether it scales 10x. We kept adding tools as we added entities, and each of them worked to a point. Orchestrating a close across all of them was the problem."
Sebastian Sandorff Jacobsen, Head of Finance and Ops, All Gravy
Applied to a shortlist, that turns into 1 question per candidate: what does the fifth entity cost. If the answer is a project, the system does not scale 10 times, whatever the feature grid says.
Who changes it after go-live
The most expensive difference between systems at this size, and the one demos never show.
Ask what happens when the business needs a new dimension, a changed approval threshold, a different revenue treatment or an extra entity 18 months from now. The answers separate cleanly into 3 groups: an admin does it, a partner does it and bills for it, or it goes on a roadmap. The middle answer is the one that quietly determines the total cost of ownership, because a system that needs a consultant for every subsequent change is a system that stops changing.
Then ask the sharper version. When something the vendor does not do turns out to matter, can the finance team build it themselves? All Gravy built its own CPQ tool and connected it to Light because its contract terms were too bespoke for standard invoicing. Dreamdata's controller wrote enough of his own automation to reach 95% of bookkeeping. Neither had to wait for a roadmap.
Implementation, measured honestly
Ask 3 specific things rather than for a timeline.
Who does the work. A vendor quoting 12 weeks with a 3-person team from the customer is quoting a different project from a vendor quoting 12 weeks with 1 part-time person.
Whether history moves. Taking an opening balance is cheaper and leaves a permanent seam in the record. Moving every transaction, reconciled year by year against trial balances, costs weeks and removes a problem the company would otherwise carry for as long as it keeps the books.
Who answers at 11pm in week 6. The realistic version of implementation support, and the thing finance leads name afterwards when asked what actually mattered.
Weak criterion
Feature coverage
Every serious vendor answers yes. The word covers systems that behave completely differently once you own them.
Strong criterion
Cost of the next change
Who adds an entity, a dimension or an approval rule 18 months from now, and whether it needs a consultant.
Strongest
What a reference customer stopped doing
Not what they like. Which specific task disappeared, and what the number was before and after.
The reference call worth having
Most reference calls produce warmth and no information. Three questions fix that.
What did you stop doing entirely? Not what got faster. What is no longer on anybody's list. At Oper Credits the answer is a half week of consolidation. At Tillo it is triaging outbound invoices, 750 to 900 of them a month. At Famly it is a 20 hour revenue process that now takes 15 minutes.
What is still broken or missing? Every honest customer has an answer. Dreamdata invoices from Light without using the contracts module, because its terms need 13 and 14 month recurrences and it has no list prices, and that is being built with them. A reference who claims nothing is missing has not used the product hard enough to be useful.
How long until it felt normal? Different question from how long implementation took, and a more predictive one.
Two criteria most lists miss
Whether the auditor has seen it. Not vendor certifications, which are about the vendor. Whether a customer has completed a statutory audit on the system and what the auditor did. Customers at $500M ARR have completed audits on Light with the auditor pulling 100% of the journal population through the API rather than sampling.
Where the work lands. Approvals in a portal get routed around. Approvals in Slack get done on the metro. This looks like a minor UX preference and it decides whether the control operates at all.
The shortlist that survives those questions is usually shorter than the one that came out of the feature grid, and it is made of different systems.
Read how All Gravy moved 6 years of history, see what Light replaces, or book a demo.