How this was made

Methodology

Every claim in this dataset is binary, checkable, cited and graded. Here is exactly what each of those words means, and where the corpus is weak.

Claims, not ratings

There are no stars and no composite score. Each capability is a single yes/no proposition about a product, scored in one of five states:

Absence of evidence is `unknown`, never `no`. A no requires positive evidence of absence — the docs enumerate the alternatives and this is not among them, or the vendor’s own support material says it is not supported. This is both an accuracy rule and a legal one: asserting a product lacks something it has is a false statement about a named business.

Evidence grades

Every claim carries a grade and, for A through E, a source URL and the date it was retrieved.

GradeSourceCells
APrimary technical documentation, API reference, regulatory filing, source code679
BVendor product documentation — the manual, help centre, admin guide7,440
CVendor pricing or feature page — first-party, but written to sell1,445
DVendor marketing claim1,075
ECredible independent third-party reporting357
FAnalyst inference. Requires a stated rationale and carries no URL4,168

A yes on a differentiator-weight claim requires grade A or B. Nothing weaker. Vendor marketing establishes claims; documentation establishes facts. A feature listed on a product page tells you the vendor wants you to believe it exists; the admin guide describing how to configure it tells you it does.

Adversarial verification

Every dossier is re-examined by an independent pass instructed to refute it, defaulting to downgrade when uncertain and permitted to mark a value upheld only when it located the supporting documentation itself. That pass challenged 3,362 values and corrected 1,042 of them — 826 downgrades against 216 upgrades, a 31% correction rate on contested cells. The systematic error it caught was always the same direction: researchers reading marketing copy as capability.

2,110 cells currently carry adversarially-confirmed evidence. The full challenge record — what was contested, on what evidence, and which way it moved — is published on each product page rather than discarded after application.

Peer groups

Comparison happens within a peer group and never across one. Each group publishes objective inclusion criteria, contestable by pull request, and each claim is scoped to the business shapes it applies to. Drive-thru timers are n/a for a taproom POS, not a failure. Denominators count only in-scope claims, so a specialist is never penalised for a capability its buyers do not want.

GroupProductsClaims in scopeParity floor
Mainstream commercial restaurant POS20278178
Enterprise & chain restaurant POS11314212
QSR & drive-thru92963
Pizza & delivery-led7311160
Bar, brewery & taproom422427
Mobile & food truck32263
Open source & self-hostable1327847
Adjacent encroachers — ordering, middleware & back-office2425578
Regional & international831256

Counting rules

The census count is generated by a published function, never authored on a record. Gross and net are reported side by side, because a defended number beats a contested one: 100 records, 93 counting toward the census. A record is excluded when it is dead or absorbed, or when it is a white-label, reseller or predecessor of another. Acquired products are never deleted — their status changes and the record persists, because the acquisition history is the point.

Sources we will not use

We do not scrape G2, Capterra, GetApp, Software Advice or Gartner Peer Insights. Their terms prohibit automated access regardless of public visibility, and importing their sentiment data would import their sampling bias with it. Capterra, GetApp and Software Advice are one Gartner Digital Markets database behind three pay-per-lead front-ends; G2 acquired Capterra in January 2026, so cross-checking one against another is no longer an independence check. They may be cited by hand for a specific proposition; they may never be evidence for a capability claim. This is enforced in code — the fetch layer throws on those hosts, and validation rejects them as evidence.

“Best restaurant POS 2026” listicles are discovery-only, never evidence. They are overwhelmingly published by competitors ranking their own categories.

Conflict of interest

This project is maintained by the authors of a POS of their own. That is a real conflict and it is disclosed here rather than buried. The mitigations are structural: the reference platform is scored by the same rubric as everyone else, is excluded by default from every ranked aggregate, is never the subject of a composite score, and is credited only for capabilities backed by shipped code — never by design documents. A test asserts this against the built output rather than the source, because the built output is what a sceptic reads.

Where this corpus is weak

Corrections

If a value here is wrong, the most useful thing anyone can do is say so with a source. Corrections arrive as pull requests against the data files and are gated by the same validation that produced them: a value change with byte-identical evidence is rejected — cite the source that shows the change, or bump the retrieval date if you re-verified. Changes requested by a vendor, what changed, and what was declined and why are all recorded publicly. That log is worth more than the corrections, because it is demonstrable independence.