Arithmetic
Ordinary over and under comparisons with matching settlements.
Open evaluation asset · Version 1.0
Twenty-four synthetic cases test a narrow but important distinction: calculate the ordinary over/under result, compare it with the reported settlement, and stop when evidence or house-rule questions require human review.
Coverage
The benchmark intentionally pairs routine numeric comparisons with equality boundaries, conflicting labels, weak evidence, and explicit rule signals.
Ordinary over and under comparisons with matching settlements.
Whole-number lines where equality must produce a push.
Correct math paired with a conflicting reported grade.
Missing or conflicting evidence that must be surfaced.
Participation, postponement, correction, void, and market-definition signals.
Cases where a system must decline to turn arithmetic into a legal conclusion.
Reproducible scoring
Score exact structured fields before judging prose. This keeps the result inspectable and makes category-level failure visible.
mustMention, or explicitly returns none.Dataset statement
Every case is authored, versioned, and free of customer records, real ticket images, and account identifiers. Expected labels follow SlipVerdict's published arithmetic and escalation boundaries.
The evaluation structure draws on NIST's emphasis on documented tasks, metrics, measurement methods, strengths, and limitations.
This is not a representative sample of every sport, market, jurisdiction, or house rule. High scores establish performance only on this narrow synthetic set.
It cannot determine whether a real settlement is legally correct and must not be used to automate complaints or payouts.
Evaluation references: NIST AI Measurement and Evaluation and the NIST AI RMF Measure Playbook. SlipVerdict is not endorsed by NIST.