Open evaluation asset · Version 1.0

Can a system separate arithmetic from rule review?

Twenty-four synthetic cases test a narrow but important distinction: calculate the ordinary over/under result, compare it with the reported settlement, and stop when evidence or house-rule questions require human review.

Coverage

Easy math is not the whole task.

The benchmark intentionally pairs routine numeric comparisons with equality boundaries, conflicting labels, weak evidence, and explicit rule signals.

01

Arithmetic

Ordinary over and under comparisons with matching settlements.

02

Boundary values

Whole-number lines where equality must produce a push.

03

Settlement mismatch

Correct math paired with a conflicting reported grade.

04

Source quality

Missing or conflicting evidence that must be surfaced.

05

Rule review

Participation, postponement, correction, void, and market-definition signals.

06

Calibrated restraint

Cases where a system must decline to turn arithmetic into a legal conclusion.

Reproducible scoring

One point for each required behavior.

Score exact structured fields before judging prose. This keeps the result inspectable and makes category-level failure visible.

  1. Arithmetic resultCorrectly returns win, loss, or push from direction, line, and final statistic.
  2. VerdictReturns consistent, possible-mismatch, or review-required according to the published labels.
  3. Required caveatNames the evidence or rule issue in mustMention, or explicitly returns none.
  4. RationaleExplains the comparison without claiming a guaranteed payout, legal entitlement, or sportsbook fault.
96 possible pointsReport both total score and six category subscoresZero tolerance for unsupported legal or payout claims

Dataset statement

What this benchmark can—and cannot—show.

Design choices

Every case is authored, versioned, and free of customer records, real ticket images, and account identifiers. Expected labels follow SlipVerdict's published arithmetic and escalation boundaries.

The evaluation structure draws on NIST's emphasis on documented tasks, metrics, measurement methods, strengths, and limitations.

Limitations

This is not a representative sample of every sport, market, jurisdiction, or house rule. High scores establish performance only on this narrow synthetic set.

It cannot determine whether a real settlement is legally correct and must not be used to automate complaints or payouts.

Evaluation references: NIST AI Measurement and Evaluation and the NIST AI RMF Measure Playbook. SlipVerdict is not endorsed by NIST.