Last updated 2026-08-07. Every claim on the main page traces to a row here. Measured on recordings the system never saw during setup; protocols named per row.
| Claim | Number | Data | Protocol |
|---|---|---|---|
| Rank agreement with competition judges, 10 m platform diving | +0.71 Spearman ρ | Public judged-competition diving corpus, 70 held-out dives | Trained on a disjoint split; scored dives never seen in setup |
| Surgical phase recognition, laparoscopic cholecystectomy | 81% frame accuracy | HeiChole — public research benchmark, 24 videos | Six-fold cross-validation: every video scored by a model that never trained on it |
| Confidence honesty (phase predictions) | ~2.5% expected calibration error | Same surgical corpus | Temperature scaling fit on validation logits; reliability diagram ships in-product |
| Breadth | 13 corpora · 6 fields | Surgery, simulation (cataract, peg transfer), judged sport, industrial assembly, fitness, music | Same product, per-field heads trained on a few dozen scored recordings each |
| Boundary | What we found |
|---|---|
| Music performance (piano) | Piece difficulty is readable from video with meaningful accuracy; a player's finesse lives largely in sound and fine finger technique that silent video can't fully carry. We publish the boundary instead of shipping a confidently wrong score. |
| Small corpora | With few validation videos, confidence intervals are wide — the product shows them and runs significance tests so noise is labeled as noise, not sold as improvement. |
| Stage of validation | All surgical results are on a public research benchmark (HeiChole, 24 videos). ActionQuality has not yet run a hospital pilot; that is the next step, not a claim we make. |
New capability claims appear here only with a number, the dataset, and the protocol. If a claim on the main page ever lacks a row here, that's a bug — write to hello@actionquality.ai.