Video skill assessment that shows its work.

ActionQuality is an AI video skill-assessment platform: it breaks a recording — an operation, a sim session, a competition run — into its steps and scores each one on your rubric, with every score one click from its moment on video. Your experts keep the final say.

Honest numbers. Humans in charge. Feedback in minutes — not weeks later.

actionquality.ai · laparoscopic cholecystectomy
ActionQuality scoring a laparoscopic cholecystectomy: the framed encoder window beside the source recording, a color-coded step timeline against ground truth, and GOALS rubric scorecards with the instructor's mark beside the model's.
ActionQuality scoring big-air snowboarding: a step timeline and an action-quality score against the judges' mark.
ActionQuality scoring a piano performance: pianist skill level and song difficulty scored against ground truth.
ActionQuality scoring 10m platform diving: a step timeline and an action-quality score against the judges' mark.
One system, any procedure — the same step timeline, rubric scorecard, and clickable evidence, whether it's surgery, sport, a music performance, or industrial assembly.
21 dayswhat an expert panel took to review 50 trainee videos, in a published study
Same daytimeline + scorecard ready as soon as processing finishes
1 clickfrom any score to the exact moment on video — see the evidence yourself
Any rubric1–5 clinical scales, 0–110 judging scores, 0–9 grading ladders — same product
Measured results

Numbers we can defend in a room

On held-out recordings, ActionQuality reaches +0.71 Spearman rank correlation with competition judges on 10 m platform diving (70 dives) and 81% six-fold cross-validated frame accuracy for surgical phase recognition on laparoscopic cholecystectomy (public HeiChole benchmark, 24 videos). Every figure below was measured on recordings the system never saw during setup — cross-validated where the corpus allows it, with confidence intervals and significance tests built into the product itself.

+0.71Spearman rank correlation with competition judges on 10 m platform diving — 70 held-out dives from a public judged-competition corpus — a level academic systems typically report only after per-sport fine-tuning
81%frame-level accuracy locating surgical phases on HeiChole, a public research benchmark of 24 real laparoscopic cholecystectomies — six-fold cross-validated, so every video was scored by a model that never trained on it
~3 ptsgap between stated confidence and reality after calibration — when it says it's 80% sure, it's right about 80% of the time
13corpora onboarded across six domains — surgery, simulation, judged sport, industrial assembly, fitness, music — one product; a new field is set up by training on a few dozen of your scored recordings, never by rebuilding the system
How it works

How video assessment works: four steps, no technical team.

Define your procedure

The steps in order, and how each rated step should be scored — in your terms, on your scale.

Add a recording

From an operating room, a sim lab, a competition, a workbench, or a gym.

Get the scored timeline

ActionQuality finds the steps, scores each rated one against your rubric, and flags anything out of order or unusual — every result linked to its moment in the footage.

You stay the judge

Click any step to watch it, accept or correct in seconds, and watch the quality line climb week over week.

Features

Everything between footage and a defensible scorecard

See the procedure
A colour-coded step timeline laid against the reference version of the procedure

Automatic step timeline

Drop in a recording and get back a timeline of the procedure: which step happened, when, and for how long — laid directly against the reference version so gaps and detours are obvious at a glance.

A video player scrubbed to a specific moment, linked from the timeline

Click straight to the moment

Every step on the timeline is a link into the footage. Click it and the video jumps there; scrub the video and the timeline follows. No more hunting through an hour of tape for the thirty seconds that matter.

The ingestion recipe panel with a one-click Run pipeline button

Score new sessions against a set-up procedure

Setting a procedure up takes a corpus — a few dozen labeled sessions to train on, with more held back to validate the scores. Once it's set up, every new recording scores against it in one click: new trainee, new session, same rubric, straight into the dashboard — nothing to reconfigure.

Score with evidence
Rubric scorecards showing the model's score beside the instructor's mark

Rubric scorecards, side by side

Each rated step gets a score on your own rubric, displayed right next to the instructor's mark. Agreement and disagreement are visible, not buried in a spreadsheet.

Validation checks flagging steps with their exact time ranges

Anomalies, flagged with timestamps

Steps done out of order, or running suspiciously long or short, are flagged automatically — each flag carries the exact time range so a reviewer can watch the moment and decide.

The review queue's per-report confidence column, shown right in the interface

Confidence you can see

Every prediction carries a confidence level, shown right in the interface. A shaky call looks shaky — it doesn't wear the same face as a sure one.

Honest numbers
The corpus split into 18 train / 6 val, with quality measured on the held-out val set

Graded on recordings it has never seen

Every quality figure we publish comes from videos held back from setup — the same standard you'd hold a trainee to. No grading on the practice test.

A comparison led by a one-word WORSE verdict: the significant drops highlighted with saturated borders, within-noise metrics dimmed and marked ns

Noise is labeled as noise

Every comparison opens with a one-word verdict — Better, Worse, or Within noise — with the changes that actually matter named beside it. Differences too small to trust are grayed out and marked not significant, and every number ships with its uncertainty range — so luck doesn't get to masquerade as progress.

A frame-accuracy quality trend with its confidence band across frozen snapshots, oldest to newest

Quality tracked over time

Trend views follow both step-finding accuracy and skill agreement across every run — each point a frozen snapshot you can name and return to, with its uncertainty band drawn in — so drift shows up as a line on a chart, not a rumor.

Humans in charge
The review queue ranked least-trusted first, with accept/flag controls

A review queue that respects expert time

Results the system trusts least surface first for a human check. Accept or flag in seconds — reviewer attention goes exactly where the machine is least sure.

Edit ground-truth controls for correcting scores in the report

Correct the record, keep the history

Instructors can fix step boundaries or scores right in the report. Every correction is checked against the procedure, saved to history, and credited to a human — the machine never silently overwrites people.

The configurable procedure pipeline: Ingest, Embed, Phase, Segment, Skill, Assess, Report

Your procedure, your rubric

Define the steps and the scoring scale for any procedure — a 1-to-5 clinical scale like OSATS or GOALS, a 0-to-110 judging scale, a 0-to-9 grading ladder all work in the same product. A new procedure is a definition, not a new project.

Use cases

From surgical training to sports judging — wherever skill is a sequence a camera can see

Surgical education

Automated surgical skill assessment, without burning faculty time

ActionQuality provides automated, video-based surgical skill assessment for residency and simulation programs. Structured, video-based skill assessment is now required in accredited US general-surgery programs, and the bottleneck is quantified: in one published study, a panel of five attending surgeons took 21 days to review 50 trainee videos. ActionQuality turns each recording into a step timeline and an OSATS- or GOALS-style rubric scorecard with clickable evidence, so faculty start from the moments that matter. Surgical results to date are on the public HeiChole research benchmark; clinical pilots are the next step.

Surgical simulation

Score every sim session, same day

Roughly 300 accredited simulation programs worldwide run sessions built to be recorded — and no published survey even measures how much of that footage gets formally assessed. ActionQuality handles classic sim tasks like peg transfer and cataract procedures: it segments each session into its steps, checks the order, and puts every session on the same quality dashboard, so throughput is no longer capped by reviewer calendars.

Sports judging & coaching

A second opinion that moves with the judges

Trained on real judged-competition footage across multiple sports, ActionQuality ranks routines in strong agreement with the judges' own scores — on 10 m platform diving, +0.71 rank agreement across 70 competition dives it never saw. Coaches get an instant, consistent read on every attempt; athletes get evidence, not vibes.

Industrial assembly

Catch the skipped step before it ships

Assembly work is a sequence with a right order and expensive wrong ones. ActionQuality builds a step timeline from training footage and flags out-of-order or missing steps with the exact moment attached — turning tribal-knowledge coaching into a reviewable record.

Fitness & physio

Every session, automatically logged by exercise

Workout recordings are segmented into their exercises — the most reliable segmentation of any domain we've tested on held-out sessions — so a coach or physiotherapist sees a clean, per-exercise timeline of each session and can track a client's work across weeks — without scrubbing through a single video.

Music pedagogy

An honest instrument for practice video

On piano performance footage, ActionQuality can read the difficulty of the piece being played with meaningful accuracy from video alone — and our own testing shows a player's finesse lives largely in sound and fine finger technique that silent video can't fully carry. We publish that boundary instead of papering over it, so teachers know exactly what the tool can and cannot see.

Why trust it

We grade ourselves the way you'd grade a trainee

Most tools show you their best day. We show you the measurement.

Held-out, always

Every quality figure we show was measured on recordings the system had never seen during setup — official held-out test sets, not highlight reels.

Luck gets named

When a difference could just be chance, we say so — in gray, in writing. A standard statistical test decides when a difference is too small to trust.

The fair number wins

Where two ways of measuring disagree, we report the fair one — even when the fair way makes us look worse. Measured per performance, not per clip, the signal is real but modest — and we say so.

Traceable to footage

Every score is traceable to video — click it and watch the moment yourself. The instructor's mark is shown bar-for-bar beside the system's score.

The honesty ledger. A standing page of measured capabilities and measured limits — including the domains where video alone isn't enough. What works today, what doesn't yet, and exactly how we measured both. It ships with the product, because knowing where the edge is, is part of what you're buying.
Why now

Video-based assessment is mandated. Expert time isn't scaling.

Structured, video-based skill assessment is now a requirement in accredited US general-surgery training, simulation centers are built around recorded sessions, and in other high-stakes fields — aviation among them — recurrent simulator-based evaluation is already law. The recordings exist. The expert hours to review them don't. ActionQuality closes that gap without taking the decision away from the expert.

FAQ

Fair questions, straight answers

Do we need a technical team to use this?

No. You describe your procedure in plain terms — the steps and the scoring scale — and the dashboard is built for instructors and program directors: timelines, scorecards, a review queue, and plain-language explanations on every number.

Will this replace our expert reviewers?

No — it multiplies them. ActionQuality proposes; your experts validate. The least-trusted results surface first, every score links to its moment on video, and a human decision always wins over a machine one. That division of labor keeps your experts in charge — and removes the wait.

How do we know the scores are any good?

Because we grade ourselves the way you'd grade a trainee: on material the system never saw during setup. Every figure ships with its uncertainty, comparisons mark differences that could be noise, and small samples are labeled as small samples. When a number is shaky, the dashboard says so — out loud.

What kinds of procedures does it handle?

Anything with a repeatable sequence of steps a camera can see. Today that spans surgery and surgical simulation, judged sports routines, industrial assembly, fitness workouts, and instrument practice — and adding a new procedure means writing a definition, not commissioning a new product.

What happens when video alone isn't enough?

We tell you. Our favorite example is piano: the system can read a piece's difficulty from video, but a player's finesse lives largely in the sound — so we publish that limit instead of shipping a confidently wrong score. Knowing where the edge is, is part of what you're buying.

Can instructors correct the system?

Yes, right in the report — step boundaries and scores alike. Corrections are checked against the procedure, kept in history, and credited to the human who made them. Your experts' judgment becomes part of the permanent record, not a sticky note on top of it.

How fast is feedback?

In a published study, an expert panel took 21 days to turn around 50 trainee videos. With ActionQuality, a recording's timeline and scorecard are ready as soon as processing finishes, so reviewers start from evidence the same day the video is captured.

Where does our video go — and who can see it?

Your recordings stay yours. For pilots we process video on EU-hosted infrastructure, with on-premises processing available for operating-room footage. Recordings are used only to assess your sessions — never to train models for anyone else — and are deleted on a schedule you set. Clinical footage should be de-identified before upload (no patient identifiers in frame or filename), and we sign a GDPR data-processing agreement before any footage moves. The surgical numbers on this page were measured on HeiChole, a public research dataset — not on customer video. This website itself runs no analytics and sets no cookies.

Is this a medical device?

No. ActionQuality reviews recordings after the fact to assess the operator's skill — it plays no role during a procedure and makes no claims about patient outcomes or care decisions. It sits in the same category as an instructor reviewing tape, not clinical software.

Can it score OSATS or GOALS automatically?

The scorecard is rubric-agnostic — OSATS, GOALS, or your program's own scale plug into the same product, and every automated score is shown beside the instructor's mark with a human review queue on top. Our surgical validation so far is phase timelines on the public HeiChole dataset (81% six-fold cross-validated frame accuracy); what is and isn't yet validated is published in the honesty ledger.

Does it work with our rubric?

Yes — the scorecard adapts to whatever scale your field uses. A 1-to-5 clinical rating, a 0-to-110 judging score, and a 0-to-9 grading ladder all run through the same product, unchanged.

See a real case scored, live.
Timeline, rubric, evidence.

A live walkthrough of the step timeline and the rubric scorecard against the instructor's marks — plus what it takes to set up your own procedure from a corpus of your sessions. We're inviting the first surgical and simulation pilot programs — your video, your rubric, a signed data-processing agreement, and a scored timeline of your own sessions as the outcome.

Who's behind this

ActionQuality is built by Marek Pawlowski, an EU-based engineer. Where it stands today, stated plainly: the system is validated on public research benchmarks and real judged-competition footage across six fields — it has not yet run a hospital pilot. If your program wants to be the first, that's exactly the conversation we're inviting: hello@actionquality.ai.