ActionQuality is an AI video skill-assessment platform: it breaks a recording — an operation, a sim session, a competition run — into its steps and scores each one on your rubric, with every score one click from its moment on video. Your experts keep the final say.
Honest numbers. Humans in charge. Feedback in minutes — not weeks later.




On held-out recordings, ActionQuality reaches +0.71 Spearman rank correlation with competition judges on 10 m platform diving (70 dives) and 81% six-fold cross-validated frame accuracy for surgical phase recognition on laparoscopic cholecystectomy (public HeiChole benchmark, 24 videos). Every figure below was measured on recordings the system never saw during setup — cross-validated where the corpus allows it, with confidence intervals and significance tests built into the product itself.
The steps in order, and how each rated step should be scored — in your terms, on your scale.
From an operating room, a sim lab, a competition, a workbench, or a gym.
ActionQuality finds the steps, scores each rated one against your rubric, and flags anything out of order or unusual — every result linked to its moment in the footage.
Click any step to watch it, accept or correct in seconds, and watch the quality line climb week over week.

Drop in a recording and get back a timeline of the procedure: which step happened, when, and for how long — laid directly against the reference version so gaps and detours are obvious at a glance.

Every step on the timeline is a link into the footage. Click it and the video jumps there; scrub the video and the timeline follows. No more hunting through an hour of tape for the thirty seconds that matter.

Setting a procedure up takes a corpus — a few dozen labeled sessions to train on, with more held back to validate the scores. Once it's set up, every new recording scores against it in one click: new trainee, new session, same rubric, straight into the dashboard — nothing to reconfigure.

Each rated step gets a score on your own rubric, displayed right next to the instructor's mark. Agreement and disagreement are visible, not buried in a spreadsheet.

Steps done out of order, or running suspiciously long or short, are flagged automatically — each flag carries the exact time range so a reviewer can watch the moment and decide.

Every prediction carries a confidence level, shown right in the interface. A shaky call looks shaky — it doesn't wear the same face as a sure one.

Every quality figure we publish comes from videos held back from setup — the same standard you'd hold a trainee to. No grading on the practice test.

Every comparison opens with a one-word verdict — Better, Worse, or Within noise — with the changes that actually matter named beside it. Differences too small to trust are grayed out and marked not significant, and every number ships with its uncertainty range — so luck doesn't get to masquerade as progress.

Trend views follow both step-finding accuracy and skill agreement across every run — each point a frozen snapshot you can name and return to, with its uncertainty band drawn in — so drift shows up as a line on a chart, not a rumor.

Results the system trusts least surface first for a human check. Accept or flag in seconds — reviewer attention goes exactly where the machine is least sure.

Instructors can fix step boundaries or scores right in the report. Every correction is checked against the procedure, saved to history, and credited to a human — the machine never silently overwrites people.

Define the steps and the scoring scale for any procedure — a 1-to-5 clinical scale like OSATS or GOALS, a 0-to-110 judging scale, a 0-to-9 grading ladder all work in the same product. A new procedure is a definition, not a new project.
ActionQuality provides automated, video-based surgical skill assessment for residency and simulation programs. Structured, video-based skill assessment is now required in accredited US general-surgery programs, and the bottleneck is quantified: in one published study, a panel of five attending surgeons took 21 days to review 50 trainee videos. ActionQuality turns each recording into a step timeline and an OSATS- or GOALS-style rubric scorecard with clickable evidence, so faculty start from the moments that matter. Surgical results to date are on the public HeiChole research benchmark; clinical pilots are the next step.
Roughly 300 accredited simulation programs worldwide run sessions built to be recorded — and no published survey even measures how much of that footage gets formally assessed. ActionQuality handles classic sim tasks like peg transfer and cataract procedures: it segments each session into its steps, checks the order, and puts every session on the same quality dashboard, so throughput is no longer capped by reviewer calendars.
Trained on real judged-competition footage across multiple sports, ActionQuality ranks routines in strong agreement with the judges' own scores — on 10 m platform diving, +0.71 rank agreement across 70 competition dives it never saw. Coaches get an instant, consistent read on every attempt; athletes get evidence, not vibes.
Assembly work is a sequence with a right order and expensive wrong ones. ActionQuality builds a step timeline from training footage and flags out-of-order or missing steps with the exact moment attached — turning tribal-knowledge coaching into a reviewable record.
Workout recordings are segmented into their exercises — the most reliable segmentation of any domain we've tested on held-out sessions — so a coach or physiotherapist sees a clean, per-exercise timeline of each session and can track a client's work across weeks — without scrubbing through a single video.
On piano performance footage, ActionQuality can read the difficulty of the piece being played with meaningful accuracy from video alone — and our own testing shows a player's finesse lives largely in sound and fine finger technique that silent video can't fully carry. We publish that boundary instead of papering over it, so teachers know exactly what the tool can and cannot see.
Most tools show you their best day. We show you the measurement.
Every quality figure we show was measured on recordings the system had never seen during setup — official held-out test sets, not highlight reels.
When a difference could just be chance, we say so — in gray, in writing. A standard statistical test decides when a difference is too small to trust.
Where two ways of measuring disagree, we report the fair one — even when the fair way makes us look worse. Measured per performance, not per clip, the signal is real but modest — and we say so.
Every score is traceable to video — click it and watch the moment yourself. The instructor's mark is shown bar-for-bar beside the system's score.
Structured, video-based skill assessment is now a requirement in accredited US general-surgery training, simulation centers are built around recorded sessions, and in other high-stakes fields — aviation among them — recurrent simulator-based evaluation is already law. The recordings exist. The expert hours to review them don't. ActionQuality closes that gap without taking the decision away from the expert.
No. You describe your procedure in plain terms — the steps and the scoring scale — and the dashboard is built for instructors and program directors: timelines, scorecards, a review queue, and plain-language explanations on every number.
No — it multiplies them. ActionQuality proposes; your experts validate. The least-trusted results surface first, every score links to its moment on video, and a human decision always wins over a machine one. That division of labor keeps your experts in charge — and removes the wait.
Because we grade ourselves the way you'd grade a trainee: on material the system never saw during setup. Every figure ships with its uncertainty, comparisons mark differences that could be noise, and small samples are labeled as small samples. When a number is shaky, the dashboard says so — out loud.
Anything with a repeatable sequence of steps a camera can see. Today that spans surgery and surgical simulation, judged sports routines, industrial assembly, fitness workouts, and instrument practice — and adding a new procedure means writing a definition, not commissioning a new product.
We tell you. Our favorite example is piano: the system can read a piece's difficulty from video, but a player's finesse lives largely in the sound — so we publish that limit instead of shipping a confidently wrong score. Knowing where the edge is, is part of what you're buying.
Yes, right in the report — step boundaries and scores alike. Corrections are checked against the procedure, kept in history, and credited to the human who made them. Your experts' judgment becomes part of the permanent record, not a sticky note on top of it.
In a published study, an expert panel took 21 days to turn around 50 trainee videos. With ActionQuality, a recording's timeline and scorecard are ready as soon as processing finishes, so reviewers start from evidence the same day the video is captured.
Your recordings stay yours. For pilots we process video on EU-hosted infrastructure, with on-premises processing available for operating-room footage. Recordings are used only to assess your sessions — never to train models for anyone else — and are deleted on a schedule you set. Clinical footage should be de-identified before upload (no patient identifiers in frame or filename), and we sign a GDPR data-processing agreement before any footage moves. The surgical numbers on this page were measured on HeiChole, a public research dataset — not on customer video. This website itself runs no analytics and sets no cookies.
No. ActionQuality reviews recordings after the fact to assess the operator's skill — it plays no role during a procedure and makes no claims about patient outcomes or care decisions. It sits in the same category as an instructor reviewing tape, not clinical software.
The scorecard is rubric-agnostic — OSATS, GOALS, or your program's own scale plug into the same product, and every automated score is shown beside the instructor's mark with a human review queue on top. Our surgical validation so far is phase timelines on the public HeiChole dataset (81% six-fold cross-validated frame accuracy); what is and isn't yet validated is published in the honesty ledger.
Yes — the scorecard adapts to whatever scale your field uses. A 1-to-5 clinical rating, a 0-to-110 judging score, and a 0-to-9 grading ladder all run through the same product, unchanged.
A live walkthrough of the step timeline and the rubric scorecard against the instructor's marks — plus what it takes to set up your own procedure from a corpus of your sessions. We're inviting the first surgical and simulation pilot programs — your video, your rubric, a signed data-processing agreement, and a scored timeline of your own sessions as the outcome.
ActionQuality is built by Marek Pawlowski, an EU-based engineer. Where it stands today, stated plainly: the system is validated on public research benchmarks and real judged-competition footage across six fields — it has not yet run a hospital pilot. If your program wants to be the first, that's exactly the conversation we're inviting: hello@actionquality.ai.