Start

Validation roadmap

The claims in this system that are argument rather than evidence, written as study designs — including, for each one, the result that would prove it wrong.

Open — no studies completed IEC 62366-1 Framework · ISO 9241-11

Why this page exists

A design system cannot be clinically validated — validation is a property of a finished device in its intended use, and it belongs to the manufacturer. What a design system can do is be honest about which of its rules are established and which are reasoned, and be specific about what would settle the difference.

The evidence classes on the Status page separate Computed and Standard claims — which are verifiable now — from Reasoned claims, which are defensible arguments with no data behind them. This page is the plan for the third category.

The point for a manufacturer

Every study below produces evidence you need for your own IEC 62366-1 file regardless of whether you use these components. These are not questions about this design system — they are open questions in clinical AI interface design, and someone has to answer them.

The studies

ClaimStudy designWould falsify itMethod
Annotation off by default reduces anchoring on the model Randomised within-subject: clinicians read matched ECG sets with annotation on versus off, scored against adjudicated truth No difference in agreement with truth, or slower reads with no accuracy gain Formative
Priority-then-waiting beats confidence as a default sort Simulated queue with seeded cases; measure time to first review of the highest-acuity patient under each sort order Confidence sort surfaces high-acuity cases as fast or faster Simulated use
Symmetric agree/disagree yields honest override data A/B on control weighting; compare override rates against blinded adjudication Override rate is unchanged by presentation — friction was never the driver Field
Withholding a score outside validated scope beats reporting a low one Vignette study on paced-rhythm and reduced-lead cases; measure inappropriate reliance Clinicians discount a low score as effectively as they treat "—" as absent Formative
No bulk acknowledge reduces unread acknowledgement Simulated use under interruption load; measure acknowledge-without-open rate Rate is unchanged, or total missed findings rise from the added workload Simulated use
Staged row insertion prevents wrong-patient selection Simulated use with timed arrivals during target acquisition No reduction in mis-selection versus live insertion Simulated use
Read-only locking prevents shared-account workarounds Field observation across shifts; count shared sessions and time to resume Shared accounts persist at the same rate Field
The 12 mm gloved touch floor is the right threshold Target-acquisition testing, gloved and ungloved, across the deployed display range Error rate is flat between 10 and 12 mm, or still rising above 12 mm Bench
Priority words alongside hue improve recognition at distance Recognition task at 1 m, 2 m and 4 m, greyscale and colour, with and without the priority word Hue alone is recognised as accurately at every distance Bench
A hard quality gate beats a hedged negative result Between-subject with lay recipients: a hedged negative ("no disease detected, image quality limited") versus an explicit "not assessable — referred". Measure what each group believes was found and what they intend to do next Both groups understand equally, or the hedged wording produces no false reassurance Formative
A system-imposed retake cap reduces forced passes Field comparison across sites with and without an operator-controlled stopping decision; compare ungradable rates and downstream grader-confirmed inadequacy Ungradable rates and grader-confirmed inadequacy are unchanged — the cap was never what drove the behaviour Field
Stating what was not assessed changes what recipients believe Between-subject comprehension test on a negative screening result, with and without an equal-weight scope statement; measure beliefs about unscreened conditions Recipients hold the same beliefs either way, or the scope statement reduces comprehension of the primary finding Formative
A symptom-based route back outperforms "consult your doctor if concerned" Vignette study: recipients presented with interval symptoms after each wording; measure whether and how quickly they would seek care No difference in recognition or intended urgency Formative
A persistent suppression count preserves trust in a quiet ward Simulated shift with seeded suppressed events; measure how often clinicians independently verify monitors, with and without the count on the clinical surface Verification behaviour is unchanged — the count was never what clinicians were uncertain about Simulated use
Outcome-linked preview changes configuration decisions Administrators given the same rule change with a parameter diff versus a preview naming delayed alarms that preceded deterioration; measure approve/reject and stated reasoning Decisions are the same under both presentations Formative
Bedside announcement of rule changes reduces misattributed fault reports Field comparison before and after announcement is introduced; measure fault reports filed against monitoring in the 48 h after a configuration change Fault report rate is unchanged, or the announcement itself generates more Field
Stating displacement keeps triage honest about its cost Simulated reading session with a promoting device; measure whether readers notice and act on studies pushed down, with and without "moved down n places" on the row Displaced studies are read at the same time either way — the statement changed nothing Simulated use
Committing an impression before marks preserves independent reading Within-subject: readers record an impression first versus seeing marks immediately; compare agreement with adjudicated truth on cases the device marks wrongly No difference on falsely-marked cases — commitment did not protect the first impression Formative
Permitting logged deviation beats prohibiting it Field comparison of a hard block versus a logged one-click deviation; measure recorded deviation rate and out-of-product reading Blocking produces no out-of-product reading and no loss of record Field
A fixed clinical axis prevents noise being read as trend Within-subject: clinicians shown matched series on auto-fitted versus clinically fixed axes; measure judgements of whether meaningful change occurred Judgements are unaffected by scaling — readers already discount the axis Formative
Stating measurement variability reduces over-investigation Vignette study on surveillance findings with sub-threshold change, with and without the variability figure; measure intended next action Intended actions are unchanged — the figure is read but not used Formative
An owned, overdue-ordered open-loop list closes more loops Field comparison against usual practice; measure count and age of unclosed findings per clinician over two quarters Open-loop age and count are unchanged, or the list is not maintained Field
Preserving superseded outputs supports fair decision review Reviewers assess past decisions with the original model output preserved versus re-scored in place; measure judgements of whether the decision was reasonable Judgements are the same either way — the preserved record changed nothing Formative
An empty dose field prevents confirmation-reflex errors Simulated dosing with a deliberately miscalculated suggestion, pre-filled versus empty-with-one-tap-to-fill; measure how often the wrong dose is committed Error rates are the same — the second action does not create a check Simulated use
Age shown with equal prominence prevents acting on stale readings Within-subject: home-use display with age subordinate versus co-equal to the value; measure decisions taken on readings older than a stated threshold Decisions are unaffected by the prominence of age Formative
Reconciled follower alerts reduce unnecessary escalation Field comparison of raw versus reconciled carer alerts; measure contacts made to subjects for events already resolved Unnecessary contacts are unchanged — followers act on the alert regardless Field
Natural frequencies are understood better than percentages by lay readers Between-subject comprehension test of the same risk expressed as a percentage, a 1-in-n ratio, and a natural frequency with a constant denominator Comprehension and risk judgements do not differ by format Formative
Version in persistent chrome cuts time-to-identify during support Timed task: clinicians asked which version they are running, with the version in the masthead versus only in an about screen Time to answer is unchanged, or users look in the about screen regardless Formative
Binding the major version to the UDI-DI is legible to users Comprehension test: given two release notes, can users say whether they are on the same regulatory device as before Users cannot tell either way — the convention carries no meaning to them Formative

Sequencing

Volume of study matters less than order. Roughly in the order that would move the most:

  1. Formative think-aloud with a handful of emergency physicians on the triage worklist and ECG review. Cheapest, fastest, and would move more of this list than any amount of further writing. Five to eight participants is enough to find the structural problems.
  2. Bench testing for the two threshold claims — touch target and recognition at distance. No clinicians required, no ethics approval, days of work.
  3. Simulated use under realistic interruption load. This is where the alarm, worklist and insertion claims live, and it is the closest analogue to the summative evaluation a manufacturer will eventually run.
  4. Field observation last. The only method that will settle the authentication behaviour, because the shared-account workaround only appears when nobody is watching for it.

How results will be handled

Working on this

If you are a manufacturer, a clinical research group or a human-factors team with access to representative users, any single row in the table above is a self-contained piece of work with a clear result. The most useful contributions in order:

Contact routes are on the overview. Nothing on this page implies an existing partnership or completed study; the status pill above says exactly where this stands.

Do's and don'ts

Do

Read this page as a description of work that has not happened, and plan your own summative evaluation accordingly.

Your IEC 62366-1 obligations do not wait on anyone else's study, and would not be discharged by one if it existed.

Don't

Cite a planned study as though it were a completed one, or a roadmap as though it were a result.

Nothing on this page has been run. A roadmap in a submission that reads like evidence is the kind of overstatement that costs credibility on everything else in the file.

Do

Reuse the study designs, task scenarios and measures as a starting point for your own protocol.

This is the most useful thing on the page. Designs are cheap to share; results are not transferable between devices anyway.

Don't

Assume a result obtained on this system's components would transfer to your device.

Usability is an outcome of a specific user, goal and context. Change the device, the users or the environment and the result has to be re-established.

NotJustAnyMed.Tech Design System · Validation roadmap · v0.1
No studies have been completed. All claims referenced here are currently Reasoned.