EU AI Act and MDR require AI models in medical devices to be supported by third-party statistical evidence — dataset representativeness, stratified performance, calibration, drift detection. Self-assessment does not satisfy a notified body. And the deadline is closer than it looks.
Backward-plan from December 2027: NB queue + technical file prep + evidence generation = start date is now. Remediation of a data representativeness or calibration gap adds 6–12 months.
Internal teams can run the statistics. What they cannot produce is third-party evidence. Notified bodies systematically discount first-party performance claims — the same way they require independent audit for QMS. The value of the certificate comes from who signed it, not from the measurement itself.
AI Act and MDR Annex II require documented evidence of performance under real distributional conditions, calibration across subgroups, and ongoing drift monitoring — all traceable to an identifiable responsible party. Internal documentation meets neither the independence nor the traceability standard a notified body applies.
An independent evidence package — SHA-256 audit manifest, regulatory cross-references, structured findings — produced by a physicist with 15 years in medical devices. The package is designed to be submitted as-is into a technical file, not summarised or re-formatted by your team.
The validation library runs in your infrastructure on your data. No patient data, device data, or proprietary training sets are transmitted or stored externally. What RE:MARK produces is the evidence report — the output, not a copy of the input. This satisfies both medtech data governance and GDPR requirements for sensitive health data.
Structured inventory of the training and evaluation dataset: class distribution, demographic and site stratification, label quality assessment, and representativeness analysis relative to the intended population. Flags gaps between training distribution and intended use population before they become a notified body finding.
Performance metrics decomposed across clinically relevant subgroups — demographic strata, device configurations, acquisition sites. Bootstrap confidence intervals on all reported metrics. Reveals subgroup gaps that aggregate accuracy hides: a model with 93% overall accuracy and 71% on a key subgroup is a regulatory problem.
Reliability diagrams, Expected Calibration Error (ECE), and overconfidence analysis. A model whose confidence scores do not match empirical probabilities fails the risk management standard even if its accuracy is acceptable — clinicians using AI-assisted outputs rely on calibrated confidence, not raw scores.
Statistical tests for covariate and concept drift between training and deployment distributions, and across time windows in post-market data where available. Documents the evidence base for the post-market surveillance plan required under MDR and AI Act Art. 72.
Conformal prediction sets with guaranteed marginal coverage — a distribution-free, assumption-minimal method that produces statistically valid uncertainty bounds even when model assumptions are not met. Documents the uncertainty quantification requirement under AI Act Art. 9 in the strongest available form.
Identify the applicable regulatory pathway (AI Act risk class, MDR device class, IVDR), the intended use population, the evaluation dataset available, and the notified body's known question pattern for this device category. Scoping produces the evidence plan before any analysis begins.
The five-module library executes in your controlled environment. Each module produces a structured output with a hash-signed audit log. Findings that indicate gaps — dataset imbalance, subgroup underperformance, miscalibration — are reported with severity and recommended remediation path.
All module outputs compiled into a single PDF evidence package with regulatory cross-references, SHA-256 audit manifest, and a structured findings summary. Format matches what notified bodies expect to see in a technical file annex. One document, ready to submit.
You are preparing for MDR conformity assessment or responding to a notified body's AI-specific questions during surveillance. You need third-party statistical evidence that is structured, traceable, and immediately insertable into the technical file — not a consulting report you have to translate yourself.
You have a Class IIa or IIb device with an AI component. Your team built the model; no one on staff has produced EU AI Act evidence packages before. You need a fixed-price engagement with a defined output, not an open-ended consulting relationship, and you need it before the timeline closes.
You hold SaMD companies with AI models in their products. Unresolved EU AI Act compliance is a valuation and exit risk — buyers and notified bodies will surface it. An evidence package produced now, before a transaction, removes that risk from the data room and from the deal conversation.
You manage the regulatory pathway but do not have a statistical validation team in-house. RE:MARK produces the evidence package; you integrate it into the technical file alongside the clinical evaluation and QMS documentation. Fixed scope, defined output, no open billing.
A regulatory affairs consultant maps compliance gaps against a checklist and tells you what is missing.
A data science consultant writes the code and hands you the numbers.
Neither produces third-party statistical evidence in the format a notified body expects to find in a technical file.
The combination that makes this work is specific: a physics-trained experimenter who has spent 15 years
inside medical device development — at Boston Scientific, across EMEA, at the intersection of clinical data
and regulatory submissions — and who has since built the statistical validation tooling from scratch.
The evidence is produced by someone who understands both what the regulation requires
and what the numbers actually mean.
Fixed-price engagement · Scope agreed before we start · No open-ended billing · Data stays in your environment
A 30-minute call to understand your device, your timeline, and whether this engagement is the right next step. No commitment. No proposal before we have understood the situation.
Evaluating an industrial AI or medtech target for investment?
RE:MARK Technical Due Diligence & Integration Advisory →