Assessment

Model Validation & Deployment

An independent check on a model you already rely on - calibration, evaluation leakage, shortcut learning, drift - and a plan for running it in production: serving, monitoring, retraining and what it costs per month.

Your model reports 94% accuracy. Somebody is about to make a decision worth real money on the strength of that number, and nobody in the building can currently tell you whether it is true - not because anyone was careless, but because the ways a model quietly looks better than it is are specialised, unintuitive, and almost never checked.

Nor, usually, is the other half of the question: what it actually takes to run this thing reliably, who notices when it degrades, and what the monthly bill looks like.

We check both, independently, on a model you already have - including one built by somebody else. You get a ranked list of everything inflating your confidence, the evidence for each, a plan for putting it into production, and the harness to keep watching after we leave.

Two ways in

Validation & Deployment Scan

€4,950 excl. VAT. Fixed price, typically two to four weeks. One model and one evaluation dataset - or just its predictions. The standard checks: evaluation leakage, calibration, an honest baseline, up to three slices. Plus a production-readiness checklist covering packaging, hosting, monitoring gaps and the obvious cost and security problems. Ends in a report with a straight verdict: go, fix first, or rethink.

Validation & Deployment Review

€24,500 excl. VAT. Fixed price, typically eight to twelve weeks. One model and its pipeline, with the full battery below and a deployment plan written for your setup. The calibration and monitoring harness is handed over to run in your CI, and one retest is included within three months.

Timelines start once the model, the data and the access are in place. If the evaluation data turns out not to support a verdict, that is itself the finding - written up with what is missing - rather than a project that stalls.

What we look for

Calibration

When the model says 90%, is it right nine times out of ten? Reliability curves and calibration error, measured on held-out data. Most deployed models are miscalibrated and have never been tested for it - which makes every confidence threshold downstream arbitrary.

Evaluation leakage

Duplicate specimens split across train and test, images from one sample landing in both folds, a time-ordered process evaluated with a random split. Endemic in scientific ML, and the single most common reason a number does not survive contact with production.

Shortcut learning

Is the model reading the physics, or the scanner, the background, the technician who prepared the batch? We check attributions against what a domain expert says should matter - and we have the domain experts to ask.

Slice performance

Aggregate accuracy hides the failures that cost money: the rare defect class, one instrument, one site, one operator. We break performance down where it actually matters to you.

Distribution shift

Do today's inputs still resemble the training set, and does the system notice when they stop? Silent degradation is the most common way a good model stops being a good model.

Honest baselines

Was a simple model ever tried? Sometimes logistic regression on four features matches the network, and that is worth knowing before you maintain the network for five years.

The deployment half

A model that survives the checks still has to be run by somebody, usually the team that is already busy. The review tier ends with a written plan for that, specific to your situation rather than a generic architecture diagram:

  • Where it runs. Cloud, on-premise or at the edge, decided by where the data is allowed to be and what latency the decision needs - not by what we happen to like deploying.
  • How it is served. Batch, API or embedded in an existing system; what the throughput has to be; what happens when it is unavailable.
  • Monitoring and drift. What is measured in production, what the thresholds are, and who gets told. Drift that nobody watches for is just a later surprise.
  • Retraining cadence. How often, on what data, and what has to be true before a new version replaces the old one.
  • What it costs. A monthly running-cost estimate, so the decision to deploy is made with the operating bill visible rather than discovered in the first invoice.
  • Security and rollout. The obvious exposure checks, then a staged rollout with the risks named and something to fall back to.

The infrastructure behind all of this is our cloud and deployment practice - we write the plan as the people who would otherwise have to implement it.

How it runs

  1. 1

    Scoping call

    One model, one decision it feeds. What it predicts, what happens when it is wrong, what you already measure. Half an hour, and it fixes the price.

  2. 2

    Handover

    The model or its predictions, the evaluation data, and how the splits were made. Everything can stay on your infrastructure if it has to.

  3. 3

    Analysis

    Calibration, leakage, slices, shift, attributions and baselines - run as a standard battery, plus whatever your domain specifically demands.

  4. 4

    Report and readout

    Findings ranked by how much each one inflates trust in the model, each with its evidence and a concrete remedy - and, on the review tier, the deployment plan. Presented to your team, not emailed as a PDF.

  5. 5

    Harness handover

    The calibration and monitoring checks are yours to keep and to run in CI, so the next model gets the same scrutiny without us.

What you do with the findings

A review that only tells you what is wrong is half a service. Most findings have a concrete, bounded fix, and we will quote it separately - or hand it to your team with the code:

  • Miscalibration → conformal calibration wrapped around the existing model, giving intervals with a guaranteed error rate. No retraining required.
  • Weak slices → targeted data collection and retraining where it pays, rather than more data everywhere.
  • Shortcut learning → the honest conversation about whether the model can be salvaged, and what a defensible version needs.
  • Drift blindness → monitoring wired into the deployment, so the next degradation announces itself.

None of these are long programmes either. Each is a bounded, separately quoted piece of work, and you are free to hand the findings to your own team instead.

Not included in the review itself: implementing the fixes, retraining, and carrying out the deployment. Those are projects, priced separately once the review says what they need to be - which is also why the review can be fixed price. How your data is handled during the work is your choice at intake; the three levels are set out with the other assessments, and the analysis can run entirely on your infrastructure.

Who this is for

R&D groups with a model heading toward a real decision. Quality and process teams running inspection or prediction on a line. Anyone whose model is about to be pointed at something regulated, audited or expensive - and anyone who has inherited a model whose original authors have left.

It is deliberately independent: we are happy to review a model we did not build, including one built by another supplier. That is usually when it is worth the most.

Where to go next