A clinical decision support company needed a model that flags early sepsis risk from vitals and lab data streaming out of hospital EHR systems. Getting the model right was, honestly, the easier half. Getting a team that could operate inside HIPAA, FDA guidance on software as a medical device, and hospital IT politics, that was the hard half.
A false negative here isn't a bad recommendation. It's a missed sepsis case. That single fact reshapes what "good" looks like in a candidate. Pure Kaggle-competition talent, however sharp, tends to optimize for aggregate accuracy. Clinical model teams need people who ask about class imbalance, alert fatigue, and what happens when a nurse starts ignoring the tool because it cried wolf too many times.
The company also needed a research lead credible enough to sit across the table from hospital clinical informatics committees. An executive search problem, stacked directly on top of a technical hiring problem.
Two tracks ran side by side. Executive search for a Head of Clinical AI Research required peer-reviewed publication history in clinical ML, prior experience presenting model validation data to an IRB or equivalent body, and comfort being the named signatory on FDA submission documentation. Technical hiring covered the build team: ML engineers with time-series and EHR data experience specifically, since general tabular skill doesn't transfer cleanly to irregularly sampled clinical vitals. Data scientists needed a background in survival analysis and clinical epidemiology, not just standard classification metrics. One AI researcher focused purely on interpretability, because clinicians will not act on a black-box score without a reason attached.
We rejected several strong-looking resumes for one reason: the candidates couldn't explain how they'd validate a model across hospitals with different charting practices. A model tuned on one hospital's EHR quirks can fail quietly at a second site. It's the single most common reason clinical AI pilots die after a promising first deployment. It became our number one screening question.
The Head of Clinical AI Research, once seated, insisted on something the founding team hadn't originally planned for: a standing weekly review with a practicing ICU nurse who had no stake in the model's success and every incentive to point out when an alert felt wrong. That review caught an alert-threshold miscalibration two weeks before the second hospital's go-live, well before it could have become a real incident.
By the time the third validation site came online, the technical team had built enough internal tooling around cross-site validation that adding a new hospital's data no longer required a bespoke engineering effort. That reusability, more than any single model improvement, is what let the company credibly tell its next hospital prospect that onboarding would take weeks, not quarters.
Head of Clinical AI Research, sourced from an academic medical center
4 ML/data science hires across two hiring cycles
Deliberately unrushed
3 hospital systems, up from 1