AI & Data Science Consultant
Maya Malamud
Malamud AI Advisory · Madrid · English, Spanish and Hebrew
I help HealthTech teams build clinical AI that works in production and that clinicians can trust: architecture, annotation, and evaluation.
I was once told there was no future in building statistical models for medicine, that I thought too far outside the box.
They were wrong. That is now my superpower.
Maya Malamud, AI & Data Science Consultant
AI & Data Science Consultant
Worked and taught with
Who I Work With
Seed-Stage HealthTech Founders
You have a clinical problem and early data, but no AI roadmap yet. I help you build the right architecture from day one, before expensive mistakes get baked in.
Roadmap and Proof of ConceptSeries A CTOs Scaling Up
Your pilot worked. Now you need production-grade pipelines, a team, and a system that holds up under clinical and investor scrutiny.
Architecture and TeamTeams Whose Annotation Costs Don't Scale
Your annotation budget grows with data volume, and your clinicians spend months labeling by hand. I redesign the workflow with active learning and a similar-cases view, so experts work from similar past cases instead of a blank page. One client went from 3 experts over 6 months to 1 semi-expert over 2 days.
Annotation EfficiencyTeams Overspending on LLMs
An agent handles every task, and the cost and latency show it. I design a cascade that sends each case to the cheapest layer that can handle it and escalates to a large language model (LLM) only when confidence is low.
Cascading AITeams Burned by Hallucinating AI
Your generative AI (GenAI) system passed all evals and still got something wrong in production. I build the detection and guardrail layer that catches what benchmarks miss.
Trustworthy GenAIResearchers with Small, Complex Data
Standard ML needs thousands of samples. I specialize in building personalized models from ultra-small clinical datasets, where most data scientists walk away.
Ultra-Small DataHave other needs not covered here?
Every clinical AI challenge is different. Reach out and I will put together a tailored approach for your specific situation.
Services: The HealthTech Bridge
Cascading AI Architecture
Deterministic logic first, then classic machine learning (ML) triage, then confidence-gated escalation to a large language model (LLM) only when the case actually needs it, with a teacher-student loop that retrains the simpler, cheaper model underneath.
Architecture Read more →- Deterministic logic layer before any model runs
- Classic ML triage for the majority of routine cases
- LLM escalation gated by confidence, not by default
- Teacher-student retraining loop to keep the pipeline improving
- Built for teams tired of "throw an LLM at everything"
Structural Label Noise Audit
Diagnosing where label noise in your training data is structural, meaning systematic rather than random, before a team spends a quarter fixing the wrong problem. In my women's cardiac case study, heart attacks recorded as anxiety at intake became false ground truth for any model trained on them.
Data Quality Read more →- Separating structural label noise from random annotator error
- Tracing noise back to source: protocol, annotator, or edge case
- Prioritizing fixes by downstream model impact
- Case study: women's heart attacks mislabeled as anxiety, traced from intake labels back to clinical notes with natural language processing (NLP)
- The fix: bias-aware annotation, validation split by patient subgroup, and a conflict layer that flags when label and evidence disagree
- Preventing wasted retraining cycles on the wrong root cause
Fractional AI Lead
One-person senior AI leadership: roadmap, hands-on proofs of concept, engineering standards, solo or with a vetted freelancer network. Includes recruiting and mentoring your first in-house hires.
Leadership Read more →- Full data science roadmap from zero
- Hands-on proof of concept build, solo or with vetted freelancer network
- Engineering standards and code review processes
- Recruiting, screening, and mentoring first in-house hires
- Clean handoff when you are ready to scale independently
Expert-in-the-Loop Annotation
Optimizing the medical labeling lifecycle by reducing clinician fatigue. For one client, active learning and a similar-cases view, where the AI shows how past cases were annotated, replaced a team of 3 experts working 6 months with 1 semi-expert working 2 days.
Data Ops- Active learning to surface the most valuable samples first
- Similar-case view: the AI shows how comparable past cases were annotated and highlights what the new case shares with them and where it differs
- Annotators decide with precedent in front of them instead of starting from a blank page
- Case study: 3 experts over 6 months cut to 1 semi-expert over 2 days
Trustworthy GenAI in Clinical Settings
Designing evaluation loops and hallucination-detection systems that catch what standard benchmarks miss. A confident, hallucinating model in a clinical pipeline is not a prompt engineering problem. It is an architecture problem.
Trust & Safety Read more →- Custom evaluation frameworks beyond standard benchmarks
- Hallucination detection and confidence scoring
- Deterministic guardrails for out-of-scope queries
- Monitoring loops for production drift detection
- Architecture review for cascading vs LLM-maximalist designs
Go Deeper
Short pages on the topics I work on most, each with its own FAQ.
Structural label error
Why systematic mislabeling survives more data and retraining, with my case study on women's heart attacks recorded as anxiety.
Read the page → Generative AI evaluationFailure-to-Metric
How I design custom evaluation metrics from real production failures, with Refusal Rate and Hallucination Specificity Score.
Read the page → ArchitectureCascading AI architecture
Rules, classic machine learning, and confidence-gated LLM escalation, and why the LLM should retrain the cheaper model beneath it.
Read the page → Working togetherFractional AI lead
Part-time senior AI leadership: who it suits, what I do, how an engagement runs, and when I am not a good fit.
Read the page → TalksSpeaking
Upcoming and past talks, topics, and a short bio for conference organizers.
Read the page →How I Work
- 1
Strategy call
You describe the clinical problem, the data you have, and what success has to look like. I tell you honestly whether AI is the right tool yet.
- 2
Diagnose
I look at your data, your labels, and your real failure cases against your key performance indicators (KPIs), so the plan targets the actual problem.
- 3
Build
I write the roadmap and build the proof of concept or pipeline, solo or with a vetted freelancer network for larger scopes.
- 4
Hand off
I set engineering standards, help recruit and mentor your first data hires, and step back when your team can run it.
Case Studies and Experience
From Data Mountains to Investor-Ready Proofs of Concept
Transforming complex, unstructured clinical data into validated ML and GenAI proof-of-concepts for HealthTech startups. I help founders and CTOs go from an idea on a whiteboard to a working prototype that investors can actually evaluate.
Trustworthy AI and Engineering Rigor
Delivering reliability through custom evaluation loops and hallucination-detection systems. Applying Industrial Engineering principles to optimize the AI lifecycle, shortening annotation rounds and improving model throughput, cost-efficiency, and clinical-grade reliability.
Fractional Leadership and Team Building
Embedding as a senior AI lead to define data science roadmaps and engineering standards from the ground up. Recruiting and mentoring first data hires, and managing a vetted freelancer network, so startups get the right expertise at every stage without full-time overhead.
Structured Clinical Query Design
Drafting structured prompts for a biomedical research platform, turning open clinical questions into tabular, citation-backed answers researchers can actually check against the source.
Custom GenAI Evaluation Design
Built the Failure-to-Metric method around two proprietary metrics, Refusal Rate and Hallucination Specificity Score, for teams whose standard benchmarks don't catch what actually goes wrong in production.
Fetal Growth Anomaly Detection
Developed a personalized monitoring model for fetal growth tracking and early anomaly detection, built entirely from routine ultrasound measurements. It estimates a growth trajectory for each fetus and monitors later measurements against it, using Statistical Process Control (SPC), a method from industrial process monitoring. I am now extending it as a co-investigator with The Open University of Israel and Beilinson Hospital's Women's and Maternity Hospital.
Structural Label Error in Women's Cardiac Diagnosis
Using a large clinical database, I identified women whose confirmed heart attack contradicted their intake label. Clinical notes showed the cardiac signal was there from intake, so the error came from labeling, not from the presentation. The result is a detection and correction method that may apply to other clinical areas with embedded diagnostic bias.
Mentoring and Teaching
I mentored data science school student teams to Best Project awards in 2024 and 2025, and I volunteer as a team mentor on nonprofit tech projects. I bring the same habits to reviewing code and coaching first hires inside client teams.
Talks and Writing
Speaker bio, topics, and all talks →
Don't Let Your RAG Improvise in the ICU
Your eval passed. Your model is confident. Somewhere in that RAG pipeline (retrieval-augmented generation, where the model pulls in outside documents before answering) it just made up a drug dosage. Why that happens, how to catch it, and what trustworthy GenAI looks like when the stakes are real.
▶ Watch on YouTubeUpcoming Talks
The Evolution of the Data Scientist: From Agentic Euphoria to Cascading AI Architecture
Putting an agent on every task ran into cost and latency. I show a three-tier cascade (deterministic logic, classic ML triage, confidence-gated LLM escalation) and how to upgrade it so the LLM retrains the simpler and cheaper model beneath it. I also cover why the confidence threshold is an economic cost-and-risk decision, not a technical setting.
Full abstract
For a while, the instinct was to put a "cool" agent responsible for every task. Then the bills and latency came in. Most of that work could've been handled far more cheaply by classic ML.
That's why production teams are shifting to a three-tier cascade: deterministic logic, classic ML triage, and confidence-gated LLM escalation. I'll show this architecture, then how to upgrade it.
Most teams stop at escalating to the LLM only when needed. The real improvement is a teacher-student loop where LLM output on hard cases retrains the classic model beneath it. The LLM's job isn't just to answer what the cheap tier can't; it ensures the cheap tier needs it less next time.
I'll also cover the trap teams fall into: treating the confidence threshold like a technical setting rather than an economic cost-and-risk trade-off that can build or break user trust.
You'll leave with a blueprint for this hybrid cascade and a pattern for using LLMs to teach your models, not just do their work.
Structural Label Error: A Case Study in Women's Cardiac Misdiagnosis
, Health Care, Rehabilitation, and Innovation Conference
Women's heart attacks are often recorded as anxiety at intake, and models trained on those labels inherit the mistake. I show how I traced the error to its origin in clinical notes, and a three-step way to detect and correct it.
Full abstract
Background: Women's acute myocardial infarction (heart attack) is frequently misdiagnosed as anxiety at intake, delaying treatment. These errors can become embedded in clinical datasets, teaching downstream models the same incorrect label as ground truth.
Objectives: To identify structural label error, a systematic mislabeling pattern tied to patient subgroups, in intake diagnoses for female cardiac patients, and to develop a method for detecting and correcting it.
Methods: Using a large clinical database, I identified female patients whose confirmed heart attack diagnoses contradicted their intake label. I applied natural language processing (NLP) to clinical notes to trace the discrepancy to its origin, then developed a three-step correction approach: bias-aware annotation, disaggregated validation, and a conflict layer flagging disagreement between label and evidence.
Results: Clinical notes showed cardiac signal was present from intake in the miscoded cases, indicating the error originated in labeling rather than in the clinical presentation itself. This pattern was consistent enough to function as a false ground truth for any model trained on it.
Conclusions: Structural label error in women's cardiac diagnosis is systematic, not random, and can be identified and corrected using the proposed annotation approach. The same method may apply to other clinical domains where training labels carry embedded diagnostic bias.
Personalized Fetal Growth Monitoring: From Statistical Process Control to a Clinical Model
, TLV Lifesciences, Research Track
A fetus can drift away from its own growth trajectory while every measurement stays inside the population range. I present a proof of concept that builds an individual trajectory from early ultrasound and monitors later measurements against it, and what it takes to make that usable in a real clinical system.
Full abstract
Fetal growth monitoring today relies mostly on population reference curves, which can miss meaningful changes in an individual fetus's growth trajectory even when measurements stay within the accepted range. I'll present a proof-of-concept project that estimates an individualized growth trajectory for each fetus from early ultrasound measurements and monitors later measurements against it, using ideas adapted from Statistical Process Control (SPC), a method originally developed for industrial process monitoring.
The approach builds on earlier published research in Quality Engineering, and Dr. Diamanta Benson, Head of the Industrial Engineering and Management Program at The Open University of Israel, and I are now extending it as co-investigators with Beilinson Hospital's Women's and Maternity Hospital, using richer longitudinal data and current statistical and machine learning methods. I'll talk about what it actually takes to make an idea like this usable in a real clinical system, and the main obstacles along the way.
The Failure-to-Metric Method: Building Evals Your Platform Can't
, AI Native Week
Building a custom eval metric is easy now. Knowing which metric serves your real business goal is the hard part. I start from a real production failure and design the metric around the moment things went wrong, with two working examples, Refusal Rate and Hallucination Specificity Score, from high-stakes clinical workflows.
Full abstract
Building a custom eval metric is easy now; every platform lets you do it. The harder part, and the one that actually matters, is knowing which metric will serve your real business KPI.
I call this the Failure-to-Metric method. Start from a real production failure, not a benchmark. Decide if it's even worth a dedicated metric. If it is, design it around the moment things went wrong, not just where they ended up: did the system refuse when it should have, and how costly was the answer when it didn't.
I'll walk through how to recognize which failures actually matter and connect them to the KPI they're quietly damaging, using examples from high-stakes clinical workflows, where a wrong answer isn't just a bad user experience, it's a patient safety issue.
Then I'll show two metrics built this way. A Refusal Rate metric, scoring whether the system correctly declines to answer when it should (higher is better). A Hallucination Specificity Score, scoring how dangerous the answer was when it didn't refuse (lower is better).
You'll leave with the method, two working examples, and a way to decide which of your own production failures deserve a metric.
Past Talks and Articles
Modeling for Ultra-Small Data
hayaData 2025
How personalized modeling delivers clinical-grade insights when traditional Big Data approaches fail.
Watch Talk →The Caveman Skill
LinkedIn Article
A prompt engineering technique for clinical AI, with unit economics arguments and a deployment decision framework.
Read Article →
Fixing the HealthTech ROI Gap
Article
Why Data ROI is an overlooked metric for HealthTech founders, and how to fix it before it's too late.
Read Article →
Building Modular AI Architectures
Article
Shifting from monolithic prompting to modular agent architectures for reliable, production-ready enterprise systems.
Read Article →
Publications
- Shore, H., Benson-Karhi, D., Malamud, M. & Bashiri, A. (2014). Customized Fetal Growth Modeling and Monitoring - A Statistical Process Control Approach. Quality Engineering, 26(3), 290-310.
- Benson-Karhi, D., Shore, H. & Malamud, M. (2017). Modeling Fetal-Growth Biometry with Response Modeling Methodology (RMM) and Comparison to Current Models. Communications in Statistics - Simulation and Computation, 47(10).
Fractional AI Lead
Senior AI leadership for startups that need to move fast, without the overhead of a full-time hire.
Roadmap & Architecture
Defining your data science strategy, selecting the right models, and laying the engineering foundations your future team will inherit.
Hands-On Execution
Building proofs of concept and production-ready pipelines directly, solo or with a vetted network of specialized freelancers for larger scopes.
Team Building & Handoff
When you're ready to hire in-house, I help recruit and mentor your first data talent. Then hand off a team that can run independently.
FAQ
1. Who is Maya Malamud and what does Malamud AI Advisory do?
I am an AI and data science consultant based in Madrid, with 15 years of experience in machine learning (ML) and statistics. Through Malamud AI Advisory I work with early-stage HealthTech teams on clinical AI: cascading pipeline architecture, expert-in-the-loop annotation, custom evaluation of generative AI, and fractional AI leadership. I work in English, Spanish and Hebrew.
2. Why hire an AI advisor instead of a standard software development agency?
I push back on scope that doesn't earn its place, so what gets built actually serves your business and clinical goals. Sometimes the right answer is a simpler model, or no model yet.
3. For whom is Malamud AI Advisory not a good fit?
I am not a good fit for startups who want AI to work like magic without the data work underneath it, or projects that don't prioritize clinical/business validity. I focus on founders who value engineering rigor and strategic growth over quick, unscalable hacks.
4. How do I know if my AI pipeline has a hallucination risk before I ship it?
Standard benchmarks often miss it. I run a short diagnostic against your actual failure modes, not generic test sets, to surface where a confident model is likely to make something up before it reaches a patient or clinician. I start from a real production failure and design the metric around it, a method I call Failure-to-Metric. Two examples are Refusal Rate and the Hallucination Specificity Score.
5. How do I stop a clinical chatbot or RAG system from hallucinating?
I treat it as an architecture problem, not a prompt problem. Retrieval-augmented generation (RAG) means the model pulls in outside documents before it answers, and it can still make things up. I add deterministic guardrails for out-of-scope questions, confidence scoring so the system can refuse when it should, evaluation metrics built from your own failures, and monitoring in production to catch drift.
6. What is a cascading AI architecture, and why not send everything to an LLM?
A cascade sends each case to the cheapest layer that can handle it: deterministic rules first, then a classic ML model for triage, and a large language model (LLM) only when confidence is low. A teacher-student loop lets the LLM's answers retrain the simpler and cheaper model beneath it. Where to set the confidence threshold is an economic decision as much as a technical one.
7. Why has my model's accuracy stopped improving, and is it a modeling problem?
Not always. Sometimes the ceiling is structural label noise, systematic errors in how the training data was labeled, not random ones. I diagnose which one you have before you spend a quarter retraining against the wrong target.
8. What is structural label error?
It is systematic mislabeling tied to a patient subgroup, as opposed to random annotator mistakes. In my case study, women's heart attacks were often recorded as anxiety at intake in a large clinical database. Clinical notes showed the cardiac signal was present from intake, so the error came from the labeling, and a model trained on those labels learns the mistake as ground truth. I address it with bias-aware annotation, validation split by subgroup, and a conflict layer that flags when a label and the evidence disagree.
9. How do you reduce annotation cost for clinical data?
Active learning surfaces the most informative cases first. For each one, the AI shows how similar past cases were annotated and highlights what the new case has in common with them and where it differs, so the annotator works from precedent instead of a blank page. For one client this replaced a team of three experts working six months with one semi-expert working two days.
10. Can you build models when I only have a small dataset?
Yes, when the problem fits. I build personalized models from ultra-small clinical data. One example is an individualized fetal growth trajectory monitored with Statistical Process Control (SPC), a method from industrial process monitoring. I am a co-investigator on the proof-of-concept extension of that work with Dr. Diamanta Benson of The Open University of Israel and Beilinson Hospital's Women's and Maternity Hospital.
11. What is "Fractional AI Leadership" and do I need it?
It's part-time senior leadership for seed-stage startups that need an AI strategy before they can afford or justify a full-time Chief AI Officer. If you're struggling to define your AI roadmap or hire your first data science team, this is for you. I work solo, or with a vetted freelancer network for larger scopes.
12. What does an engagement look like?
It starts with a strategy call. From there I do the smallest piece that answers your question: a diagnostic such as a label noise audit or a hallucination-risk review, a proof of concept, or ongoing fractional leadership. I work solo, or with a vetted freelancer network for larger scopes. I agree scope and cost with you after the first call.
13. How do you ensure AI investments actually deliver ROI for Seed/Series A startups?
By starting from the simplest model that works, not the most exciting one. Every dollar spent on compute or talent has to earn its way back.
14. How do I get a healthcare AI pilot into production?
I look for the usual blockers first: unclear key performance indicators (KPIs), noisy or systematically wrong labels, and an architecture that sends every case to the most expensive model. Then I define the KPI, audit the labels, and build the simplest pipeline that meets it, with monitoring in place. When you are ready to hire, I help recruit and mentor the team that runs it.
15. When is it too early to bring in a Data Science team?
If you haven't defined your core KPIs (key performance indicators) or identified the specific business problem AI is meant to solve, it's too early for a full-time team. I help you build the prototype and strategy first, saving you months of expensive "exploratory" hiring.
16. How do you help with technical talent acquisition?
I don't just recruit. I help you build the team, from the hiring plan through mentoring your first hires, so you start with high engineering standards and a clear mission.
17. Do you speak at conferences or run workshops?
Yes. Upcoming talks are listed under Talks and Writing, and I also give talks and workshops for healthcare and engineering teams. To discuss one, message me on WhatsApp or book a strategy call.
A confident, hallucinating model in a clinical pipeline is not a prompt engineering problem. It is an architecture problem.
Maya Malamud, AI & Data Science ConsultantLet's Build Something Trustworthy
Whether you're a founder scaling soon, or a CTO looking to enrich your team, let’s talk strategy.
AI & Strategy Advisory
For HealthTech startups and companies looking for fractional AI leadership or architecture review.
Book a Strategy CallSpeaking & Workshops
For healthcare conferences and teams, looking to enhance their AI capabilities.
WhatsApp Chat