Shipping AI Providers Trust: Two Features, One TPM Playbook
I put two AI features into the hands of 200+ clinical users, and I owned more than the timeline: I decided what the AI actually did. I set the thresholds, rules, and model selection for evaluation suites that flag and interpret lab results, and for a dashboard that reads a clinic's practice health, then drove the rollout across engineering, clinical, and product. This is where AI product judgment and program discipline meet.
Two jobs, held at once: shape the AI, ship the program
On both features I carried the product requirements and the delivery at the same time. I defined what the AI did, the thresholds it judged against, the rules it followed, and which model fit the job, and I ran the program that turned those decisions into something live: a charter, the dependencies between engineering, clinical, and product, a phased rollout, and the stakeholder alignment to keep everyone moving on one plan. The two halves below show that pairing in practice.
What the AI did
The product layer: requirements, thresholds, rule logic, and model selection. These were product decisions, and I owned them with engineering rather than handing off a spec.
How it shipped
The program layer: charter, cross-functional dependencies, phased rollout, validation gates, and the stakeholder cadence that kept engineering, clinical, and product aligned to launch.
AI-assisted evaluation suites
AI that analyzed raw lab data, flagged abnormal results, and surfaced diagnostic interpretations for providers and patients.
Raw results, expert interpretation, a manual bottleneck
Lab panels produced large volumes of raw values, but the value to a provider or patient is the interpretation: what is abnormal, what it means, and what to do next. That read depended on expert time, which made it slow and uneven. Two patients with similar results could get different depth of explanation depending on who reviewed them.
The opportunity was to put a consistent, fast first-pass interpretation in front of providers and patients, while keeping clinical judgment in the loop where it mattered.
Own the timeline and the requirements
My mandate covered both sides. I had to define what the AI flagged and how it interpreted results, the thresholds, the rules, and the right model for the job, and I had to drive the program that built and shipped it across engineering, clinical, and product. Getting the AI logic right and getting the rollout right were the same job.
Shaping the logic, then driving it to live
What the AI did
- ruleThresholds & rules: worked with engineering and clinical to define what counted as normal, abnormal, or escalate, so flagging was explainable, not a black box.
- psychologyModel selection: weighed model options against accuracy, interpretability, and the cost of a wrong call, and picked the fit for a clinical context rather than the flashiest option.
- diversity_3Two audiences: shaped the interpretation so a provider got clinical depth and a patient got plain-language meaning from the same underlying result.
How I drove it
- account_treeCross-functional execution: aligned engineering, clinical, and product on one charter and mapped the dependencies between them so handoffs did not stall.
- verifiedValidation gates: clinical sign-off on flagging behavior was a gate, not a courtesy, before anything reached users.
- rocket_launchPhased rollout: expanded access in stages, watching behavior at each step before widening, on the way to 200+ users.
Raw lab data
Panel values arrive from the lab system as the input to evaluation.
Evaluation engine
Threshold & rule layer: normal / abnormal / escalate
Model: interpretation selected for clinical fit
Clinical review gate: human-in-the-loop on escalations
Surfaces
Provider view with clinical depth; patient view in plain language.
Step 1
Read result
Compare each value to its threshold band.
In range
Normal
Interpretation surfaced, no flag raised.
Out of range
Flag
Abnormal marker with plain-language meaning.
Critical
Escalate
Routed to clinical review before it reaches users.
Faster, consistent reads, with a human on the risky ones
200+
users on rollout
Faster
first-pass interpretation
Consistent
reads across patients
Verified outcome: the evaluation suites rolled out to 200+ users, giving providers and patients a faster, more consistent first read of lab results while a clinical-review gate kept judgment on the calls that needed it.
Clinic analytics dashboard
A provider-facing dashboard where AI analyzed how a clinic was performing, from financial signals to patient interaction, and explained what it meant.
Plenty of data, no intelligent read
Providers ran on signals scattered across systems: financial performance on one side, patient interaction on the other. Seeing the numbers was possible; understanding what they meant for the health of the practice was not. The story was buried in spreadsheets and disconnected reports.
The opportunity was a single view that did not just display the numbers but read them: surfaced what was working, flagged what was not, and pointed at why.
Deliver a dashboard that explains, not just displays
I led delivery of a provider-facing dashboard where AI analyzed practice performance end to end, from financial signals like revenue, collections, and reimbursement trends down to patient interaction metrics like visit volume, follow-up rates, and engagement, and turned them into an intelligent read on business health. The bar was insight, not a prettier report.
From scattered signals to an explained view
What the AI did
- mergeSynthesized signals: combined financial and patient-interaction data into one read of practice health instead of two disconnected views.
- lightbulbSurfaced insight: explained what was working and what was not, in words, rather than leaving the provider to infer it from charts.
- trending_upFlagged trends: called out movement worth attention so a slipping metric did not hide inside a busy dashboard.
How I drove it
- account_treeData dependencies: coordinated the financial and interaction data sources across teams so the dashboard had one trustworthy feed.
- rocket_launchPhased delivery: shipped in stages, validating each data domain before layering on the AI read.
- handshakeStakeholder alignment: kept product, engineering, and the provider-facing teams aligned on what "useful" meant.
Financial signals
Revenue
Collections
Reimbursement
Patient interaction
Visit volume
Follow-up rate
Engagement
AI insight: revenue is steady, but follow-up rate is trending down and is the likeliest drag on repeat visits. Worth a look at scheduling.
Trend flag: follow-up rate down vs. prior period.
An at-a-glance read on practice health
Outcome: providers got a single, explained view of business health across financial and patient-interaction signals, with the AI surfacing what was working, flagging what was not, and pointing them at the trends worth acting on, instead of leaving the story buried in disconnected reports.
Where AI product thinking met TPM execution
Across both features, the through-line was the same: I made the AI product decisions and I ran the program that shipped them. Three moments capture that intersection.
I specified the model, not just the milestone
Thresholds, rule logic, and model selection were product decisions, and I owned them with engineering. The schedule mattered, but so did what the AI actually decided, and I shaped both.
Human-in-the-loop by design
For diagnostic AI, safe failure is the feature. I built clinical-review gates into the flow so the system accelerated the routine reads and deferred the risky ones to a clinician.
From numbers to narrative
The dashboard's value was the AI explaining practice health, not displaying it. Delivering that meant coordinating the cross-team data dependencies that made one trustworthy read possible.
What I'd do differently
Across both features, what I would change next time.
- check_circleDefine the evaluation metrics before the model. I would lock how we measured a "good" flag, precision against the cost of a miss, even earlier, so model selection was scored against an agreed bar from day one rather than refined in flight.
- check_circleInstrument trust, not just usage. For both features I would ship feedback capture with v1, was an insight acted on, was a flag overridden, so the AI's value was measured by provider trust, not only adoption.
- check_circleBring clinical review into design earlier. The review gate worked, but pulling clinical stakeholders into the threshold and rule design from the start would have shortened the validation loop before rollout.