Clinical notes, prior-auth, and coding are begging to be automated, but nothing can touch a public model under HIPAA. We make private inference the default and gate the rest.
Every health system is running the same experiment right now: clinicians paste into consumer chatbots because the sanctioned tools are slower than the unsanctioned ones. That is PHI crossing a boundary it can never come back across, and the fine print of a public API is not a business-associate agreement.
The uncomfortable math is that healthcare's highest-value AI workloads (ambient documentation, coding support, prior-auth triage) are also its most sensitive. Waiting is not neutral: shadow usage grows while the policy committee meets. The answer is not to ban the behavior; it is to make the compliant path the fastest one.
Private small language models changed this calculus. Clinical summarization does not need a frontier model. It needs a specialized model that lives inside your HIPAA boundary, tuned to your note formats, with an audit trail an OCR investigator can walk through.
On-prem and private SLMs for clinical summarization, coding, and prior-auth triage. PHI stays inside your boundary and nothing leaves for a public API.
Least-privilege access per department, immutable audit trails, and mandatory human sign-off on anything clinical or diagnostic.
A TCO audit scoped to a HIPAA boundary, then a private intelligence stack that keeps sensitive workloads local and cloud reserved for the non-sensitive tail.
Private SLMs on owned or colocated hardware for anything touching PHI; frontier models only through a gate that strips and logs.
Least-privilege per department and role, so the pharmacy model does not read oncology notes.
Immutable logs of every prompt, model, and output touching patient data: evidence, not assurances.
Anything clinical or diagnostic routes through a human before it acts. Non-negotiable by design.
A break-even model for on-prem clinical AI versus per-token pricing at your actual note volume.
HIPAA compliance is a property of the deployment, not the model. Private inference inside your boundary, access controls, audit trails, and a BAA where a vendor is involved: that combination is achievable today with private SLMs.
For narrow, high-volume tasks like summarization, coding suggestions, and triage, specialized small models routinely match or beat general frontier models, because the task is narrow and the context is yours. We prove it with a blind bake-off on your own documents before anything deploys.
The four-minute readiness assessment scores your data, security, and governance posture, and the TCO calculator models on-prem versus cloud at your volumes. Both are free and run in your browser.