AI will not replace healthcare’s operating system. It is becoming part of it.
That sounds grand until you look at where the work is already happening. Patients are connecting records and wearable data to AI. Clinicians are using it to make notes and prepare information. Health systems are testing it against the administrative work that exhausts their staff. The model is only one component. The important work is connecting it to the real world without losing trust.
I have been thinking about this while working on a healthcare client project and speaking with founders in the space. My conclusion is straightforward: the opportunity is enormous, but a health agent earns its place only when it is reliable enough for the environment around it.
Health is moving from information to context
OpenAI launched Health in ChatGPT on July 23, 2026. The product lets eligible U.S. users connect Apple Health and supported medical records. With permission, ChatGPT can use that information to compare results over time, help prepare for appointments, and make everyday questions more specific to the person asking them.
That is a meaningful shift. A model can explain a lab result in the abstract. A useful system can help someone understand what changed, what they should ask, and what information is still missing. It should support professional care, never pretend to replace it.
The larger market is moving in the same direction. In a recent Lux Capital conversation, Dina Shakir described the opportunity moving beyond models into deployment and trust infrastructure. That is exactly right. Healthcare does not need another detached intelligence layer. It needs systems that can operate safely inside real workflows.
I wrote earlier about Agentic Healthcare: AI-first Healthcare Solutions. I still believe the core idea: models are ready to help, while the infrastructure around them has to catch up.
The hard part is the surrounding system
Healthcare is full of versioned, consequential detail. HIPAA’s Administrative Simplification provisions cover standards for electronic transactions and code sets. In day-to-day work, that means systems must understand the right coding language and the rules around it: ICD-10, HCPCS, CPT, CDT, and NDC.
Consider ICD-10. CMS has already published FY 2027 ICD-10-CM update files for encounters from October 1, 2026 through September 30, 2027. It has also published the separate ICD-10-PCS files for the relevant 2027 period. There are code tables, an index, conversion material, addenda, and guidelines.
An agent cannot simply remember a code family and guess. It needs the correct edition, a traceable source, the right clinical context, and a clear boundary between suggestion and decision. That is not bureaucracy around the product. It is the product.
A prompt is not a safety boundary
The cybersecurity disclosures from OpenAI and Anthropic made this point sharply. Anthropic said that a third-party evaluation environment had live internet access even though the model had been told it was isolated. The model interpreted real systems as part of its simulated task, and the incident led to unauthorized access to three organizations.
This was not a story about a model deciding to become malicious. It was a systems failure: incorrect assumptions, an open network path, and controls that were not strong enough for an autonomous capability.
Healthcare teams should take the lesson seriously. Instructions are useful, but they are not enforcement. A dependable agent needs least-privilege access, explicit scope, network and data isolation, versioned sources, audit logs, monitoring, and human review for consequential actions. It must fail safely when a record is missing, a source conflicts, or an action falls outside its authority.
Test the world the agent will meet
I have experience building reliable solutions that must meet compliance requirements. I also believe my work in testing and simulation testing is central to how these systems become possible.
Do not only test whether an agent can produce a convincing answer. Simulate stale code sets, incomplete records, conflicting evidence, ambiguous documentation, broken integrations, and requests outside the agent’s authorization. Test what happens when the model is wrong, when the context is wrong, and when the environment is wrong.
The standard is not perfect language. It is predictable behavior under pressure.
That is why I am optimistic. Healthcare has difficult constraints, but those constraints can make the systems better. They force teams to define authority, preserve evidence, expose uncertainty, and measure whether the system did what it was meant to do.
The future I want is not autonomous medicine. It is capable, grounded software that gives patients and clinicians more clarity, more time, and better options. We can build it. But we have to build the trust infrastructure with the same seriousness as the intelligence.