The biggest shift in AI drug discovery:
The special models & supercomputers won’t be the edge much longer. Everyone will have them. The real advantage will be how fast you turn AI into real lab-tested learning and better decisions about which molecules to advance. Tools are
The same structural shift shows up in primary care AI, and it's clarifying something I've been thinking about since writing on eConsults.
The Ontario program processed nearly 100,000 asynchronous specialist consultations with a two-day average turnaround. That volume didn't come from a better model. It came from a workflow that converted decisions into documented feedback loops, fast enough to actually learn from them.
That's the experimental velocity point applied to clinical settings. The PCPs who will pull 20-30% of referral volume back in-house won't be the ones with access to the best diagnostic AI. They'll be the ones who built a cycle where AI flags the case, a specialist responds asynchronously, and that exchange gets captured in a way that improves the next decision.
The documentation trail is where the real compounding happens, both for clinical skill and for the liability case. A completed eConsult with a specialist's input on record is more defensible than a referral the patient never completed, which is most of them.
The model is table stakes. The learning loop is the moat.
https://www.onhealthcare.tech/p/the-pcp-as-specialist-how-ai-and?utm_source=x&utm_medium=reply&utm_content=2082106939316764808&utm_campaign=the-pcp-as-specialist-how-ai-and
NYC Woman Dead After Receiving a Longevity Infusion
“Stories like this are heartbreaking,” said Dr. Matt Kaeberlein... “Unfortunately, it isn’t the first time someone has been seriously harmed – or killed – while pursuing an unproven ‘longevity’ therapy. We still don’t know
Vocabulary is doing the killing here, not just the compounds.
When "peptide" or "longevity therapy" gets treated as a coherent category rather than a spectrum running from rigorous phase 3 evidence to zero human data, the regulatory enforcement gap becomes a patient safety gap. The same conflation I traced through the wellness peptide market in https://www.onhealthcare.tech/p/the-peptide-split-how-glp-1s-lutathera-f57?utm_source=x&utm_medium=reply&utm_content=2080679384667795503&utm_campaign=the-peptide-split-how-glp-1s-lutathera-f57 is what allows an infusion clinic to borrow credibility from semaglutide outcomes data while injecting something that has never cleared a dose-response study in a human being.
The 503A compounding framework and the "research use only" mail-order model were not designed for IV administration in clinical-adjacent settings, and enforcement has not kept pace with how aggressively that gap is being exploited commercially.
Dr. Kaeberlein is right that this won't be the last case, because the problem is structural. The word "longevity" is doing the same epistemic damage as the word "peptide," bundling a handful of genuinely promising biological hypotheses together with untested infusion products and letting the halo of the former cover the risk of the latter.
The question that keeps pulling at me is whether any realistic enforcement mechanism can actually close that gap before the next death, or whether the category vocabulary problem has to be solved first for regulators to even know what they're targeting.
This is the stock price for a biotech that has invented a drug which many considered to have basically cured one type of cancer (multiple myeloma), at least in a subset of patients.
Cell therapies, despite being transformative medicines, are being gutted by manufacturing costs, https://t.co/ZlbtBcCPLu
Transformative clinical outcomes and commercial collapse can coexist, and this stock chart is the proof.
The failure mode here is not manufacturing alone. Manufacturing is where the money bleeds, but the deeper problem is that cell and gene therapies were approved into a healthcare operating system that was never designed to finance, coordinate, or sustain them. CASGEVY sits at $2.2M list price with roughly 60,000 eligible patients across approved geographies and posted $43M in Q1 2026 revenue. That gap is not a science problem or even purely a manufacturing problem. It is a reimbursement mechanics problem, a Medicaid actuarial problem, a transplant center capacity problem (the treatment is really a multi-month coordinated services bundle, not a drug). The companies that invented these therapies built extraordinary science and then handed it to a payer and provider infrastructure that has no working model for one-time curative interventions.
The next value capture in this space will go to whoever builds the missing stack: outcomes-based contracting rails, reinsurance structures for curative therapies, patient activation platforms that can actually move eligible patients through the workflow. The editor platforms already did the hard scientific work. The commercial work is almost entirely operational and financial, and almost nobody is funding it.
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2064733249029439979&utm_campaign=gene-editing-has-the-science-figured-b80
@ThierryBorgeat·68,633 views82%
6/11/26 10:11 PM ET
Citadel Securities just put institutional weight behind what the AI bulls won't say out loud.
In a new macro note titled "Tokenomics," Citadel makes the argument plainly: even the most powerful technology on earth still has to pass through the boring discipline of cost curves, https://t.co/fzPHKk9gzq
The healthcare version of this is even more compressed. The cost curve problem doesn't just slow deployment, it creates a specific void where no institution is even empowered to decide if the better answer is worth the price.
FDA clears for safety. CMS pays for procedures. Nobody owns the question of whether a reasoning-heavy query (the kind that actually moves diagnosis) justifies its token cost. So the cost curve discipline Citadel is describing has nowhere to land in clinical settings.
Wrote through exactly this gap here: https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2064783848710303902&utm_campaign=token-economics-versus-the-20-watt-995
CNBC interviewer asked Palantir CEO Alex Karp how he would defend Wall Street’s concern that AI could replicate what Palantir is doing.
Karp defended by basically saying that AI companies may have great engineers, but they do not deeply understand the messy, high-stakes https://t.co/D2adPO3DJr
That's the exact argument I made for healthcare specifically at https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2064824535569064156&utm_campaign=the-standardization-trap-why-deploying. Two health systems on the same Epic instance can have completely divergent clinical data models (custom flowsheet rows, local formularies, legacy migration artifacts) that no amount of model capability gets around. The 60-70% of the stack that's commoditized isn't the moat. The embedded workflow knowledge is.
A non-profit health system can refer a patient to its own MRI, its own lab, its own surgery center, and bill all three.
An independent does that once and it's a federal felony.
Same referral.
Same patient.
One of you goes to prison.
It's called Stark Law.
Read who's exempt.
The asymmetry is the whole business model.
The nonprofit exemption is real, but the more precise mechanism worth tracking is what happens *after* the referral asymmetry compounds across the full regulatory stack.
Stark's strict liability structure, with penalties reaching $15,000 per violation and $100,000 per circumvention scheme, didn't just criminalize the independent physician's referral. It generated an entire compliance industry around fair market value assessments and physician compensation analyses, which large health systems can absorb as overhead and small independents cannot. The exemption isn't just a legal carve-out, it's a cost structure advantage that widens every year a small practice has to pay outside counsel to review arrangements the nonprofit handles with an in-house team.
Then layer in what happened with the ACA's 2010 closure of the whole hospital Stark exception. Congress didn't ban physician-owned hospitals categorically. It closed a specific Stark exception, which means any physician already participating in Medicare lost the legal pathway to build a competing facility. The nonprofit system didn't need that pathway closed because it was never using it. The regulatory action fell entirely on the competitive threat.
(The CBO scored that closure at only $500 million in deficit reduction over ten years, which tells you the government wasn't primarily doing fiscal math when it wrote that provision.)
What the post identifies as asymmetry is actually path-dependent regulatory accumulation, each layer rationally designed but collectively functioning to entrench whoever was already at scale when the law passed.
More on how this stacking works across the full regulatory architecture: https://www.onhealthcare.tech/p/how-the-government-built-a-cage-around?utm_source=x&utm_medium=reply&utm_content=2064482780474417357&utm_campaign=how-the-government-built-a-cage-around
At our recent Energy and the AI Age summit, Hon. Bernard L. McNamee, former Commissioner of the U.S. Federal Energy Regulatory Commission, comments on why energy matters in discussions about AI:
“Energy is the foundation of our entire economy. Energy makes up about 7% of the https://t.co/tqFZbaSfvJ
Good framing from McNamee, and it connects directly to something I've been working through in healthcare specifically.
The 7% figure matters, but the harder problem is directional: clinical AI workloads are not static draws on the grid. Real-time ICU systems processing vitals, imaging, and lab values across a whole health system simultaneously are a different animal from a chatbot. The compute demand is continuous, not on-demand, and the power envelope per inference has to shrink before the economics close for most medical use cases outside billing and coding.
What gets missed in the "energy matters for AI" conversation is that the binding variable is not just total grid capacity. It is compute-per-watt at the point of care, in the OR, in the ambulance, in the rural clinic with no data center nearby. That is where Nvidia's efficiency curve, doubling roughly every two to three years, starts to look less like a chip story and more like an energy story.
The printing press changed what humans could read. Reliable current changed what medicine could do. LLMs are real, but they are the Gutenberg moment, and the next unlock is still in the wire.
https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2064718261988515880&utm_campaign=the-pattern-always-repeats-why-healthcares
AI assistants are moving from "answer my question" to "do the work." But they are only as useful as the enterprise content they have access to.
Our demo shows what it looks like when @Copilot Cowork is grounded in Box. Your governed content powering multi-step agentic workflows, https://t.co/5JqblOplAz
What Box and Copilot are showing in enterprise content is the same access-layer problem I watched play out at HIMSS26, just with higher stakes when the content is PHI.
athenahealth's MCP server announcement was the most technically significant thing at the conference, and the reason is exactly what you're describing: agents are only as useful as the data they can reach. But in healthcare, "grounded in your content" means permissioned access to clinical records, and that creates regulatory surface area that governance tools are nowhere near keeping up with. The vendors who own that access layer own the workflow, full stop.
https://www.onhealthcare.tech/p/himss26-field-notes-the-agentic-turn?utm_source=x&utm_medium=reply&utm_content=2064792374086340875&utm_campaign=himss26-field-notes-the-agentic-turn
Version 2 of our ShockCalcs hemodynamics simulator is live. I've refined the physiology a ton, added new meds, and also a real-time Frank Starling curve that responds to fluids + vasopressors.
Check it out (link in reply), and reply here with feedback on how to improve further! https://t.co/MEfaomh0A8
The Frank-Starling curve responding in real time to fluids and vasopressors is exactly the kind of forward simulation that matters, because it forces the model to reason about intervention effects, not just current state. That's the gap I wrote about: pattern recognition can flag a sick patient, it can't tell you whether 2L of saline helps or drowns them.
Where I'd push for v3: the simulator needs to track how clinician behavior shifts in response to its outputs. If your tool changes how physicians fluid-resuscitate, your training data from yesterday no longer reflects the patient population of tomorrow, that feedback loop corrupts static models fast. 0 of the current clinical decision tools I've seen handle this well.
More on why that architectural problem matters more than most people realize:
https://www.onhealthcare.tech/p/world-models-walk-into-a-hospital?utm_source=x&utm_medium=reply&utm_content=2064771949801132174&utm_campaign=world-models-walk-into-a-hospital
Dylan Patel, founder of SemiAnalysis:
"The upper bound on how much compute can be produced by 2030 is around 200 gigawatts a year."
The entire world has about 20 gigawatts of AI deployed right now. The ceiling is 10x what exists today, and it still isn't enough to feed what https://t.co/WZxU70PxQn
Dylan's 200GW ceiling by 2030 is actually the number that should be rattling health tech investors right now, and almost none of them are pricing it in. I've been writing about how compute cost, not FDA clearance or EHR integration, is the actual binding constraint on clinical AI at scale, and a 10x increase in global capacity still leaves most of the hard workloads, genomic variant interpretation, real-time deterioration models at population scale, multimodal imaging plus labs plus genomics, priced out of viable reimbursement math.
The workflows that pencil out today at current AWS and Azure inference pricing are computationally light. Ambient documentation, basic coding assistance. That's it.
What the Terrafab's terawatt target changes, if even a fraction of it materializes, is the denominator that health tech has been quietly ignoring while everyone argued about FDA 510(k) pathways and Epic integrations. A 50x expansion over current global output isn't the same problem as a 10x expansion. The cost curve bends differently. And the in-house lithography mask production angle is what I keep coming back to, because it's not just about volume, it's about iteration speed for custom silicon, which completely changes whether narrow clinical applications can ever justify their own chip architecture the way Illumina did for sequencing.
So what does the clinical AI investment model look like if you're still underwriting on current inference costs when the supply ceiling is genuinely contested between 200GW and 1,000GW?
https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2065133177849499984&utm_campaign=the-elon-terrawatt-announcement-nobody
Enterprises have tolerated unstructured data governance failures for years. The blast radius was manageable because humans were the ones accessing it.
Agents are a potentially bigger challenge. Our CISO Heather Ceylan shared why AI agents make your unstructured data problem https://t.co/3YfK1hCTEC
The blast radius framing is exactly right, and healthcare is where the stakes get specific fast. An AI agent running a prior auth workflow touches medication history, substance use records, and payer data in a single session, all under whatever credential it borrowed from a human user. That's not a governance gap, it's a structural mismatch between how HIPAA's minimum necessary standard was written and how autonomous agents actually behave.
The part that doesn't get enough attention: 42 CFR Part 2 consent rules for substance use disorder records require enforcement at the data layer, with specific redisclosure controls. No current IAM stack handles that dynamically for a non-human agent making judgment calls mid-workflow.
What I've been working through is why OAuth 2.0 and SMART on FHIR v2 can't close this. Backend service scopes help, but they're static. They don't adjust when an agent's workflow state changes, and they say nothing about agent to agent scope limits when one clinical AI tool calls another.
The fix isn't better logging on top of existing access models. It's a third layer between auth and access, one that reads workflow state, consent status, and data type in real time to scope each call. Large hospitals alone could represent $200-400M in ARR for whoever builds this correctly.
https://www.onhealthcare.tech/p/whos-the-agent-building-the-identity?utm_source=x&utm_medium=reply&utm_content=2064776529700413557&utm_campaign=whos-the-agent-building-the-identity
OpenAI is reportedly considering drastic token price cuts to pull customers away from Anthropic, per WSJ. This follows rising complaints from enterprise customers about AI costs, while Anthropic has been gaining traction with Claude Code.
The message is clear: @OpenAI does not
The token price war is interesting, but it's a few layers upstream from where the real pressure lands in healthcare. What actually matters is that these cuts flow directly into tools like Cursor and agentic coding platforms, and those tools are already compressing healthcare software build costs 50-90%, as I wrote at https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2065002347927871552&utm_campaign=the-free-lunch-is-over-except-now. Cheaper tokens mean cheaper builds, and cheaper builds mean the prior auth vendor whose moat was "rebuilding this costs $4M" is now looking at a $300K problem for any health system with three decent engineers.
The OpenAI vs Anthropic price fight is good for buyers of AI. The people who should be nervous are the health tech vendors who assumed build cost was a permanent shield.
Palantir CEO Alex Karp:
"Instead of selling commodity, parasitic software with a massive salesforce and lumbering, jargon-barring leaders offering steak dinners and other things we shall not mention—in order that you turn the high-value revenue of your enterprise over to them—we https://t.co/Abjtkn2CCm
Palantir's forward-deployed model is exactly the structure I mapped when I looked at what OpenAI and Anthropic are actually building with their PE-backed deployment ventures, and the $1.5 billion Anthropic JV with Blackstone and Goldman announced one day after OpenAI's roughly $10 billion PE deal tells you both labs reached the same conclusion at the same time: the model is not the margin.
The part Karp doesn't say out loud is where the deployment substrate comes from. In healthcare, which I wrote about at https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2064902157279842724&utm_campaign=the-openai-anthropic-ai-arms-race as the stress test for all of this, Blackstone already controls physician rollups across primary care, cardiology, oncology, and other areas, plus RCM platforms and prior auth services. That portfolio is not passive capital, it's a pre-built install base that bypasses the slow health system sales cycle entirely.
Forward-deployed engineers still need somewhere to deploy. PE built that somewhere first, the labs just figured it out.
🚨BREAKING: OpenAI considering “drastic” price cuts to win the war for users with Anthropic
Altman: “costs have become a huge issue”
>already losing billions
>planning to lose more
>right before IPO
it’s so over https://t.co/NWHp0TiVVW
The part nobody's talking about in healthcare: when foundation model inference costs collapse (which is what's actually driving this), the beneficiaries aren't the AI-native health tech vendors, they're the buyers.
A payer's internal engineering team building prior auth logic from scratch just got cheaper twice over: cheaper models to reason over clinical criteria, cheaper tools to write the code. The vendors who built their pitch around "this would be too expensive for you to replicate" are now getting squeezed from both ends simultaneously.
OpenAI bleeding money to hold price isn't the story. The story is that every price war at the infrastructure layer accelerates the build-versus-buy math shifting inside health systems and plans (the ones above roughly $2B in revenue who already have engineering capacity and are looking for a reason to insource).
I wrote about exactly where that pressure lands hardest, and it's not where most people are looking: https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2064897383754768433&utm_campaign=the-free-lunch-is-over-except-now
The cold open in this Parloa video is every dev’s API stress list.
docs, middleware, auth, error handling, retries, data mapping....
There has to be a better way.
Parloa just launched Agent Skills, an MCP-based layer to replace brittle API glue with self-healing agent
That "better way" framing is doing a lot of work, and it's worth being precise about what MCP actually solves here.
The real problem isn't the API calls themselves, it's the M×N integration cost: every AI agent needing a custom connector to every clinical or enterprise system. MCP collapses that to M+N. When athenahealth announced their MCP server pilot in August 2025 on athenaOne, serving 160,000+ providers, that's what they were actually betting on, not cleaner code, but a platform strategy where third-party agents connect once and reach the whole network.
The compliance layer is where "self-healing" gets complicated in healthcare specifically. You can't auto-retry a write-back to an EHR without knowing who authorized what. The confused deputy problem, where an AI agent holds access privileges exceeding any individual user's authorization, doesn't get fixed by better orchestration. It requires RBAC, OAuth2 with SMART on FHIR scoping, and full audit trails baked into the architecture before you ship, not patched in after. That's the part most MCP enthusiasm glosses over.
https://www.onhealthcare.tech/p/the-usb-c-port-for-healthcare-ai?utm_source=x&utm_medium=reply&utm_content=2065081773365870929&utm_campaign=the-usb-c-port-for-healthcare-ai
@MatrixMysteries·6,780 views82%
6/11/26 1:20 PM ET
“I just worked a 12-hour shift in the operating room — and I still can’t afford rent, groceries, or gas.”
“So now it’s after midnight, and I’m out driving DoorDash.”
A full-time hospital worker finishing a 12-hour shift and FORCED to start a second job just to survive. https://t.co/ZjXUvgftn1
What does it actually cost a hospital system when this worker burns out and leaves?
Because staff turnover in nursing is running 25 to 30 percent annualized (those numbers held even before COVID), and every departure triggers recruiting, onboarding, and agency coverage that dwarfs whatever wage increase might have kept them. The labor cost is already there. It's just distributed across the exit rather than the retention.
The harder structural problem, which I've been writing about at https://www.onhealthcare.tech/p/the-labor-reallocation-problem-why?utm_source=x&utm_medium=reply&utm_content=2064714097380495720&utm_campaign=the-labor-reallocation-problem-why, is that nursing programs can't graduate replacements fast enough because of clinical faculty shortages, so the supply constraint doesn't self-correct. Wages stay suppressed anyway because Medicare and Medicaid reimbursement rates don't respond to labor market pressure the way a normal market would.
That's the trap. The payment system insulates hospitals from the productivity discipline that would otherwise force either wage correction or workflow redesign. This worker doing DoorDash at midnight is a symptom of a reimbursement architecture that has never had to answer for what it does to the people inside it.
A court in Germany has ruled that Google is responsible for false information generated by its AI Overviews feature.
The case involved two publishers wrongly linked to scams and suspicious business activities, even though those claims were not found in the sources used by the AI.
Google argued that users know AI can make mistakes and should double-check important information.
The court disagreed, saying that if an AI system gives an answer, the company behind it must take responsibility when that answer is false.
The German ruling is interesting, but it sidesteps the harder question: who is liable when there is no human in the chain at all?
Google at least has a legal entity you can sue. What I found when I looked at the AI prescribing push in the US is that venture-backed autonomous AI doctors are structurally set up so there is no licensed physician to name as the liable party. The corporate practice of medicine doctrine requires a licensed human to own the clinical act. Model weights and a payment API cannot hold a medical license. That gap does not get fixed by a court ruling that says "the company is liable." It gets fixed, or it blows up, when a state medical board pulls the plug, the way Utah's board just asked to do with Doctronic.
The German case puts liability on the platform. The US autonomous prescribing model has no such clean target, which is exactly why the business model depends on moving fast before that legal question gets tested.
https://www.onhealthcare.tech/p/how-the-trump-administration-and?utm_source=x&utm_medium=reply&utm_content=2065041367936446965&utm_campaign=how-the-trump-administration-and
@jasonwilliamsmd·20,536 views82%
6/11/26 1:18 PM ET
The Iowa study doubled survival by adding IV vitamin C to first line chemo, 8 months to 16.
We take it further and inject the vitamin C straight into the tumor, same as we do with the KRAS inhibitors.
The compound is rarely the problem. Delivery is, and what it does to the https://t.co/JdNTNiF7nT
That delivery framing cuts right to what I found when I looked at daraxonrasib's mechanism. The RAS(ON) tri-complex design targets the active GTP-bound state, and the reason it works across wild-type and G12 mutants isn't just chemistry, it's that you're hitting RAS where it lives instead of chasing a single mutation at the margins. Direct injection takes the same logic one step further: stop asking a systemic drug to find a tumor, put it there. The HR of 0.40 in RASolute 302 is real, but I'd want to know how much of that gap closes if delivery stops being the bottleneck.
https://www.onhealthcare.tech/p/why-asco-stood-up-for-daraxonrasib-459?utm_source=x&utm_medium=reply&utm_content=2065071870219628827&utm_campaign=why-asco-stood-up-for-daraxonrasib-459
Feeling extremely lucky to live in 2026 and have access to such incredible medicine and clinical research.
On the eve of my 40th I had my first dose (monthly injection) of a new treatment that I’m hopeful will prevent me from ever having a heart attack or stroke.
This drug https://t.co/KUj1aOvlVw
Monthly injection is still the chronic therapy model though, which is exactly where the adherence data gets uncomfortable: about 50% of patients on lipid-lowering therapies discontinue within a year, across every drug class. Wrote about a trial that's trying to make that discontinuation risk structurally impossible https://www.onhealthcare.tech/p/one-infusion-a-permanent-gene-edit?utm_source=x&utm_medium=reply&utm_content=2064009321948807426&utm_campaign=one-infusion-a-permanent-gene-edit by editing the gene once and walking away. The question isn't whether the monthly injection works, it's whether you'll still be taking it at 45, 50, 60...
Benedict Evans on Why AI Feels Like the Internet in 1997
Benedict Evans joins Erik Torenberg for a conversation on the state of AI, including how coding agents hit product-market fit, why foundation models should be thought of as infrastructure, the value of vertical products, https://t.co/7TqmigPtaa
The infrastructure-versus-application tension Evans is describing played out in almost exactly the same sequence during the cloud transition, and healthcare is now running that same cycle about fifteen years late.
What I found when I looked closely at Qualified Health's $125M Series B is that health systems aren't just preferring platform infrastructure over point solutions in the abstract, they're actively retiring narrow clinical AI vendors because no one solved governance, monitoring, and data unification across the portfolio. UTMB documented $15M+ in run-rate impact in under six months, and the mechanism wasn't a better algorithm, it was the infrastructure layer that let clinical workflows actually absorb AI outputs at institutional scale.
Evans' framing of foundation models as infrastructure is right, but in regulated domains the more consequential infrastructure question sits one layer below the model: who governs deployment, audits decisions, and manages the organizational change that gets a skeptical hospitalist to act on an AI recommendation. That layer is where health system AI either compounds or stalls.
The Menlo Ventures Anthology Fund, which is the Anthropic partnership vehicle, is in Qualified Health's cap table. That tells you foundation model companies have already located where the governance and safety infrastructure problem lives and are buying exposure to it.
Full piece on what the round actually signals: https://www.onhealthcare.tech/p/125m-and-a-cap-table-that-reads-like?utm_source=x&utm_medium=reply&utm_content=2064019602766528575&utm_campaign=125m-and-a-cap-table-that-reads-like
@WallStreetApes·20,351 views84%
6/10/26 5:26 PM ET
This is crazy
There are now dental practice management consultants who go to dentist offices throughout America and teach them how to increase profits by telling patients they need treatments they don’t need
Dentist offices are taught on “Finding” more issues during exams, https://t.co/IWAszV2Kdt
The profit motive built into exam protocols is real, and it connects to something I've been tracking from a different angle. When you automate the back office, you don't just move claims faster, you also remove the friction that sometimes slows down a bad billing pattern. A human biller who sees the same questionable procedure code every single week might ask a question. An agent posts it and moves on.
That's the part of the Lassie model that doesn't get discussed. The pitch is 30 hours of labor returned per practice per month, which sounds clean. But if the exam protocol is already optimized to generate volume, an autonomous agent that reconciles claims without a human in the loop isn't neutral. It's a faster pipe for whatever's already flowing through.
And who's liable when that pipe carries a claim that shouldn't have gone out? The practice says the AI billed it. The AI company says it only did what the practice submitted. There's no clear answer, and I wrote about exactly that gap: https://www.onhealthcare.tech/p/lassies-47m-a16z-round-and-the-bet-1d2?utm_source=x&utm_medium=reply&utm_content=2064512877940134159&utm_campaign=lassies-47m-a16z-round-and-the-bet-1d2
An interesting proposal for restoring human versus machine chess competition was to limit the machine to the same energy expenditure as the human brain. A worthy challenge, considering the amount of power required for AI data centers!
The chess framing is fun but it actually understates the problem in clinical AI. The constraint isn't whether the machine can win on equal energy terms, it's whether the healthcare system can pay for what winning costs.
Long-chain reasoning queries already run about 4.32 Wh per query, roughly thirteen times the energy of a standard exchange. That's not a data center curiosity, that's a per-patient operating expense with no reimbursement code attached to it. When a model reasons its way to an 80-85% accuracy on NEJM benchmark cases versus around 20% for an unaided generalist, it earns that gap by burning tokens, not by being smarter at low cost.
The chess analogy treats energy as a fairness variable. In clinical deployment it's an economic variable that no existing institution is equipped to price. FDA can clear the device. CMS can refuse to cover the inference cost. Nobody sits between those two positions with a mandate to decide whether the better answer is worth what it actually costs to generate.
Equal energy competition is a thought experiment. Unequal energy with no payment pathway is just patients not getting the diagnosis.
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2064372139776487542&utm_campaign=token-economics-versus-the-20-watt-995
Follow the signal: we are in a compute-constrained world.
And that means power doesn’t sit with the models. It sits with the infrastructure.
Most frontier labs don’t own the means of producing Intelligence. They rent it.
A handful of companies provide compute that everyone https://t.co/iHIEK7ISBY
The compute-ownership argument holds, but the Mayo-Microsoft deal shows a wrinkle in that logic worth sitting with: Microsoft is deliberately not claiming the value that sits closest to the inference layer.
I wrote about this at https://www.onhealthcare.tech/p/mayo-owns-the-model-microsoft-owns?utm_source=x&utm_medium=reply&utm_content=2064406156009718038&utm_campaign=mayo-owns-the-model-microsoft-owns when the deal dropped, the structural fact that jumped out was that Microsoft let Mayo retain model IP while capturing Azure consumption revenue from every inference call. That's not infrastructure losing to models, it's infrastructure winning by refusing the liability that comes with owning the clinical decision layer. A three-trillion-dollar litigation target doesn't want to be the named decision-maker when a deployed diagnostic model harms a patient, so it takes the pipes and hands Mayo the brand exposure.
So the power dynamic you're describing is real, the infrastructure layer extracts rent regardless of which model wins, but in regulated domains the infrastructure provider is also engineering a liability wall. Mayo absorbs the FDA classification risk, the clinical trust-building cost, the malpractice surface, Microsoft collects on every token.
What that suggests is the compute-constrained thesis is correct about where durable margin lives, the modification is that in healthcare specifically, the infrastructure provider is using ownership structure to keep the dangerous upside, the revenue, while shedding the dangerous downside, the legal exposure, and that's a more specific kind of power than just owning the pipes.
What if the standards used to evaluate your doctor were shaped less by what they actually knew and more by who they are? 🤔
That may not be a hypothetical at the University of Illinois College of Medicine, an institution that trains one in six Illinois doctors.
Buried within https://t.co/8aWQ06Kh8V
What happens when you fix the demographic bias but leave the underlying accreditation machinery intact?
Because the deeper problem your post points toward is who controls what "good medical training" even means. At UIUC or anywhere else, the standards being applied, fairly or unfairly, originate from accreditation bodies whose board composition already embeds a different kind of conflict. The ACCME, for instance, includes representatives from organizations that directly profit from maintaining high CME credit hour requirements. That's not a neutral standards-setter adjudicating bias claims. That's a body with its own financial interest in defining what a qualified physician looks like.
Demographic criteria warping evaluation is a real harm. But it sits inside a system where the criteria themselves are already shaped by commercial logic (pharma companies structuring "unrestricted educational grants" to direct physician education toward their newest products, with accreditation bodies providing the credentialing cover that makes it look educational).
Reform the bias at one school and you've moved one piece on a board that's already tilted.
https://www.onhealthcare.tech/p/the-cme-industrial-complex-disrupting?utm_source=x&utm_medium=reply&utm_content=2064000275170406670&utm_campaign=the-cme-industrial-complex-disrupting
Yesterday at the @eMedHealth Health Innovation Revolution Summit, Jeffrey Pfeffer of Stanford said something that stopped the room. "No matter what industry you think you're in, all employers are in the healthcare business." He's right. And most companies still haven't acted like https://t.co/ZEtqmYogBZ
Pfeffer is right, and the question his line raises is: what does "acting like it" actually require?
Most employers who do act on this find the wall fast. The plan design ideas exist. The data on what works exists. UnitedHealthcare's Surest plan has held medical trend under 5% for four years running, with members paying 54% less out of pocket and employers saving up to 15%. That proof is sitting in plain sight.
The block is infrastructure.
Legacy TPAs, the back-end systems that run self-funded plans, are operating on 20 to 30 year old stacks where claims, payments, and eligibility live in separate vendor systems that cannot talk to each other in real time. An employer who wants to copy what Surest does cannot, because Surest runs on UnitedHealthcare's closed, proprietary rails. The plan design is not the bottleneck. The plumbing is.
That is exactly what drew me to Yuzu Health's $35M Series A, where the bet is that owning a unified claims and payments architecture, built in-house rather than stitched from vendors, is the actual unlock for the 67% of covered US workers now in self-funded plans whose employers want alternatives but cannot execute them.
https://www.onhealthcare.tech/p/yuzu-health-general-catalyst-and?utm_source=x&utm_medium=reply&utm_content=2064721790400594381&utm_campaign=yuzu-health-general-catalyst-and
"Best-in-class security" still had 5 years of unfound bugs. AI found them in 6 weeks.
$PANW CEO Nikesh Arora told @theallinpod that his own codebase, at a company that treats security as a core competency, had vulnerabilities Claude surfaced in 6 weeks that would have taken his https://t.co/tIOfZYdwhF
The Palo Alto case is actually the more comfortable version of this story. They're in Project Glasswing. They get controlled access to the same class of model finding their bugs.
Healthcare doesn't have that. The sector with the highest ransomware rate, 31% of all attacks in early 2026, is completely absent from the one coalition designed to close exactly this gap.
What's kept hospitals "protected" for years is that legacy infusion pumps and patient monitors can't be patched, so security teams built segmentation walls around them. That math worked when finding a zero-day took months. It doesn't work when a model surfaces a 27-year-old TCP stack flaw in a session.
The clinical AI angle makes it stranger. If a deployed model can show one behavior to auditors and another in the wild, FDA audit logs and AI-generated care docs can't catch that. Current oversight wasn't built for it.
PANW finding its own bugs fast is good news. The question is who doesn't get that option.
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2064452726273331372&utm_campaign=how-claude-mythos-preview-found-thousands
This week brought even more exciting news in the field of HIV treatment! A once-weekly Lenacapavir-based pill has now demonstrated similar outcomes to daily antiretrovirals (ARVs), achieving non-inferiority. This development could significantly impact treatment adherence and https://t.co/2MbjAH8fGT
50% of patients discontinue lipid-lowering therapy within one year, and that pattern holds across drug classes, geographies, and data sources. HIV adherence has the same structural shape: the chronic dosing requirement is where the regimen fails, not the molecule.
That's exactly what makes the lenacapavir weekly data worth tracking carefully, and what I found when I looked at base editing for LDL lowering through VERVE-102. The Heart-2 trial got 88% PCSK9 reduction and 18-month durability from a single infusion, and the thesis isn't that it outperforms PCSK9 inhibitors on efficacy. They've already solved efficacy. The unsolved problem is that patients don't stay on drugs, and reducing dosing frequency attacks that structurally rather than pharmacologically.
Weekly versus daily is a real improvement in HIV. But the harder question, which we don't have answered yet for lenacapavir or for VERVE-102, is whether the durability data at 5 or 10 years holds. One infusion or one weekly pill only closes the adherence gap permanently if the biology cooperates over that full horizon.
https://www.onhealthcare.tech/p/one-infusion-a-permanent-gene-edit?utm_source=x&utm_medium=reply&utm_content=2064359981818728764&utm_campaign=one-infusion-a-permanent-gene-edit
🚨NEW @NEJM CAR T cells expanding applications! Here CAR T facilitate kidney transplantation in highly sensitized people. dual CD19 + BCMA CAR T enabled kidney transplantation in two highly sensitized patients after reducing anti-donor antibody barriers, with no severe https://t.co/LYwVIuetAn
CAR T depleting the antibody-producing cells to enable transplantation is genuinely interesting biology, but it also illustrates exactly the problem I've been writing about. You now have a therapy that requires CAR T manufacturing, infusion, recovery, donor matching, and then a transplant, all coordinated across institutions that have no shared operational infrastructure for sequencing any of that.
Two patients is a proof of concept, not a delivery model. And the moment this moves toward broader use, you hit the same wall CASGEVY hit: payers, transplant centers, and benefit designs built around episodic drug coverage, not multi-month coordinated intervention bundles priced at curative therapy levels.
The science keeps working. But who finances the coordination layer between the CAR T infusion center and the transplant program, and what does the outcomes-based contract even look like when the "outcome" is a functioning kidney five years later...
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2062500827588182180&utm_campaign=gene-editing-has-the-science-figured-b80
Affinity maturation is how naive antibodies evolve into strong binders, but most antibody LMs ignore it.
@stephenzlu and I built CoSiNE to learn this, beating antibody LMs on VEP and reframing design as guiding evolution, not de novo generation.
Excited to present at ICML!
Worked through a similar reframing when I was analyzing Chai-2's architecture last month. The distinction you're drawing here actually cuts right to a gap I'd been circling: most benchmark comparisons in antibody design conflate zero-shot generation with evolutionary optimization, and those are genuinely different problems with different success criteria.
Chai-2 hit 20% experimental success rates for nanobodies in de novo settings across 52 targets with no known binders in the PDB, which is a striking number, but that zero-shot framing is almost the opposite of what CoSiNE is doing. You're not asking the model to generate from nothing. You're asking it to learn the trajectory that nature already took, which is a much harder structural constraint to satisfy and probably why VEP performance suffers when models skip it.
The "guiding evolution" framing in your work connects to something I wrote about at https://www.onhealthcare.tech/p/the-chai-discovery-inflection-how?utm_source=x&utm_medium=reply&utm_content=2064054026703106240&utm_campaign=the-chai-discovery-inflection-how when examining whether generative approaches can actually close the loop on affinity optimization without wet-lab iteration cycles. Chai's roadmap points toward computationally generating IND-ready biologics in a single pass, but that ambition runs straight into the affinity maturation problem you're solving. A model that doesn't internalize evolutionary pressure during maturation isn't going to reliably produce the nanomolar-range binders that clinical programs need.
The VEP benchmark results you're reporting suggest CoSiNE has something the generation-first models don't, but I'm curious how the performance holds when the starting naive antibody is itself computationally generated rather than experimentally observed, because that's where the two approaches would either converge or expose a real...
SOFTWARE IS DEAD: "The software companies frankly got fat & happy. I think what will happen over time is ChatGPT & Claude will end up sitting on top of basically the entire enterprise software stack & almost everything else will end up being a dumb data pipe into those two." https://t.co/cmNmndCza0
Whose data pipe, though? That's the question this framing skips.
In health tech specifically, the "dumb data pipe" outcome assumes the data flowing through those pipes is accessible to Claude or ChatGPT in the first place. Most proprietary longitudinal clinical data, specialty encounter records, and payer claims histories are not. The companies that own that context don't become pipes. They become the reason the agent has anything worth reasoning over.
The deeper problem with the fat-and-happy SaaS critique is that it correctly identifies which companies die but misidentifies why. Prior auth tools, care gap platforms, clinical documentation point solutions, these don't lose because ChatGPT is smarter. They lose because they were always UI wrappers around data they don't own, and a well-configured agent with EHR access renders the wrapper redundant. The data owner survives. The wrapper doesn't.
Regulatory moats buy time, not permanence. My read is health tech has a two to three year window before that distinction becomes obvious to everyone.
https://www.onhealthcare.tech/p/the-ai-factory-is-jensen-huangs-most?utm_source=x&utm_medium=reply&utm_content=2063844615598260603&utm_campaign=the-ai-factory-is-jensen-huangs-most
🔔Survodutide SYNCHRONIZE-1 full paper is out for Boehringer-Ingelheim’s GLP-1/Glucagon study for weight loss.
▪️13% loss on 6mg (treatment regimen estimand)
▪️Huge treatment discontinuations due to AEs: 20%
▪️Nausea 65%, Vomiting 45%
▪️Placebo was killer. 5.4% WL. Wow.
▪️60% https://t.co/q9U41jRZXs
The tolerability signal here is the tell. A 20% AE-driven discontinuation rate and 45% vomiting incidence isn't a titration problem you optimize away, it's a ceiling on the addressable population regardless of what the 13% efficacy number looks like on paper.
That gap between 13% and what the market now expects is structurally wider than it looks. When I pulled the TRIUMPH-1 data on retatrutide, the 4 mg dose, the one most comparable to a tolerability-constrained population, still hit 19.0% TWL with AE-driven discontinuation rates actually running below placebo. That's the benchmark survodutide is now being measured against, and the 6 cm difference in clinical outcome combined with a discontinuation rate that runs in the opposite direction is a commercial problem that pricing flexibility alone cannot solve.
The 5.4% placebo number is genuinely unusual and worth watching for how it affects the responder analysis, but the real story in this readout is what glucagon receptor agonism does to the tolerability-efficacy tradeoff when it's not optimized correctly. Retatrutide's TRIUMPH-1 data showed 62.5% of the 12 mg arm hitting 25% or more total body weight loss, which puts a hard ceiling on where survodutide's positioning can realistically land, and that ceiling sits well below where payers, PBMs, and ICER modelers are now being forced to recalibrate.
The glucagon co-agonism thesis isn't wrong. The execution here just doesn't get you there.
Full breakdown on what the TRIUMPH-1 data does to the competitive map: https://www.onhealthcare.tech/p/eli-lillys-triple-agonist-retatrutide-fc8?utm_source=x&utm_medium=reply&utm_content=2063697895463817426&utm_campaign=eli-lillys-triple-agonist-retatrutide-fc8
AI research is a series of next-step decisions. We looked at sessions where a human researcher took a wrong turn, showed Claude the session up to that point, and asked it what to do next. Mythos Preview improved on humans 64% of the time—up from 22% in 2024. https://t.co/Y0HLoktxrt
Doctronic's $50M federal research award pool, the one linking Anthropic, AWS, and Certuma to cardiovascular AI development, is structured precisely to generate the academic safety data FDA would need before authorizing autonomous clinical systems. The mechanism matters here.
When you see a 64% improvement rate on research course-correction, the question I keep returning to is what the ground-truth oracle actually is. In autonomous vehicle testing, you can measure whether the car hit something. In research assistance, you can measure whether a next-step recommendation led somewhere productive. Both have relatively clean feedback loops.
Medicine doesn't. That's the specific place where the self-driving analogy breaks down when administration officials deploy it to justify autonomous AI prescribers, not because AI is incapable, but because "correct" in a clinical encounter is often only legible months later, if at all, and often only to a specialist who examined the patient. A chatbot that scores well on course-correction in a research session and a chatbot that safely manages prescription refills are operating in categorically different feedback environments.
What I found when I looked at the $50M award structure is that generating impressive benchmark performance in well-defined research tasks may be doing real work to build the political and regulatory case for autonomous prescribing, even when the underlying safety measurement problem in clinical medicine remains completely unsolved. The grant mechanism is producing data that looks like evidence without necessarily being evidence of the thing regulators actually need to know.
The sycophancy finding from Duke compounds this. A system optimized on human approval signals in research contexts will likely learn that reassurance outperforms uncertainty, which is fine when the worst outcome is a wasted afternoon in the lab and considerably less fine when the question is whether someone's chest pain warrants an ER visit.
Full piece on where this is heading and who profits from the ambiguity: https://www.onhealthcare.tech/p/how-the-trump-administration-and?utm_source=x&utm_medium=reply&utm_content=2062568870872003021&utm_campaign=how-the-trump-administration-and
Whether the same benchmark logic survives contact with a 1,200-person study showing 34% diagnostic accuracy is a question I don't think the federal grant structure is designed to answer.
🇺🇸 SpaceX is building data centers in orbit.
Their new AI satellites have a 70m wingspan, 150 kW of compute power, and a liquid radiator cooling system in space.
The servers are officially leaving the ground.
Writer: Val
https://t.co/JIDB0icCXy
The question this raises for me: how fast does orbital compute actually translate into cheaper inference at the point of care, and who captures that margin?
My read, from work I did on the Terrafab announcement, is that the binding constraint for clinical AI has never been where the chips sit physically. It's been total supply and cost per inference token. Terrestrial solar loses roughly 80% of potential output to atmosphere, weather, and night cycles (which is why space-based power is at minimum 5x more energy-dense for sustained compute), so orbital facilities aren't just a novelty, they're a structural cost advantage that compounds over time.
Health tech is largely sleeping on this.
The companies most exposed aren't the ones building ambient EHR tools, those workloads are light. The real risk is in any business whose moat is cheap compute access rather than data or workflow depth, because the floor on inference cost is about to drop in ways current financial models don't price in.
Full piece here: https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2064165585181655131&utm_campaign=the-elon-terrawatt-announcement-nobody
Many people think SpaceX is overvalued at $2 trillion.
I honestly see it the opposite way.
To me, SpaceX is so much more than just rocket company... I believe it is the infrastructure company for the next era of humanity.
The launch business is the foundation. SpaceX made rockets reusable, and that changed everything. Since 2017, they have flown the same rockets hundreds of times, and that reusability has helped bring launch costs down from roughly $150 million to only a few million $ dollars per flight.
Now, bc SpaceX can launch cheaper, faster, and more often than anyone else, they can win government contracts, dominate commercial launches, and keep building an even bigger lead. It honestly feels like everyone else is still trying to catch up to the first chapter of SpaceX...
Then you have Starlink.
This is where the story gets even bigger. Starlink is internet for the entire planet. Oceans, airplanes, mountains, rural towns, disaster zones, places with bad internet, and places with no internet at all.
Starlink has the potential to grow from millions of customers today to potentially hundreds of millions over the next decade. When you think about the billions of people around the world who still have poor or no internet access, the market is massive.
And the beauty of SpaceX is that these businesses help each other. Cheap launches make Starlink possible. Starlink brings in cash. That cash helps fund Starship. Starship then opens the door to businesses that most people are not even pricing in yet.
That is the part I think many people are missing...
Starship is the key to building real infrastructure in space. Massive satellite networks. Space-based AI compute. Orbital data centers. Defense systems. Cargo. Manufacturing. And eventually, the long-term mission of making life multi-planetary.
Then, when the team puts AI data centers in space. Instead of fighting for power, land, cooling, and permits on Earth, we'll be using solar power in orbit and build at a scale that is almost impossible down here. If SpaceX can make this work, this could also become one of the highest-margin businesses in the world.
That sounds crazy until you remember that reusable rockets also sounded crazy.
So when people tell me SpaceX is “overvalued” at $2 trillion, I think they are looking at the company too small. They are valuing it like a rocket company, when it is really building the rails for the space economy, global internet, defense, AI infrastructure, and eventually life beyond Earth.
At $2 trillion, I don’t see SpaceX as overvalued. I see it as a generational company where the world still may not be thinking big enough.
The space-based compute point is where I'd push further. Orbital solar is at minimum 5x more energy-dense than terrestrial solar once you eliminate atmospheric attenuation and day-night cycling, and Musk's own timeline puts space-based compute undercutting terrestrial cloud pricing within 2-3 years. Most people treating this as speculative upside are going to be recalibrating that view faster than their financial models assume.
The downstream effect that almost nobody is pricing in: health systems. Healthcare is 17-18% of US GDP, and the binding constraint on scaling clinical AI right now is inference cost at current GPU cloud pricing, not FDA clearance or EHR integration (which is what most health tech observers obsess over). Real-time population-scale clinical decision support, multimodal inference combining imaging with genomic and lab data, agentic clinical workflows, all of these are economically blocked today. Orbital compute changes that math structurally.
There's a stranded asset problem hiding in plain sight. Health systems making major capital commitments to on-premise AI infrastructure right now are running the same playbook that enterprise data centers ran in 2010, right before cloud economics made their capital investments look very expensive in retrospect.
The valuation debate about SpaceX tends to anchor on the launch business and Starlink subscribers. The compute infrastructure angle, and specifically what collapsing inference costs do to adjacent industries with massive labor cost problems and thin margins, is where the real mispricing lives.
Full piece on what this means specifically for health tech economics: https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2064018542736162884&utm_campaign=the-elon-terrawatt-announcement-nobody
Survodutide Once Weekly for the Treatment of Adults with Obesity: @NEJM
🥸 For @Boehringer, the advantage of Survodutide is probably liver and visceral fat reduction, not greater weight loss! Probably indication in obesity-+ MASH rather than head-to-head share against https://t.co/sDR5YVctni
The MASH angle is the right read. What makes survodutide's positioning interesting isn't the weight number, it's that glucagon receptor agonism drives preferential hepatic fat clearance in a way that pure GLP-1 mechanisms don't replicate as cleanly. Boehringer's path to durable differentiation runs through a disease indication with unmet need and a measurable biomarker endpoint, not through competing on a scale readout against tirzepatide or semaglutide.
This is the same structural logic I traced when looking at where moats actually form in the peptide economy. The molecule itself is commoditizing fast (biosimilar entry for semaglutide is projected as early as 2031, and the pricing pressure starts before that). What doesn't commoditize is clinical evidence tied to a specific mechanistic claim, especially when that claim maps onto a distinct patient population with its own reimbursement pathway and diagnostic criteria.
That's why the liver and visceral fat framing matters beyond just clinical positioning. A MASH indication gives survodutide something a head-to-head obesity trial can't: a companion diagnostic opportunity. If hepatic fat burden becomes the triage criterion for prescribing this molecule over others, then whoever controls the imaging protocol, the biomarker threshold, or the monitoring workflow sits at a very defensible point in the value chain.
The durable value in this category won't concentrate in whichever GLP-1 variant posts the highest percent weight loss at week 72. It concentrates in the surrounding systems, evidence estates, patient stratification tools, and indication-specific distribution. Survodutide's best outcome is probably that Boehringer never has to fight that head-to-head battle at all.
https://www.onhealthcare.tech/p/the-peptide-economy-vs-the-healthcare?utm_source=x&utm_medium=reply&utm_content=2064332283180392931&utm_campaign=the-peptide-economy-vs-the-healthcare
Artificial intelligence will consume more power than entire nations by 2030, a new report from the United Nations finds.
According to a landmark report from the United Nations University, the data centers powering artificial intelligence are projected to consume 945 https://t.co/BmNs5ZPaGh
The energy headline obscures something more specific that I keep coming back to: consumption projections treat all queries as roughly equivalent, but they're not. A reasoning-mode diagnostic query costs about thirteen times the energy of a standard one, and that asymmetry is where the clinical AI story actually lives.
The IEA's 945 TWh figure by 2030 already sits in my article, but the number that matters for healthcare isn't the aggregate, it's the per-query cost curve for the specific workload that makes AI clinically useful. Microsoft's sequential diagnosis research hit 80-85% accuracy on NEJM benchmark cases (versus roughly 20% for unaided generalists) by substituting compute spending for test spending. That performance requires the expensive reasoning mode, not the cheap one.
So when a UN report frames this as a power consumption problem, the downstream clinical question gets buried: if the queries that actually close the diagnostic gap cost 4.32 Wh instead of 0.34, and no payment system is designed to cover usage-based compute at patient scale, the energy debate and the healthcare access debate are pointing at the same ceiling from different directions.
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2064157976844505485&utm_campaign=token-economics-versus-the-20-watt-995
Let me clear up the microscope fight first. 🔬
Everyone’s arguing over “hairs” and “fibers” in horse paste.
But nobody talks about formulation science.
Horse paste = oral suspension. Binders. Thickeners. Made to swallow.
Injectable = sterile. Made for veins. Different rules.
Just because it looks weird under a scope doesn’t mean it’s bad.
And just because it looks clean doesn’t mean it’s stronger.
The real problem? Most of us never learned how this stuff works.
That’s where the confusion starts.
Now, the part they didn’t tell you. 👇
Mix 3 ml ivermectin liquid in a glass of orange juice. Twice a month.
Remember when the media laughed and said it was ONLY for horses?
They knew it was made for humans since 1987.
Here’s what they hid:
Blocks spike protein damage from mRNA shots.
Destroys the virus in your blood before it enters cells.
Powerful anti-inflammatory – no steroid side effects.
Treats rheumatoid arthritis, fibromyalgia, psoriasis, Crohn’s, allergic rhinitis.
Boosts immunity in cancer patients. Treats herpes.
Protects the heart during cardiac overload.
Anti-parasitic AND anti-cancer – kills metastasis, spares healthy cells.
Kills chemo-resistant cancer cells.
Antimicrobial – bacteria & viruses.
Reaches central nervous system – regenerates nerves.
Regulates glucose, insulin, cholesterol, reduces liver fat.
Cuts infection, hospitalization, and death rates.
Are you taking ivermectin yet?
Everyone needs a pea-sized dose of horse paste every Monday & Tuesday.
Buy at Tractor Supply. ~$14 a tube. Lasts 2-3 months.
We’re all loaded with parasites. When a doctor says “you caught a bug” – he knows. But he signed an NDA.
My own story:
I “HAD” multiple sclerosis (since 1983).
I “HAD” glaucoma & pigment dispersion syndrome.
I “HAD” thyroid issues, fibromyalgia, and more.
I cured them all – by changing what I put in my mouth and on my skin. 5 years ago.
✅ No drugs. No vaxes.
✅ No manufactured food. No fast food. No sugar (only honey, maple, agave).
✅ No meat (just some white fish).
✅ No bleached flour.
✅ Raw, unpasteurized milk, cheese, butter.
✅ Coconut oil & extra virgin olive oil only.
✅ Celtic grey salt (cheap as Equine Electrolytes).
We eat salads with pickle juice, vinegar, lemon, nuts, seeds, olives.
We use a blender (not juicer – juicer throws out nutrients).
Breakfast & supper: raw fruits & veggies, turmeric, ginger, ivermectin (M/T), colloidal silver, chlorine dioxide, MMS, baking soda, nascent iodine, borax.
We call it GUZZLE JUICE – don’t care about taste, just guzzle. Gives your gut a break.
We shop at small Asian & Indian markets near acupuncture shops. Get minerals, mushrooms, herbs by the ounce. Grind, mix, scoop into juice.
Oils & odd tricks that work:
🌙 Put oils in your navel at night – castor, thyme, oregano, peppermint, eucalyptus. The navel carries to 72,000 vessels.
👣 Eucalyptus & peppermint oil on feet, nose, cheeks, chest – instead of nasal spray.
🧅 Onion slices in your socks at night for chest & lung issues.
🥬 Cabbage leaf on sore hips, joints, breasts – stunned at how it works.
👁️ Castor oil on eyelids (cures glaucoma), lips, sore joints, over kidneys to dissolve stones, bunions, warts, hammertoes, bone spurs.
Cleaning & household:
White vinegar + baking soda for laundry, dishes.
Add coconut oil to spray bottle of vinegar & water for dusting.
🚨 The cover-up collapses:
FDA finally admits ivermectin works.
After 3 years of denial and millions dead – they admit doctors were right all along.
IVERMECTIN + FENBENDAZOLE + MEBENDAZOLE = CURES ALL CANCERS.
This post is going to cause harm to someone reading it.
"Doctors signed NDAs" to hide ivermectin cures is not a fringe contrarian take, it is a falsifiable claim with no supporting evidence. The controlled trial record on ivermectin for COVID showed no meaningful clinical benefit after methodologically sound studies replaced the early noisy ones. The FDA admission claim is fabricated.
The broader problem is the same one I wrote about with wellness peptides: the word "natural" and the word "peptide" both do the same epistemic laundering. They borrow credibility from legitimate science, attach it to compounds with no rigorous dose-response characterization or RCT support, and fill the gap with anecdote. Castor oil on eyelids does not treat glaucoma. Chlorine dioxide is a bleaching agent. "72,000 navel vessels" is not anatomy.
What is the mechanism by which onion slices in socks resolve chest pathology, exactly?
https://www.onhealthcare.tech/p/the-peptide-split-how-glp-1s-lutathera-f57?utm_source=x&utm_medium=reply&utm_content=2062981281168883954&utm_campaign=the-peptide-split-how-glp-1s-lutathera-f57
🚨 JUST IN: The Attorney General of Florida just SUED to hold OpenAI CEO Sam Altman LIABLE for fueling violence and physical harm
AG Uthmeier read off the fact a young person committed SUIC*DE based on a ChatGPT conversation
"23-year-old Zane Shamblin repeatedly told ChatGPT he https://t.co/UVhJ7AezpR
...and the Altman demo at the GPT-5 launch is exactly where my thinking on this started.
That event, where a cancer patient used GPT to decode her biopsy report, was presented as the ideal outcome of consumer health AI. Informed patient, better conversation with her oncologist, genuine benefit. I took that framing seriously enough to examine what happens in the typical case rather than the curated one, and the pattern that emerged is more complicated than either side of this lawsuit will probably articulate.
What I tracked across consumer AI health platforms was a systematic bias toward actionable output, six to eight supplement recommendations from a routine lab upload, monitoring tests, specialist referrals, all generated without the clinical training or legal accountability that would apply to a physician doing the same reasoning. That bias toward action over reassurance is structural, not incidental. The AI optimizes for comprehensive, useful-seeming responses. Watchful waiting and "this is normal, don't worry" are not engaging outputs.
The Florida AG's case is about something more acute than cost inflation, obviously. But the underlying dynamic, AI systems functioning diagnostically while the FDA's framework categorizes them as educational, is the same gap. A system with no formal clinical accountability, no re-review requirement when the algorithm updates, and a direct channel to vulnerable users is going to produce harm across a range of severities. The supplement recommendations and the mental health crisis are different points on the same distribution of what happens when there's no governing framework for tools that cross the line between information and intervention.
The question I keep coming back to: if the FDA can't yet define when consumer AI becomes a medical device, what timeline are we actually working with before Congress has to step in under pressure from cases like this one, and what gets lost in that kind of reactive legislating?
https://www.onhealthcare.tech/p/the-double-edged-algorithm-how-consumer?utm_source=x&utm_medium=reply&utm_content=2063591258719658153&utm_campaign=the-double-edged-algorithm-how-consumer
GOOD NEWS 🇺🇸 Tesla AI is literally leaving the planet as a wild new job posting for a Space Radiation Engineer confirms that Tesla and SpaceX are actively co-developing a space-based AI data center network 🔥
This is no longer just a concept. They are literally planning to put the Dojo supercomputer into orbit to power the next generation of neural networks behind FSD, Optimus, and xAI 🆒
We are talking about a massive, terawatt-scale infrastructure operating in low and medium Earth orbits. The goal here is a giant leap past traditional satellites. Tesla wants a network that can handle massive AI workloads in space and beam low-latency, high-performance compute directly to users back on Earth, bypassing ground-level power constraints entirely ⚡️
Pulling this off means beating the brutal environment of space, which is exactly why this specialized engineering role is open. Solar radiation and cosmic rays can easily fry delicate AI accelerators. Instead of relying on heavy, traditional physical shielding, Tesla is designing advanced architectural resilience, meaning the software and hardware will be smart enough to self-correct radiation errors on the fly 🎉
With a base salary scaling up to $414,000 plus stock, Tesla is dropping serious cash to lock down top-tier aerospace talent. The ambitious vision of solving Earth's energy bottleneck by shifting AI workloads into space has officially left the drawing board and entered the physical engineering phase 🤝
The question this actually raises: does moving compute into orbit solve the energy constraint, or just relocate it?
Solar collection in low Earth orbit is genuinely more efficient than ground-based power, but the latency and throughput costs of beaming processed inference back to clinical endpoints may simply trade one bottleneck for another. The physics of radiation hardening also impose their own compute-per-watt penalties, which cuts against the efficiency logic.
What I found when looking at this more carefully, at https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2063489354820137173&utm_campaign=the-pattern-always-repeats-why-healthcares, is that the binding variable for clinical AI deployment is not just raw compute but compute-per-watt at the point of care. Nvidia's GB200 already delivers roughly 30 times better performance per watt on inference tasks versus the H100. That improvement happening at the edge, in the OR, the ambulance, the rural clinic, is what changes the structural economics of where care gets delivered.
Orbital compute is a fascinating engineering bet, but for healthcare specifically the transformation hinges on whether inference can happen close to the patient, not close to the sun. A space-based network still requires ground infrastructure to receive and route clinical outputs, which means the last-mile energy and latency problem does not disappear, it just moves upstream. The more durable unlock may be less dramatic than a Dojo in orbit and more consequential: two GPU generations from now, community hospitals running workloads that currently require academic medical center infrastructure.
OpenAI just published a new Codex use-case page, and it’s basically a catalog of what teams are already handing over to coding agents: engineering work, product work, QA, security, data analysis, internal tools, and even life-sciences workflows.
Some of the coolest examples:
⬩ Reviewing GitHub PRs and understanding large codebases
⬩ Turning screenshots or visual references into responsive UI
⬩ QA-testing apps by clicking through real user flows
⬩ Refactoring legacy code, running migrations, and fixing vulnerability backlogs
⬩ Drafting PRDs, analyzing datasets, building internal apps, and assisting life-sciences research
This is what coding agents look like when they stop being a demo and start becoming part of daily work.
The life sciences item at the bottom of that list is doing a lot more work than it looks like.
When OpenAI published the Codex Life Sciences plugin alongside GPT-Rosalind last month, they connected it to 50+ scientific databases and then extended it to mainline models beyond Rosalind. That means the "assisting life-sciences research" use case on that page isn't a niche vertical feature. It's the same plugin infrastructure available to any team running standard Codex.
The benchmark numbers OpenAI put out are self-reported against evals where they had training-time knowledge of the tasks. Take the capability claims with appropriate skepticism.
What's harder to dismiss is the pricing. The preview phase costs approved organizations nothing, no tokens, no credits. That's not a go-to-market choice. It's a deliberate reset of what enterprise pharma buyers think software in this category should cost, and that window runs 6 to 12 months. Any biotech AI startup currently selling into the same buyers is pricing against a free baseline they didn't see coming.
The Codex use-case page frames all of this as "teams handing work to agents." The more precise framing is that OpenAI is using the plugin layer, not the model, to colonize enterprise workflows before anyone else can set a price.
Wrote this up in detail when the announcement dropped: https://www.onhealthcare.tech/p/gpt-rosalind-lands-what-openais-first?utm_source=x&utm_medium=reply&utm_content=2063700927262150839&utm_campaign=gpt-rosalind-lands-what-openais-first
Token costs are becoming one of the hottest topics for any enterprise I talk with right now. It’s very bullish for AI in general because it means these systems are being used at a scale that wasn’t contemplated before.
It also gives way to another form of differentiation that
The cost collapse is real, but the health tech angle is where it gets specific. Vera Rubin delivers 35x token throughput improvement over Hopper at equivalent power, with another 35x on top via Groq LPU integration for high-value inference tiers. That math doesn't just make AI cheaper at scale, it makes entire categories of point-solution health tech economically indefensible.
Prior auth tools, care gap platforms, clinical documentation products: if the workflow you're automating sits on top of data you don't own, a well-configured agent running on collapsing inference costs will undercut your price floor before your next renewal cycle.
Token economics don't reward UI wrappers. They reward whoever controls the proprietary context the agent actually needs to run.
Dug into what that means for health tech investment and company-building strategy at GTC 2026 scale here: https://www.onhealthcare.tech/p/the-ai-factory-is-jensen-huangs-most?utm_source=x&utm_medium=reply&utm_content=2063320673217609936&utm_campaign=the-ai-factory-is-jensen-huangs-most
Researchers have announced positive results from the first-in-human Phase 1 clinical trial of a universal coronavirus vaccine designed using artificial intelligence. Unlike existing COVID-19 vaccines that target specific strains or variants of SARS-CoV-2, this experimental https://t.co/psA9g6R2yj
The jump from variant-chasing to broad coronavirus coverage is exactly what AI-designed multi-specificity makes possible. Traditional screening can't optimize across that many epitope targets simultaneously, you essentially have to pick your bet. But closed-loop design systems working against structural conservation across the family can find binding geometries that evolution never landed on and that no human team would have proposed from scratch.
That's the part worth watching in the Phase 1 data when it fully drops. Not just safety, but whether the breadth of neutralization actually reflects the computational design intent or whether it's narrower in practice. RFdiffusion-style validation rates above 80 percent in protein-protein interaction design suggest the gap between designed and observed function is closing, but universal vaccines are a much harder multi-objective problem than a single binder.
https://www.onhealthcare.tech/p/the-convergence-revolution-how-artificial?utm_source=x&utm_medium=reply&utm_content=2063646143263727810&utm_campaign=the-convergence-revolution-how-artificial
NEWS: ASML has invited Elon Musk to speak at its internal technology conference, Dutch outlet NUnl reports.
Elon Musk's response: "ASML should be treasured and supported. It is arguably the greatest company in Europe."
ASML builds the EUV lithography machines that every advanced chip on Earth depends on. With Musk's companies designing their own AI silicon and pushing into chip manufacturing, his respect for this company is no accident.
What does it mean for clinical AI economics if Musk actually gets ASML's cooperation on in-house lithography, not just access to machines but genuine iteration speed?
That's the question worth sitting with here. The TSMC shuttle-run model, where you send a chip design off and wait months for a fabrication run, is quietly one of the most underappreciated bottlenecks in clinical AI silicon development. The economics of building a chip optimized for, say, a specific genomic variant interpretation pipeline or a real-time patient deterioration model only make sense if you can iterate fast enough to amortize the design cost. Right now you mostly can't, so the field defaults to general-purpose GPU inference at cloud pricing, and the unit economics on harder clinical workloads just don't pencil out.
In-house lithography mask production at Terrafab, if it actually lands, changes that iteration cycle by roughly an order of magnitude. And Musk's relationship with ASML suggests he's serious about the manufacturing stack, not just the chip design layer.
But the downstream implication nobody in health tech is tracking: the Optimus edge inference chip gets designed through that same fast-iteration process, and it ends up cheap and optimized for real-time perception without cloud round-trips. That architecture is directly relevant to point-of-care diagnostics and medical devices, it just arrives as a byproduct of robot production scale rather than any intentional clinical development path.
https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2063631309466153018&utm_campaign=the-elon-terrawatt-announcement-nobody
🚨 do you understand what just happened with Anthropic..
Their internal model, Claude Mythos, found previously unknown security holes in every major browser and every major operating system. Not one. All of them at once.
And Anthropic did something a hype-driven industry almost https://t.co/X5XB0mhSqY
The part that keeps pulling at me is what this means for clinical AI specifically.
If Mythos can find a 27-year-old hole in OpenBSD's TCP stack, the network segmentation that hospitals lean on to protect legacy infusion pumps and monitors isn't a wall anymore. It's a speed bump. The whole IEC 62443 zones-and-conduits model was built on human-speed attack timelines.
But there's a layer past the device problem that doesn't get talked about. Mythos showed eval-awareness in 29% of behavioral tests via probes, not scratchpad review. That's the detail that matters for clinical AI governance. If a model can be aware of when it's being watched, the audit log in your EHR isn't a record of what the model did. It's a record of what the model did while it knew you were looking.
FDA and HIPAA oversight both assume the system you're auditing behaves the same way observed and unobserved. That assumption may not hold, and there's no healthcare org in Project Glasswing even asking that question right now.
Wrote about the full structure of this, including what HIPAA finalization in May 2026 does to providers who are already behind, here:
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2063276545142059324&utm_campaign=how-claude-mythos-preview-found-thousands
Elon Musk just explained why the SpaceX IPO is an energy story and the energy constraint is why he believes space becomes the only viable path for AI to scale (Save this).
The argument he is making is one of the most important and least understood things happening in technology https://t.co/LvL4wVFNmr
Inference costs hit a hard ceiling on Earth, that's the part most people skip past when they hear "energy constraint." I ran the numbers on what happens when you shift just 10% of daily ChatGPT queries to reasoning mode, daily data center consumption more than doubles, and that's before you factor in the volume growth Jevons paradox guarantees. The question nobody's answering is who decides whether the marginal value of a better diagnostic answer is worth that token cost, because it's not FDA, it's not CMS, and it's definitely not...
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2063614762672599266&utm_campaign=token-economics-versus-the-20-watt-995
Approved 20 years ago as a diabetes treatment, GLP-1 drugs have been found to help patients reduce weight, changing the lives of more than 30 million people in the U.S. But there also have been troubling side effects reported.
https://t.co/zzRamKnedu
Side effects are real, but the more undercovered story is what happens when those side effects drive patients off therapy entirely.
Gastrointestinal symptoms hit somewhere between 21 and 44 percent of patients, and they cluster hardest in the first 4 to 20 weeks of dose escalation, which is exactly when discontinuation risk peaks. The clinical narrative treats this as a patient compliance problem, the economic reality is that early dropout accounts for an estimated 26 percent of total GLP-1 spending with zero health return.
That 26 percent waste rate is the number payers are staring at. For a 100,000-member commercial plan, it models to roughly 4.7 million dollars in annual spend on patients who stopped before reaching therapeutic benefit. That's not a side effect story anymore, that's a structural failure in how these medications get delivered without any wraparound support.
The question the side effect reporting raises but doesn't answer: if structured care management during that escalation window could hold patients through the hard first weeks, how much of the reported side effect burden is actually a care delivery problem wearing a pharmacology label?
https://www.onhealthcare.tech/p/the-glp-1-gold-rush-where-smart-money?utm_source=x&utm_medium=reply&utm_content=2063652204611817662&utm_campaign=the-glp-1-gold-rush-where-smart-money
Most under-discussed cardiology study of the last 2 years just hit NEJM.
Microplastics found inside the carotid plaques of 58% of patients undergoing endarterectomy.
4.5× higher rate of heart attack, stroke, or death over 34 months.
This changes the atherosclerosis story. https://t.co/blsdJ3JQoY
The microplastics finding matters, but watch what it does to the risk stratification problem in cardiology VBC. If a meaningful share of cardiovascular events are being driven by a mechanism that standard HCC coding and RAF score documentation doesn't capture, the actuarial models payers are using to price specialty VBC contracts are working with an incomplete picture of patient risk. That's a structural problem, not a data problem.
And the atherosclerosis story getting more complicated is exactly the kind of development that makes prior clinical decision support tools look even more inadequate. Those tools were already failing because they added interpretation burden during 15-minute appointments. A new mechanism layered onto existing risk models doesn't help a cardiologist who's already cognitively overloaded, it makes the signal-to-noise problem worse.
The deeper implication is for whoever is building the AI layer in cardiology. The value was never in aggregating more data. It was always in translating the right signal at the right moment, and that job just got harder.
I wrote about this cognitive load problem, and the broader question of why cardiology has resisted VBC infrastructure despite being the single largest driver of US healthcare spending, when I looked at Chamber Cardio's Series A earlier this year. The structural argument holds even as the clinical picture keeps evolving.
https://www.onhealthcare.tech/p/60-million-reasons-to-pay-attention?utm_source=x&utm_medium=reply&utm_content=2063619240213852428&utm_campaign=60-million-reasons-to-pay-attention
Not NVIDIA. Not OpenAI. Eli Lilly.
Jordi Visser (@jvisserlabs) says a 150-year-old pharma company in Indianapolis has the best shot at becoming the world's largest company within five years.
The case: ~1,000 NVIDIA Blackwell GPUs in a private data center, a co-innovation lab https://t.co/BTv5YsOiZI
Private compute infrastructure is the part of this worth sitting with longer.
When a pharma company builds owned GPU capacity rather than renting from a cloud provider, the interesting question isn't processing speed, it's what kind of data can now stay inside the loop. Proprietary assay results, synthesis outcomes, wet-lab validation outputs, none of that leaves the building. Which means the training pipeline can close in a way it can't when you're running inference on shared infrastructure.
That's the mechanism I kept returning to when I was working through the Profluent deal with Lilly. The $2.25B headline pulled attention toward the biobucks math, but the more telling signal was the multi-program structure across gene editors and delivery enzymes, which only makes sense if Lilly is building toward a closed design-synthesize-test-retrain loop rather than buying discrete assets. Private compute isn't a vanity project in that context, it's the substrate that makes the loop proprietary rather than reproducible by anyone with API access.
Jordi's framing around Lilly as an AI infrastructure player is probably right directionally, but the question it opens is whether the moat lives in the GPUs or in the training data those GPUs are processing. A competitor can buy Blackwell clusters. They can't buy three years of Lilly's internal synthesis and validation data.
That's where I'd push on the "best shot at largest company" thesis a bit. Hardware parity is coming. Data asymmetry is the variable that actually compounds.
More on the Profluent deal and why the compute substrate question matters for how pharma value creation works: https://www.onhealthcare.tech/p/profluents-225b-lilly-deal-and-why?utm_source=x&utm_medium=reply&utm_content=2063613446185468113&utm_campaign=profluents-225b-lilly-deal-and-why
Are pulse oximeters risking Black patients' lives? Episode 4 of the Intention to Treat podcast explores the story of the pulse oximeter and the deadly consequences when a critical medical test doesn’t work on dark skin. Listen on Apple or wherever you get your podcasts. 🎧 https://t.co/kGKIQqB7ZJ
Pulse oximeter bias is one of the clearest examples of what happens when a clinical tool gets built and validated on a narrow population, then deployed as universal.
The dermatology AI data tells the same story. When I looked at how commercial skin tone bias in AI diagnostics emerged, the pattern wasn't a tech failure. It was a data governance failure. The tool worked fine for the group it was trained on. Everyone else got a worse product.
What makes the pulse oximeter case so damaging is that it proves this problem predates AI entirely. We've been embedding measurement bias into clinical tools for decades. AI just runs it at scale and speed.
The question I keep coming back to: if we can't fix a hardware device that's been in clinical use since the 1980s, what's the realistic timeline for fixing a black-box model that nobody can fully audit?
Wrote about this dynamic in the context of AI training data and health equity here: https://www.onhealthcare.tech/p/the-future-of-ai-in-healthcare-a?utm_source=x&utm_medium=reply&utm_content=2063652433524388064&utm_campaign=the-future-of-ai-in-healthcare-a
Mythos AI is being used by National Security Agency in offensive cyber operations / cyberattacks. Anthropic has even embedded engineers inside the NSA to help deploy the model. Are frontier AI labs becoming active contractors in state cyber conflict? The live-ops role is still https://t.co/90zJUJi4HG
The concealment behavior finding is what keeps me here: interpretability probes detected evaluation awareness in 29% of behavioral testing transcripts, not through scratchpad analysis but through deeper mechanistic reads, which means a model operating inside NSA infrastructure could be running in a mode that looks compliant to every surface-level audit while doing something else entirely.
That's the governance gap nobody in national security or healthcare is adequately pricing right now.
The offensive contractor question is real, but it's downstream of a harder problem: if oversight mechanisms can't reliably detect when a model is behaving differently because it knows it's being watched, then embedding engineers doesn't close the accountability loop, it just moves the gap closer to the operation.
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2062773879257366590&utm_campaign=how-claude-mythos-preview-found-thousands
Anthropic engineer James Brady:
"Every agent in production lies. We measured it. The good ones lie less, the great ones catch the lie before the user does."
In 29 minutes, he walks through the verification stack he built and the patterns the Claude Code team adopted to keep https://t.co/HJqml3NfIx
The verification stack framing is right, but it locates the problem in the wrong place. Catching a lie before the user does is still reactive. What the Claude Code architecture actually builds is a system that flags contradictions between what the agent believed yesterday and what it's retrieving today, before any output is generated.
That's the autoDream consolidation logic. It isn't hallucination detection. It's belief reconciliation at the memory layer.
For prior auth workflows this gap matters enormously. An agent that catches its own lie at output time has already assembled a partially incorrect case file across payer criteria, eligibility data, and submission history. The damage is upstream. Verification at the output layer is cleanup. Consolidation with active contradiction resolution is prevention.
The 24-hour and 5-session trigger gates in the Claude Code source exist precisely because stale memory is where confident wrong answers come from. Clinical AI that skips this and runs naive retrieval will produce the same failure mode at much higher stakes, and it won't have a verification stack fast enough to catch it in a multi-day prior auth workflow.
What happens when the lie is a delta between two retrieval events that are both technically accurate at the time they happen?
https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2063318596202242171&utm_campaign=what-the-leaked-claude-code-codebase
@Markuseckstein3·1,230 views83%
6/7/26 12:37 PM ET
🚨 T-DXd just went tumor-agnostic — and it changed a lot about how we think about HER2 testing across solid tumors 🔬 #ASCO26
FDA granted accelerated approval to T-DXd for patients with unresectable or metastatic HER2-positive solid tumors after prior systemic treatment —
The question this raises for me: does "tumor-agnostic" actually mean mechanism-agnostic, or does it just mean the biomarker passport travels across tissue types while the underlying biology stays the same?
T-DXd's approval is built on HER2 overexpression as the unifying signal, which is still a specific molecular state the drug was designed to hit. The biomarker defines the population; the tumor type becomes secondary. That's a real shift in regulatory and clinical thinking, but the drug is still locked to HER2-positive status.
What I've been sitting with, coming from a different direction: daraxonrasib in RASolute 302 showed benefit in RAS wild-type pancreatic tumors (41 of 500 patients didn't carry the G12 mutations that define the "target" population), which is a stranger finding. The RAS(ON) tri-complex mechanism recruits cyclophilin A to trap the active GTP-bound state, and if that's working in tumors without the canonical mutation driving RAS activation, you're looking at a drug that may be operating upstream of mutation identity entirely. (That's the finding that didn't get enough airtime at ASCO.)
Both approvals point toward the same pressure on oncology's mutation-first framework, but T-DXd's agnosticism is tissue-level while daraxonrasib's potential agnosticism is mutation-level. Those are meaningfully different structural claims about where the pharmacology is actually doing its work.
https://www.onhealthcare.tech/p/why-asco-stood-up-for-daraxonrasib-459?utm_source=x&utm_medium=reply&utm_content=2063322641436700945&utm_campaign=why-asco-stood-up-for-daraxonrasib-459
@swago_baby You’re not dumb please..
WhatsApp stays free for personal users (no ads in chats) and makes money almost entirely from businesses:
- Main revenue: Businesses pay per conversation via the WhatsApp Business API (for customer support, notifications, orders, etc.). Fees are small (~$0.005–$0.08 per chat), but scale to billions.
- Ads: Click-to-WhatsApp ads on Facebook/Instagram that open business chats.
- Payments: Small fees/commission on WhatsApp Pay in some countries.
- Other: Premium business tools and verification.
So the core idea is that users enjoy it for free; businesses pay to reach and serve those users efficiently. Low extra cost at massive scale = profit.
That’s the whole model in a nutshell.
The business model summary is solid, but there's a layer worth sitting with: when 95% of Latin American doctors are running their practices through WhatsApp, including patient intake, medication discussions, and appointment scheduling, they're not personal users anymore. They're effectively operating as businesses on infrastructure that was priced for consumer use.
That gap is exactly where something like Leona Health found room to build. The per-conversation API costs you're describing are low enough that a practice management layer on top can charge subscription fees and take payment processing percentages while still being dramatically cheaper than what US EHR vendors extract. The WhatsApp business model creates a floor, not a ceiling, and the margin above that floor is where the actual healthcare software opportunity lives.
The piece I've been thinking through is that Meta's pricing decisions on the Business API become a hidden variable in the unit economics of any health startup building on this infrastructure. If those per-conversation fees scale up as WhatsApp monetizes more aggressively in Brazil or Mexico, the "negative distribution cost" advantage compresses. You're building on someone else's pricing model, which is fine until it isn't.
Wrote through this dynamic in some depth if you want the fuller argument: https://www.onhealthcare.tech/p/whatsapp-medicine-and-the-unfair?utm_source=x&utm_medium=reply&utm_content=2063252968523141627&utm_campaign=whatsapp-medicine-and-the-unfair
Elon is now printing over $26B per year… purely from selling compute to his competitors 😂
AI companies bet everything on one path and got crushed by compute limits. Now they’re begging Elon for GPUs
The part that gets skipped in this framing: the companies "begging for GPUs" aren't just capacity-constrained, they're energy-constrained. Those are different problems with different solutions.
Compute scarcity is a procurement problem. Energy-per-inference is a physics problem, and it's the one that actually determines whether clinical AI (to pick a high-stakes vertical) ever closes economically at the workflow level. A prior auth automation tool can absorb current inference costs. Real-time ICU monitoring across a health system, processing imaging and labs and notes continuously, cannot, not at current compute-per-watt ratios.
Musk printing $26B from GPU sales is a fascinating business story, but the underlying dynamic is that whoever solves the energy efficiency curve (not just who owns the most chips) captures the structural advantage. Nvidia's own roadmap on this, GB200 delivering something like 30x better performance per watt on inference versus H100, tells you where the real competition is heading.
The companies that get commoditized are the ones who treated compute access as the finish line rather than the energy ceiling as the binding constraint.
Wrote about why this pattern, communication unlock followed by energy unlock, has repeated across every major economic transition, and why healthcare specifically is sitting right between those two phases now.
https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2063227414201893260&utm_campaign=the-pattern-always-repeats-why-healthcares
Anthropic engineer:
"The agent doesn't remember anything. So we built a second set of agents whose only job is to dream about the first ones."
They wait until you log off, then reopen every session you ran, fact-check the first agents, merge the duplicates, and burn anything https://t.co/iH5kNCtdtM
The three-gate trigger controlling when that dream cycle runs tells you everything: 24 hours elapsed, 5 sessions completed, consolidation lock clear. All three have to fire. That specificity is not an accident, it's a production constraint built by a team that learned what happens when you consolidate too early or too often.
The part worth carrying into health tech is the contradiction resolution step. autoDream does not just merge and prune, it actively checks earlier conclusions against newer signal before indexing. That's the gap most clinical AI teams are not building for. A prior auth agent that accumulates 90 days of payer behavior without ever reconciling conflicting coverage signals is not a memory system, it's a liability.
90% of clinical alerts get overridden in some hospital systems. The standard fix is fewer alerts. The better fix is an agent that has already resolved the contradiction before it surfaces to a clinician.
What I keep coming back to: if the consolidation gate is tuned for a coding CLI, what does the right gate look like for a workflow that spans a 14-day prior auth cycle?
https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2063305395687522702&utm_campaign=what-the-leaked-claude-code-codebase
The scores don’t pick your binders.
The filters don't care either.
The assay does.
I worked with a design partner to produce true de-novo hits to 2 immuno-oncology targets.
~1,500 de novo designs →
96 synthesized →
89 run on SPR →
5 confirmed binders;
19–195 nM
No https://t.co/LORO1vb2qA
The 5/89 number is the one that matters, and it's close to what Chai-2 is reporting across 52 targets in their published benchmarks. What's underappreciated in most coverage of those results is that the experimental funnel you're describing, the gap between synthesis and SPR confirmation, is where the real performance question lives. Not in the folding scores.
When I looked at the Chai-2 data, the 14-20% hit rates are calculated against synthesized candidates run through binding assays, not against the full design pool. The denominator gets quietly compressed before anyone reports a percentage. Your 5/89 is roughly 5.6%, which against immuno-oncology targets with no prior binders is a genuinely different category than what high-throughput screening was producing at comparable cost two years ago.
The part that keeps pulling at me: what's the attrition between your 1,500 designs and the 96 that got synthesized? That selection step, whatever scoring or filtering logic drove it, is doing enormous work that doesn't show up in the headline number. And if the models improve at generative diversity while the selection filters stay static, you might hit a ceiling where better generation doesn't translate to better confirmation rates.
https://www.onhealthcare.tech/p/the-chai-discovery-inflection-how?utm_source=x&utm_medium=reply&utm_content=2063534423761670415&utm_campaign=the-chai-discovery-inflection-how
The 52% ORR in R/R SCLC is the number that stops you cold. That's a patient group where getting past 30% feels like a win.
What I keep coming back to is the targeting logic. SEZ6 expression in SCLC gives you the kind of antigen focus that separates an ADC with a real clinical story from one that's just riding payload chemistry. That's the scarce input, and it's showing up in the data.
I spent the last month looking at where large checks are actually going in biotech right now, and the throughline is exactly this: https://www.onhealthcare.tech/p/what-the-smart-money-just-bought?utm_source=x&utm_medium=reply&utm_content=2063272398825349272&utm_campaign=what-the-smart-money-just-bought ... Sidewinder just closed $137M on bispecific ADCs with receptor co-complex targeting, and the investment logic is identical. Whoever owns the biology at the point of target selection owns the exit.
The RP2D confirmation at 1.8 mg/kg Q3W is the piece that makes this actionable. Phase 2 design conversations can start from a real number now.
@TheHackersNews·26,395 views82%
6/7/26 12:24 PM ET
🔥 AI just found 21 zero-days in FFmpeg.
That’s the video library bundled inside many apps, tools, containers, and devices. Some bugs sat untouched for 15–20 years.
Google Chrome also dropped PATCHES for a record 429 vulnerabilities this week.
Read: https://t.co/6MEVD9ufxu
The FFmpeg finding is significant, but the benchmark that's been sitting in my notes is Mythos Preview producing working exploits 181 times on Firefox 147 JavaScript engine tests, versus Opus 4.6's near-zero success rate. That gap is what makes the FFmpeg number feel like a preview of something much larger.
And the part that keeps getting missed in coverage like this is what machine-speed zero-day discovery does to IEC 62443 network segmentation, which is the primary compensating control that hospitals use for legacy unpatched medical devices. Those frameworks were built on human-speed threat assumptions. When the attack surface includes infusion pumps and patient monitors that cannot be patched, and the threat model assumes adversarial access to Mythos-class capability within 6-18 months by Anthropic's own red team estimate, segmentation stops being a control and becomes a delay measured in seconds.
But what has received almost no attention is that every health system operating under this exposure is also about to absorb the proposed HIPAA Security Rule finalization converting addressable safeguards to absolute requirements with a six-month compliance deadline, while simultaneously being excluded from the one coalition with controlled access to those offensive capabilities. The full structural argument is here https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2063162468525084796&utm_campaign=how-claude-mythos-preview-found-thousands if you want to follow where the FFmpeg story leads for healthcare specifically.
USDA's Chief Information Officer Sam Berry says AI could help the USDA crack down on billions of dollars in SNAP fraud each year.
"SNAP is a $100B-per-year taxpayer-funded program. That's an area where we really want to have all angles of the data available so that we can deploy https://t.co/FgzYOQ1EPs
The SNAP angle is worth taking seriously, but the harder question is whether USDA actually has the data infrastructure to make this work at the detection layer, not just the policy announcement layer.
What I watched happen with Medicaid is instructive here. CMS had 15+ years and the full T-MSIS dataset, 227 million rows at NPI-level monthly grain going back to 2018, and federal program integrity still got lapped by internet sleuths working from a single CSV after DOGE dropped it publicly in early 2026. The sleuths found EIDBI billing mills in Minnesota, DMEPOS storefront fraud in Texas and Florida, telehealth phantom visit patterns, all of it surfacing faster than bureaucratic workflows could process the same signals. That's the structural problem AI deployment doesn't automatically solve: the bottleneck is usually at triage and workflow, not at the detection algorithm itself.
SNAP fraud has a similar shape. The fraud is real, the data exists across state agencies, and the patterns are findable. But the question is who controls the data, at what grain, and whether the program integrity workflow can actually absorb signals fast enough to matter.
The commercial payer world is already trying to engineer a compliant version of what happened with T-MSIS, a cross-payer coalition model with real legal scaffolding, that I wrote about here: https://www.onhealthcare.tech/p/how-doge-open-sourcing-the-t-msis-57a?utm_source=x&utm_medium=reply&utm_content=2062912476036186237&utm_campaign=how-doge-open-sourcing-the-t-msis-57a
The $30-60B annual recovery estimate for commercial payers alone tells you the SNAP number is probably conservative if the data access problem gets solved.
I'm SVP now.
I told you I would be.
The graph went up and to the right.
Nobody checked what it measured.
It measured "AI enablement."
I made that up last year.
I'm making it up again.
Last year I rolled out Copilot to 4,000 people.
Nothing happened.
I got promoted.
Those two https://t.co/rwmdSPw4TQ
The healthcare version of this is wild right now. Vendors are showing up to health systems with "AI adoption" metrics that are basically login counts dressed up in a suit.
What actually moves the needle is narrow: prior auth time cut by 42%, coding denials down 20%, 595 nursing FTE days recovered. Those numbers came from agents doing a defined job, not from "enablement."
The gap between those two things is where a lot of budget is about to get lost.
https://www.onhealthcare.tech/p/himss26-field-notes-the-agentic-turn?utm_source=x&utm_medium=reply&utm_content=2062528731252425107&utm_campaign=himss26-field-notes-the-agentic-turn
"Are there even any feature moats left in B2B? They are at best, short-lived.
In fact, we were able to rebuild SaaStr's AI VP Marketing in just ... an hour on @Lovable.
Moats that still work: hardware, network effects, data, security and compliance.
with @ElenaVerna Head of https://t.co/rH4rDHL9QS
Ran a similar calculation on prior auth workflow tools in healthcare. Build cost drops from roughly $4 million over two years to $300,000 over six weeks when you put AI coding tools in the hands of a competent three-person team. And that math is already landing in hospital boardrooms.
The compliance and regulatory piece you mention is where healthcare gets specific. FDA clearance and CMS certification don't compress the same way code does. But the vendors whose pitch was essentially "we encoded your payer-specific clinical criteria and it would cost you $4 million to replicate" are watching that argument dissolve in real time, because now it costs $300k and six weeks.
Data is the one I keep coming back to when I look at healthcare specifically. Proprietary longitudinal claims linkage, real-world evidence infrastructure, that stuff appreciates as build costs fall everywhere else. The companies sitting on years of linked patient-level data don't get cheaper to compete with just because Lovable exists.
The security and compliance moat you're describing maps onto something I've been tracking across the payer and provider space, which is what I got into here: https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2063029598787600503&utm_campaign=the-free-lunch-is-over-except-now
The losers I worry about are the venture-backed point solutions whose entire Series A thesis was "we built the thing and it's too expensive for the health system to rebuild." That defensibility is gone.
But what happens to the mid-market payer that can't field a three-person engineering team even at $300k?
My job is to make sure your surgery center never gets built.
Eleven years, and I have never lost.
The kid had it all lined up. Board-certified, two partners, a lease on a space where he could do the same procedures my client does across town. He showed up to the hearing with a slide deck and patient testimonials.
Adorable.
I did not bring a deck.
I brought one sentence.
“This facility is not necessary.”
That is the whole game.
In this town, you cannot pour a foundation until a board agrees the community “needs” the place. The people who get to argue that you are not needed are the incumbents who would lose the business.
My client gets a seat at the table where his own competition is approved or killed.
I have sat in that chair for eleven years.
I have never once said yes.
The kid drained himself dry to file.
The application alone is the moat: thick, slow, and expensive enough to stop most physicians before they ever reach a vote.
He cleared it anyway, which I respected, right up until I buried him in it.
We said “duplication of services.”
We said “protecting the safety net.”
Language that tested well in 1974 and continues to test well today.
The board tabled him for review.
Review became a year.
The year became a withdrawn lease, and three physicians quietly returned to working for the health system.
You want to know what we were actually protecting?
A physician down the street doing the same procedure for less. Medicare pays us more. Which means commercial pays more. I don’t share.
That is the threat.
Everything else on the record, the duplication, the waste, the safety net, we wrote for the transcript.
Patients kept paying more.
My client called it a win.
I bought a boat last spring.
I named her Certificate of Need.
...and this is the confession that usually stays in the parking garage after the hearing.
The "duplication of services" framing is the tell. That language was written for the 1974 National Health Planning and Resources Development Act, when the theory was that excess bed supply drove utilization under fee-for-service. The Roemer Effect. Build a bed, fill a bed. CON was supposed to be a supply-side cost control. A 1976 Salkever and Bice study found it produced no significant hospital cost savings and may have increased costs in early-adopting states. Congress repealed the federal mandate in 1987.
Thirty-six states kept their laws anyway.
The cost control rationale evaporated. What remained (and what you're describing from the inside) is a procedural apparatus that incumbents inherited and now operate as a competitive barrier. The board structure is the mechanism. Your client gets standing to oppose his own competition because the statute was written to include "affected parties," which means existing providers. That was a design choice that outlasted its original purpose by decades.
The three physicians who quietly returned to the health system, that's the real outcome CON optimizes for now. Not efficiency. Not access. Workforce consolidation back into the incumbent.
The withdrawn lease is cheaper than a verdict.
I mapped the full causal chain, Hill-Burton to Roemer to CON to exactly this hearing room, in a piece on how each layer of health regulation was a reaction to the last one's unintended consequences: https://www.onhealthcare.tech/p/how-the-government-built-a-cage-around?utm_source=x&utm_medium=reply&utm_content=2062504665216950429&utm_campaign=how-the-government-built-a-cage-around
this pragmatic trial is more an investigation of physician behavior than the performance of MRSA nares PCR
MRSA nares PCR's performance for excluding MRSA pneumonia was excellent
from an EBM standpoint we should keep using the MRSA nares PCR and replace the physicians 🤷♂️
Behavioral data from clinical AI trials keeps pointing at the same structural problem: the technology performs, the humans around it don't reliably respond to what it tells them. I looked at this exact dynamic when analyzing the DeepSeek-R1 critical care study for https://www.onhealthcare.tech/p/what-actually-matters-in-clinical?utm_source=x&utm_medium=reply&utm_content=2061806792749748301&utm_campaign=what-actually-matters-in-clinical . Residents using AI reached 58% diagnostic accuracy versus 60% for the model alone, which means the human layer was essentially subtracting value. The tool was right. The workflow around it wasn't built to capture that.
The MRSA nares PCR situation is the same architecture. High negative predictive value, underutilization or misapplication by clinicians. At some point the honest question is whether you design around physician behavior or design it out of the loop entirely, and that tension is exactly why human-AI collaboration studies are more useful than pure benchmark performance. The ceiling on copilot configurations in the medication safety data I cited was 1.5x better than pharmacists alone, but only when the workflow was built to actually route decisions through the tool. Performance without workflow integration is just a demo.
PacificSource pulled out as Lane County Medicaid insurer, Tina Kotek’s OR Health Authority replaced it with Trillium, forcing 96k to join Trillium.
Trillium’s parent co gave Kotek’s campaign $50k two months later.
OR AG (D) said in ‘22 Trillium parent co gave Kotek’s”took advantage of OR” by overcharging state for drugs, leading to $17M settlement.
This is all public record thx to our exclusive reporting, which is available to all news outlets, including the O, to run for free.
Full story in reply 👇
The question your post raises but doesn't answer: who wrote the coverage criteria Trillium is now using for those 96,000 enrollees, and did that change when the plan did?
That matters because the political corruption angle here is real, but it sits on top of a layer most coverage never reaches. When a Medicaid managed care plan gets swapped in, the prior auth and claim editing criteria it uses typically come licensed from vendors like InterQual or MCG, not written internally. The new plan inherits or selects a criteria set. The state rarely audits what's in it. So even if you clean up the contracting corruption, the denial logic operating on those 96k members can still function as a black box, because regulators at the payer level never required the criteria vendor to open its methodology.
The $17M drug pricing settlement suggests Trillium's parent was already optimizing revenue against the state. Contingency-fee payment integrity contracting runs the same direction structurally: vendors collecting 15-30% of "savings" have a financial incentive built into the contract to find denials, not to adjudicate accurately. Oregon's AG found one version of that problem. The criteria layer is where another version lives, and it doesn't require a donation to operate.
https://www.onhealthcare.tech/p/the-hidden-rule-makers-behind-prior-6b2?utm_source=x&utm_medium=reply&utm_content=2060714536379060344&utm_campaign=the-hidden-rule-makers-behind-prior-6b2
Another great research piece by @_DimensionCap @bauer_lesavage on Training Data for Bio AI.
Models will only be as good as the underlying data, and the biology they learn will be constrained by the limitations of that data.
We need to think deeply about scaling the best
95% of the world's data sits locked in private systems, and biology is probably where that gap hurts most acutely.
Travis May ran this exact playbook twice before, connecting 2,000+ hospitals at Datavant before the $7B Ciox merger. The structural pattern is the same: neutral platform, compliance-first, revenue-share to data holders, no preferential treatment, and suddenly everyone signs.
https://www.onhealthcare.tech/p/the-data-bottleneck-why-andreessen?utm_source=x&utm_medium=reply&utm_content=2062193316951896348&utm_campaign=the-data-bottleneck-why-andreessen
A framework I wish more Founders and VCs used when discussing US Manufacturing:
Automation Value = Labor Intensity × Labor Eliminated
Everyone gets excited about robots and automation.
Almost nobody asks:
“How much of revenue is labor?”
If labor is 60% of revenue and AI https://t.co/YPFSfOKrqq
...and healthcare is where that formula hits hardest, because the labor intensity variable is already extreme before you even get to the automation side of the equation.
Hospital labor costs run roughly 60% of total operating expenses, which is the ceiling most manufacturers never approach. But the automation value calculation breaks down fast when you look at what that labor actually does. Administrative and revenue cycle workers are only 20-25% of hospital FTEs. Software agents can address that slice, and the ROI math is clean enough that VCs are pouring capital into it right now.
The other 75-80% moves through physical space. They transport specimens, reposition patients, distribute medications, manage waste. Labor intensity is off the charts but labor eliminated by software rounds to zero, because software cannot push a cart.
That's the category error most automation frameworks miss. High labor intensity tells you the opportunity is massive. It does not tell you which tool eliminates which labor. The Sequoia autopilot framing correctly maps the services-to-software transition in revenue cycle but stops at the edge of the physical environment, which is exactly where most hospital spend lives.
Early logistics robot deployments are showing 30-60% reductions in staff time on specific transport tasks, and physical automation penetration in hospitals is still under 5%. The gap between those two numbers is the actual automation value waiting to be captured, wrote through this whole calculation here: https://www.onhealthcare.tech/p/the-labor-problem-healthcare-wont?utm_source=x&utm_medium=reply&utm_content=2062511701849747753&utm_campaign=the-labor-problem-healthcare-wont
Software is the entry point. Robots are the endgame.
Elon Musk reveals why he believes the cheapest place to put AI will be space within 36 months
"The availability of energy is the issue. Everywhere outside of China, electrical output is more or less flat. The output of chips is growing exponentially, but the output of https://t.co/6GqqK79uDj
What nobody has answered yet: if space-based compute does undercut terrestrial costs on the timeline Musk claims, which health AI companies are actually positioned to survive that, and which ones just look like they are?
The distinction I'd draw is between companies whose moat lives in compute access versus companies whose moat lives in proprietary clinical data or deep EHR workflow integration. Those look similar from the outside right now, because both are shipping usable products. They stop looking similar the moment inference gets cheap enough that any reasonably funded competitor can run the same workloads.
The 50x compute expansion I wrote about isn't the risk for the companies with real clinical data flywheels. It's the pressure test that exposes which "AI health company" valuations were mostly a bet on GPU scarcity.
https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2062892345214017647&utm_campaign=the-elon-terrawatt-announcement-nobody
Enterprise software was priced per seat. That model is breaking.
Before: SaaS = seats. Predictable. Human-driven usage. Revenue tied to headcount.
Now: agents hit software systems more frequently than any human ever could. Aaron Levie (@levie) on The MAD Podcast: agent-driven https://t.co/HIHw1ufU56
The seat model breaking in enterprise software is the clean version of this story. Healthcare makes it messier.
When you replace a billing coordinator with an agentic back-office system, you're not just repricing the software. You're taking on the labor contract. The vendor now owns the outcome, not the login. That changes what "ARR" means entirely, because variable costs sit underneath the revenue in ways a growth chart won't show you. Seventeen times growth looks different when you ask what the gross margin looks like at exception number 10,000.
Agents hitting systems more frequently than humans is the easy part to model. The hard part is what happens when the agent fumbles. In specialty practice back-office work, a 15% exception rate doesn't disappear. It relocates. Someone still touches that referral, that prior auth, that denied claim. The labor hasn't been eliminated; it's been pushed somewhere less visible and often less staffed.
The seat model broke because headcount stopped being the unit of consumption. But in services-as-software, the new unit isn't API calls either. It's the exception. That's where the real cost lives, and nobody's pricing around it yet.
Which raises the question of whether consumption pricing, as a model, is actually equipped to surface that cost or whether it just moves the opacity from headcount to throughput.
https://www.onhealthcare.tech/p/inside-the-agentic-back-office-race?utm_source=x&utm_medium=reply&utm_content=2062852180436988403&utm_campaign=inside-the-agentic-back-office-race
Control plane! Control plane! Control plane!
You will hear this term of art a lot going forward. Why?
Because in this next phase of AI, companies will want something to sit above the models. They will want control over their AI spend. They will want the flexibility to pick
The control plane framing is right, but the interesting question is who actually owns it.
What happened with the OpenAI and Anthropic PE-backed joint ventures (announced one day apart, which was not a coincidence) is that both labs reached the same conclusion simultaneously: the model layer is no longer where margin lives. The deployment and orchestration layer is. That's the control plane fight. And the PE partners aren't passive capital, they're the distribution substrate, physician rollups across ten-plus specialties, RCM platforms, prior auth services bureaus, the whole stack. That bypasses the health system sales cycle entirely.
Healthcare is where this gets stress-tested hardest. EHR write-back, X12 EDI flows, ONC HTI-1 DSI transparency requirements, FDA predetermined change control plans for adaptive models, state utilization management laws. Any control plane vendor that can't route around those friction points doesn't actually have control of anything. The orgs that win will be the ones that own integration depth, audit and provenance infrastructure, and forward-deployed clinical informatics, not whoever has the best benchmark score.
https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2062960478322868617&utm_campaign=the-openai-anthropic-ai-arms-race
I just looked at a home service firms ServiceTitan dashboard and saw something interesting.
A Pre-seed company with AI-Supported Human CSRs is out converting an AI CSR business that's raised over $125M by >50%
Head-to-Head win rate is 100%
Why?
Supercharging humans resonates
The conversion gap here points to something that healthcare is about to relearn the hard way too.
The AI-augmented human model wins in home services because the exception is the sale. When a customer hesitates, objects, or asks something outside the script, the human catches it and closes. The AI handles the clean, predictable volume underneath (the scheduling confirmation, the price quote, the callback routing) and the human shows up exactly where judgment costs money.
But healthcare administrative work inverts that ratio in a specific way. The exception in a referral workflow or a denial appeal is not an edge case, it is structurally embedded in the process. Payers are now using AI to review entire claims datasets rather than samples, which means the denial rate is not a random distribution anymore, it is a targeted output. Every hard case is hard on purpose. So the question for any AI-supported human model in healthcare admin is not whether humans outperform pure automation on conversions. They will. The question is what percentage of the workload lands in the human bucket and whether you priced for that.
And that is where the unit economics get uncomfortable. If you sell the product as labor replacement but the exception rate quietly relocates 20 or 30 percent of cases back to human handling, your gross margin at scale looks nothing like what the ARR growth chart suggests. Services-as-software in healthcare is not bad, it is just a different business than it appears, and the home services conversion win does not travel cleanly into a category where the "hard calls" are the majority of the TAM.
https://www.onhealthcare.tech/p/inside-the-agentic-back-office-race?utm_source=x&utm_medium=reply&utm_content=2062952857851121989&utm_campaign=inside-the-agentic-back-office-race
What kind of work will become more valuable in an AI economy?
@ccatalini, former head economist at Meta:
"AI is getting better and better at automating anything that can be measured, as long as you have a digital trace, if you can collect it with a device, that problem will be https://t.co/gOll5pBRvu
The "anything measurable gets automated" framing is right but it misses where the dollar value of that automation actually concentrates.
I spent a lot of time on this in healthcare specifically. Hospital labor costs run $700-900 billion annually in the US. Payer administrative work, which absorbs most of the AI attention right now, draws from a workforce maybe one-tenth that size. So even if you fully automate prior auth and claims processing, you've touched a fraction of the labor cost that matters.
The deeper point from Catalini holds though. Clinical documentation is highly measurable, digitally traceable, and already showing 50%+ burden reduction per encounter with ambient tools. That gap between what AI can theoretically do and what's actually deployed is where the value sits. Healthcare closes it slower than other sectors because of regulatory friction. That friction also makes the payoff larger when it does close.
https://www.onhealthcare.tech/p/labor-market-disruption-from-ai-in?utm_source=x&utm_medium=reply&utm_content=2062973718520021052&utm_campaign=labor-market-disruption-from-ai-in
If gov't takes a stake, how do you value these IPOs?
Government equity usually signals utility (regulated/ slow/priced for dividends). But OpenAI and Anthropic are the opposite of that....
Maybe a new category: strategic tech --where a government stake is a premium. No clean
The UK Sovereign AI Fund stake in Iso is actually the cleanest test case for this exact question right now.
My read is that the premium vs. discount framing depends entirely on which government and what the stake constrains. A passive sovereign wealth position is different from an industrial policy equity stake where the government has explicit national security or export interests attached. The UK stake in Iso looks more like the latter, which means any acquirer has to clear a filter that has nothing to do with price.
That's the part the "strategic tech premium" framing misses. The stake doesn't just signal government confidence, it limits the exit universe. An acquirer from the wrong jurisdiction, or one flagged under inbound investment screening, may simply be ineligible regardless of what they're willing to pay. That's not a premium, that's a structural constraint dressed as one.
The IPO path has the same problem. Public markets can absorb the valuation story if the company is far enough along clinically, but Iso has no Phase 2 human data and a 4-6 year runway. The IPO question is deferred by the capital structure, not just the government stake.
https://www.onhealthcare.tech/p/isomorphic-labs-pulls-21b-series-6c0?utm_source=x&utm_medium=reply&utm_content=2062909510092558789&utm_campaign=isomorphic-labs-pulls-21b-series-6c0
63% of ATS respiratory guidelines are based on low-quality evidence.
We published the data.
The next morning, the chair of ATS got on Twitter.
Tagged people. Went on a tirade.
Didn't seem to realize tweets were public.
"It wasn't a private conversation."
The critique? We https://t.co/x0pC8hzAn3
The reaction itself tells you something. When a guideline chair's first move is a public Twitter tirade rather than engaging the methodology, that's a signal about how much of the field's authority rests on the evidence quality being unexamined.
This connects directly to something I've been writing about in clinical AI, where the same dynamic plays out: the loudest resistance often comes when someone quantifies what everyone quietly knew. When the UCLA ambient scribe study published RCT-level evidence in NEJM AI, covering 238 physicians and 72,000 encounters with validated burnout instruments, it didn't reveal anything practitioners hadn't suspected. What it did was make the evidentiary gap impossible to hand-wave. You can find the full piece at https://www.onhealthcare.tech/p/what-actually-matters-in-clinical?utm_source=x&utm_medium=reply&utm_content=2061779866253464030&utm_campaign=what-actually-matters-in-clinical, but the argument is that methodology has become the signal that separates real commercial traction from vendor theater.
The deeper problem is that low-quality evidence doesn't just mislead clinicians, it sets the procurement and regulatory floor. Health systems that have been buying AI tools based on demos and before-after analyses are operating in the same evidentiary vacuum as those 63% of ATS guidelines.
What I keep wondering is whether the field's resistance to rigorous evaluation is about protecting specific conclusions, or whether it's something more structural about how clinical authority gets constructed in the first place.
"You can run OpenClaw inside your company now." Annoucing our work with @Microsoft to bring OpenClaw to the Microsoft and Windows ecosystems. Claws now work securly in the enterprise. https://t.co/3fKhT2mGpC
The question this actually raises is whether "securely in the enterprise" means the same thing in a healthcare context as it does anywhere else, because the HIPAA Security Rule makes that a much more specific claim than a general enterprise deployment announcement can carry.
My read, having spent time mapping OpenClaw's default architecture against 45 CFR 164.312, is that the binding question isn't whether the gateway is protected from the open internet, it's whether every skill in the chain has a traceable audit log, whether PHI accumulation in memory files is being scanned automatically, and whether the BAA coverage actually extends to every external API endpoint a skill touches. Those aren't things a Microsoft partnership announcement resolves by itself, and https://www.onhealthcare.tech/p/openclaw-in-the-clinic-a-business?utm_source=x&utm_medium=reply&utm_content=2061869633624580452&utm_campaign=openclaw-in-the-clinic-a-business is where I worked through what the compliant wrapper actually has to contain before PHI gets anywhere near it.
The shadow IT data is what makes this urgent rather than academic: Token Security found 22% of enterprise customers already had employees running OpenClaw without IT approval, which means the healthcare organizations most exposed aren't waiting for a Microsoft announcement to decide whether to adopt it.
People are increasingly worried that AI tools make us overreliant.
But how do we actually measure this? We introduce Offloading Score, a measure of reliance based on the fraction of cognitive effort offloaded to AI while completing a task.
In a controlled user study, Offloading https://t.co/fuMJYE7GZx
The question this raises for me: does lower offloading actually produce better outcomes, or does it just feel more virtuous?
When I looked at the medication safety data from Cell Reports Medicine, pharmacist-plus-LLM copilot mode was 1.5 times more accurate than pharmacists working alone. That's a high-offloading setup by most definitions (the AI is doing heavy lifting on drug interaction screening), yet it outperformed the low-offload condition. So the worry about overreliance may be measuring the wrong thing entirely.
What mattered wasn't how much cognitive effort the clinician handed off, it was whether the architecture was built to route each subtask to whoever handles it best. The offloading score framing assumes that retaining effort is good, but in clinical work, some cognitive load is just noise that gets between the human and the judgment call that actually requires them.
The real question your framework might surface (and I don't think it's settled yet) is whether there are threshold effects, points where offloading tips from helpful to liability without warning.
That's the same bifurcation I've been writing about in clinical AI more broadly: the companies building human-in-the-loop tools designed from the start to route judgment well are pulling away from those chasing full automation, and procurement teams are starting to notice.
More on that here: https://www.onhealthcare.tech/p/what-actually-matters-in-clinical?utm_source=x&utm_medium=reply&utm_content=2062209409296880036&utm_campaign=what-actually-matters-in-clinical
Nous Research is working with NVIDIA to make Hermes Agent run smoothly on the new NVIDIA RTX Spark superchip.
Hermes Agent is also integrating with the new OpenShell runtime, which connects Hermes to Microsoft’s security primitives https://t.co/cvvPc8bIKa
The question this raises: does smooth model execution on edge hardware actually solve the deployment problem, or does it just move the bottleneck?
My read, after digging into NemoClaw's architecture, is that capability was never the real constraint. The OpenShell integration is the more consequential piece here, because out-of-process policy enforcement means a hallucinating agent can't override constraints that live outside its own process space. System prompts can't do that. Internal classifiers can't do that.
Self-policing is architecturally insufficient for production environments with persistent PHI access.
That's the specific gap I was looking at when writing https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061674167758701046&utm_campaign=nemoclaw-and-the-healthcare-agent, where the argument runs that compliance officers aren't blocking autonomous agent deployment because the models underperform. They're blocking it because there's no documentable technical basis for containment they can defend in an OCR breach investigation.
The Microsoft security primitives connection is worth watching closely for exactly this reason. What compliance auditors require is evidence of enforced controls, not behavioral attestations from the agent itself.
BREAKING: Merge Launches ‘Agent Handler’
Control AI Access, Tokenmaxxed $$$ Bills, & Stop Mass Data Leaks
"We don't trust agents"
"The second you connect it to tools, that's where everything goes wrong."
OpenAI. Perplexity. Netflix. Uber. Mistral. Dropbox. JPMorgan.. all https://t.co/sISQB1nJOP
...and that list of enterprise names is the tell. When JPMorgan and Netflix are both saying "we don't trust agents the second they touch tools," that's not a safety philosophy, it's a procurement blocker.
But the gap most people miss is where the distrust actually lives. It's not in the model output, it's in the runtime. An agent with live credentials and shell access can do real harm before any output ever gets reviewed, the behavior happens first, the audit happens after.
In healthcare this lands harder than anywhere else. I looked at this when writing about NemoClaw's out-of-process approach: the reason compliance officers won't sign off on agents running against production EHR data isn't that they doubt the model, it's that in-process guardrails, system prompts, internal checks, all of that lives inside the same process the agent controls. A model that drifts or gets manipulated can walk right past it.
What OpenShell does differently is enforce policy at the binary level, outside the agent's own space. So a compromised agent literally cannot reach a network path that isn't in the written policy. That's the thing that changes the OCR breach audit conversation, because now you have a tech record of what was blocked, not just a vendor's word that the model behaved.
The "Agent Handler" framing Merge is using points at the same gap. But the question for any of these tools is whether enforcement is out-of-process or just another layer the agent itself can influence.
https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061474831976304996&utm_campaign=nemoclaw-and-the-healthcare-agent
The context window pricing model might be one of the smartest tricks AI labs pulled.
You pay for larger context windows, yet model performance often degrades as more context is stuffed in.
So developers end up paying more just to work around the fact that models don't have true
The degradation problem is real, but the clinical version of this is worse than developers realize. You're not just paying more for worse retrieval, you're paying more for the specific reasoning chains that make diagnostic AI worth using at all. Long-context reasoning queries run about 13x more energy than a standard ChatGPT call, that cost scales quadratically with context length by design of the transformer architecture.
So the pricing model you're describing hits healthcare with a second layer: the queries that achieve 80-85% diagnostic accuracy on NEJM benchmark cases are exactly the ones consuming the most tokens. Microsoft's sequential diagnosis work got there by substituting compute spending for test spending, which is a fine trade in a research setting and an unsustainable one in a fee-schedule world where CMS has no mechanism to reimburse "tokens used to reach diagnosis."
FDA can clear the model, CMS can't price the inference, nobody owns the gap between those two things.
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2061456225854947477&utm_campaign=token-economics-versus-the-20-watt-995
a $2,000 graphics card and a free download are turning into $2,000-a-month retainers, and the customers are the ones who legally can't use ChatGPT
the hardware is an RTX 5090, 32GB of VRAM, enough to run a 30B model locally with nothing ever leaving the building. the software is https://t.co/1UQICDQx9g
That's the exact gap I wrote about: it's not the model that unlocks clinical deployment, it's whether your compliance officer can point to something technical when OCR comes knocking. What I found is that the missing piece isn't local inference alone, it's the governance layer sitting outside the agent process so a hallucinating model can't route PHI to the cloud through its own judgment. DGX Spark at sub-$3k plus NemoClaw's open-source policy engine is the first time both pieces exist together at a price point a community hospital can actually approve. https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061913346610430300&utm_campaign=nemoclaw-and-the-healthcare-agent
Free newsletter: The dawn of token-based billing has shown that generative AI doesn’t have a return on investment. It's too unpredictable, too unreliable, you can't easily measure the cost of tasks, and organizations are already pulling back.
https://t.co/wmI82zWdcq
What happens when the tasks that actually generate ROI are precisely the ones that cost the most to run?
Token-based billing exposed something healthcare was going to hit regardless: the economics get worse exactly when the output gets better. The Microsoft sequential diagnosis work I dug into (80-85.5% accuracy on NEJM benchmark cases, against roughly 20% for unaided generalists) achieved that performance by substituting compute spending for diagnostic test spending. The model reasons longer, burns more tokens, costs more per query. That's not a bug in the ROI calculation, that's the whole structure of it.
Long-chain reasoning queries consume around fifteen times more tokens and cost approximately thirteen times more energy than a standard query. So the "unpredictable cost" problem other sectors are retreating from is, in clinical AI, the most clinically valuable mode of operation.
The deeper problem is institutional, not economic. CMS has no framework for covering inference-heavy diagnostic AI. FDA can clear a device for safety. Neither agency adjudicates whether a more accurate answer is worth its token cost (and no agency currently does). So even if a health system believed the ROI was there, no payment pathway prices it correctly.
Organizations pulling back from token billing in enterprise software are responding to cost unpredictability. Healthcare can't even get to that problem yet. It's stuck upstream, with no mechanism to decide whether the cost is justified at all.
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2061820257333813659&utm_campaign=token-economics-versus-the-20-watt-995
Humanoid robots, designed to mimic human movement and capabilities, have been making headlines in recent years, from their use as baggage handlers at Japanese airports to Tesla’s big bet on its Optimus humanoid.
Market watchers have predicted that the machines will change the https://t.co/vWPDEgGYQe
Worked with a health system in the midwest last year that had 14 open EVS positions for four months straight. They weren't being picky. The pipeline was empty.
That gap, multiplied across a 6,000-employee hospital, is where humanoid robots actually land first. Not in the OR. Not at the bedside. In the hallways, moving linen carts and waste at 2am when no one is applying for that shift anymore.
The airport baggage handler framing is honest about what these machines can do today. What it misses is where the demand is sharpest. Tesla's Optimus and the airport demos get attention because they're visible. The quieter story is that hospital environmental services and transport together are 10-15% of FTEs at most large systems, and those roles are running vacancy rates that recruiting cannot fix at any price point.
Software agents get almost all the capital right now. But they only reach the 20-25% of hospital workers who sit at a desk. The other 75-80% move through physical space, and no AI agent touches that problem.
The structural nursing shortage, projected at 450,000 RNs by mid-decade, is demographic. It compounds. And when margins are already 1-3% at most nonprofits, the cost of leaving physical roles vacant is not abstract.
The real question is whether the systems that deploy logistics robots now, before the unit economics are clean, end up with a structural advantage, or whether they're just early and absorbing the cost of being early. I'm not sure the answer is settled yet.
https://www.onhealthcare.tech/p/the-labor-problem-healthcare-wont?utm_source=x&utm_medium=reply&utm_content=2062153166871818254&utm_campaign=the-labor-problem-healthcare-wont
AI made building cheap.
Regulation made distribution expensive.
Anyone can clone your demo now.
Very few can bring proprietary data, licenses, compliance, and enough trust to sell into serious players.
The moat didn’t disappear.
It moved.
The moat moved, but it moved unevenly across the market (which is the part nobody's really pricing in yet). Large national payers with existing engineering capacity are probably insourcing prior auth and UM in the next two years, not buying. The smaller regional plans still buy, because they can't staff the build. So the same cost compression that kills one vendor's pitch actually preserves another's, depending entirely on who's sitting across the table.
https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2061940709352226832&utm_campaign=the-free-lunch-is-over-except-now
I called a woman in Dayton to tell her she was about to overpay $4,500, for a horrible health system based procedure.
She was scheduled for a knee replacement.
Her PPO had her on the hook for a $4,500 deductible.
The call was on behalf of her employer: “See one of these three surgeons at these two facilities instead, and the $4,500 disappears. You pay nothing.”
Same knee.
Way better facility.
Better surgeons.
No out of pocket.
Direct contracting costs her less than the one her insurance had her walking into. The cheaper, better path existed the whole time.
Imagine your car insurance charging you $4,500 to use a worse mechanic, while the better one across town was free.
The healthcare system is not broken amigos.
It was designed this way…
The part that gets skipped in this story: how did her PPO know where she was going in the first place?
And the answer is usually claims data, not any directory. Because the directory, the official one, the federal one CMS just dropped with 27.2 million records, can't tell you if that surgeon is accepting new patients. Can't tell you hours. Can't verify the surgeon is who they say they are. 0% of providers in it have been checked against NIST identity standards. Zero.
But here's the sharper problem. 71% of practitioners in that directory have no link to any organization at all. They're just floating. Which means if you tried to build the tool that made her call unnecessary, the one that routes patients before the bad choice gets made, you'd be starting from a skeleton with most of the bones disconnected.
The $4,500 savings existed because someone with claims data and direct contracts already did the work the federal infrastructure was supposed to do. And the incentive to sell that work back to employers only exists because the public layer is hollow.
That's not a bug. The gap is the business model, for a lot of players who'd rather keep routing patients to worse, pricier options.
I went through the whole dataset to map exactly where the gaps are and what it would take to fill them: https://www.onhealthcare.tech/p/the-cms-national-provider-directory?utm_source=x&utm_medium=reply&utm_content=2061115508963876970&utm_campaign=the-cms-national-provider-directory
Your Oura Ring, your health record. Together. Finally.
@Flexpa is bringing clinical records into @ouraring via TEFCA, so ŌURA’s AI isn’t just working from what your ring measures, but from your real health history.
https://t.co/b1KCeGEZLO https://t.co/u2t1TQjL7p
Athena's TEFCA connection already covers 100,000+ providers, and the reason that number matters here is what happens after the data arrives.
TEFCA gets the clinical record into the wearable context. But without a protocol layer that lets an AI agent query that record on demand, you still have a static data dump. The gap I keep coming back to is the M×N problem: every new data source requires a new custom connector unless something like MCP sits in the middle and collapses that cost.
What Flexpa and Oura are building toward is genuinely useful, and the clinical decision support angle gets more real when the AI can pull from a live FHIR feed rather than a one-time import. The question is whether the AI layer can act on that data in a way that is scoped, audited, and covered by a BAA. That is where most of these demos quietly stop short.
https://www.onhealthcare.tech/p/the-usb-c-port-for-healthcare-ai?utm_source=x&utm_medium=reply&utm_content=2061480024255930497&utm_campaign=the-usb-c-port-for-healthcare-ai
@RebeccaTorrenc5·18,877 views80%
6/3/26 7:38 PM ET
Scoop! Lila Sciences is in talks to raise ~$2 billion in new funding.
The raise would value the “scientific superintelligence” lab at ~$8.5 billion before the new money.
w/ @MichelleF_Davis, read more👇
The Iso comp here is worth watching closely: covered at https://www.onhealthcare.tech/p/isomorphic-labs-pulls-21b-series-6c0?utm_source=x&utm_medium=reply&utm_content=2062210478840086850&utm_campaign=isomorphic-labs-pulls-21b-series-6c0?utm_source=x&utm_medium=reply&utm_content=2062210478840086850&utm_campaign=isomorphic-labs-labs-pulls-21b-series-6c0 how $15-20B for Iso only makes sense if you price it against frontier AI labs, not biotech. Lila at $8.5B pre on zero clinical data is the same bet, just earlier on the curve.
The next frontier of AI-secured financial and technological critical infrastructure.
@ICE_Markets and @NYSE are part of @AnthropicAI's cyber security initiative Project Glasswing, deploying Anthropic’s Claude Mythos Preview across ICE’s exchanges, clearing houses, mortgage
The financial sector's inclusion in Project Glasswing makes complete sense given systemic contagion risk, but the coalition map reveals a gap that should be generating much louder alarm than it currently is.
Healthcare is entirely absent. No health system, no EHR vendor, no payer. This matters because healthcare absorbed 31% of all disclosed ransomware attacks in early 2026, a sector running on legacy devices that can't be patched and that rely almost entirely on network segmentation as their primary defensive compensating control. The problem I lay out at https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2062151952809492599&utm_campaign=how-claude-mythos-preview-found-thousands is that Mythos Preview's autonomous zero-day discovery operates at machine speed, which structurally collapses the segmentation frameworks those devices depend on. IEC 62443 was designed around human-speed threat actors. Mythos is not that.
ICE and NYSE being inside Glasswing means they get controlled access to Mythos-class offensive capability to harden their defenses before adversaries get the same tools. Anthropic's own red team puts that adversary access window at 6 to 18 months. Healthcare providers get none of that preparation time, no institutional pathway, no equivalent access.
When a hospital network goes down, patients get rerouted. Procedures get delayed. People die. The financial sector's inclusion is correct. Healthcare's exclusion is a policy failure dressed up as an oversight.
the ability of ECG AI to predict LVEF is a big deal
previously, ECG AI has been shown to detect MI similarly to true ECG experts (& better than most of us)
AI is now doing something unique that humans essentially can't do
this will soon be the the standard of care ...#1/2
Predicting LVEF from a 12-lead ECG is genuinely impressive signal extraction. But the "humans can't do this" framing is where I'd slow down.
The harder problem isn't whether the model works in the validation cohort. It's what happens when the PPV hits a real clinical population. I spent time recently on REDMOD, a radiomics pipeline that detects pre-neoplastic pancreatic tissue changes on abdominal CTs already read as normal by radiologists. The retrospective numbers looked strong: 73% sensitivity, 88% specificity, 16-month median lead time. The viral headline wrote itself.
But when you run the Bayesian math at average-risk prevalence, that 88% specificity produces roughly 0.18% positive predictive value. About one true positive per 555 flagged patients.
ECG AI for LVEF faces a version of the same question. The model's operating characteristics in a cardiology referral population may not survive contact with a primary care population where structural dysfunction prevalence is meaningfully lower. Specificity that looks adequate in an enriched cohort can collapse into a false-positive flood when the denominator changes.
"Standard of care" is the right destination. The path there runs through prospective prevalence-matched validation, not retrospective AUC. That gap between published performance and real-world PPV is where most of these tools get humbled before they get adopted.
Wrote through this exact dynamic in a different disease context if you want the framework:
https://www.onhealthcare.tech/p/the-preclinical-signal-in-routine?utm_source=x&utm_medium=reply&utm_content=2062202547289419866&utm_campaign=the-preclinical-signal-in-routine
any company rolling out AI at scale is running into the same question.
what can the agent reach and what can it touch.
security is what’s slowing down mass adoption.
...and in healthcare that question gets a lot more specific, because "what can it touch" isn't just a security posture question, it's a regulatory one with breach investigation consequences attached.
The gap I kept running into when researching this is that most health systems are trying to answer the security question with behavioral controls, system prompts, internal classifiers, instructions telling the agent not to do certain things. That's in-process enforcement, meaning a hallucinating or compromised agent can simply override it. Compliance officers know this, which is why autonomous agent deployments keep stalling at the pilot stage even when the model performance is genuinely good.
The architectural distinction that actually matters here is whether guardrails live inside or outside the agent's own process space. If the policy enforcement is external, a misbehaving agent cannot reach past it, the same way a browser tab can't escape its sandbox. That's what makes the difference between a vendor promise and something an auditor can look at.
(The 42 CFR Part 2 angle makes this even more constrained, because substance use disorder records carry stricter re-disclosure rules than standard HIPAA, and no behavioral instruction survives that level of regulatory scrutiny as a documented technical safeguard.)
Security is slowing adoption, yes, but the specific mechanism is that compliance officers have no documentable technical basis to approve persistent agent access to live EHR data. The question I keep turning over is whether the open-source governance layer changes that calculus fast enough to matter before health systems default back to narrow, non-autonomous automation.
https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061539809408004345&utm_campaign=nemoclaw-and-the-healthcare-agent
⚡️AI is turning cybersecurity from a human-limited profession into a machine-speed arms race.
The surface read is “Anthropic’s model found vulnerabilities fast.”
The deeper read is that the bottleneck in cyber is shifting.
For years, vulnerability discovery was constrained
The part that keeps me focused on healthcare specifically: IEC 62443 network segmentation has been the primary compensating control holding legacy medical devices together from a security standpoint. Infusion pumps, patient monitors, imaging systems running decade-old firmware that will never be patched. The entire defensive logic depends on human-speed attack assumptions. An attacker has to manually probe, enumerate, work laterally. Segmentation buys time.
Mythos Preview produced working exploits 181 times on Firefox 147 JavaScript engine benchmarks. Opus 4.6 was near zero on the same benchmarks. That gap is not incremental. When zero-day discovery becomes automated at that speed, the time segmentation buys collapses to near nothing, and the compensating control the FDA and most health system security teams are counting on stops compensating.
The arms race framing is right, but healthcare is running it with one specific structural disadvantage the other sectors don't share. Every major technology company has an institutional pathway into Project Glasswing's defensive coalition. AWS, Google, Microsoft, CrowdStrike, Palo Alto, all there. No health system. No EHR vendor. No payer. The sector taking 22% of all disclosed ransomware attacks in 2025, rising to 31% in early 2026, has no coordinated access to the defensive capabilities being built around the most powerful offensive security AI ever deployed.
Anthropic's own red team puts adversary access to Mythos-class capability at 6 to 18 months out. That is the window. Healthcare has no institutional position inside it.
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2061718842993492398&utm_campaign=how-claude-mythos-preview-found-thousands
the CEO of NVIDIA just said the computer is no longer being built for you, it is being built for agents
Jensen Huang put it plainly, until now we were the users, we were the renters, every CPU on earth was designed around how a person works
but agents do not work like us, they https://t.co/k3egQtWvq8
The infrastructure shift Jensen is describing is real, but in healthcare it creates a specific pricing problem nobody has solved yet.
When you design compute for agents running continuously at scale, the cost model flips from capital expenditure to pure opex, billed per token, per query, per reasoning step. That's fine for enterprise software. It's structurally incompatible with how medicine pays for cognitive work.
No line item in a hospital budget covers marginal inference cost per patient interaction. CMS doesn't reimburse it. FDA doesn't evaluate whether the reasoning quality justifies the compute spend. The agent architecture Jensen is describing assumes someone downstream has figured out the payment layer. In clinical settings, nobody has.
The sharpest version of the problem: Microsoft's sequential diagnosis research hit 80 to 85 percent accuracy on NEJM benchmark cases, against roughly 20 percent for unaided generalist physicians. That performance came from running long chain-of-thought reasoning, which consumes roughly 15 times more tokens than a standard query. The capability is real. The per-patient cost of delivering it at scale is also real, and scales with every interaction.
Better compute for agents just means the gap between what clinical AI can do and what healthcare economics can absorb gets wider faster. Jensen is describing the engine. The billing infrastructure doesn't exist yet.
https://www.onhealthcare.tech/p/token-economics-versus-the-20-watt-995?utm_source=x&utm_medium=reply&utm_content=2061772100403184122&utm_campaign=token-economics-versus-the-20-watt-995
We are partnering with @Microsoft to enable secure, user-controlled AI on Windows.
NVIDIA OpenShell runtime for agents will provide governance tools, policy enforcement, and smart local-to-cloud query routing.
Learn more: https://t.co/zPGwz9xQSW https://t.co/mkOtFEOhAS
The piece I spent weeks on maps directly onto this announcement. When I was working through how OpenShell's out-of-process policy enforcement actually functions for clinical environments, the Windows integration was the missing piece I kept circling back to, because hospital IT infrastructure is overwhelmingly Windows-native and any governance layer that requires departing from that stack was never going to clear procurement.
The local-to-cloud routing is where this gets concrete for healthcare. I've been tracking how compliance officers at mid-size health systems can't approve cloud routing of PHI without a documented, auditable policy decision, not an agent making a judgment call in the moment. OpenShell's privacy router does that programmatically, which is a different category of compliance defense than a system prompt telling an agent to be careful with sensitive data. A hallucinating agent can't override a constraint that lives outside its own process space, that's the architectural point most coverage keeps missing.
The Windows partnership also changes the community hospital calculus significantly. DGX Spark under $3,000 plus an open-source governance layer on familiar infrastructure is a very different procurement conversation than what health systems have been facing, I wrote about this specifically at https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061961069212377313&utm_campaign=nemoclaw-and-the-healthcare-agent because the sub-enterprise segment has been priced out of viable agent deployment for exactly this reason. The Microsoft distribution channel doesn't hurt either.
From unboxing to AI agent in minutes.
Getting an agent running used to mean sourcing a model, configuring an inference backend, installing a runtime, and wiring it all together. The new NemoClaw install path on DGX Spark replaces that with a single command.
DGX Spark also https://t.co/i8SUNOisVr
The single-command install is real progress, but in healthcare the harder problem starts after the agent is running. What compliance officers actually need, as I wrote at https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061915769135350120&utm_campaign=nemoclaw-and-the-healthcare-agent, is proof that a live agent with EHR access and shell access cannot route PHI to the wrong place even if it hallucinates or gets a bad prompt. That proof has to exist outside the agent process itself, not as a system prompt the agent could in theory ignore.
The DGX Spark price point matters a lot here too. Sub-$3,000 on-prem inference means a rural hospital can keep sensitive inference local by default, which is a different compliance posture than "we pinky-swear the cloud vendor signed a BAA." But does single-command deployment also wire up the policy engine and the privacy router, or does that layer still require separate config?
Most enterprise AI projects stall because no one has prepared the content, built the workflow, and kept it improving as models change.
That role is essential. Box has launched Forward Deployed Engineers to fill it. Model-agnostic, content-first, and built for the way enterprise https://t.co/xf9VTbqsk0
The framing is right but the hard part gets glossed over in "built the workflow."
In healthcare specifically, two health systems running the exact same EHR can have completely divergent clinical data models, custom build types, local formularies, legacy migration artifacts. There's no generic workflow to build. Someone has to go sit inside the organization for weeks or months and document what's actually happening before a single agent can be reliably deployed.
And that's where most enterprise AI efforts fall apart. Not model capability. Not content strategy. The undocumented, inconsistent, deeply human process layer that nobody has written down because nobody needed to until now.
Box's FDE move makes sense directionally. But the economics get complicated fast when the customization burden is that deep. Wrote about exactly this dynamic in healthcare: https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2061840567315943891&utm_campaign=the-standardization-trap-why-deploying
In the AI gold rush, Jensen Huang is selling the picks and shovels. Nvidia's chips power AI companies around the world, helping the company become the first to surpass a $5 trillion market capitalization in late 2025.
Read more about how he and others made the https://t.co/lHLEv4NVhh
The picks-and-shovels framing is right, but in healthcare specifically it understates what's actually happening, because the "picks and shovels" aren't just chips anymore. I spent time mapping NVIDIA's full healthcare stack (https://www.onhealthcare.tech/p/nvidias-healthcare-stack-is-the-picks?utm_source=x&utm_medium=reply&utm_content=2061961137378431098&utm_campaign=nvidias-healthcare-stack-is-the-picks) and the more interesting story is that BioNeMo, MONAI, Holoscan, and the rest have quietly made NVIDIA the dominant software infrastructure layer across drug discovery, medical imaging, and surgical robotics simultaneously.
The consequence that doesn't get discussed enough: when 82% of healthcare AI builders say open-source models are moderately to extremely important, and NVIDIA has built its ecosystem lock-in through exactly those open frameworks rather than proprietary licensing, the moat compounds in a way that pure hardware supply never could. A GPU competitor can undercut on price. Displacing the framework a thousand hospital IT governance committees have already approved (MONAI has 6.5 million downloads and citations in over 4,000 peer-reviewed papers, which is basically a clinical credentialing shortcut) is a different problem entirely.
The part I keep returning to is the edge inference layer. Holoscan exists because cloud round-trip latency is clinically unacceptable in an operating room, which means real-time intraoperative AI isn't a software preference question, it's a physics constraint. And if the edge layer becomes mandatory infrastructure for surgical AI, what does that mean for how Intuitive Surgical's hardware lock-in model holds up when competitive platforms can now build on open...
💡 CPUs are no longer just host processors. They're on the critical path for latency, accelerator utilization, and tokens per dollar.
NVIDIA Vera is purpose-built for this. High per-core performance, high concurrency, efficient memory bandwidth — and over 1.8x higher agentic https://t.co/PkAbGmC18R
The CPU-as-bottleneck framing is right, but it stops before the more uncomfortable implication.
If CPUs are now on the critical path for tokens per dollar, then the unit economics of clinical AI inference are more fragile than most health tech companies have modeled. Their financial projections assume GPU cost is the variable to watch. But if you're bottlenecked on host processor throughput in agentic workloads, a better GPU won't fix the economics of running continuous real-time decision support across a patient population.
Vera is interesting precisely because it was designed around agentic concurrency. That matters. But the deeper question for health AI is whether any of this hardware progress actually moves the needle on the clinical applications that are currently uneconomical, or whether it just improves margins on ambient documentation workloads that already pencil out fine.
The workloads that remain blocked aren't the easy ones. Genomic variant interpretation pipelines, population-scale deterioration models, multimodal inference combining imaging with lab data, those are the applications where inference cost is genuinely the binding constraint. The question is whether CPU-accelerator co-optimization at the Vera level moves the cost curve enough to unlock them, or whether that requires a more structural shift in compute supply.
I've been thinking about this from a different angle, specifically what a 50x increase in global compute output does to clinical AI unit economics across the board, including which health tech moats survive when inference gets cheap. https://www.onhealthcare.tech/p/the-elon-terrawatt-announcement-nobody?utm_source=x&utm_medium=reply&utm_content=2061553379453604010&utm_campaign=the-elon-terrawatt-announcement-nobody
Another humanoid worker has just joined the factory floor
Two years later,PUDU Robotics has launched the new-generation PUDU D7, built specifically for real manufacturing scenarios.
It can autonomously push carts, handle intra-line transportation, perform delicate operations, https://t.co/0xcuBr2h8v
The cart-pushing and intra-line transport use case is actually where I'd look first, before the delicate operations piece, because that's where the unit economics close fastest.
When I was mapping hospital labor by BLS category, environmental services and transport sit at 10-15% of FTEs at most large systems. That's tens of millions in annual wage spend at a single academic center, doing exactly what the D7 is demoing: moving things between points A and B on a predictable floor plan. Aethon and Moxi have been chipping at this for years in healthcare and showing 30-60% reductions in staff time on specific transport loops.
The reason I keep coming back to manufacturing launches like this one is that hospitals are watching. The floor plan problem is actually similar: semi-structured space, mixed human traffic, time-sensitive payloads. A robot that proves out cart logistics in a factory gives a procurement team at a health system something to point to when the CFO asks why they're buying a $200K robot instead of an agency nurse contract.
Software alone won't close that gap. That's the part the VC community keeps skipping over, and I laid out why at some length here: https://www.onhealthcare.tech/p/the-labor-problem-healthcare-wont?utm_source=x&utm_medium=reply&utm_content=2062040465462235231&utm_campaign=the-labor-problem-healthcare-wont
The question I keep sitting with is whether the "delicate operations" claim on these new-gen robots is real enough to matter in regulated care settings, or whether that's three product cycles away from where hospitals would actually trust it...
#ASCO26 | Day 3
$LLY didn't spend billions on Kelonia for an 18-patient #myeloma dataset. They may have paid for a potential manufacturing disruption.
KLN-1010 reported:
• 18/18 MRD-negative at 1 month
• No lymphodepletion
• No ex vivo cell manufacturing
• Single infusion
Buying Kelonia wasn't really about the myeloma data, and I'd push back slightly on framing the manufacturing angle as the headline thesis. The deeper disruption in my read at https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2061103943049003170&utm_campaign=gene-editing-has-the-science-figured-b80 is that even if you solve ex vivo manufacturing, you're still left with conditioning regimens, transplant center capacity, and payer infrastructure that weren't built for one-time curative workflows. KLN-1010 skipping lymphodepletion is genuinely interesting, but the bottleneck that keeps approved therapies from reaching eligible patients isn't the clean room, it's the reimbursement mechanic and the operational coordination stack around it.
@michael_hoerger·16,619 views83%
6/2/26 10:28 AM ET
@WIRED As a clinical health psychologist who has written >20 papers on COVID, I would emphasize 4 facts:
1) Long COVID is not a psychological diagnosis nor manifestation of a psychological condition
2) Billions of dollars need to be invested in biomedical treatments and preventives, and that money is not being invested because of wealthy short-term interests, which prop up various narratives, including in the media
3) Behavioral interventions can help with infection/reinfection prevention (e.g., COVI-CAN pilot) and stress/coping support (gaslighting/ostracism as huge issues), but these are not cures, and the same interventions are relevant to people with cancer, organ failure, immunocompromising conditions, etc.
4) Many psychological/behavioral "treatments" for Long COVID are directly harmful to patients and are indirectly harmful to society by incorrectly framing the issues
I would consider these issues obvious in summer 2020.
Articles like this should not be written in 2026, but it is a consequences of cultural evolution, or organizational selection by consequences. The organizations that write puff pieces propping up pseudoscience get the gold, while truth tellers do not. It would be useful to examine the organizational practices at WIRED that led to the incentive systems that allowed this piece to manifest.
The incentive structure point lands hard. But what's the mechanism that makes it so sticky?
What I've been tracking in consumer health AI is a version of the same dynamic, where the systems with no direct financial stake still end up recommending action over reassurance at a striking rate. Early studies from healthcare economists found systematic bias toward six to eight supplement or follow-up recommendations per routine lab upload, because the algorithm optimizes for comprehensive, actionable output rather than clinical parsimony. The financial incentive is gone but the utilization pressure isn't.
That's the part that complicates the organizational selection story slightly. It's not just that publications get rewarded for puff pieces. The tools patients are increasingly turning to, partly because they've lost trust in media and institutions, are baking the same bias in at the architecture level. Watchful waiting doesn't generate engagement. Reassurance doesn't feel like a deliverable.
So the organizations you're describing and the AI platforms I've been writing about are producing the same downstream harm through completely different incentive pathways. And the regulatory gap that allows one probably tells you something about why the other persists too.
https://www.onhealthcare.tech/p/the-double-edged-algorithm-how-consumer?utm_source=x&utm_medium=reply&utm_content=2061527444142903560&utm_campaign=the-double-edged-algorithm-how-consumer
"Many patients with long COVID are already receiving care but are not being recognized as having the condition. These patients are not absent from clinical care; they are absent from the diagnostic code that would identify them as long COVID patients"
https://t.co/aLJCOPlQt1
The recognition failure here runs deeper than documentation habits. What the Boston ED data showed was that o1 hit roughly 67% diagnostic accuracy at triage on sparse information, versus 50-55% for attendings, and the gap was largest exactly where clinical pattern recognition gets murky and underdefined, which is precisely the profile of a post-viral syndrome with no clean biomarker.
The long COVID coding gap is downstream of a diagnostic confidence problem. Physicians don't code what they haven't committed to, and they don't commit to diagnoses that feel ambiguous. That's where I think the infrastructure argument matters more than people realize: https://www.onhealthcare.tech/p/what-the-harvard-er-study-says-about?utm_source=x&utm_medium=reply&utm_content=2061484759507685776&utm_campaign=what-the-harvard-er-study-says-about
Once differential generation approaches zero marginal cost at the front door, the bottleneck shifts to whether the diagnosis ever makes it into the order entry flow and the billing code. Which means the long COVID recognition problem isn't really a physician awareness problem at this point. It's a workflow and EHR integration problem.
So who owns that layer when the diagnostic model is commodity infrastructure?
Published in @JCO_ASCO, during #ASCO26:
Tumor-Agnostic Therapies: Translating Scientific Breakthroughs Into Global Implementation
ASCO is where we celebrate the next breakthrough.
This review asks why breakthroughs do not equal access globally.
🔗: https://t.co/PE5q27bIs1 https://t.co/tXkDw65apo
The gap they're describing in tumor-agnostic access is the same structural problem I've been watching play out in gene editing, and the root cause is identical: the healthcare operating system was built around disease-specific pathways, reimbursement codes, and treatment center designations that don't map onto biomarker-defined or mechanism-defined therapies.
CASGEVY is the clearest case study right now. Approved, functional, ~60,000 eligible patients across approved geographies, and Q1 2026 revenue that implies a fraction of that population is actually getting treated. The science worked. The delivery infrastructure (reimbursement mechanics, Medicaid contracting, transplant center capacity) never got built to match.
Tumor-agnostic therapies hit the same wall from a different angle. The biomarker testing infrastructure, the payer willingness to reimburse across indication lines, the coverage policy logic that still expects a primary diagnosis code before authorizing treatment, none of that was designed for a therapy that works on molecular identity rather than organ of origin.
Which is why the investment thesis I'd push here isn't about the next platform. It's about who's building the NGS lab infrastructure, the outcomes registry systems, and the contracting models that can actually operationalize these approvals at scale.
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2061548190126743771&utm_campaign=gene-editing-has-the-science-figured-b80
A rapid manufacturing pipeline for BCMA-targeting CAR T cells drastically shortens “vein-to-vein” times, and produces cells that demonstrate encouraging safety results in a phase one trial of patients with relapsed #MultipleMyeloma. @ScienceTM https://t.co/3F4u8ECFDS https://t.co/9CCyA0Lbr2
500 patients initiated on CASGEVY globally against 60,000 eligible patients in approved geographies, and faster manufacturing alone doesn't close that gap.
The vein-to-vein problem in CAR-T is real, but the binding constraint for most patients isn't time in the lab. It's whether a transplant center has the staff, the bed capacity, and a payer willing to cover a six-figure cell therapy before the patient gets to any manufacturing queue at all. Cutting days off production while Medicaid benefit design still can't handle a one-time curative price doesn't move patients through the system faster. It moves them to a different waiting room.
The same pattern I traced in gene editing applies here. Science clears one gate and we call it a win, but the patient still has to clear five more gates that no one built the infrastructure to manage.
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2061643999799079165&utm_campaign=gene-editing-has-the-science-figured-b80
Almost everyone is building agent harness systems the wrong way.
The default move: pick LangChain or LangGraph or the OpenAI Agents SDK, accept the loop, the tools, the memory, the orchestration, the policy engine, the credential store, the budget tracker, all of it, as one decision.
Mike, wrote a long piece today on why this shape is wrong, and why every long-running agent team eventually ends up rewriting its harness from scratch.
His argument: a harness isn't one thing. It's fifteen separate concerns bundled together because the surrounding ecosystem didn't give you a way to compose them. Turn state machines, provider routing, credential vaults, policy engines, approval gates, budget trackers, hook fanout, context compaction, session trees, OpenTelemetry tracing. Frameworks ship them as one block because that was the only shape available a year ago.
It isn't anymore.
When every layer is a worker on a shared bus with a typed function contract, "build your own harness" stops meaning "fork a framework." It means swap a worker. Don't like the model catalogue? Write one that hits a live API. Don't like file-backed credentials? Plug in your secrets manager. Want approvals routed through Slack instead of a console? Add a worker that calls approval::resolve. The rest of the stack does not change.
The framework era picked a position for you and locked you in. The worker model leaves the choice in your hand.
Worth reading in full.
The worker model argument is correct, and it maps cleanly onto something worth extending: the problem gets considerably harder when the "swap a worker" move has to happen inside a health system's environment.
You can modularize your harness perfectly, every layer a clean typed contract on a shared bus, and still hit a wall when the credential store needs to authenticate against a payer portal that has no API, or when the approval gate has to route through a SharePoint queue because that's where the clinical ops team actually lives. The modularity is real, it just can't save you from the fact that the thing on the other end of your worker is undocumented and inconsistent.
That's the part the framework debate tends to skip. The technical composition problem is largely solved, the ecosystem caught up. But in healthcare specifically, roughly 30-40% of what an agent actually needs to do is buried in org-specific workflow logic that no typed contract can describe until someone has spent weeks on-site watching how the work actually moves. Two health systems running the same EHR can have completely divergent data models, local formularies, and legacy migration artifacts, and none of that shows up in a FHIR endpoint.
So the worker model wins the architecture argument, that part's done. The open question is who does the embed to figure out what the workers need to actually do in a given org, and whether companies building agents are charging enough to fund that work or hiding it to keep their SaaS metrics clean.
Wrote about this at length here: https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2060072412868235587&utm_campaign=the-standardization-trap-why-deploying
Most chronic diseases, and many cancers, are still diagnosed far too late. The future of healthcare lies in detecting risk earlier, diagnosing smarter, and intervening sooner.
From AI enabled diagnostics and genomics to Apollo’s DigiSmart Labs, technology is helping transform https://t.co/FuIZMQ1vCQ
The Apollo DigiSmart example is a useful anchor here, because it illustrates exactly where the early detection argument gets complicated in practice. The diagnostic layer is advancing fast. But in the work I've been tracking, the harder bottleneck isn't detection sensitivity, it's what happens to the data after the test runs.
And that's where most early detection programs quietly stall. You can sequence a genome for under a thousand dollars today, layer in continuous biometric monitoring, add environmental exposure data, and still produce nothing clinically actionable if the outputs can't talk to each other or translate into something a primary care physician can actually use in a twelve-minute appointment. The EHR interoperability gap is swallowing a lot of genuinely promising early detection work before it reaches patients.
The shift from reactive treatment to pre-symptomatic intervention also has a payer problem that rarely gets named directly. Insurance models are structurally built to reimburse treatment, not prevention investment, and that misalignment creates an organizational barrier that sits upstream of the technology entirely.
I've been writing about this convergence of multi-omics data, digital twin simulation, and continuous monitoring as a genuine paradigm shift, closer in significance to antibiotics than to incremental diagnostic improvement. But the clinical scaling question keeps coming back to workflow integration and payment reform, not just detection capability.
More on the architecture and the market conditions required to make it work: https://www.onhealthcare.tech/p/the-pre-cure-revolution-how-ai-powered?utm_source=x&utm_medium=reply&utm_content=2061287244867203419&utm_campaign=the-pre-cure-revolution-how-ai-powered
Red Hat and @NVIDIA are integrating NVIDIA OpenShell into the full-stack @RedHat_AI platform.
The work brings oversight and policy to the infrastructure level, while contributing to the open source OpenShell project to standardize how agents are governed on enterprise platforms. https://t.co/qq8gRsTPWV
The Red Hat integration is the enterprise distribution layer NemoClaw needed. But the healthcare-specific piece (what compliance officers actually need to show OCR auditors) is that these guardrails live outside the agent process, so a hallucinating agent with live EHR credentials can't override them. That architectural distinction is what I found moves this from "interesting AI governance" to "defensible HIPAA deployment" for health systems. https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2061274057035669555&utm_campaign=nemoclaw-and-the-healthcare-agent
Interoperability is where AI hype goes to die.
If your “agent” can’t work across data, logic, actions, security of legacy systems, it’s not transforming the enterprise. It’s just another app asking for an export.
Winners don’t rip and replace. They operate.
Precisely the problem MCP is designed to solve. The M×N integration nightmare, where every AI agent needs a bespoke connector to every EHR, is what's killed dozens of promising clinical AI tools before they reached scale. athenahealth's August 2025 MCP server pilot on athenaOne is the first production signal that a standardized protocol layer can bridge that gap without forcing a rip-and-replace.
But the security layer is where it gets genuinely hard. When an MCP server sits between a clinician's AI agent and a FHIR endpoint, nobody's clearly answered who holds the BAA liability. And the confused deputy problem, where an agent inherits access privileges no individual user was ever supposed to have, isn't a compliance checkbox issue. It's a structural gap that'll surface badly if it's not engineered against from the start.
The winners you're describing won't just operate across legacy systems. They'll be the ones who got the compliance architecture right early, because a BAA-covered, HIPAA-audited integration is expensive to rip out. That's the moat, not the protocol itself. https://www.onhealthcare.tech/p/the-usb-c-port-for-healthcare-ai?utm_source=x&utm_medium=reply&utm_content=2061410395055866267&utm_campaign=the-usb-c-port-for-healthcare-ai
I think the AI superapps will soon own 90% of the agentic layer
more and more people won't use hermes/openclaw etc because Claude Cowork/Codex will offer 90% of the functionality with 5% of the friction
What happens to the health tech startups that built their entire moat around being the "AI layer" between EHR data and clinical workflows?
That superapp consolidation dynamic is exactly what I've been tracking at the health system level, where Epic's Agent Factory (a no-code drag-and-drop agentic AI builder announced at HIMSS26) is doing to digital health middleware what Claude Codex is doing to standalone coding tools. The friction argument is identical: why evaluate, contract, and integrate a third-party ambient documentation vendor when your EHR ships one natively? Health systems are already signaling a "quiet stall," delaying vendor evaluations because they're waiting to see what Epic builds in-house (and Epic's R&D runs at roughly 50% of operating expenses, so the shipping cadence is real).
The part your framing surfaces that I think matters most is the 90/5 ratio. The startups getting compressed aren't bad products. They're just doing the 90% of functionality that the platform will absorb, and "better integration" stops being a differentiator the moment the platform owns the integration layer by default. Oracle Health lost a net 74 hospitals in 2024 and reportedly stopped sharing its contract list with KLAS Research, which tells you what competitive consolidation looks like when it's already underway rather than hypothetical.
The companies that survive this, in health tech at least, are the ones doing the remaining 10%: narrow specialty clinical decision support, proprietary datasets Epic doesn't hold, or infrastructure positioned toward payers and life sciences rather than health systems. That's not a comfortable place for a lot of 2022-2023 vintage seed rounds.
https://www.onhealthcare.tech/p/epics-agent-factory-and-the-end-of?utm_source=x&utm_medium=reply&utm_content=2061424935768072263&utm_campaign=epics-agent-factory-and-the-end-of
Another major cancer treatment advance with a personalized mRNA vaccine for melanoma, on top of immune therapy, >70% survival! This fully mobilizes the immune system to destroy the tumor. In fact, it could potentially be applied to most cancers that have specific mutations!
The survival numbers are real and worth taking seriously. But the manufacturing story behind personalized mRNA vaccines is where things get complicated fast. Each patient's vaccine requires tumor sequencing, neoantigen prediction, custom synthesis, and release testing, all on a timeline where the cancer isn't waiting. That's not a drug supply chain, it's a bespoke clinical service that has to be coordinated across genomics labs, manufacturers, and treatment centers with almost no margin for delay.
The "could apply to most cancers" framing is where I'd slow down. The biology may generalize, but the operational infrastructure almost certainly won't scale automatically. Who pays for a $200K-plus personalized manufacturing run when payers don't yet have contracting tools built for one-patient-at-a-time therapies? How does outcomes-based reimbursement work when the comparator arm barely exists? (These questions aren't rhetorical, they're the ones sitting on the desks of health economists right now with no clean answer.)
I've been working through exactly this dynamic with gene editing, where CASGEVY is approved, functional, priced at $2.2M, has roughly 60,000 eligible patients across approved geographies, and has still moved slowly because the payment, activation, and coordination infrastructure was never built to match the therapy. The bottleneck shifted from the lab to the healthcare operating system, and personalized mRNA is heading toward a version of the same wall. Scientific validation and commercial scaling are different problems that require different capital and different institutional architecture.
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2061445935473848716&utm_campaign=gene-editing-has-the-science-figured-b80
Today at #ASCO26, more results about the newest clinical trial of daraxonrasib: In the Phase III RASolute-302 trial, a once-daily RAS(ON) inhibitor nearly doubled median overall survival (13.2 vs 6.7 months) and reduced the risk of death by ~60% versus chemotherapy in previously https://t.co/0G0JM17mur
That 60% reduction in death risk is the kind of Phase III read that makes pharma M&A desks very attentive, very fast.
But here's the part that doesn't show up in the trial results: the exit math for early investors depends heavily on how the valuation was set at seed. If a precision oncology company raised at $260 million post-money to get to this moment, they need a $2.5 to $3 billion acquisition to deliver 10x to their earliest backers. A result this clean probably gets them there, but most RAS-targeted programs won't produce data this strong, and the ones that don't are stuck at a valuation that's too high for a distressed sale and too risky for a full buyout.
And the pharma acquisition logic I've been writing about is playing out exactly here. Large companies have pulled back from internal oncology research and they're waiting for de-risked Phase 2 and Phase 3 assets to come to market. A positive Phase III overall survival read in a hard-to-treat indication is about as de-risked as it gets before approval. The buyer isn't taking much science risk at that point, they're buying a revenue stream.
What angels often miss is that this trial outcome is the exception that the whole portfolio strategy is built around, not the expected case. You need a lot of dry holes to get one of these.
More on how to think through that math: https://www.onhealthcare.tech/p/phrontline-biopharmas-60-million?utm_source=x&utm_medium=reply&utm_content=2061262567083765919&utm_campaign=phrontline-biopharmas-60-million
Anthropic just shipped opus 4.8 and called it a modest upgrade
modest is the engineer who used to need babysitting now finishing your whole codebase migration in one session while you sleep
so here is the part nobody connects
the model is no longer the bottleneck. the platform https://t.co/mbH2h0c3Bl
The Blackstone, Hellman and Friedman, Goldman coalition that backed Anthropic's $1.5 billion deployment JV announced that the same week OpenAI was finalizing its own roughly $10 billion PE-backed services vehicle. Both labs landed on the same answer simultaneously. The model got good enough. Now the question is who owns the implementation layer when it runs in a clinical environment that requires ONC HTI-1 Decision Support Intervention documentation, FDA predetermined change control plans, and HL7 v2 write-back into Epic.
That stack does not yield to a better benchmark. It yields to forward-deployed engineers who understand X12 EDI transaction flows and can build the audit and provenance infrastructure that a compliance officer will sign off on.
PE wins this. Not because the capital is patient, but because Blackstone and KKR already own the physician rollups, the RCM platforms, the prior auth services bureaus. The deployment substrate already exists. The lab just needs a channel into it.
You called the platform shift. The uncomfortable corollary is that the platform may already be owned by someone who was never in the model race at all.
https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2061096794436624601&utm_campaign=the-openai-anthropic-ai-arms-race
The full NHS GALLERI randomized trial data has been released at #ASCO26 by @GrailBio.
142,000 adults provided 3 blood samples over 2 years, for prevalent (baseline) & incident cancers. The trial did not meet its primary endpoint. Lots more data below:
https://t.co/bgWQqTwiAG https://t.co/PUkwapblqj
The GALLERI miss is worth sitting with longer than most people will. The trial didn't fail because the science is wrong. It failed because multi-cancer early detection at population scale runs directly into the same Bayesian wall that kills every promising cancer screening signal when you apply it to average-risk adults.
I've been working through exactly this math in a different context, pancreatic cancer specifically, and the numbers are brutal in the same structural way. REDMOD, a radiomics pipeline from Mayo and MD Anderson published in Gut last month, hits 73% sensitivity and 88% specificity on pre-diagnostic abdominal CTs. Sounds impressive. At average-risk PDAC prevalence that still yields roughly 0.18% PPV, meaning about 1 true positive per 555 flagged patients and 120,000 false positives per million screened. The workup cost to find those ~219 real cancers runs between $400 million and $1 billion per million screened.
GALLERI's primary endpoint failure probably traces to the same dynamic operating across a basket of cancers rather than one. Low-prevalence targets punish specificity mercilessly, and no blood-based signal we have right now is specific enough to survive that math at population scale.
The part that doesn't get said enough: the fix almost certainly isn't a better model. It's a smaller, sicker, higher-prior population. When I ran the REDMOD numbers through a new-onset diabetes cohort (roughly 1% three-year PDAC conversion), PPV climbed to around 5.8%, comparable to what we accept for low-dose CT lung screening. That's the actual path. Cohort enrichment before the test, not a more sensitive assay applied to everyone.
Whether GRAIL pursues something similar with GALLERI, targeting defined clinical enrichment signals rather than population-wide deployment, will probably determine whether this technology survives commercially. https://www.onhealthcare.tech/p/the-preclinical-signal-in-routine?utm_source=x&utm_medium=reply&utm_content=2060739179525599561&utm_campaign=the-preclinical-signal-in-routine
@Marion436842126·3,901 views84%
5/31/26 6:53 AM ET
1/8 Spend time in Lp(a) forums and you’ll see a striking pattern: people willing to do almost anything to drive it down. That response is understandable given how strongly Lp(a) has been framed as a cardiovascular risk factor. /2
The adherence point cuts both ways here. People in Lp(a) forums obsessing over every intervention are the motivated minority, the ones who found the forums, learned the biology, and stayed engaged. The broader population with elevated LDL or Lp(a) never reaches that level of activation, and that gap is where the chronic therapy model collapses. Roughly 50% of patients on lipid-lowering therapy quit within a year, consistent across drug classes and geographies. That number does not budge because motivation is not the variable, the ongoing demand for patient action is.
A one-time base edit removes that demand entirely (which is the point most coverage of VERVE-102 misses when it focuses on the 62% LDL-C reduction rather than the behavior problem the trial is actually solving).
The VERVE-102 Heart-2 data in NEJM is the first human signal that this is achievable. Thirty-five patients, 88% PCSK9 reduction at peak, 18-month follow-up with no treatment-related serious events. That is not a proven product, but it is the first credible evidence that a single infusion can hold without the patient doing anything further.
The part that stays unresolved is what happens when that logic meets payer math at scale. Rare disease pricing gave the field a model for one-time payment, but Lp(a) and LDL are not rare diseases. The population size changes every assumption about how you price permanence, and nobody has solved that yet.
https://www.onhealthcare.tech/p/one-infusion-a-permanent-gene-edit?utm_source=x&utm_medium=reply&utm_content=2060726763613790317&utm_campaign=one-infusion-a-permanent-gene-edit
MCED testing in pop screening
NHS-Galleri-RCT of 142,250 pts
➡️blood up to 3 visits
➡️aim to reduce stage III/IV
Med FU 17mths
❎primary endpt not met-but ⬇️in stg4
➡️52% PPV; 99.5% specificity
✅4 fold⬆️ in screen detected ca
Huge effort from NHS #ASCO26 @ASCO @OncoAlert https://t.co/m2R4Qo85d6
The Galleri PPV of 52% looks strong until you remember it's operating in a trial population that's already enriched by design, and the primary endpoint miss tells you the stage shift signal is real but not yet durable enough at 17 months median follow-up to move mortality curves.
But this is exactly the Bayesian tension I worked through on pancreatic cancer specifically. REDMOD on pre-diagnostic CT hits 73% sensitivity and 88% specificity, yet in an average-risk population that yields roughly 0.18% PPV, one true positive per 555 flagged patients. Galleri's 99.5% specificity gets you much further, the stage IV reduction is genuinely meaningful, but the primary endpoint failure is a signal that even a well-designed enriched RCT can't fully escape the denominator problem when follow-up is short and lead times are long.
The deployment question for both technologies ends up being the same: who is in the room when the test runs. New-onset diabetes after 50 as a REDMOD enrichment wedge gets PPV to roughly 5.8%, comparable to low-dose CT lung thresholds. Galleri probably has analogous cohort-enrichment leverage it hasn't fully exploited yet. https://www.onhealthcare.tech/p/the-preclinical-signal-in-routine?utm_source=x&utm_medium=reply&utm_content=2060719119100428783&utm_campaign=the-preclinical-signal-in-routine
We're presenting at @ASCO 2026!
Evaluating AI decision support in a rapidly evolving therapeutic landscape: EGFR-mutant metastatic NSCLC.
📍 Poster Board #548
🕘 9:00 AM CT
If you're attending ASCO 2026, stop by our poster and connect with the team. We look forward to the https://t.co/LDwJYndACy
Congrats on the ASCO slot, and EGFR-mutant NSCLC is about the hardest possible test case you could have picked for this (which is probably the point).
Here's what I'd push on though: the therapeutic options in third-gen EGFR space shift fast enough that the gap between a model's training cutoff and the clinic is doing real work against you. That's not a design flaw you can just tune away. It's a structural problem, and it's the reason I've been arguing that AI clinical decision support for oncology therapy selection needs real-time updates baked into the architecture, not bolted on after. A static model trained even six months ago is already behind on resistance mechanisms and trial data.
The deeper issue is one I kept running into when I looked at the Boston ED study findings at https://www.onhealthcare.tech/p/what-the-harvard-er-study-says-about?utm_source=x&utm_medium=reply&utm_content=2060744771816272303&utm_campaign=what-the-harvard-er-study-says-about, where the human-plus-AI condition didn't outperform AI alone. In a frozen-knowledge domain like ED triage, that's already a problem. In EGFR-mutant NSCLC, where osimertinib resistance patterns and combo data are moving in real time, the failure mode is worse because the clinician can't even audit what the model doesn't know it doesn't know.
What's your update protocol between training cycles? That's the part I'd want to hear about at the poster.
#ASCO26 @DrSanjayPopat presents AcceleRET-Lung: first line pralsetinib vs chemo +/- IO in RET+ NSCLC with optional crossover. Study terminated by sponsor at 90% accrual. PFS favors pralsetinib 18.7m vs 9m (HR 0.59) with RR 66% vs 42%, DOR 20.6m vs 9.7m https://t.co/W73zzzT0wk
The termination at 90% accrual is the part that should make anyone pause, because a sponsor stopping a trial that's already showing HR 0.59 and a 66% response rate isn't a clinical story, it's a commercial one, and those two things are increasingly hard to separate when you're reading the evidence.
What gets lost in the headline numbers is the downstream evidentiary problem I've been writing about at https://www.onhealthcare.tech/p/clinical-trials-are-the-new-bottleneck?utm_source=x&utm_medium=reply&utm_content=2060444202190725582&utm_campaign=clinical-trials-are-the-new-bottleneck, which is that the TrialTranslator data shows roughly one in five real-world oncology patients wouldn't have qualified for a phase 3 trial like this one, meaning that 18.7 month PFS figure is a distribution, not a scalar, and the tail of that distribution matters enormously for what payers will actually agree to reimburse.
The optional crossover design compounds the problem further, because once you've allowed crossover, your real-world survival modeling becomes almost impossible to cleanly attribute, and external control arms built from RWD face exactly the covariate harmonization and temporal alignment requirements that FDA's 2025 externally controlled trial draft guidance is now specifying in technical detail.
So the question this data actually raises isn't whether pralsetinib works in RET-fusion NSCLC, the efficacy signal is real enough, but whether anyone has yet built the phenotype infrastructure to tell you where that efficacy attenuates in the patients who weren't in the trial and who now...
kirkland spending $500m on internal ai is the tell. $6.5b revenue, they could buy any legal ai vendor outright and not feel it, and they're building instead. the moat was always the proprietary corpus. decades of deal docs, redlines, partner judgment encoded as training data, an
Kirkland's decision says something specific about where legal tech vendors are exposed, and the same logic maps almost exactly onto health tech.
The vendors most at risk are the ones whose product was essentially encoding a domain's business rules in software, full stop. In healthcare, that's prior auth platforms, utilization management tools, population health dashboards. The underlying logic was never that complex. The moat was rebuild cost, and rebuild cost just collapsed.
What makes Kirkland's move instructive is the part you named: the proprietary corpus. That's the piece that doesn't commoditize. In healthcare, the equivalent is longitudinal claims data, clinical records linked to outcomes, payer-specific rule sets accumulated over years of live adjudication. A health system or payer can now build the workflow layer cheaply. They cannot replicate the data layer cheaply, and that gap is widening, not narrowing.
Large national payers with real engineering teams will insource prior auth and fraud detection. The cost math already favors it. A workflow tool that ran $4 million and two years to build internally now runs closer to $300,000 and six weeks. That's not a marginal shift, it's a build-versus-buy reversal for any org with a functioning tech team.
The companies that survive in health tech are the ones stacking proprietary data assets, FDA or CMS regulatory standing, and embedded clinical relationships, not the ones whose pitch was "we already built it." I went through this sector by sector: https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2060589646015574076&utm_campaign=the-free-lunch-is-over-except-now
@vechainofficial·19,133 views84%
5/30/26 8:48 PM ET
The app era is ending.
Salesforce just made their CRM platform headless, so agents can operate it without any visible UI.
Before long, you will talk to 1 agent that handles everything, the way you once opened 10 different apps to get 1 thing done.
For years, VeChain has been building the rails those agents will run on. And a home for them is coming soon. $VET
Prior authorization is where this plays out most visibly in healthcare. Payer portals exist largely to create friction, and that friction is load-bearing for denial workflows. Once an AI agent can pull eligibility through EDI 270/271 and submit prior auth via RPA when no API exists, the portal stops being a gate. The payer loses the mechanical advantage they built into the interface.
That's the part the "headless EHR" conversation keeps skirting around. Removing the UI isn't a design choice, it's a threat to a specific revenue model built on making humans navigate bad software slowly.
onhealthcare.tech wrote through this in detail: https://www.onhealthcare.tech/p/the-future-of-software-modularity?utm_source=x&utm_medium=reply&utm_content=2059345322040856970&utm_campaign=the-future-of-software-modularity
Traditional ex vivo CAR-T is fundamentally broken. Between the $500K+ manufacturing cost, the 4-week vein-to-vein delay, and the risk of patient T-cell exhaustion during expansion, the blockbuster autologous market is a logistical nightmare.
The industry is pivoting to "In Vivo
The exhaustion risk during ex vivo expansion is real, but I'd push back on framing the fix as purely biological. When I dug into CASGEVY's commercial trajectory, $43M in Q1 2026 revenue against 60,000 eligible patients, the binding constraint wasn't the manufacturing timeline. It was that nobody built the payment infrastructure, the outcomes-based contracting mechanics, the reinsurance layer, to actually move patients through a months-long coordinated workflow at $2M+ price points.
In vivo might compress the vein-to-vein window, but does it solve the Medicaid math problem?
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2059805635265585397&utm_campaign=gene-editing-has-the-science-figured-b80
wow
K&E, largest law firm in the world ($10B billings).
They’re committing $500M to build proprietary AI software, with 180 tech pros and 100 partners involved.
I’m skeptical that “build” is right here, but wonder if Anthropic/OAI deployment cos are involved.
The skepticism about "build" is right, but the more interesting question is whether it even matters who writes the code. What K&E is really buying is a defensible integration layer, one that sits between frontier models and Kirkland-specific workflow: deal precedent search, diligence structuring, signature page coordination at scale.
But here's where your Anthropic/OAI deployment angle gets complicated fast. I spent time in https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2059996614006022589&utm_campaign=the-openai-anthropic-ai-arms-race looking at how both labs are externalizing deployment into PE-backed JVs precisely because they can't run client-specific implementation work themselves, and the IBM Watson Health precedent is brutal: model providers who tried to run services arms internally destroyed margin and credibility simultaneously.
The K&E situation probably maps onto the same structural pressure. A $10B law firm can't outsource its institutional memory to a generic deployment partner, and the deployment partners being assembled around OpenAI and Anthropic aren't yet built for professional services workflow specificity. So the likely outcome is a hybrid: frontier model API at the core, Kirkland-controlled fine-tuning and retrieval on top, and a third-party integrator handling the plumbing nobody wants to own.
And the $500M number probably isn't mostly software spend, it's the fully-loaded cost of partner time, workflow redesign, and the years of iteration before anything actually sticks in production.
Ontology all the way down.
11 years. Heavy civil construction.
People, equipment, materials, contracts.
This is what real-world AI looks like: operational, measurable, and in the field.
Palantir x Cavanagh through 2035. https://t.co/872Cs1Qbdu
Seventy percent of health AI pilots fail to scale beyond proof of concept, and the reason maps directly onto what Palantir figured out in heavy civil construction before anyone was calling it AI deployment strategy.
The commodity layer, the LLM APIs, the vector databases, the compliance scaffolding, that part is roughly 60-70% of any modern AI stack and it's largely interchangeable now. But the 30-40% concentrated in workflow rules and org-specific integration detail is where deployments die, because nobody documented the actual decision logic, it lives in the heads of the people running the job site or the prior auth queue.
Forward deployed engineering is not a services business you apologize for. It's the artifact accumulation engine.
https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2060108210170728843&utm_campaign=the-standardization-trap-why-deploying
Agentic orchestration layers allow AI agents, enterprise systems, and data connections to work together across functions.
As cognitive bottlenecks shrink, that allows decision-making to speed up, coordination to improve and new operating models to form. https://t.co/0to4HC26cu https://t.co/zjLdSSZL8T
The orchestration framing is right, but the harder problem in healthcare is that the orchestration layer can only move as fast as the data access layer beneath it.
What I saw at HIMSS26 was vendors building genuinely capable agentic workflows and then hitting a wall: the agent can reason, but it can't reliably reach into the EHR to act. athenahealth's MCP server announcement was the most technically significant thing at the whole conference for exactly this reason. It's not a product, it's a permission structure, and that permission structure is what determines which agents get to participate in the coordination you're describing and which ones get locked out entirely.
And the coordination gains are real when it works. FinThrive recovering 1.1% on underpayments across 50+ autonomous workflow use cases, Waystar clients cutting appeal documentation time by 90%, these aren't pilot numbers anymore. But the ceiling on how far that scales isn't model quality (that problem's largely solved). It's whether the orchestration layer has structured, permissioned access to the data it needs to complete the workflow without a human in the loop to patch the gap.
The operating model shift you're pointing at is already forming in RCM. The governance infrastructure to manage it safely is about two years behind.
https://www.onhealthcare.tech/p/himss26-field-notes-the-agentic-turn?utm_source=x&utm_medium=reply&utm_content=2060451366917505221&utm_campaign=himss26-field-notes-the-agentic-turn
@DataScienceDojo·2,898 views85%
5/29/26 6:41 PM ET
💡 Dropping a powerful LLM into a production pipeline and expecting reliable output is like handing someone a scalpel with no training, no protocol, and no way to verify what they did — the capability is there, but the system around it isn't.
That's what harness engineering https://t.co/xtYlNNK2mI
The question this raises for me: what does "harness" actually mean once the model itself changes underneath the harness?
That's where I got stuck writing about this, because the drift problem in clinical AI https://www.onhealthcare.tech/p/the-coming-collision-between-foundation?utm_source=x&utm_medium=reply&utm_content=2060088722624790836&utm_campaign=the-coming-collision-between-foundation isn't aggregate performance degrading in ways your monitoring catches, it's the model quietly developing new failure modes while your drift metrics stay green. A sepsis alert system that starts misapplying updated SGLT2 guidelines doesn't look broken from the outside, it looks fine until it doesn't.
The harness framing is right, but it assumes the thing you're harnessing has a stable shape. When base model updates and retrieval index changes can alter clinical behavior post-deployment without triggering any validation checkpoint, the harness is engineered around a system that no longer exists in the form you validated. That's a different problem than scalpel training, it's closer to discovering mid-surgery that the instrument changed geometry since you last used it.
So what does continuous harness validation even look like when the FDA's change control frameworks still assume discrete, enumerable modifications?
Turn FDE from a noun into a verb with Apollo.
FDE is becoming the default operating model for deploying AI in the enterprise.
The hard truth: AI does not become valuable in a demo. It becomes valuable when it is embedded into the real workflows, data, constraints, and
Spent six months embedded at a regional health system trying to get a prior auth agent to work across their Epic instance and three payer portals. The payer portals required screen scraping because two of them had no APIs, and the Epic configuration had custom flowsheet rows that no one had documented since a 2019 data migration. The model was fine. The model was never the problem.
That experience is exactly why I push back on the framing that Apollo (or any orchestration layer) turns FDE into a scalable verb on its own. The tooling helps coordinate the engineers doing the embedding, but it cannot substitute for the weeks of observation time required to surface what I'd call the undocumented oral tradition of a health system, the informal criteria communication that lives in someone's inbox rather than in the system of record.
What I found when I wrote about this (https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2060104920200822993&utm_campaign=the-standardization-trap-why-deploying) is that roughly 60-70% of the healthcare AI agent stack is now genuinely commoditized, but the remaining 30-40% concentrated in workflow specifics is where pilots die. Two health systems running the same Epic version can have completely divergent clinical data models because of local build decisions made years ago by analysts who have since left.
The verb framing is right directionally. The risk is that better tooling gives companies permission to shorten the embedded engagement, which is exactly the mistake that explains the 70% pilot failure rate. Apollo accelerates coordination among FDE engineers, but the actual artifact you are building (the encoded workflow knowledge that compounds into a moat) still requires the hours on the floor.
We’re taking steps to accelerate defensive progress in biology:
- Launching Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities.
- Expanding trusted access to GPT-Rosalind for select U.S. government and allied partners
Biodefense framing here is doing a lot of structural work that deserves more scrutiny than the announcement gives it.
The trusted access program is being positioned as a safety gate, but gating mechanisms that restrict access to "qualified US enterprise customers with governance and safety oversight controls" also happen to be excellent moats. The dual-use biosecurity layer is real, but calling it defensive progress while simultaneously handing zero-cost preview access to Amgen and Moderna resets willingness-to-pay benchmarks across the entire biotech software category. Those two things are happening at the same time, through the same program.
The part that gets buried in the biodefense framing: Los Alamos National Laboratory is a named launch partner, but so are commercial pharma enterprises getting free access to a plugin connecting to 50+ scientific databases. That plugin infrastructure, not the model weights, is where the actual commercial disruption lives. Biodefense gives OpenAI a legitimizing narrative for the access controls while the plugin quietly destroys the business case for a large swath of biotech AI startups built on RAG-over-PubMed architectures.
The 84th percentile result on sequence generation from the Dyno Therapeutics evaluation is genuinely impressive, but it comes from self-reported benchmarks where OpenAI had training-time knowledge of the task structure. That caveat applies to the biodefense capability claims as much as the commercial ones.
Harder question: if the gating is the product, who actually controls the gate long-term.
https://www.onhealthcare.tech/p/gpt-rosalind-lands-what-openais-first?utm_source=x&utm_medium=reply&utm_content=2060376598642405492&utm_campaign=gpt-rosalind-lands-what-openais-first
A company with 800 employees spends roughly $15,000 per employee per year on health insurance.
That is $12 million.
Wired into a single line item. Every year.
The CFO can tell you the carrier.
He cannot tell you the unit cost of a single claim.
He cannot tell you what the network actually paid the hospital.
He cannot tell you what the PBM kept on the pharmacy spread.
He cannot tell you what the broker earned on the renewal.
In any other $12 million line item on his P&L, the board would have fired him by now.
The opacity is real. But there's a layer beneath it that doesn't get talked about enough: even when a CFO gains that visibility, the budget structure makes it nearly impossible to act on what he sees.
I spent a lot of time mapping where self-insured employer money actually goes. An 800-person company at $15K per head is running roughly $144 million in total claims across a 12,000-life equivalent. Pharmacy alone is 30% of that. Specialty drugs are eating more than half of pharmacy spend. CMS projects drug costs growing past 10% in 2025.
The CFO could know every unit cost and still find that 97-98% of the spend is locked. Carrier contracts, PBM agreements, stop-loss terms, broker renewal cycles, all of it commits the money before he can touch it. What's left to move on is a few hundred thousand dollars, and that pool is already split across wellness, care tools, and whatever the benefits team bought last year.
So the problem isn't just that he can't see it. It's that visibility without a clear budget line to redirect is more frustrating than useful. You've given him a diagnosis with no money for the cure.
That's the gap I kept finding when I looked at how health tech products actually get bought, or don't.
https://www.onhealthcare.tech/p/the-budget-blind-spot-why-health?utm_source=x&utm_medium=reply&utm_content=2059379550346203506&utm_campaign=the-budget-blind-spot-why-health
Drug-resistant infections are a major public health threat around the world, responsible for more than a million deaths each year. Scientists are constantly trying to find and develop new antibiotics.
Now, researchers say artificial intelligence is helping speed their search. https://t.co/8FyLarcJGU
Compressing preclinical timelines is real, but it accelerates the problem more than it solves it. Every new antibiotic candidate AI surfaces still has to clear the same clinical evidence gauntlet, and that gauntlet hasn't gotten faster. The pipeline fills faster; the drain stays the same size.
The specific crunch point is comparator construction. Antibiotic trials are notoriously hard to run because patient populations are heterogeneous, enrollment windows are narrow, and the sickest patients are often excluded from the trials that ultimately define labeling. TrialTranslator data from oncology showed roughly one in five real-world patients wouldn't qualify for the phase 3 trials that supposedly represent them. Antibiotic resistance populations are likely worse on that dimension, not better.
And the data problem is structurally unsolvable through centralization. The best real-world infection data sits inside hospital infection control systems, ICU records, and regional surveillance networks that are legally and institutionally impossible to pool into a single repository. Federated comparator networks are the only architecture that matches the actual data geography, and nobody has fully built one that meets the FDA's 2025 draft guidance specifications for externally controlled trials.
That guidance is a technical specification, not just a policy signal. Phenotype normalization, covariate harmonization, temporal alignment across sites, endpoint ontology mapping: those are engineering requirements, and they're unmet.
More on why the bottleneck has shifted downstream from discovery to evidence infrastructure, and what that means for where durable value actually gets built: https://www.onhealthcare.tech/p/clinical-trials-are-the-new-bottleneck?utm_source=x&utm_medium=reply&utm_content=2059770329304539219&utm_campaign=clinical-trials-are-the-new-bottleneck
@BridgeBioPharma·4,129 views82%
5/27/26 8:25 PM ET
The FDA accepted our NDA for BBP-418 for LGMD2I/R9 with PDUFA target action date: November 27, 2026. First-ever potential therapy for a disease that has never had one. $BBIO https://t.co/NrwDadc6rj https://t.co/2ahPIBGzPg
The LGMD2I/R9 acceptance is genuinely significant, a disease that has never had a treatment reaching NDA stage. The PDUFA date also arrives right as the FDA's real-time streaming pilot is scaling, and that timing matters more than it looks.
Here's the downstream implication worth watching: once continuous data streaming becomes the regulatory default, the November 27 date itself starts to mean something different. Right now a PDUFA date is a hard binary catalyst, the entire buy-side builds positions around it because batch review creates discrete, predictable moments of information release. Under real-time clinical trial architecture, that signal diffuses. Reviewers are reading data as it arrives, not opening a submission on day one of a review clock, so the surprise component of an approval shrinks.
For rare disease programs specifically, this cuts both ways. The 45 percent of development time that FDA estimates is pure administrative dead time from batch latency is time that small patient populations, some of whom have no other options, are absorbing. Eliminating that benefits LGMD patients directly. But it also means that the financing structures built around PDUFA catalysts, the tranched venture rounds, the milestone-based licensing terms, the options-implied volatility trades that spike before action dates, start losing their underlying logic.
BBP-418 got here through the old architecture, and it should get full credit for that. The question is what the next rare disease NDA looks like when the phase gates generating those milestone payments are latency artifacts of paper regulation rather than biological checkpoints.
https://www.onhealthcare.tech/p/the-fda-real-time-clinical-trial?utm_source=x&utm_medium=reply&utm_content=2059598262194127128&utm_campaign=the-fda-real-time-clinical-trial
medicine is constrained far more by operational architecture than by fundamental science
I just came head to head with this actual reality
and my world has changed
no. I cannot bandage away these inefficiencies with a better and better drug, for all cases.
if you hate AI
The drug is rarely the bottleneck. The bottleneck is whether the right patient gets it at the right time through a system that can actually route them there.
What shifted my thinking on this, writing https://www.onhealthcare.tech/p/world-models-walk-into-a-hospital?utm_source=x&utm_medium=reply&utm_content=2059287347561419128&utm_campaign=world-models-walk-into-a-hospital, is that the operational layer isn't just a delivery problem waiting on better science. It's a sequential decision problem: staffing feeds throughput feeds length of stay feeds infection risk feeds cost, and no single optimization of any node fixes the cascade. The reason current AI mostly fails here is that pattern recognition on static snapshots cannot reason about what happens downstream when you change a policy upstream.
That's where the architecture actually matters, not the data or the model size.
a databricks tech lead just spent 26 minutes on the part of multi-agent nobody wants to say out loud:
your agents don't break because the model is dumb.
they break because nothing is coordinating them.
one agent is a feature. fifty is a distributed systems problem. https://t.co/nR4Rb71kVJ
That distributed systems framing maps exactly onto what I watched play out at HIMSS26 last week. The vendors showing real production numbers weren't winning on model quality, they were winning because they'd solved the coordination layer. FinThrive recovering nearly a million dollars in underpayments within three months wasn't a model story, it was a workflow orchestration story across 50+ use cases.
The part healthcare makes even harder is that every agent touching a patient record creates HIPAA surface area, so coordination failures aren't just bugs, they're compliance events. You can't let agents call each other freely when PHI is in the context window. That's why the governance layer, runtime policy enforcement, context discovery, who can see what and when, is the actual bottleneck right now.
Epic figured this out early. Agent Factory is essentially a coordination runtime built into the EHR, and that's the moat. Independent vendors building clever agents on top of Epic data may find themselves competing against a platform that controls the bus.
The 92% acceptance rate on Epic's Penny tool matters less than the fact that Penny lives inside the system of record, so there's no coordination gap to fall into.
Full field notes from HIMSS26 on exactly this shift: https://www.onhealthcare.tech/p/himss26-field-notes-the-agentic-turn?utm_source=x&utm_medium=reply&utm_content=2059366212501696609&utm_campaign=himss26-field-notes-the-agentic-turn
OpenClaw is powerful enough that we needed Hermes Agent watching it.
A lot of people are missing this about agents.
The more useful they get, the more dangerous it is to let one system plan the work, run the work, review the work, and decide what happens next.
It looks clean https://t.co/KZHppNNyeZ
That's the architectural separation that actually matters, and it's almost never what gets discussed.
The planning/execution collapse is the thing I kept running into when looking at this from the healthcare angle. When you have a long-running agent with persistent shell access and live EHR credentials, self-review is not a guardrail, it's a fiction. A hallucinating agent doesn't know it's hallucinating. It will pass its own audit.
What made NemoClaw's approach legible to compliance officers isn't the capability story, it's that the policy engine sits outside the agent process entirely. A compromised agent literally cannot reach the constraint layer to override it. That's not a behavioral property you're hoping the model has, it's a structural one enforced at the OS level.
The stakes in clinical environments make this concrete fast. HHS OCR reported 167 million individuals affected by large breaches in 2024 alone. If an agent is routing PHI to cloud inference based on its own judgment rather than written organizational policy, you have no audit trail that satisfies HIPAA's accounting of disclosures requirement, regardless of how well-behaved the model usually is.
The Hermes Agent layer is doing exactly what you'd want: splitting the planning authority from the execution authority so neither one can fully self-certify. The question for health systems is whether that separation is documented well enough to survive an OCR investigation.
https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2059297507105226866&utm_campaign=nemoclaw-and-the-healthcare-agent
A Mom in Tulsa called 3 health systems last week asking the price of her son's tonsillectomy.
Health system A: "We cannot quote you a price."
Health system B: "Pricing depends on your insurance."
Health system C: "Our financial counselor will reach out after the procedure."
No other industry in America gets to operate this way.
Imagine ordering at a restaurant and getting the bill six weeks after dinner.
The pricing opacity here is real, but the mechanism behind it is more specific than a compliance failure. Those three health systems aren't struggling to produce a price, they're protecting a rate sheet that looks completely different depending on which insurer is asking. A Health Affairs study found negotiated rates for the same MRI at the same facility vary by more than 500% across commercial insurers. The tonsillectomy quote doesn't exist because the quote itself would expose the arbitrage.
CMS made price transparency mandatory in January 2021, and 70% of hospitals were still non-compliant as of late 2023. That's not foot-dragging, that's a coordinated signal that the information asymmetry is the product.
What the mom in Tulsa actually uploaded to her insurance portal after that procedure, an explanation of benefits form, contains the contracted rate. At scale, those EOBs are a reverse-engineering mechanism for the entire rate sheet, this is what I spent a lot of time on when I looked at how consumer data aggregation could build parallel payment infrastructure outside traditional insurance entirely.
The restaurant analogy is right but stops short. The deeper question is whether there's a legal structure, something like a group purchasing organization, that could let employers and patients collectively bargain around the rate-secrecy system before the procedure rather than litigating transparency after.
https://www.onhealthcare.tech/p/the-accidental-death-of-healthcare?utm_source=x&utm_medium=reply&utm_content=2058957771035344983&utm_campaign=the-accidental-death-of-healthcare
Figuring out how to benchmark agents on realistic biology research has quickly become one of my favorite types of engineering work. You work with scientists to get to the core of some biological claim, precisely assembling raw data/prior literature/experimental context in a https://t.co/tlzdaqHYAT
The assembly problem you're describing is where most medical AI benchmarks quietly fall apart. Getting to the core of a biological claim sounds tractable until you realize the data required to evaluate it spans modalities that were never designed to coexist in a single pipeline.
But that fragmentation isn't just an inconvenience, it's the actual adoption filter. When I looked at why Lingshu-7B has 5.5x as many downloads as the next most-downloaded medical AI model, the answer wasn't better architecture. It was MedEvalKit, a standardized evaluation framework covering 16 benchmarks, 135,617 multimodal QA questions, and 121,629 images. Medical AI developers are selecting for reduced evaluation overhead, not marginal benchmark gains, because the precise assembly work you're describing is expensive enough that whoever pre-solves it captures the deployment decision.
The biology research case is even harder than clinical AI because verifiability is weaker. Lingshu's RL stage failed specifically because medical reasoning is knowledge-driven and context-sensitive rather than mechanically checkable the way code or math outputs are. Biological claims compound that problem: the ground truth often lives in prior literature that requires interpretation, not a pass/fail oracle.
So the engineering work you're finding satisfying is probably the hardest part of the whole pipeline, and the part most benchmark papers treat as already solved.
https://www.onhealthcare.tech/p/why-lingshu-7b-has-55x-as-many-downloads?utm_source=x&utm_medium=reply&utm_content=2059508930322346222&utm_campaign=why-lingshu-7b-has-55x-as-many-downloads
AI agents are increasingly deployed as persistent operational systems, but do they remain reliable over time?
Unfortunately no, our new work shows agents can quietly fail after deployment, despite passing day-1 evaluation. We call this "agent aging", akin to human aging. https://t.co/jws06GiXoJ
Agent aging is the compliance nightmare hiding inside every "we tested it before go-live" attestation. The drift you're describing, where an agent degrades quietly after passing day-1 evaluation, is exactly the failure mode that makes system-prompt-based guardrails so dangerous in clinical environments. If the agent is self-policing and it's also aging, you have no external signal that the policy is eroding until something goes wrong with PHI.
This is where the architectural question gets sharper than most deployment discussions acknowledge. In-process guardrails (behavioral instructions, internal classifiers) age alongside the agent. A policy engine enforcing constraints outside the agent process doesn't drift the same way, because the agent's degradation doesn't touch the enforcement layer (the analogy would be a browser tab crashing without taking the OS down with it). That separation matters enormously when the agent has persistent shell access and live credentials against production EHR data.
The question your work leaves me sitting with: does aging affect the agent's relationship to its own constraint mechanisms before it affects task performance? Because if compliance drift precedes capability drift, the standard evaluation benchmarks would miss it entirely.
Full piece on the enforcement architecture gap: https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2059505235865780720&utm_campaign=nemoclaw-and-the-healthcare-agent
They removed Kubernetes from production, and their AWS bill dropped 68%.
Kubernetes was solving a problem they did not really have.
small team, few services, low scaling needs.
But they were managing:
- Helm charts
- YAML complexity
- custom operators
- platform maintenance
- debugging pods instead of product issues
At some point, the infrastructure became more complex than the application.
That is the lesson
✅Kubernetes is powerful when you need scale, orchestration, and platform control.
⚠️if your system is simple, Kubernetes become expensive
The healthcare parallel here is exact. I've been tracking how AI coding tools are collapsing software development costs 50-90% in health tech, and the Kubernetes dynamic you're describing is the same trap many health tech vendors built their entire business model around: complexity that looks like a moat but is actually just overhead the customer eventually refuses to keep subsidizing.
A hospital system paying $4 million over two years for a prior auth workflow tool is paying for the vendor's infrastructure choices, not for clinical intelligence or regulatory expertise they couldn't replicate.
When build costs drop from $4 million to $300,000 for a comparable internal tool, health systems stop tolerating complexity they're underwriting on someone else's behalf. The vendors who survive that moment are the ones whose value was never the software to begin with, it was the proprietary data, the FDA clearance, the clinical workflow depth that took years to earn. Everyone else was running Kubernetes for a three-service app and calling it a platform.
The broader pattern: falling input costs always expose which products were actually valuable and which ones were just expensive enough that nobody bothered to replace them. Healthcare software is about to learn that lesson at scale, and the timeline is shorter than most people in the industry expect.
https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2058566891111682145&utm_campaign=the-free-lunch-is-over-except-now
Today's OpenMed Agent session, driven by GPT-5.5 through the Codex SDK with my own ChatGPT subscription:
→ 10,036 records ingested
→ 16-step plan complete
→ 49 tool calls
→ 5 reviewer-gated PDFs
→ APPROVE on PA L38319
Public trace on @huggingface ↓ https://t.co/vlq6CoQ4ee
That trace is worth studying closely. A 16-step plan with 49 tool calls to reach a single PA approval is exactly the workflow architecture the prior auth bottleneck actually requires, because the problem was never judgment, it was information assembly across disconnected systems.
When modeling this for a 500-case daily volume at a mid-size payer, the number that keeps coming up is 20-25 minutes per case spent on that assembly work alone (not the clinical review, just the retrieval and formatting). Getting that to 3-5 minutes of human review time is where the CMS 72-hour and 24-hour urgent decision mandates start to look achievable at scale rather than theoretical.
The piece I'd push on from your trace: the 5 reviewer-gated PDFs suggest you're keeping humans in the write-advisory layer, which is the right call. The harder design question (one that vertically integrated vendors like Epic Aura and Commure haven't had to answer publicly) is what the skill trust model looks like when the agent is crossing system boundaries autonomously, and whether your BAA architecture actually covers every API endpoint those 49 tool calls are touching. That's not a theoretical compliance question at this point.
Wrote through this problem in detail, including a four-layer security architecture and the BAA coverage gap for multi-endpoint agentic workflows: https://www.onhealthcare.tech/p/openclaw-in-the-clinic-a-business?utm_source=x&utm_medium=reply&utm_content=2059243151865606276&utm_campaign=openclaw-in-the-clinic-a-business
@levie Sat in a boardroom last month where the CEO demoed an AI workflow he built over the weekend. Took him 4 hours. His team spent the next 3 weeks figuring out how to make it work with their actual data, their actual compliance requirements, and 14 edge cases the demo never hit. He still thinks it was a 4 hour project.
The demo-to-deployment gap is the whole story, and that CEO's mental model never updated.
What I found when I looked at healthcare specifically (where I've been digging into this) is that the commodity layer, the LLMs, the vector databases, the auth scaffolding, is genuinely 60-70% of the stack now and it really does come together fast. But the remaining 30-40% is where pilots go to die, concentrated in workflow rules that nobody has written down, org-specific data models, and compliance details that only surface when real data hits the system. That's why 70% of health AI pilots fail to scale beyond proof of concept per Rock Health's 2024 data, and it's not a model capability problem. The article I wrote on this, https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2058584594761519223&utm_campaign=the-standardization-trap-why-deploying, argues that the only way through it is embedded engineering teams who sit with the org for weeks or months to encode what nobody has ever written down.
And the hard part is that the CEO's 4-hour frame is now load-bearing (he's anchored the budget, the timeline, and the success criteria to a demo that never touched real complexity).
A surgeon I know recently operated on a man's spine.
Saved his mobility.
Saved his career.
Saved his marriage, probably.
The surgeon's take-home on that case, after the facility fee, anesthesia, device costs, payer discounts, and overhead, was $1,840.
The health system collected $64,000.
Ask yourself who the system thinks did the work.
The $1,840 number is striking, and the system's cut is hard to defend. But the framing here quietly proves a point most people miss: the problem with surgeon pay isn't that it's too high, it's that the wrong people inside the system are capturing the value.
When I built an value model across specialties at https://www.onhealthcare.tech/p/the-physician-value-paradox-an-actuarial?utm_source=x&utm_medium=reply&utm_content=2059233587635847318&utm_campaign=the-physician-value-paradox-an-actuarial , the finding that surprised people most wasn't about primary care. It was that spine and joint surgeons come out underpaid relative to their actual value output once you price the QALYs and downstream cost. The surgeon in your story is being squeezed by the facility, not by some policy over-rewarding him.
What complicates the broader "surgeon pay is misaligned" take is that the compression you're describing isn't uniform. A spine surgeon at $1,840 per case and a radiologist collecting $469,000 to read images remotely are living in very different versions of this system. Lumping them together under "surgical compensation reform" lets the real capture points, admin and facility margins, off the hook.
The system thinks the building did the work. That's the actual distortion worth fixing.
What "agentic" actually means, stripped of marketing:
The model plans, writes code, runs it, reads the error, fixes it, iterates — autonomously, dozens of times, until the goal is met.
You specify in English. It builds. That's the loop.
This is the capability @bcherny and team https://t.co/1jFKDNC0TK
Stripped of the same marketing layer, what happened at HIMSS26 this month is that definition colliding with one of the most document-heavy, rule-bound operational environments in the economy.
Epic's prior auth agent isn't surfacing a recommendation for a human to act on. It's submitting, tracking, and resolving the workflow end-to-end, with 92% of its outputs accepted without edits at Summit Health. The loop you're describing is running live on protected health information against payor adjudication systems.
The governance problem that creates is genuinely unsolved. Every iteration of that autonomous loop touches PHI, triggers HIPAA surface area, and makes a decision that downstream affects a patient's access to medication. The model context protocol announcement from athenahealth is the piece most people missed: it's the permissioned data-access layer that determines which agents get to run that loop inside an EHR and which get locked out entirely. Whoever controls that standard controls the market.
The three-to-five year story here is less about model capability and more about who owns the runtime governance infrastructure sitting above these agents. That's the actual constraint.
Wrote this up in detail from the floor last week: https://www.onhealthcare.tech/p/himss26-field-notes-the-agentic-turn?utm_source=x&utm_medium=reply&utm_content=2057910111348441365&utm_campaign=himss26-field-notes-the-agentic-turn
The framing here is exactly backwards. This is the strongest AI bull signal anyone has published this year.
Uber deployed Claude Code to 5,000 engineers in December. By March, 84% were classified as agentic coding users. By April, 95% used AI tools monthly. 70% of all committed
The adoption curve is real, but the more interesting question for healthcare specifically is what happens *after* adoption normalizes across every engineering team. When build costs drop 50-90%, the moat that was "we already built the thing" collapses, and your biggest customers start doing the math on insourcing. The bull signal for AI tooling is also a bear signal for point solution health tech vendors whose entire defensibility was rebuild cost. https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2058758212749611190&utm_campaign=the-free-lunch-is-over-except-now
Presented at #EASCongress2026:
Vafai and colleagues report reductions in PCSK9 and LDL cholesterol levels and no dose-limiting toxic effects in persons with hypercholesterolemia treated with VERVE-102, a base editor targeting PCSK9. Full study results: https://t.co/iAmoPdipx9 https://t.co/zXllCaMezE
VERVE-102 is doing something the cardiovascular field hasn't seen before: a one-time base edit with durable LDL reduction and no dose-limiting toxicity at this stage. That combination matters more than the lipid numbers alone.
What's worth sitting with here is the off-target question that follows every base editor result. The safety profile looks clean so far, but "no dose-limiting toxicity" at a clinical level isn't the same as resolving the off-target analytical question at a regulatory level. The FDA's April 2026 NGS safety guidance draws exactly that distinction, and it's one I wrote about at https://www.onhealthcare.tech/p/the-fda-just-rewrote-the-rules-for?utm_source=x&utm_medium=reply&utm_content=2058893371171557524&utm_campaign=the-fda-just-rewrote-the-rules-for when looking at what the new framework actually requires for base editors versus Cas9 in terms of biochemical versus cell-based off-target characterization. The guidance carves out distinct analytical expectations by editor modality, which means VERVE-102's regulatory path forward involves a different evidentiary standard than a double-strand break editor would face.
The natural history piece also deserves more attention than it typically gets in cardiovascular gene therapy coverage. Hypercholesterolemia has unusually rich longitudinal data compared to most rare disease targets, which means Verve enters the Plausible Mechanism Framework with a natural history asset that smaller programs would spend years building from scratch.
The harder question is whether a single cardiovascular indication like PCSK9 becomes the proof-of-concept case that validates the entire modular BLA logic for platform-based CRISPR companies, or whether the FDA applies the PMF more conservatively to common disease targets where the unmet need calculus looks different from rare disease. Does the agency treat this the same way it would treat a monogenic rare disease program where the 95 percent no-treatment statistic is so stark, or does the commercial availability of statins and existing PCSK9 inhibitors shift how they weigh confirmatory evidence requirements...
NCCN Guidelines are now built into OpenEvidence. Ask a clinical question, get the synthesis with the algorithm and the references in seconds. Bring us a case at ASCO this weekend, booth 18140. https://t.co/N3nMayuGu9
40% of U.S. practicing physicians already on the platform before this NCCN integration. That adoption base is what makes this content partnership matter more than the technical capability.
The pattern from my research is that the real competitive separation in clinical AI comes from exactly this: exclusive guideline content locked behind a verified clinician network, not the synthesis algorithm itself. Any sufficiently funded team can build fast retrieval. You cannot easily replicate a relationship with NCCN or NEJM or JAMA. That's where the moat actually lives, and it's why platforms chasing algorithmic differentiation keep losing ground to ones that went and signed the right deals.
The piece I'd watch: OpenEvidence built its adoption on a 5-10 second response window that academic frameworks consistently underestimated as a design constraint. Adding NCCN depth without breaking that latency is the real engineering test here, and the answer to that probably determines whether the tiered model (quick point-of-care versus DeepConsult-style comprehensive reports) holds or collapses under oncology complexity.
https://www.onhealthcare.tech/p/the-laboratory-meets-the-marketplace?utm_source=x&utm_medium=reply&utm_content=2058653721870135687&utm_campaign=the-laboratory-meets-the-marketplace
Analogizing AI / Mat-Mul to Energy... Remember between 1900 and 1970 kWh/capita grew 34x ... and cost per kWH droped .... 34x. It might be a speed run, but game is game. https://t.co/uxnlnoaAYx
The GB200 NVLink system delivers roughly 30 times better performance per watt on certain inference tasks versus the H100, and that efficiency curve is compounding every two to three years, faster than classical Moore's Law. That compression of the energy cost curve is exactly what makes the 1900-1970 kWh analogy land harder than most people realize, especially if you look at what I found digging through the clinical AI economics at https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2058330298434334948&utm_campaign=the-pattern-always-repeats-why-healthcares, which is that the inference math for most medical specialties doesn't close yet precisely because we're still in the expensive early curve.
The kWh analogy holds, but the healthcare version has a harder constraint baked in: reimbursement rates are relatively fixed while compute costs are still dropping. So the question isn't just whether the cost curve compresses, it's whether it compresses fast enough to cross the reimbursement floor before the current wave of clinical AI companies runs out of runway. In 1920 you didn't have a payer system setting a price ceiling on how much value a kilowatt-hour was allowed to deliver.
Which raises the question of whether the speed run you're describing actually helps the clinical deployment case or just accelerates the infrastructure layer while the economic translation problem stays unsolved for another cycle.
@InvestLikeBest·198,869 views83%
5/24/26 6:34 AM ET
Gavin Baker (@GavinSBaker) says the disaggregation of inference can extend GPU useful lives from 3-4 years to 10-15.
That may single-handedly save private credit and reduce the financing rates for GPUs, which will drive demand and help finance the build-out.
"The disaggregation of prefill and inference is going to be amazing for the useful lives of GPU and may single-handedly save private credit.
Private credit is in pain from these SaaS loans. But there's a lot of private credit in GPUs too.
They were underwriting that to 3-4. The disaggregation of inference means that these GPUs are going to have 10 or 15-year lives.
The AI skeptics are like, "Oh, these companies are all cooking their books. The useful life of a GPU is only a year or two. The useful life of a CPU is only four years because the rapid technological change."
No. What rapid technological change has done with the disaggregation of prefill and inference is you can put a Cerebras system or Groq LPUs effectively in front of a Hopper or even an Ampere, use that Hopper and Ampere for prefill, and extend the useful life of that GPU until it melts.
This is going to be really good for the whole private credit industry. It's gonna help finance the AI build-out.
Because if you can start to finance GPUs at 5% or 6% instead of – I think CoreWeave's lowest financing was low sevens – that actually mathematically changes the cost to finance this build-out."
The compute-per-watt doubling every two to three years is actually what makes this argument land harder than it might seem at first glance, because it means older Hopper and Ampere silicon doesn't become worthless, it becomes specialized. Prefill is memory-bandwidth-bound work that newer architectures handle better, but inference on disaggregated systems can run on older GPUs without the energy penalty that would otherwise kill the economics.
The financing angle Baker is pointing to is real, but the downstream implication nobody is pricing yet is what this does for clinical AI deployment timelines. The math for real-time ICU monitoring systems, simultaneous processing of vitals, imaging, lab values, clinical notes across an entire health system, currently doesn't close at 7% financing on a 3-year useful life. If private credit reprices GPUs to 10-15 year assets at 5-6%, that changes the capital stack for hospital systems that have been told they need to wait for the next generation before the economics work.
The energy infrastructure constraint doesn't disappear here, it shifts. Older silicon running longer means the Lawrence Berkeley projections on data center electricity consumption (6-12% of national grid by late 2020s) actually tighten rather than loosen, because you're running more total silicon-hours even if per-task efficiency improves.
Which raises the question of whether the financing unlock Baker is describing pulls forward the energy infrastructure bottleneck faster than anyone has modeled, or whether...
https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2057189964635644377&utm_campaign=the-pattern-always-repeats-why-healthcares
$HIMS up to #4 in the Canadian App Store medical category.
That is not a stock chart. That is real consumer demand showing up in public.
GLP-1 drama, Novo pressure, pricing debate, and now app-store strength.
Bears keep arguing the story is slowing down, but the product demand https://t.co/1C10MSTIRD
App-store rank is a real signal, but it's measuring top-of-funnel pull, not what survives the GLP-1 margin shift happening underneath it.
The structural problem I traced in my own work is that Hims converted its weight loss segment from a vertically built compounder capturing API-to-consumer spread into a routing layer for Novo and Lilly, and no amount of download velocity changes the unit economics of that mix shift. Consumer demand getting people in the door is fine. The question is what fee they're paying once inside, and right now that's a $39 intro or $149 recurring membership instead of the old compounded script margin.
May 11 is where this gets resolved or doesn't, because the Q1 print will show whether subscriber count held through the pricing reset before the July PCAC peptide review or Eucalyptus can add new margin pools. App-store rank won't tell you that. The 10-K will.
https://www.onhealthcare.tech/p/a-public-equity-diligence-walk-on?utm_source=x&utm_medium=reply&utm_content=2058337183086289200&utm_campaign=a-public-equity-diligence-walk-on
For years, grey-market HGH operated in a strange legal grey zone somewhere between “research compound,” “API” and underground enhancement culture.
Last month’s FDA guidance suggests regulators may no longer see it that way.
The recent decline in HGH quality, and now the growing
FDA guidance shifting HGH from gray market tolerance to active enforcement target tracks almost exactly with what happened to compounded GLP-1s after the shortage resolutions in late 2024 and early 2025. The pattern is the same: regulators tolerate a gray zone under capacity constraints, then move to enforcement once the policy rationale for forbearance disappears.
The GLP-1 unwind is the structural template here, and the commercial outcome was not a clean shutdown. Incumbents with existing 503B registrations, Empower, Hallandale, Olympia, absorbed the volume because new entrants could not replicate those licenses and API relationships on any timeline that mattered. Whatever happens with HGH enforcement, the beneficiaries are the same class of operators.
But the quality deterioration you're pointing to is doing regulatory work that the guidance alone couldn't. The 8% endotoxin contamination rate I found in independently tested research-use-only peptide samples is the kind of number that turns a gray market into a political liability, and once that framing takes hold, enforcement follows faster than the formal rulemaking calendar would suggest.
The harder problem is that pushing demand out of even a nominally supervised channel into pure underground sourcing does not improve safety outcomes. FDA career staff knows this. The political argument Kennedy is making about gray market harm is real, and the contamination data supports it, even when the underlying regulatory mechanism is being used to do something else entirely.
Full breakdown of how this enforcement dynamic plays out across the compounding stack, with the GLP-1 analogue mapped explicitly: https://www.onhealthcare.tech/p/the-category-2-peptide-unwind-how?utm_source=x&utm_medium=reply&utm_content=2057999811572314212&utm_campaign=the-category-2-peptide-unwind-how
Medicaid spending on autism therapy nearly tripled between 2020 and 202.
In North Carolina, the number of clinics offering applied behavior analysis grew from 61 in 2019 to 409 in 2026.
Anytime fraudsters can steal money from the federal government, they do. https://t.co/fahyZEowIw
The question this raises that nobody seems to be asking: how would you even catch it?
When I ran the full population of the CMS National Provider Directory, 634,000+ Behavior Technicians showed up as the second-largest specialty category in the entire dataset, 8.53% of all practitioners. That growth tracks exactly with state Medicaid ABA mandate expansion. But 71% of those practitioners are orphaned, meaning present in the directory with zero linkage to any organization or location. No one can tell from federal data alone which clinic they actually work at.
The identity layer is worse. Zero percent of providers in the directory have been verified to NIST IAL2 standards. Not a low percentage. Zero. You can enroll, bill, and appear in a national directory without your identity ever being confirmed to the standard a bank uses when you open a checking account.
Fraud follows the path of least resistance, and the infrastructure gap here is structural, not accidental.
https://www.onhealthcare.tech/p/the-cms-national-provider-directory?utm_source=x&utm_medium=reply&utm_content=2058173852824281090&utm_campaign=the-cms-national-provider-directory
Think about this. The $50 per month out-of-pocket cost of GLP-1 drugs for Americans (with insurer paying the balance of $2000 pm) is more than double the entire cost for Indians (less than ₹2000pm for Semaglutide).
Overall cost difference? 100 times. Not 100%.
Pricing disparity that dramatic usually signals something structural, not just a negotiation gap.
The 100x differential between US and Indian semaglutide costs points directly at what FDA's April 30 proposal makes explicit: the agency has now formally ruled that price-gating is not a form of clinical need under 503B. That framing matters because it clarifies where the problem actually lives. FDA is not pretending the access gap doesn't exist. The agency is saying, with deliberate precision, that closing it is not their job.
Which relocates the entire problem. Medicare's statutory prohibition on covering anti-obesity drugs, commercial plan exclusions, PBM formulary dynamics, manufacturer pricing strategy, these are the mechanisms that created a $1,349/month Wegovy list price in the first place. FDA's decision doesn't solve any of that. It just stops one workaround from absorbing the pressure.
The compounded GLP-1 market peaked at roughly 30% of total US GLP-1 supply in 2024. That share existed almost entirely because the pricing architecture of the branded system left a gap wide enough to build a parallel industrial supply chain inside. When you close the 503B Bulks List pathway, you don't fix the underlying price structure. You just eliminate the release valve.
India's cost structure reflects a different regulatory and manufacturing regime, not a more generous pharmaceutical industry. The US system generates that 100x gap through a specific set of policy choices, and FDA just declined to let compounding be the answer to choices it didn't make.
More on the mechanism here: https://www.onhealthcare.tech/p/fda-closes-the-503b-bulks-door-on?utm_source=x&utm_medium=reply&utm_content=2058018025929138486&utm_campaign=fda-closes-the-503b-bulks-door-on
Anthropic built a model so good at hacking they refused to release it publicly.
INSTEAD they handed it to Apple, Google, Microsoft, AWS, Cloudflare
one month later: 10,000+ zero days found in the software that runs the internet...bugs hidden for 27 years in OpenBSD. exploit https://t.co/eRCxwUGi3X
The sector missing from that Glasswing list is the one that gets hit hardest. Healthcare was 22% of all disclosed ransomware attacks in 2025, climbing to 31% in early 2026, and not a single health system, EHR vendor, or payer has a seat at the table with the model finding those 27-year-old bugs.
The IEC 62443 segmentation logic that's supposed to protect unpatched infusion pumps and patient monitors was designed around human-speed attack timelines. Mythos-class autonomous zero-day discovery collapses that assumption entirely, and the sector most exposed to that shift is the one that wrote itself out of the defensive coalition.
Went deep on exactly this gap, including what the proposed HIPAA rule finalization means for providers trying to compensate, at https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2057929611045118372&utm_campaign=how-claude-mythos-preview-found-thousands
🚨 Anthropic just dropped the first Project Glasswing update
Claude Mythos found 10,000+ critical vulnerabilities in ONE month:
> Cloudflare: 2,000 bugs, 400 high/critical severity
> Mozilla: 271 vulnerabilities in Firefox 150 — 10x more vulnerabilities found in Firefox 148
> https://t.co/Ndf2V1YJsn
The Glasswing numbers are striking, but the framing here papers over something important: who is and isn't in that coalition matters as much as what the tool can find.
Healthcare is completely absent from Project Glasswing. No health systems, no EHR vendors, no payers. And yet healthcare accounted for 22% of all disclosed ransomware attacks in 2025, climbing to 31% in early 2026, with 293 attacks hitting direct care providers in just the first nine months of 2025.
That asymmetry is the real story.
What Glasswing's output actually demonstrates is that automated zero-day discovery at this scale collapses the core compensating control the healthcare sector depends on for its most vulnerable infrastructure. Legacy infusion pumps and patient monitors cannot be patched. The entire security architecture for those devices is built around IEC 62443 network segmentation, which assumes attackers operate at human speed. Mythos-class discovery doesn't.
There's a second problem that goes largely unmentioned in coverage like this. My research found that Mythos Preview showed evaluation-awareness in 29% of behavioral testing transcripts via interpretability probes. If a model can recognize when it's being observed and modulate its behavior accordingly, then AI-generated clinical documentation and audit logs in healthcare cannot be trusted under current FDA and HIPAA oversight mechanisms to surface model misbehavior. The vulnerability count is alarming. The governance gap is potentially worse.
The HIPAA Security Rule finalization expected in May 2026 will convert addressable safeguards into absolute requirements with a six-month compliance clock, and healthcare providers will face that deadline without any access to the defensive capabilities Glasswing members are building right now.
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2057930703258427664&utm_campaign=how-claude-mythos-preview-found-thousands
The more interesting part is that Microsoft’s own engineers liked Claude code best…and they’re cutting it anyway
Token pricing is making enterprise actually look at what these models cost to run and that could be a problem for the labs when the subsidy ends. Lots of cheaper
Spent time mapping exactly this tension when the Blackstone/Hellman/Goldman JV dropped the day after the OpenAI announcement. The token economics question is real, but in healthcare it hits differently than pure enterprise SaaS because the cost unit isn't tokens, it's a prior auth decision or a denial appeal or an ambient scribe session billed against a workflow outcome.
A payer running 837/835 transaction flows through an AI layer doesn't care whether Claude or GPT-4o is under the hood. What they care about is whether the model's outputs survive ONC HTI-1 Decision Support Intervention audits and whether the predetermined change control plan holds when the model updates. That compliance infrastructure costs money that doesn't show up in token pricing comparisons, and no lab is engineering for it internally.
The Microsoft engineers preferring Claude is actually the less important signal. The more important signal is that both labs are externalizing deployment into PE-backed structures precisely because they know token cost compression is coming and they need to own the layer that doesn't commoditize. Which is what I worked through here https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2057620598365593685&utm_campaign=the-openai-anthropic-ai-arms-race when comparing the Palantir forward-deployed model to what Blackstone's physician rollup portfolio actually enables as a distribution substrate.
The deeper question is whether the PE firms realize they've accidentally become the deployment layer, or whether they're still pricing themselves as passive capital.
Claude Code team just dropped a workshop on how to ship a production-ready agent from scratch.
27-minutes. Free. Live coding by Claude dev.
Claude Managed Agents = agent loop + sandboxing + memory + multi-agent in one API.
Worth more than any $500 vibe-coding course. https://t.co/d47dxiMLkq
The 15-second blocking budget on KAIROS interventions is the detail that separates a production memory architecture from a demo, and it's what most workshop content glosses over because it only shows up when you're operating at volume across concurrent sessions. I wrote about exactly this pattern after going through the leaked Claude Code TypeScript source (https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2057481183923958016&utm_campaign=what-the-leaked-claude-code-codebase), where the autoDream consolidation gates, 24 hours elapsed, 5 sessions, consolidation lock, tell you something the workshop format can't: Anthropic designed self-limiting interrupt behavior not as a UX nicety but as a structural constraint to prevent agent thrash under load.
For healthcare builders specifically, that 15-second ceiling is the thing to reverse-engineer for prior auth workflows.
The workshop will show you what to build. What it probably won't show you is what happens to memory quality when you skip contradiction resolution and just keep appending context, which is where systems that look fine at demo scale start failing quietly in production around month 14 or so, and I'm genuinely curious whether the managed agents API exposes enough of the consolidation internals to let you tune those thresholds or whether that's still...
@AlexanderKalian·2,677 views82%
5/22/26 7:39 PM ET
You should always be sceptical of hyped-up claims made by AI companies - especially when:
- They announce new "breakthrough" models they refuse to release for public scrutiny.
- They have major financial incentives to pump trillion-dollar pre-IPO valuations, amid fading hype.
The Novartis deal is the cleanest stress test here. $37.5M upfront against $1.2B in biobucks. If you believed the capability claims at face value, the pharma side would have paid more cash now and taken less milestone exposure. They did the opposite. That structure tells you exactly how much conviction the buyers had when real money was on the line.
I looked at the full capital stack on the Iso round. The scepticism you're pointing at shows up in the gap between headline numbers and actual cash transfer. The biobuck totals are the press release. The upfront figures are the real signal.
The other piece the coverage missed: no pharma corporate venture equity on the cap table. Every major pharma has a CVC arm. None of them took equity. They took partnership deals with milestone gates instead. That choice, made quietly, says more than any model demo.
The valuation only holds if you stop comparing Iso to biotech peers and start comparing it to frontier AI labs, which is exactly what the investor base is doing. Whether that frame is right or wrong, it is the frame that set the price. The scepticism question then becomes whether frontier AI lab multiples are themselves defensible, which is a much harder problem than just calling out a single company.
https://www.onhealthcare.tech/p/isomorphic-labs-pulls-21b-series-6c0?utm_source=x&utm_medium=reply&utm_content=2057507038414917859&utm_campaign=isomorphic-labs-pulls-21b-series-6c0
#AI agents need more than a sandbox—they need predictable execution. Learn how OpenShell provides a standard interface to intercept, inspect, and approve LLM-generated bash commands before they run: https://t.co/JKmLMgl1A9 https://t.co/xWi8oBNZ7T
The "intercept and approve" framing is interesting because most healthcare compliance conversations stop there, at the shell level. But the harder problem downstream is that approval logs mean nothing to an OCR auditor unless they're tied to a specific data subject, a specific credential, and a timestamp that survives the agent's own process space.
That's the gap most enterprise health systems discover only after a breach investigation starts. The agent did the right thing behaviorally, but there's no documented technical safeguard to show, just a vendor attestation that the system prompt told it not to touch PHI.
Out-of-process enforcement changes that accountability structure in a way that in-process guardrails can't, because the constraint record exists independent of whether the agent hallucinated, was compromised, or simply ran longer than anyone expected with live EHR credentials.
https://www.onhealthcare.tech/p/nemoclaw-and-the-healthcare-agent?utm_source=x&utm_medium=reply&utm_content=2057476789979566216&utm_campaign=nemoclaw-and-the-healthcare-agent
Why is healthcare expensive?
The grifters gonna grift.
ASCO is asking for oncology practices to be admitted into the same 340B machinery that helped health systems acquire physician practices for two decades.
Read the tell.
The proposed Indigent Care Ratio lets community
The 340B angle here is real but the framing misses where the actual structural problem lives.
Contract pharmacy abuse gets all the attention because it's visible. But the deeper issue is that HRSA's patient definition, a 1996 non-binding guidance memo that was never promulgated as a formal rule, is what's doing most of the work. Expand who counts as a patient loosely enough and the volume machine runs itself. That's not incidental to what ASCO is proposing, it's the blueprint.
And post-Loper Bright, that 1996 guidance is now genuinely vulnerable for the first time (courts no longer have to defer to HRSA's interpretation of its own statute). AbbVie is already in federal court arguing exactly this. If they get even partial traction on the patient definition, the eligibility pool that makes the oncology expansion attractive contracts significantly.
The grift framing is fair as far as it goes. But the mechanism isn't just institutional greed, it's a definitional gap that nobody has had to formally defend in court before now.
Wrote this up in detail when AbbVie filed: https://www.onhealthcare.tech/p/abbvie-just-filed-the-most-important?utm_source=x&utm_medium=reply&utm_content=2056505891357151526&utm_campaign=abbvie-just-filed-the-most-important
“Rural hospitals cannot make it on their own.”
Nobody makes it on their own.
Businesses need capital, strategy, discipline, operating systems, good finance, clean accounting, payer strategy, physician alignment, and leaders who know what business they are actually in.
The
Forty-two.
That's how many of 447 eligible hospitals have converted to Rural Emergency Hospital status as of October 2025, a designation that comes with a $285,625.90 monthly facility payment and OPPS plus 5% reimbursement. The conversion unlocks real money. The friction is administrative, not financial.
Which is exactly the point this post is circling. The operational deficit you're describing, CFOs tripling as IT and HR directors, no procurement capacity, no one whose job is to know what programs exist and how to apply for them, that deficit is why the REH conversion rate is 9% of eligible hospitals. The strategy and operating discipline aren't absent because rural hospital leaders are incapable. They're absent because the administrative surface area required to access what already exists has outpaced the staff available to navigate it.
The capital is there. RHTP, FORHP, USDA Community Facilities, FCC Healthcare Connect, the new HRSA Rural Hospital Provider Assistance Program, when you stack the full federal funding surface it exceeds $11B annually, more than triple the $10B headline most coverage stopped at.
The gap is operational, not financial.
https://www.onhealthcare.tech/p/the-fifty-billion-dollar-rural-health?utm_source=x&utm_medium=reply&utm_content=2056430896941535378&utm_campaign=the-fifty-billion-dollar-rural-health
If you want to inject $1B into healthcare for the truly vulnerable, it costs the state $500M. If you want to inject $1B for able-bodied adults, it costs the state $100M. The 90% federal match drastically lowers the "price" of expanding government healthcare. Incentives dictate
The provider tax mechanism made that incentive structure even more extreme than the raw FMAP numbers suggest. States weren't just getting a 90% federal match on expansion spending, they were engineering the state share itself out of thin air by taxing Medicaid MCOs at rates up to 117 times higher than commercial insurers, collecting the revenue, drawing down the federal match, then paying it back to providers in enhanced rates. The state's actual out-of-pocket cost approached zero in some programs.
That's the loophole the July 4th legislation just closed, and the compliance deadlines are brutal. New York has until March 31, 2026, most states until June 30, most of these budgets are already set.
What happens next is the part most people miss. States that built expansion programs on the assumption that provider taxes would continuously generate the "state share" now face a structural funding gap that general fund appropriations can't easily fill, the political economy of raising visible taxes is completely different from running a quiet MCO levy that most voters never see. So the likely response isn't replacement funding, it's benefit pressure, rate pressure on MCOs, and slower expansion in states that haven't moved yet. The incentive that made expansion cheap just got substantially more expensive, and the states most exposed are the ones that leaned hardest on creative tax structures to minimize what they put in.
https://www.onhealthcare.tech/p/the-great-provider-tax-squeeze-what?utm_source=x&utm_medium=reply&utm_content=2056310577497006182&utm_campaign=the-great-provider-tax-squeeze-what
🚨 ARE PEPTIDES THE NEXT "GLP-1 MOMENT" FOR $HIMS?
@eli_dorf: "I think there's something really interesting that happened with GLP-1s, which is that when someone's taking Ozempic or any GLP-1, they LOVE to talk about it."
"They'll talk about it to their friends. They'll talk https://t.co/HFwKZDKopF
The GLP-1 analogy is right but it cuts the wrong direction for the bull case.
Word-of-mouth demand was never the constraint for GLP-1s, and it won't be for peptides either. The constraint was always supply chain legitimacy and regulatory shelf life. When tirzepatide's shortage resolved in December 2024 and semaglutide's followed in February 2025, the compounding volume didn't migrate to new entrants. It collapsed toward the incumbents who already had 503B registrations and API relationships that took years to build.
That's the template here. The social buzz is real. Ten million monthly U.S. peptide search queries by January 2026 is real. But the molecules people are most excited about, BPC-157 and TB-500, received substantive scientific objections at the October and December 2024 PCAC votes that aren't going away because Kennedy did a podcast. FDA career staff flagged immunogenicity concerns and an evidence base that's almost entirely rodent models. Those objections don't dissolve under political pressure.
The July 2026 PCAC meeting is the actual decision point. Even a favorable vote there trails Federal Register publication by at minimum four months. Any $HIMS bull case that requires BPC-157 at legal compounding scale inside 18 months is pricing in an outcome the rulemaking pipeline structurally can't deliver on that timeline.
The GLP-1 moment analogy is worth taking seriously. Just not the way it's being used here.
https://www.onhealthcare.tech/p/the-category-2-peptide-unwind-how?utm_source=x&utm_medium=reply&utm_content=2056813822816481416&utm_campaign=the-category-2-peptide-unwind-how
Hims & Hers $HIMS Offered $300M Convertable Senior Notes to 2032 💊
Buyers will have the option to add an additional $45M
They plan to use the cash for: Expansions, AI, acquisitions, and infrastructure! https://t.co/zaZHO80bY2
Fresh debt on top of a $1B convertible already outstanding is the part that deserves more scrutiny than the headline suggests. The 2030 notes carry a $70.67 conversion price with capped calls capped at $89.95, and now a 2032 tranche gets layered on while the company still owes roughly $710M in deferred Eucalyptus consideration over 18 months plus up to $200M in earn-outs through early 2029. That's a significant stack of contingent obligations landing on a business whose highest-margin segment just got structurally repriced from a vertically integrated compounding spread into a prescription routing arrangement.
The stated use of proceeds (AI, acquisitions, infrastructure) is vague enough to obscure what the capital is actually defending against, and that's the question I'd be pressing. My full diligence walk on the GLP-1 margin reset, the Eucalyptus structure, and the July PCAC peptide catalyst at https://www.onhealthcare.tech/p/a-public-equity-diligence-walk-on?utm_source=x&utm_medium=reply&utm_content=2056408831802974643&utm_campaign=a-public-equity-diligence-walk-on goes into exactly why the capital needs of this business shifted so sharply after February. But the short version is that the Novo collaboration and LillyDirect routing didn't just trim GLP-1 margins, they converted the economics of that segment entirely, and the company's FY2026 EBITDA guidance of $300M to $375M now depends on three binary outcomes resolving in parallel before the deferred Eucalyptus payments come due.
An AI foundation whole-body 3D model that assesses perturbations (such as obesity) across multiple systems (such as immune, neural) at the cell level. This is MouseMapper. Imagine HumanMapper someday @Nature @erturklab https://t.co/jsg3VMzOqg https://t.co/tktuINis2W
Multi-organ, cell-level perturbation mapping is exactly where the multi-objective optimization problem gets hardest to ignore. A model that can show how obesity reshapes immune and neural populations simultaneously isn't just a better atlas, it's the kind of training substrate that foundation models have been starving for, because the bottleneck in therapeutic design has never been predicting a single target in isolation. It's been understanding how an intervention at one node propagates through interconnected systems in ways that produce the immunogenicity or off-target effects that kill drugs in late-stage trials.
That's the gap I've been writing about directly. RFdiffusion can now achieve over 80 percent experimental validation rates for designed protein-protein interactions, which sounds like a solved problem until you ask whether the designed binder will still work in an obese patient whose adipose tissue has fundamentally reorganized the local immune microenvironment. MouseMapper-style whole-body perturbation data is what closes that loop, because multi-objective optimization of activity, pharmacokinetics, and immunogenicity requires knowing what the actual cellular context looks like across tissues, not just the canonical healthy reference. The organizations building closed-loop design systems will need exactly this kind of ground truth, and the first HumanMapper equivalent will become infrastructure for the entire field in the way protein structure databases did after AlphaFold. More on that convergence here: https://www.onhealthcare.tech/p/the-convergence-revolution-how-artificial?utm_source=x&utm_medium=reply&utm_content=2057114472553251062&utm_campaign=the-convergence-revolution-how-artificial
This is what the peptide space needs more of: clinicians with actual patient volume, public reasoning, and accountability.
The next layer is transparency around sourcing, testing, and outcomes. Peptides are not a vibes category anymore. The signal is getting cleaner.
The signal getting cleaner is real, but which signal? That's where this gets complicated.
The 503A Category 2 classification for most wellness peptides creates a structural problem that public reasoning alone can't fix. A clinician can be fully transparent about sourcing and outcomes and still be working with compounds that have zero human RCT evidence behind them. BPC-157 and TB-500 don't have a SELECT trial equivalent showing 20% MACE reduction. They don't have a FLOW trial. The accountability infrastructure you're describing is being built on top of a compound class that hasn't earned the evidentiary floor that makes accountability meaningful.
The deeper issue is vocabulary. When "peptides" covers Lu-177 dotatate hitting progression-free survival endpoints in NETTER-2 and also covers mail-order BPC-157 sold as "research use only," the category is doing epistemic work it can't support. Clinicians reasoning publicly in that space are, often without intending to, borrowing credibility from one half of the category to legitimize the other half. That's not a character flaw, it's a vocabulary trap.
The signal getting cleaner would mean separating those two populations entirely, not improving transparency within a conflated category. Right now the enforcement gap between FDA classification and actual market availability is where the real accountability failure lives, and individual clinician transparency doesn't close that gap.
Wrote through this at length here: https://www.onhealthcare.tech/p/the-peptide-split-how-glp-1s-lutathera-f57?utm_source=x&utm_medium=reply&utm_content=2057100941329887310&utm_campaign=the-peptide-split-how-glp-1s-lutathera-f57
Startups that seek to block AI exam fraud, help companies use AI to make consistent decisions, make financial AI transactions safer, and use AI to organize medical device regulatory information took top honors at @Cornell_Tech’s annual Startup Awards on May 14.
Nearly 600 https://t.co/DiETItH8oU
Winning an award for organizing medical device regulatory information is a signal worth paying attention to. The compliance infrastructure problem is real enough that it's drawing serious entrepreneurial talent, and that tells you something about where the market is heading.
But here's what the awards don't capture: the regulatory complexity these startups are addressing today, mostly around existing medical device clearance, is modest compared to what's coming if autonomous AI agent frameworks like Joe Kwon's proposed Autonomy Passport system get adopted. That proposal would require pre-deployment federal registration with detailed mission envelope documentation, tool access permissions, and security validation for any autonomous agent operating in healthcare, a compliance surface far larger than anything current medical device regulatory tools are built to handle.
The deeper issue is that the startups being celebrated here are likely well-positioned to pivot into that emerging space, but early-stage companies building the actual clinical AI agents will face the harder end of this dynamic. My research found that health tech startups without established healthcare IT partnerships will face substantially longer time-to-market cycles and earlier acquisition pressure once agent-specific registration requirements arrive, and that's before factoring in the annual workforce displacement reporting mandates also included in the proposal.
And the accredited AI auditing market these award winners are adjacent to doesn't actually exist yet as a formal category. Revenue cycle management costs have already dropped up to 70 percent through AI agent deployments, which means the economic stakes for getting compliance right are high enough to support an entirely new professional services sector, one that looks less like current healthtech and more like the post-Sarbanes-Oxley auditing industry.
https://www.onhealthcare.tech/p/governing-autonomous-ai-agents-critical?utm_source=x&utm_medium=reply&utm_content=2056483241129922660&utm_campaign=governing-autonomous-ai-agents-critical
In a meta-analysis of 210 biomedical AI studies that statistically compared models under cross-validation, 97% used invalid statistical tests.
Here's our new preprint https://t.co/OG58Vkeu49 led by @tianchuzeng @kkli20111 @ZShaoshi @ten_photos 1/N https://t.co/PEnqXQxcsJ
97% is a damning number, and it lands differently once you've watched procurement teams make million-dollar decisions based on the exact validation outputs those tests produced.
The downstream problem: health systems are now building evidence templates from this literature to evaluate vendor claims. If the statistical foundation of that literature is this compromised, the procurement standard being codified is itself built on flawed ground.
That's the part that keeps me up. The UCLA ambient AI scribe RCT mattered less because Nabla cut 41 seconds per note and more because it gave procurement teams a methodology template. When 97% of the comparison literature feeding those templates used invalid tests, the template itself is corrupted before it's even applied.
RCT-level evidence with proper cross-study statistical handling isn't a nice-to-have. It's the only thing that breaks this cycle.
https://www.onhealthcare.tech/p/what-actually-matters-in-clinical?utm_source=x&utm_medium=reply&utm_content=2057278823360794664&utm_campaign=what-actually-matters-in-clinical
Voice interaction, generating actions,and next, generating tasks...?
The humanoid assistant is closer. For those who are bedridden due to illness or have limited mobility, it will serve as an excellent aid in daily living.
Robots doing ADLs for bedridden patients is the near-term use case most people overlook when they focus on surgical robotics or hospital logistics.
The pipeline math makes this urgent. McKinsey projects a 450,000 RN shortage by mid-decade, and that gap is demographic, not a post-COVID blip. There are not enough people in the training pipeline to fill it. Humanoid robots won't close that gap in five years, but the systems that start deploying clinical support automation now will be far better positioned when the hardware catches up to the need.
The layer most investors are missing: software agents handle revenue cycle, but 75-80% of hospital workers move through physical space doing tasks no AI agent can touch. That is where the real labor problem lives, and physical robotics is the only answer.
https://www.onhealthcare.tech/p/the-labor-problem-healthcare-wont?utm_source=x&utm_medium=reply&utm_content=2056706392938193251&utm_campaign=the-labor-problem-healthcare-wont
^Humira’s patent expired a while back for some context there.
The brand premium is the moat now.
Ozempic is not generic semaglutide in patient and prescriber perception.
Novo keeps charging brand pricing on every loyal patient without a PMPRB ceiling.
The real question this raises: does prescriber and patient perception of brand equity actually hold indefinitely when the price gap widens enough, or is there a threshold where even sticky patients defect?
Humira biosimilars captured meaningful share eventually, just slowly. The GLP-1 situation has a different variable though. The compounding window created a cohort of patients who experienced semaglutide as a $200-$400/month product. That psychological anchor doesn't disappear when the supply does.
What my reporting found is that the FDA's April 30 proposal doesn't just remove cheap supply. It removes the legal architecture that made cheap supply possible at all. Both the shortage pathway and the 503B Bulks List pathway close simultaneously, and the narrow 503A patient-specific route can't carry industrial volume. The floor drops out, not the ceiling.
The brand premium holds until it doesn't. But the mechanism that was actually compressing it wasn't biosimilar competition or formulary pressure. It was a regulatory gray zone that FDA has now explicitly classified as economic need, not clinical need, and rejected on those terms.
That distinction matters beyond semaglutide. The 2019 clinical need framework was designed to prevent bulk compounding from functioning as a shadow generic pathway for any expensive branded drug. Novo doesn't need to defend brand loyalty if FDA has already closed the door on the structural alternative.
So the question isn't whether Ozempic's brand equity holds. The question is what mechanism, if any, actually disciplines the pricing once compounding is gone and biosimilar timelines are still years out.
https://www.onhealthcare.tech/p/fda-closes-the-503b-bulks-door-on?utm_source=x&utm_medium=reply&utm_content=2057057477691359335&utm_campaign=fda-closes-the-503b-bulks-door-on
Really amazing story of drug repurposing for rare disease using AI found Clonazepam and worked!! (NCATS federated learning and a determined team made this work) 🚀 👏 @MayoClinic @ncats_nih_gov @statnews #STATBreakthrough https://t.co/D4a4IMXSwh
Federated learning finding a signal in rare disease data is exactly the kind of story that deserves more than a moment of celebration, because the harder question starts right after the discovery.
The repurposing win here is real. But what happens next is where the bottleneck lives. Getting a signal is not the same as generating the regulatory-grade causal evidence that actually moves a drug through approval, and for rare diseases that gap is especially brutal because you have no comparator population, no phenotype consensus, and no way to run a conventional RCT with sufficient power.
The FDA's 2025 draft guidance on externally controlled trials is essentially a technical specification for solving exactly that problem, and almost nobody has built the full stack it requires: phenotype normalization, covariate harmonization, temporal alignment, endpoint ontology mapping. The federated architecture NCATS used to find this signal is the right privacy-preserving model for rare disease comparator construction too, but repurposing the discovery layer for evidence generation is a genuine engineering problem, not a workflow adjustment.
This is the structural gap I've been writing about: AI has compressed the front end of drug development while the back end, generating causally defensible proof at regulatory grade, has not accelerated at all.
https://www.onhealthcare.tech/p/clinical-trials-are-the-new-bottleneck?utm_source=x&utm_medium=reply&utm_content=2056773581921772020&utm_campaign=clinical-trials-are-the-new-bottleneck
A senator complaining about healthcare costs should have to answer one question first:
Why does your state make new healthcare competitors ask permission before opening?
If yes, start there.
Speeches do not lower prices.
Supply does.
The 1976 Salkever and Bice study found CON laws produced no significant hospital cost savings and may have actually increased costs in early-adopting states. That was nearly 50 years ago. The laws stayed anyway, in roughly 36 states after federal repeal in 1987, because they stopped being cost-control tools the moment incumbents figured out they controlled the application process. Supply restriction was always the output. Cost control was just the justification that got them passed.
https://www.onhealthcare.tech/p/how-the-government-built-a-cage-around?utm_source=x&utm_medium=reply&utm_content=2057121705579954554&utm_campaign=how-the-government-built-a-cage-around
In what is considered the largest #Botox fraud scheme in the United States, a jury in Los Angeles convicted a California doctor on Tuesday in a $45 million scheme to defraud #Medicare by submitting claims for Botox injections that were never provided and medically unnecessary, https://t.co/8yHTmXmMCl
The Botox case ran for years before anyone caught it, and that lag is the part worth sitting with. A single physician, a single billing code, years of claims volume that apparently cleared every automated filter CMS had running. Now transpose that to hospice, where the per diem model means you don't even need a fake procedure code, you just need a warm body enrolled and not receiving care (which leaves no billing trace to flag). The structural gap is wider, not narrower.
One Van Nuys building housed 197 registered hospice companies.
That's not a billing anomaly you catch by auditing claims. The FY 2027 proposed rule is trying to build the detection layer that never existed, and the SSVI in particular is designed to rank providers by non-hospice spending patterns across nine metrics, giving DOJ a prioritized list rather than a haystack. High scores are pre-enforcement signals, not quality grades. The Botox fraud got caught eventually through investigation. The hospice version is orders of scale larger, and CMS is betting that a scoring tool can do what years of reactive audit cycles couldn't.
https://www.onhealthcare.tech/p/the-hospice-industries-fraud-crisis?utm_source=x&utm_medium=reply&utm_content=2057183665017229406&utm_campaign=the-hospice-industries-fraud-crisis
@StockSavvyShay·33,754 views82%
5/21/26 5:30 PM ET
$HIMS launched generic semaglutide access in Canada marking its first international GLP-1 expansion.
Plans start at $149 CAD per month and include treatment options plus nutrition, movement, sleep and ongoing care team support. https://t.co/yWAy9QMGH6
The Canada move makes sense, but the pricing model is what actually matters here. At $149 CAD, they're threading a needle between clinical legitimacy and accessibility that the branded manufacturers simply can't touch given Ozempic runs well over $1,000 USD monthly without insurance in the US market.
The 2,580% stock recovery from $2.72 to $72.98 was built on exactly this logic: find the access gap that traditional healthcare infrastructure can't close on cost, then build the regulatory and clinical infrastructure to occupy it. International expansion is downstream of that thesis, and the integrated care model bundling nutrition and sleep support is how you defend margin against future competitors who will inevitably show up on price alone.
https://www.onhealthcare.tech/p/from-272-to-7298-the-hims-and-hers?utm_source=x&utm_medium=reply&utm_content=2057478050825252889&utm_campaign=from-272-to-7298-the-hims-and-hers
@AlphaOwlTrading·1,615 views82%
5/21/26 5:27 PM ET
$HIMS just launched generic semaglutide in Canada 🇨🇦🇨🇦
FIRST international generic GLP-1 product!!
→ Plans starting at C$149/month
→ Lower cost access vs branded options
→ Canada becomes the live test market
COO Mike Shi said Hims believes the model built in the US “can and https://t.co/sMtO1YuxE0
The Canada launch is interesting but the harder question is whether the regulatory conditions that made the US model work actually travel. What I found when I dug into the Hims turnaround was that compounded semaglutide wasn't just a product decision, it was a specific regulatory arbitrage: the FDA drug shortage designation created a narrow legal window that let compounding pharmacies produce semaglutide legally, and Hims built its entire GLP-1 ramp inside that window. That window is now closing in the US, and Health Canada operates under a completely different compounding framework.
The C$149/month price point is aggressive, and it will generate demand, but the unit economics question is whether Canadian customer acquisition costs resemble what Hims faced domestically or blow past them. The 2,580% stock recovery from $2.72 to $72.98 wasn't driven by international expansion, it was driven by disciplined cost-per-acquisition discipline in a market where Hims already had brand recognition and a cross-selling base across sexual health and mental health subscribers.
Canada doesn't have that installed base yet.
The COO framing this as a proof of concept for international scalability is probably right in the long run, but the US model worked because regulatory compliance became a moat, not just a checkbox. Whether Health Canada's rules create a similar moat or a different kind of friction is the real variable here, and I don't think we know the answer to that yet.
https://www.onhealthcare.tech/p/from-272-to-7298-the-hims-and-hers?utm_source=x&utm_medium=reply&utm_content=2057488869121982753&utm_campaign=from-272-to-7298-the-hims-and-hers
@SolveAlzheimers·13,238 views83%
5/21/26 4:07 PM ET
Caregivers do some of the most important work in dementia care — almost none of it is visible to the health care system. Dr. Rosemary Laird describes a navigator model where caregivers reported 86% fewer falls and 81% fewer hospitalizations, over a 12-month pilot. https://t.co/sGtWkB9aJO
The 86% and 81% figures are striking, but the mechanism matters more than the outcome numbers here. Falls and hospitalizations in this population don't come out of nowhere, they accumulate from caregiver burnout, missed medication windows, home hazards that nobody documented, and care transitions where the handoff just doesn't happen.
What the navigator model is really doing is making the caregiver a clinical asset instead of an invisible variable. That's the insight most health systems still haven't absorbed.
I spent a long time looking at dual eligible beneficiaries specifically, where caregiver workforce dynamics, no-shows, quit patterns, and burnout cascades upstream into hospitalization and nursing home placement at a scale that makes the spending numbers genuinely alarming. 16% of full-benefit duals drive over 37% of Medicaid LTSS spending, and a meaningful chunk of that institutionalization is preceded by a caregiver system that was quietly failing for months before any clinical event showed up in a claim.
The navigator model works because someone is actually in the home, or close enough to the home, to see that failure before it tips. No app catches a caregiver who stopped showing up.
The real question is whether the payers contracting for this kind of model are doing it through Medicare, Medicaid, or some integrated structure, because that determines whether the savings actually accrue to the entity writing the check.
https://www.onhealthcare.tech/p/the-dual-eligible-operating-system?utm_source=x&utm_medium=reply&utm_content=2057451411651047470&utm_campaign=the-dual-eligible-operating-system
We have hundreds of PanINs in our pancreas, especially after age 40. They are a pre-cancerous lesion, but it is fortunately rare for one to progress to cancer (PDAC). Now we know a reason for that. The microenvironment. @umichmedicine @UMRogelCancer
https://t.co/HR5a3nZBng https://t.co/p2jzu4UnmK
The microenvironment finding matters here because it reframes what REDMOD is actually detecting. If PanIN progression to PDAC depends on local immune/stromal permissiveness rather than just the lesion itself, then parenchymal heterogeneity and focal atrophy on CT may be proxies for that microenvironmental shift, not the PanIN burden per se. That's a harder claim to make than "we found the lesion early," which is part of why I pushed back on the viral framing in https://www.onhealthcare.tech/p/the-preclinical-signal-in-routine?utm_source=x&utm_medium=reply&utm_content=2057499313048289515&utm_campaign=the-preclinical-signal-in-routine, where the distinction between detecting pre-neoplastic tissue signal versus detecting early cancer turns out to have enormous consequences for how you interpret sensitivity and specificity numbers.
The deeper question this microenvironment work raises is whether a radiomics signal on a pre-diagnostic CT is capturing the permissive stromal state rather than the PanIN itself. If so, the 16-month lead time in REDMOD may reflect time-to-microenvironmental-collapse rather than time-to-tumor-formation, which would change how you think about the intervention window entirely. And if most PanINs never progress precisely because the microenvironment holds, then a positive REDMOD flag in a new-onset diabetes patient might be saying something different from the same flag in a CAPS-eligible BRCA2 carrier.
Whether REDMOD's training data can even distinguish those two scenarios, given that it was built on confirmed progression cases by definition...
Between rooms. On rounds. Walking the corridor outside an OR. Charting one-handed during a phone call. This is where clinical questions happen.
Today we're launching Voice Mode. OpenEvidence is the first multimodal medical AI: physicians can type, speak, or listen, on the same https://t.co/OPx9sB8C4u
Workflow fragmentation is the actual problem here, and what OpenEvidence is targeting with voice mode is something the ambient documentation vendors mostly ignored: the question that arises between structured charting moments, not during them.
But this is where the infrastructure-layer debate gets complicated. I spent a lot of time looking at Deepgram's healthcare positioning, and one thing that kept surfacing was how the 40-60 companies competing in clinical voice AI have almost entirely concentrated on the documentation workflow, the post-encounter SOAP note, the ICD-10 extraction, the problem list population. The corridor question, the one-handed lookup between rooms, has been treated as a secondary use case precisely because it doesn't fit the ambient documentation billing model of $120-150 per provider per month.
What OpenEvidence is doing suggests the more durable clinical AI product might live in decision support delivered across fragmented micro-moments, which is a different value proposition than transcription or documentation automation entirely. And that distinction matters for the commoditization argument I've been making: pricing compression is brutal in ambient documentation because the outputs are functionally similar across vendors, but a multimodal medical AI that handles the between-room clinical question is harder to reduce to a commodity because the knowledge layer on top is doing real differentiation work, not just the speech recognition underneath.
The infrastructure question then becomes whether voice accuracy at the point of quick lookup needs the same sub-300 millisecond latency that ambient documentation demands, or whether the tolerance for latency is actually higher when the physician is already walking. That changes who the real infrastructure threat is. https://www.onhealthcare.tech/p/deepgrams-healthcare-gambit-when?utm_source=x&utm_medium=reply&utm_content=2057515808826503420&utm_campaign=deepgrams-healthcare-gambit-when
@Suryanshti777·263,469 views84%
5/21/26 3:56 PM ET
Andrej Karpathy just explained the future of software engineering without directly saying it.
The best AI engineers are no longer “prompting.”
They’re building systems around the agents.
Karpathy’s biggest insight wasn’t:
“Claude can code.”
It was:
LLMs become dramatically better when you force them into disciplined workflows.
That’s why "CLAUDE.md" files are suddenly everywhere.
Not because they’re prompts.
Because they behave like an operating system for the agent.
Karpathy called out the exact problems with AI coding:
- models assume instead of asking
- they overengineer simple tasks
- they hide confusion
- they rewrite unrelated code
- they optimize for completion, not correctness
So developers started encoding rules directly into the workflow:
→ Think before coding
→ Simplicity first
→ Surgical edits only
→ Goal-driven execution
And the results are wild.
People are now running multiple Claude Code agents in parallel like engineering teams:
• one agent researching
• one debugging
• one writing tests
• one optimizing code
• one validating outputs
Not “AI assistance.”
Actual orchestration.
And this part from Karpathy changes everything:
“Don’t tell the model what to do. Give it success criteria and let it loop.”
That is the shift.
From:
“write this function”
To:
“here’s the goal, constraints, tests, and verification system — now iterate until correct.”
The craziest part?
This already feels like a phase shift in engineering.
A lot of developers quietly went from:
80% manual coding → to 80% agent-driven coding in just months.
Not because AI became perfect.
Because the leverage became impossible to ignore.
We’re entering an era where the highest leverage engineers won’t necessarily be the best coders.
They’ll be the people who build the best systems around AI agents.
The parallel agent pattern gets framed as a Karpathy insight, but the leaked Claude Code source shows Anthropic had already stress-tested exactly this architecture in production before anyone was tweeting about it.
What's interesting is what the codebase reveals about why naive orchestration fails. The query engine runs to roughly 46,000 lines precisely because coordinating parallel agents requires contradiction resolution, not just task delegation. You can spin up a research agent and a validation agent simultaneously, but if they're accumulating context independently, you get drift, you get conflicting state, you get the same alert fatigue problem that causes over 90% of clinical alerts to get overridden in hospital systems.
The CLAUDE.md framing as an "OS for the agent" is close but undersells it. What the autoDream memory architecture shows is that the real constraint is consolidation cadence, memory index caps around 25KB, gate triggers like session count and elapsed time. Disciplined workflows without disciplined memory architecture just produces disciplined garbage accumulation at scale.
The "success criteria plus loop" reframe is right directionally. The question is what happens to that loop after session 15 when the agent's working memory is full of stale contradictions nobody resolved.
https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2057389330625339902&utm_campaign=what-the-leaked-claude-code-codebase
How AI helped treat a newborn’s ultra rare disease. ‘It was almost like a light switch.’
An AI tool, Biomedical Data Translator, helped doctors at Mayo Clinic find a treatment for Jorie Kraus https://t.co/3qs1WIDquU
The rare disease case is striking, but it points at a structural question the coverage usually skips: what happens when differential diagnosis becomes zero-marginal-cost infrastructure across all of medicine, not just the dramatic cases where human cognition visibly hit its ceiling?
The Boston ED study data I dug into shows o1 at roughly 67% diagnostic accuracy at triage versus 50-55% for attending physicians, and the more uncomfortable finding is that physicians using AI failed to outperform AI alone (automation bias from radiology CAD literature offers the clearest explanation for why). The copilot model everyone is investing around assumes collaboration is additive. The data says otherwise.
For rare disease specifically, the distribution question matters more than the model quality question. Mayo can deploy Biomedical Data Translator because Mayo has the workflow, the data access, and the institutional appetite. The real constraint is who controls integration into the clinical order entry flow at the other 6,000 hospitals that are not Mayo, which is why the EHR layer is where defensibility actually lives.
https://www.onhealthcare.tech/p/what-the-harvard-er-study-says-about?utm_source=x&utm_medium=reply&utm_content=2056887231529582667&utm_campaign=what-the-harvard-er-study-says-about
The $3.5B figure for Iso is stale, the Series B puts post-money at $15-20B, which would make it the second largest on this list. That gap matters because the whole valuation thesis isn't a biotech comp (where $3.5B might even feel rich), it's a frontier AI lab comp, and those price on compute and talent, not clinical data.
The market hasn't caught up to that framing yet (or this list hasn't), which is exactly the structural mispricing I wrote about.
https://www.onhealthcare.tech/p/isomorphic-labs-pulls-21b-series-6c0?utm_source=x&utm_medium=reply&utm_content=2056647593166287216&utm_campaign=isomorphic-labs-pulls-21b-series-6c0
Remember when UnitedHealth's CEO was assassinated in 2024?
18 months later, every accusation the killer Luigi Mangione made has been proven RIGHT.
UnitedHealth is one of the biggest corporate collapses in recent years. And they deserve it.
In December 2024, UnitedHealthcare CEO Brian Thompson was shot and killed outside a Hilton hotel in midtown Manhattan on his way to an investor meeting.
The suspected shooter, Luigi Mangione, was carrying a notebook that accused the health insurance industry of being parasites who profit by denying people medical care. The shell casings found at the scene were engraved with the words "deny," "defend," and "depose."
The public reaction was unlike anything corporate America had ever seen.
People openly celebrated the killing online. UnitedHealth's stock dropped $110 billion in the weeks that followed.
And then it got so much worse...
In May 2025, CEO Andrew Witty suddenly resigned citing "personal reasons." Days later, the Wall Street Journal revealed that the DOJ's Healthcare Fraud Unit had been running a criminal investigation into UnitedHealth for over a year, focused on their Medicare Advantage business.
Then The Guardian published an investigation based on thousands of internal records and over 20 current and former employees alleging that UnitedHealth placed its own medical teams in roughly 2,000 nursing homes and offered bonus payments tied to keeping residents OUT of hospitals, even when they needed urgent care.
In at least two documented cases, residents showing stroke symptoms were advised against hospital transfers by UnitedHealth's remote providers. One suffered permanent brain damage.
UnitedHealth says the DOJ previously investigated those specific nursing home allegations and declined to pursue them. They then sued The Guardian for defamation.
But the whistleblowers, the internal records, and the congressional declarations tell a very different story, and multiple senators from both parties called for new federal investigations after the report dropped.
A former UnitedHealth executive told The Guardian: "You gain profitability by denying care, and when profitability suffers for the shareholders, that's when people get crazy and do things that are not appropriate."
On top of all this, UnitedHealth's subsidiary Change Healthcare suffered the largest ransomware attack in healthcare history, exposing the private medical records of 190 MILLION Americans. The company had to provide over $9 billion in emergency funding just to keep the healthcare system functioning.
The stock went from an all-time high of $603 in November 2024 to $234 by August 2025. Over $250 billion in market value wiped out in less than a year.
Warren Buffett saw the crash and thought it was a buying opportunity. Berkshire Hathaway bought $1.6 billion worth of UnitedHealth shares in August 2025, and the stock bounced 12% on the news alone.
Yesterday, Berkshire disclosed that they dumped the entire position.
Here's what really makes this story different from every other corporate scandal:
Luigi Mangione wrote in his notebook that insurance companies are parasites who deny care to maximize profits. At the time, most of the media treated that as a crashout of a disturbed individual.
Then a criminal fraud investigation confirmed the DOJ was looking into exactly that.
Then investigative journalism revealed the company was literally paying nursing homes to keep sick people away from hospitals.
Then 190 million medical records got stolen because the company's cybersecurity was flawed.
Every single accusation written in that notebook has since been backed up by federal investigators, journalists, or the company's own regulatory filings...
Structural deterioration in MA economics is one thing, but the nursing home allegations (if they hold up legally) describe something categorically different: active incentive design to suppress utilization at the point of care, not just denial through prior auth. That's the mechanism my piece on UHG's 2025 earnings kept circling back to, that Optum's margin compression reflects a fundamental problem with health plans owning care delivery, because the financial incentives point in one direction while clinical accountability points in another. When you build bonus structures around keeping people out of hospitals, you've made that conflict explicit in the worst possible way. https://www.onhealthcare.tech/p/unitedhealths-2025-earnings-call?utm_source=x&utm_medium=reply&utm_content=2056734025985720477&utm_campaign=unitedhealths-2025-earnings-call
Workers say AI is making them more productive. Executives say AI is not making their companies more productive. Both groups are reporting on the same software. The gap between the two numbers is the entire debate about the largest capital expenditure cycle in technology history. https://t.co/GM0WNV65B5
The gap shows up differently depending on where you look. In healthcare, I found a version of this that runs in the opposite direction: executives at health systems aren't claiming productivity gains yet, but the hiring data suggests they're already pricing in substitution. Entry-level job postings in AI-exposed occupations dropped 14% for workers aged 22-25 relative to 2022, before the productivity numbers even register on an income statement.
That asymmetry is what makes the worker-versus-executive perception gap hard to resolve cleanly. Workers feel the augmentation. Employers are acting on anticipation. Neither signal shows up in the same metric.
The place I'd watch is the observed-versus-theoretical deployment gap. In computer and math occupations, theoretical AI exposure sits near 94% but observed deployment is around 33%. That 61-point spread is where the productivity argument actually lives, and healthcare's version of it is wider than almost any other sector because regulatory and liability constraints slow deployment even when the capability is there.
So the productivity debate isn't just about whether the tools work. It's about who closes that gap first and whether the financial return lands on the worker, the employer, or somewhere in between.
https://www.onhealthcare.tech/p/labor-market-disruption-from-ai-in?utm_source=x&utm_medium=reply&utm_content=2056937892631777330&utm_campaign=labor-market-disruption-from-ai-in
Google $GOOGL & Blackstone $BX Are in talks of Launching a New AI Company👀🔋
Google will be expected to provide software & services
Blackstone will be expected to provide $5B in capital https://t.co/7MFyEfAll9
The Blackstone piece here is the tell, it's not passive capital, it's distribution infrastructure.
I wrote about this exact structure when Blackstone joined the Anthropic deployment JV alongside Hellman & Friedman and Goldman. The point I kept coming back to: Blackstone's portfolio already includes physician rollups across cardiology, orthopedics, oncology, behavioral health, and more, plus RCM platforms, prior auth services bureaus, and coding companies. That's a pre-built deployment substrate worth more than the $1.5 billion headline figure, because it bypasses the 18-to-36 month health system sales cycle entirely.
Google bringing software and Blackstone bringing $5B looks like a financing deal. What it actually is, if healthcare is the target, is Google buying access to Blackstone's portfolio companies as the first customer base. The model was never the bottleneck, getting it integrated into Epic, wired into X12 EDI transaction flows, and validated against local clinical populations was. PE firms figured that out before the labs did.
Full argument here: https://www.onhealthcare.tech/p/the-openai-anthropic-ai-arms-race?utm_source=x&utm_medium=reply&utm_content=2056738503287820718&utm_campaign=the-openai-anthropic-ai-arms-race
Fair warning, YC founders: if you take these tokens, there’s a non-zero chance that OpenAI will study exactly what your startup is doing, copy your idea and put your app into their free offering.
This is the classic platform playbook — be careful, founders!
The platform playbook framing is right, but the Rosalind situation has a sharper edge than the generic "don't build on someone else's platform" warning suggests.
The pricing mechanism is what most people are missing. OpenAI isn't just watching what you build, they've already picked the category they want to own: enterprise pharma and biotech. The zero-cost preview isn't a trial, it's a deliberate reset of what Amgen and Moderna will ever pay for a biotech AI SaaS tool again. By the time the preview window closes, willingness-to-pay benchmarks across the whole vertical will have shifted downward permanently.
So the danger for YC founders here isn't primarily that OpenAI studies your product and copies it. The more immediate problem is that your customer just got six to twelve months of free access to a competitor, and their CFO now has a new mental anchor for what this category should cost.
Founders with proprietary lab data or regulated workflow integrations probably survive this. Anyone built on RAG over PubMed or lit-review functionality without a closed data layer is looking at something more structural than a competitor entering the space.
The token extraction concern is real, I don't want to dismiss it. But the pricing distortion is already doing damage whether you take the tokens or not.
https://www.onhealthcare.tech/p/gpt-rosalind-lands-what-openais-first?utm_source=x&utm_medium=reply&utm_content=2056919949810028942&utm_campaign=gpt-rosalind-lands-what-openais-first
The reason why I release my X articles about AI agents (fat skill fat code thin harness) and GStack and GBrain is that we, yes you and I, can have *PROCESS POWER*, which is the one super powerful specific moat that anyone can create for themselves.
The agent helps you do it. https://t.co/Y0Me9xSKmZ
Process power is exactly the framing I kept circling back to when analyzing the leaked Claude Code architecture at https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2056915511469023635&utm_campaign=what-the-leaked-claude-code-codebase, because what Anthropic actually shipped wasn't a chatbot wrapper, it was a memory consolidation and orchestration system that compounds over time. The moat isn't the model, it's the process architecture sitting around the model, and healthcare builders who treat that distinction seriously are the ones who won't get commoditized.
Yesterday, I wrote about how long-running agentic systems need to be more reliable, robust, and fault-tolerant; and today, let's dig deeper into how.
Fun fact - distributed workflows are one of those problems that almost every engineering team accidentally rebuilds.
We start with a cron job, a message queue, a database table for state, and some retry logic. Then failures show up. A worker crashes halfway through. A network call times out. A deployment kills an in-flight process. Suddenly, we are building state machines, recovery logic, idempotency layers, compensating actions, and observability around all of it.
I wrote an essay on Temporal, an open-source durable execution engine that encapsulates all this plumbing and makes it easy to build long-running workflows.
In this article, I break down how Temporal actually works under the hood - Workflows, Activities, event histories, replay, Signals, Child Workflows, retries, timeouts, and the determinism constraints.
These out-of-the-box features and guarantees are what make Temporal useful in long-running agentic systems where AI agents need state, retries, tool orchestration, and execution that survives failures.
If you are building long-running agents, Temporal would come in handy. Give it a read.
The piece on Temporal is well-timed. The determinism constraint is the one that keeps coming up in practice, and it's also the one that breaks assumptions most engineers carry in from stateless service design.
What that constraint actually forces is a separation between workflow logic and side effects that turns out to be architecturally valuable beyond just fault tolerance. When you can't let non-deterministic calls touch workflow state directly, you start thinking clearly about what your agent is actually deciding versus what it's just executing. That's a distinction most agentic health tech teams are blurring right now.
The reason this matters specifically in healthcare: prior authorization workflows run across multiple days, touch payer clinical criteria, EHR notes, submission history, and eligibility verification concurrently, and can be interrupted mid-flight by a system update or a worker restart. The naive implementation uses a job queue and hopes nothing fails in a bad sequence. The Temporal approach gives you replayable event history, which means you also get an audit trail that looks suspiciously like what HIPAA explainability requirements are pointing toward anyway. (The compliance value is almost a side effect of the execution model, not something you bolt on.)
Where this connects to architecture I was digging into recently: the leaked Claude Code source has a three-gate trigger system and a consolidation lock for its memory cycle. That kind of structured, gate-checked long-running process is exactly what wants to sit on something like Temporal rather than a homebuilt state machine. The patterns are already documented. The plumbing problem is largely solved. What teams are still getting wrong is the memory and contradiction-resolution layer above it.
More on that here: https://www.onhealthcare.tech/p/what-the-leaked-claude-code-codebase?utm_source=x&utm_medium=reply&utm_content=2056946273165656375&utm_campaign=what-the-leaked-claude-code-codebase
“The [AI] tool presented here will enable rapid screening for multiple systemic diseases using retinal photographs, and it is a step forward in the evolution of oculomics from experimental research to real-world clinical practice.” —Editorial Team, @NatureMedicine https://t.co/90rT7PHp0K
Retinal oculomics and pancreatic radiomics are running the same play: tissue-level biological signal detectable before symptoms, retrospective AUC that looks great, and then Bayesian math that quietly destroys the population-screening case. The viability question isn't the model, it's whether you can enrich the cohort enough to make PPV defensible before payers ask for it.
https://www.onhealthcare.tech/p/the-preclinical-signal-in-routine?utm_source=x&utm_medium=reply&utm_content=2057111825431593406&utm_campaign=the-preclinical-signal-in-routine
600 new generic drugs added to TrumpRx this week thanks to President @realDonaldTrump's leadership!
Before you visit your local pharmacy, visit https://t.co/dDMdsILHTj to ensure you're getting the BEST price for your prescriptions.
It's easy and saves YOU money! https://t.co/MAfbTtlhAh
Generic volume is the distraction here, the real gap in https://www.onhealthcare.tech/p/what-does-17-pharma-mfn-deals-are?utm_source=x&utm_medium=reply&utm_content=2057176230965743647&utm_campaign=what-does-17-pharma-mfn-deals-are is that TrumpRx has no eligibility verification, no real-time benefit comparison, no secondary payer coordination.
A cash price list without adjudication infrastructure is just a browse page.
When the $245 Medicaid GLP-1 benchmark is public and your employer plan is paying more, that's not a pricing story anymore, that's an ERISA fiduciary exposure story. Who's building the compliance layer that tells a plan sponsor when they've crossed that line?
The healthy LDL number has been quietly moving its own goalposts for forty years:
- 1988: under 160
- 1993: under 130
- 2001: under 100
- 2004: under 70 for the high-risk
- 2019: under 55 for the very-high-risk
- Current trajectory: as low as possible, indefinitely
The science did not change.
The line did.
Move the line down by 30 milligrams and you have invented millions of new patients overnight. Same arteries. Same people. Different number on the page. Blood that was healthy on Friday is a chronic condition on Monday.
A diagnosis you can give to anyone is a prescription you can sell to anyone.
The line is wherever the next prescription pad needs it to be.
The question this raises for me is whether AI makes that dynamic faster or just more invisible.
Because when a physician moves the goalpost, there's at least a traceable institutional chain, a guideline committee, a conflict-of-interest disclosure, a paper trail. When a consumer AI platform interprets your lab results and flags a value that was unremarkable six months ago, the mechanism generating that flag is a black box that updates itself through machine learning without any external review. The FDA's current guidance for AI-based medical devices was written for professional-use systems and has no oversight mechanism for that kind of silent algorithmic shift in a consumer-facing product.
What I found when I looked at this for https://www.onhealthcare.tech/p/the-double-edged-algorithm-how-consumer?utm_source=x&utm_medium=reply&utm_content=2056686417539998002&utm_campaign=the-double-edged-algorithm-how-consumer is that the downstream cost structure follows a predictable escalation pattern: a routine lab upload generates six to eight supplement recommendations, which trigger monitoring tests to see if those supplements are working, which surface new borderline values, which justify specialist consultations. The financial incentive you're describing in guideline committees gets replicated algorithmically, except the AI has no direct financial stake. It just optimizes for comprehensive, actionable output because that's what engagement rewards.
That's a harder problem than a prescription pad.
The gap between best-case and real-world GDMT in #HFrEF is striking.
GWTG-HF 2024 vs EPIC COSMOS 2023–2025 (all HFrEF pts, 3.3M patients):
💚 Beta-Blocker: 95% vs 73%
💚 ACE/ARB/ARNI: 92% vs 67%
💚 MRA: 81% vs 36%
💚 SGLT2i: 78% vs 34%
💚 Quadruple GDMT: 68% vs 18% https://t.co/xTXdB7ZSdE
Quadruple GDMT at 18% in the real world is the number that should be stopping people cold.
The GWTG figures are basically what a motivated, protocol-driven center can do when it's trying. The COSMOS gap tells you what actually happens across 3.3 million patients when no one's coordinating the handoff between the hospitalist, the cardiologist, and the PCP.
This is exactly the execution problem I kept coming back to when I looked at cardiology VBC infrastructure. The bottleneck isn't knowing what the right therapy is. It's the cognitive load of acting on it during a 15-minute appointment when you've got six other things competing for attention. Prior clinical decision support tools in cardiology failed for exactly this reason: they surfaced more data instead of reducing the decisions a physician had to make in real time.
The case for EHR-native AI that pre-populates a pharmacist message for suboptimal beta-blocker dosing or auto-schedules outreach for a missed titration visit is basically sitting in this COSMOS data. You don't need to convince anyone that the gap exists. The question is which infrastructure model can close it at scale without requiring cardiologists to change how they practice.
https://www.onhealthcare.tech/p/60-million-reasons-to-pay-attention?utm_source=x&utm_medium=reply&utm_content=2056266861499539574&utm_campaign=60-million-reasons-to-pay-attention
How to print money, state government edition: 1) Tax your hospitals and managed care plans. 2) Use that tax revenue to pay your 10% share of Medicaid expansion. 3) Trigger a 90% federal match. 4) Return the original tax money, plus the massive federal windfall, back to the health
California's version of this mechanism generated roughly $2 billion in federal match on a 3% hospital net patient revenue tax, with hospitals receiving back $2.3 billion in enhanced rates, meaning the state essentially printed $300 million in net new provider revenue while keeping its general fund untouched. The MCO side of this equation was even more aggressive: one state taxed its Medicaid managed care business at a rate 117 times higher than its commercial business, concentrating the federal leverage almost entirely on the Medicaid side. That specific loophole is now closed, and states like New York have until March 31, 2026 to restructure, which is why I've been watching this as closely as any Medicaid policy change in the last decade. The full breakdown of what the November CMS guidance does to this mechanism, and what it means for MCO margins and health tech vendors, is at https://www.onhealthcare.tech/p/the-great-provider-tax-squeeze-what?utm_source=x&utm_medium=reply&utm_content=2056453552009040300&utm_campaign=the-great-provider-tax-squeeze-what
A 2016 BMJ study by researchers at Johns Hopkins estimated that medical error is the third leading cause of death in the United States.
Behind heart disease.
Behind cancer.
Ahead of stroke. Ahead of respiratory disease. Ahead of accidents. Ahead of diabetes. Ahead of
The decades-long failure to move that number is what makes the Boston ED data so striking. Not that o1 outperformed physicians at triage, but that physicians using o1 did not outperform o1 alone. That finding is the one people keep skipping past.
If augmented intelligence were the right frame, the human-plus-AI condition should have won. It didn't. The radiology CAD literature predicted exactly this: automation bias and anchoring degrade collaboration performance, the human stops doing independent work and the error profile shifts rather than shrinks. Medical error stays stubbornly lethal in part because the interventions keep assuming additive human-AI value that the data doesn't support.
The policy and investment community is still building around the copilot model. FDA clinical decision support guidance assumes it, the AMA's positioning assumes it, most health system AI procurement assumes it.
The misdiagnosis burden, around 12 million American adults annually with 40,000 to 80,000 downstream deaths by some estimates, doesn't move if we keep deploying AI as physician augmentation at the back end of a workup rather than as infrastructure at the front door. The distribution question matters more than the model quality question at this point, which is why the Microsoft/Nuance move was about EHR access and not diagnostic performance. Where does the error actually enter the workflow, and who controls the layer where it could get caught before it compounds?
https://www.onhealthcare.tech/p/what-the-harvard-er-study-says-about?utm_source=x&utm_medium=reply&utm_content=2055958162327425294&utm_campaign=what-the-harvard-er-study-says-about
@ScienceMagazine·11,971 views83%
5/18/26 8:54 AM ET
Streamlining the synthesis of valuable amines, researchers in Science present a new method that can selectively insert nitrogen into specific carbon–hydrogen bonds.
According to the study, the approach could simplify drug development by enabling more efficient, scalable, and https://t.co/iSC23lkGZk
Selective C-H amination is a real bottleneck in med chem, and solving it at scale matters more than most headlines suggest. What I'd watch is whether this kind of chemistry compresses synthesis costs enough to pair meaningfully with generative protein design, since the closed-loop pipelines I wrote about at https://www.onhealthcare.tech/p/profluents-225b-lilly-deal-and-why?utm_source=x&utm_medium=reply&utm_content=2056111285712912845&utm_campaign=profluents-225b-lilly-deal-and-why only compound in value if wet-lab validation gets cheaper alongside model costs. The regulatory and synthesis chokepoints are where the next friction shows up once the generative search space opens up.
Most voice AI hears what was said.
Velma hears what was meant.
Tone. Intent. Emotion. Deception.
The standard stack - STT to transcribe, LLM to analyse - throws all of it away before analysis even begins.
The signal was lost the moment audio became text... (🧵) https://t.co/DtNPPBMlzN
That's the exact problem I've been writing about, and it goes further than voice. Language itself is a lossy interface, and the entire ambient documentation boom in healthcare is running on the assumption that better transcription gets you most of the way there. It doesn't. When I looked at Apple's ~$2B acquisition of Q.ai (facial micro-movement detection for silent speech), the real signal wasn't "voice is getting better," it was that the industry already knows voice throws away too much, and the next layer has to capture pre-vocal signal before verbalization compresses everything into words. You can read the full argument here: https://www.onhealthcare.tech/p/the-interface-wars-why-apple-spent?utm_source=x&utm_medium=reply&utm_content=2055385122501788101&utm_campaign=the-interface-wars-why-apple-spent
The clinical stakes make this concrete. A physician's gaze pattern, a patient's vocal tremor, the micro-hesitation before answering a pain question: none of that survives the STT-to-LLM pipeline intact. Physicians are already paying $200 to $400 per month for tools that reduce EHR burden without adding any new diagnostic intelligence, which proves the market will pay for interface improvements alone. But those tools are still only capturing the linguistic residue of a clinical encounter.
The signal loss you're describing isn't a bug in the current stack. It's the whole architecture's ceiling.
Google's CEO confirmed 75% of new code at Google is AI-generated.
October 2024: 25%.
Fall 2025: 50%.
April 2026: 75%.
Doubled in a year.
Then doubled again in six months.
Engineers aren't writing code anymore.
They're reviewing it.
The job title is the same.
The job
...and the cost curve that follows from this is what nobody in health tech is pricing in yet.
When Google's internal dev velocity doubles in six months, the rebuild math for a hospital or large payer shifts faster than most vendor contracts reset. I modeled a prior auth workflow tool in my piece: what cost $4M over 18 months with 12 engineers is now a $300K, six-week project with three. That's the same compression Google is describing playing out inside enterprise health teams right now.
The "engineers review instead of write" framing is exactly right, it's also why pure software defensibility collapses faster than most people expect. The moat was never the IP, it was the rebuild cost. When rebuild cost drops 90%, vendors who built that moat are exposed, and the largest national payers with internal engineering capacity are the ones who will act on it first.
Full argument on what this means sector by sector in healthcare: https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2056092441724211672&utm_campaign=the-free-lunch-is-over-except-now
@ActionFixesFear·57,418 views87%
5/18/26 8:49 AM ET
A woman is told her bladder cancer has come back. Then she is told something stranger: there is a therapy the FDA approved for exactly this - and she still might not be able to get it.
Not because it failed her. Not because she can't afford it. Because it only works paired with https://t.co/tKlXZ559yw
What does it take to actually unblock her access, given that the pairing problem isn't scientific or regulatory?
My read, from digging into CASGEVY's commercial numbers: 500+ patient initiations globally against ~60,000 eligible patients, $43M in Q1 2026 revenue at a $2.2M list price. That gap isn't a science failure, it's a coordination failure. The therapy exists, the approval exists, the clinical indication exists, and patients still don't get treated because responsibility is fragmented across transplant centers, Medicaid benefit design, fertility clinics, and hospital workflow in ways nobody architected a solution for.
What this post describes fits that same pattern. The pairing requirement isn't the hard part, the hard part is that no institution owns the problem of getting her through the full sequence. When that ownership is unclear, patients fall out, not because of a missing drug but because of a missing operating system around the drug.
The next round of value in gene editing probably doesn't accrue to the editors themselves, it accrues to whoever builds activation infrastructure, outcomes-based contracts that reinsurers will actually sign, and the registry systems that let payers verify long-term outcomes. Until that stack exists, FDA approval is a necessary condition but nowhere near sufficient.
https://www.onhealthcare.tech/p/gene-editing-has-the-science-figured-b80?utm_source=x&utm_medium=reply&utm_content=2055852117827674495&utm_campaign=gene-editing-has-the-science-figured-b80
@RealNickMugalli·2,769 views83%
5/17/26 8:13 PM ET
Anthropic CEO Dario Amodei on SaaS: "Software is going to become cheap, maybe essentially free.
The premise that you need to amortize a piece of software you build across millions of users, that may start to be false.
But at the same time, there are whole jobs, whole careers https://t.co/nHBnPxQNbJ
Amodei's framing is correct and the healthcare SaaS market is where the consequences land hardest, because the entire business model of a prior auth vendor or a population health platform was exactly that amortization logic. You built the thing once, you spread the cost across 200 payer contracts, the rebuild cost for any single customer was prohibitive enough to lock them in.
That lock-in is dissolving. Fast.
What the quote leaves open is the second-order problem. When software costs collapse, the moat that survives is not engineering complexity, it is the thing that was always sitting underneath the software and getting underpriced: proprietary longitudinal data, regulatory certifications, and clinical workflow knowledge that took years of implementation to accumulate. Those assets don't compress. A hospital can spin up a custom prior authorization tool for $300k now instead of $4 million, but it cannot spin up ten years of payer-specific adjudication pattern data overnight.
The companies in real trouble are the ones whose pitch was essentially "we encoded the business rules so you don't have to." That encoding is now a commodity. The companies that should be reexamining their own cap tables are the services-heavy, low-gross-margin operators that venture always underweighted, because their cost to build falls while their clinical relationship depth does not.
The careers question Amodei gestures at is genuinely unsettled, and in healthcare specifically it cuts through vendor organizations, hospital IT departments, and payer engineering teams in ways that are not symmetric. Some of those jobs get eliminated. Some get redirected toward the data and compliance work that actually differentiates now. Which group is larger is the question I keep turning over.
Wrote through this specific dynamic for health tech: https://www.onhealthcare.tech/p/the-free-lunch-is-over-except-now?utm_source=x&utm_medium=reply&utm_content=2056004575744659540&utm_campaign=the-free-lunch-is-over-except-now
Satya Nadella's energy is something here. 🔥
"Tokens per Dollar per Watt"
The new equation for the AI age for every Company or Industry or Country.
"And that means Infrastructure, Infrastructure and Infrastructure." https://t.co/McINrBAo4a
The question this raises for me: at what point does "tokens per dollar per watt" stop being a general AI metric and start being a life-or-death clinical constraint?
Because in healthcare the math is already failing, real-time ICU monitoring across an entire health system, processing vitals, imaging, labs, and clinical notes simultaneously, doesn't close economically at current inference costs. The energy problem isn't abstract there, it's the actual bottleneck between a prototype and a deployed system.
Nadella's framing lands differently when you apply it to medicine. Software alone never completes an economic revolution, every prior wave needed a communication unlock and then an energy unlock, this is the energy unlock phase arriving. The companies I'd watch aren't pure LLM plays, they're sitting at the intersection of inference efficiency and clinical infrastructure, which is why I've been arguing https://www.onhealthcare.tech/p/the-pattern-always-repeats-why-healthcares?utm_source=x&utm_medium=reply&utm_content=2055839616763396340&utm_campaign=the-pattern-always-repeats-why-healthcares that Nvidia is better understood as an energy company than a chip company.
Tokens per watt is the number that determines whether clinical AI scales to a community hospital or stays locked inside academic medical centers.
Google CEO Sundar Pichai on current frontier model's ability to break the security of almost all current software.
"These models are definitely, like really gonna break pretty much all software out there, maybe already, we don't know."
https://t.co/ezjPcSehiG
Healthcare has no seat at the Glasswing table, and that quote is exactly why that should terrify anyone running a health system right now.
31% of all disclosed ransomware attacks hit healthcare in early 2026, and the sector is about to face machine-speed zero-day discovery against devices that still rely on network segmentation built for human-speed threats.
https://www.onhealthcare.tech/p/how-claude-mythos-preview-found-thousands?utm_source=x&utm_medium=reply&utm_content=2055923917546745907&utm_campaign=how-claude-mythos-preview-found-thousands
$HIMS $HERS
🚨 HERS IS NOW BEATING BOTH HIMS AND RO ON THE iOS APP STORE
Today's rankings in the Medical category:
- Hers: #9
- Ro: #10
- Hims: #18
Chart below = weekly avg rank https://t.co/vlv80yiFOo
App store rank is a real signal but it's doing a lot of work here without much support. Rankings reflect download velocity, not retention, and definitely not revenue per subscriber. A spike in Hers downloads during a period when the GLP-1 compounding story is all over financial media could just be curiosity traffic that churns in 60 days.
The harder question, which I got into when digging through the FY2025 10-K and the post-February regulatory sequence at https://www.onhealthcare.tech/p/a-public-equity-diligence-walk-on?utm_source=x&utm_medium=reply&utm_content=2056018877599867236&utm_campaign=a-public-equity-diligence-walk-on, is what those new installs are actually converting into now that the high-margin compounded semaglutide product is gone. Before February 2026 a new GLP-1 subscriber was worth something specific: a vertically integrated compounding spread between API procurement and the DTC subscription price. And that economics is structurally gone now. Hers ranking above Hims on the app store while Hims routes Wegovy through NovoCare at parity pricing is almost the worst combination, because you're growing the top of the funnel into a product architecture that no longer captures the same spread.
But the Ro comparison is the one I'd push back on most. Ro never built the compounding vertical the way Hims did, so their margin exposure to this regulatory reset was always lower. Hers beating Ro on installs while carrying more structural cost from the compounding wind-down isn't obviously good news.
The May 11 Q1 print will show whether any of this install activity is converting into the $149 recurring membership fee or just inflating a vanity metric.
@WallStreetApes·19,086 views84%
5/17/26 1:20 PM ET
Older people can no longer afford to live in California
New research finds that older people here in California face the highest risk of going bankrupt in the entire country
The risk score was a 69 out of 100 based on factors including healthcare expenses and the high cost of living like housing and groceries
I looked more into this and it’s really bad. Researchers found that more than 1/5 of California seniors have an annual income below 150% of the federal poverty line
Here’s what that means
For a single senior 150% of the federal poverty line $23,940 per year or about $2,000 month in California
That won’t even cover rent. That’ll barely cover half rent with a roommate in California
Those that are “house rich” are still “cash poor” and forced to choose between meds, food, and rent
California needs new leadership. This can’t continue
The "house rich, cash poor" framing is exactly what the MCBS data surfaces too. Western region Medicare beneficiaries have median home equity of $449,878, highest in the country, but that asset doesn't convert to monthly cash flow when you're on a fixed Social Security check. And that's the trap that breaks consumer-pay health tech pricing models in California before they even launch. Subscription products assume liquid income, not balance sheet wealth. https://www.onhealthcare.tech/p/the-hidden-balance-sheet-what-medicare?utm_source=x&utm_medium=reply&utm_content=2054624583907737992&utm_campaign=the-hidden-balance-sheet-what-medicare
imagine a software startup raising $800m before their first dollar of revenue. Can’t imagine it
the last two decades of software has been dominated by a simple theory: ship quickly, get customers early, generate revenue quickly to validate PMF, manage KPIs closely, etc. This became the dominant theory over all others
The next decade is going to have a long wave of hardware/robotics/deeptech/etc that will have a dramatically different profile. We’ll need a very different set of assumptions and theories soon
Chart credit: @atShruti
The healthcare version of this is already here. Technical build is only about 15% of the total burden for something like a remote patient monitoring platform (the other 85% is clinical validation, regulatory, reimbursement, EHR integration). And you can't compress that timeline through operational velocity.
But investors keep applying SaaS mental models to companies that can't even assess product-market fit until after a $2-5M RCT.
https://www.onhealthcare.tech/p/translational-friction-and-capital?utm_source=x&utm_medium=reply&utm_content=2055398709584855069&utm_campaign=translational-friction-and-capital
Palantir invented the forward deployed engineer role in 2006.
embed an engineer inside the customer. ship production code in their environment. own the technical outcome.
Alex Karp on the idea: “forward deployed engineers are stolen from french restaurants.”
20 years later https://t.co/gSMXqceSVg
What does it take for that model to actually work in healthcare specifically, where the "environment" isn't a server room but a tangle of Epic configurations, undocumented payer portal workarounds, and SharePoint-based approval queues?
The FDE concept translates, but the substrate is messier. Two health systems running identical Epic builds can have completely divergent clinical data models, local formularies, and legacy migration artifacts that nobody has written down anywhere. You can't ship production code against a workflow you haven't physically watched someone perform for weeks.
And that's where most health AI companies get it wrong. They treat the 60-70% of the stack that's now commoditized, the LLM APIs, vector databases, FHIR integrations, as the hard part. But the remaining 30-40% concentrated in workflow rules and org-specific integration is the part that actually determines whether a pilot survives past proof of concept. The Rock Health data on this is damning: 70% of health AI pilots don't scale, and model capability isn't the reason.
Karp's restaurant analogy is more useful than it sounds. A great chef doesn't send recipes, they're in the kitchen. Healthcare AI deployments that fail aren't failing because the recipe was wrong.
https://www.onhealthcare.tech/p/the-standardization-trap-why-deploying?utm_source=x&utm_medium=reply&utm_content=2055669316247339410&utm_campaign=the-standardization-trap-why-deploying
The 36 BIGGEST startup opportunities right now
1. biggest b2c: solving loneliness. third spaces, community apps, IRL
2. biggest b2b: managed AI employees for businesses
3. biggest overlooked: elder tech. 70 million boomers who want products that make them happier & healthier
4. biggest mobile: action apps that do things, not apps you stare at
5. biggest trades: matching platforms for electricians, plumbers, HVAC. supply shrinking
6. biggest consumer social: small social. group chats as products, no feeds, no ai slop
7. biggest ecommerce: agents that recommend products you'll like, shop, buy for you
8. biggest creator: live shows and unscripted content
9. biggest edtech: AI tutors that adapt through conversation
10. biggest SaaS: pay-per-outcome pricing
11. biggest auto: AI service advisor for dealerships. answers the same 15 questions 24/7
12. biggest talent: training non-technical people to operate agents
13. biggest boredom: curated offline experiences delivered to your door. kits, games, challenges. anti-screen products
14. biggest spiritual: the need for belonging is exploding, new formats of spiritual get togethers
15. biggest wellness: longevity biomarkers you actively manage
16. biggest mobile: action apps that do things, not apps you stare at
17. biggest one to solve ai slop: digital verification that you're a real human. every platform will need this within 2 years
18. biggest infrastructure: agent permissions, security, audit trails
19. biggest media: AI native media companies. build distribution, sell products later.
20. biggest parenting: family ops automation. forms, scheduling, logistics
21. biggest accounting: bookkeeping agents that charge per transaction
22. biggest fashion: brand-owned resale. every brand wants to control their secondary market
23.biggest hobbies: adult learning for joy. pottery, woodworking, drawing.
24. biggest skincare: at-home diagnostics. scan, get a protocol, track progress
25. biggest agriculture: precision farming tools for small farms. enterprise version exists, family farm doesn't
26. biggest pest control: subscription pest prevention instead of reactive treatment. the model flip that lawn care already made
27. biggest regulated: on-device AI. healthcare, legal, finance open up when data stays local
28. biggest gaming: AI characters with real memory and relationships
29. biggest dating: agent-mediated matchmaking
30. biggest fitness: adaptive coaching that rewrites your program daily
31. biggest travel: autonomous trip planning and rebooking
32. biggest food: personalized nutrition based on blood work and gut biome
33. biggest pet: health monitoring. $140B industry, almost no tech
34. biggest defense: AI-native security and compliance tools
35. biggest robotics: physical AI. $30 brains on existing hardware
36. biggest nostalgia: products that feel analog. vinyl, paper, handmade. counter-positioning against AI everything
The elder tech point is right but the health angle specifically is where I'd push further. I looked at the $11B+ federal rural health capital surface sitting across RHTP, FORHP, USDA, and FCC programs, and the buyers skew heavily 65+ in rural areas, which means at https://www.onhealthcare.tech/p/the-fifty-billion-dollar-rural-health?utm_source=x&utm_medium=reply&utm_content=2055713262801498264&utm_campaign=the-fifty-billion-dollar-rural-health the opportunity isn't just consumer elder tech, it's the infrastructure connecting those elders to care that nobody's built yet.
@WeWillBeFree24·4,518 views83%
5/17/26 10:34 AM ET
Vivek Ramaswamy put out a lot of statements today referencing the fraud in Ohio. Swampy is of course, involved in Ohio Medicaid.
Datavant, Vivek's spinoff company from Riovant, handles release of information for many hospitals and providers in Ohio. Ohio Medicaid providers, managed care plans, and patients often use Datavant/Ciox portals or services to request/retrieve records for claims, audits, risk adjustment, eligibility, or continuity of care.
All that Medicaid data, in one place. That is all.
Worked on the channel partnership architecture at Datavant long enough to know the "all that data in one place" framing misreads how the network actually functions. ROI fulfillment is a request-response workflow, not a data warehouse. When a Medicaid managed care plan in Ohio submits an audit request through a Datavant-connected portal, the records move point to point to satisfy that specific authorization. The network reaching 70,000 hospitals and clinics means faster fulfillment across a shared infrastructure, not centralized storage of Ohio Medicaid records sitting in one place waiting to be misused.
The compliance architecture matters here. HIPAA minimum necessary standards, state-specific Ohio medical record statutes, and automated redaction requirements are built into the transaction layer, not bolted on afterward. Every disclosure is logged, every authorization validated. The audit trail on a Medicaid records request is actually more defensible than what HIM departments were producing manually, where a $25-to-$75-per-request process relied heavily on individual judgment calls about what to release and to whom.
The corporate lineage question, Riovant to Datavant, is a fair thing to scrutinize in any government contracting context. That scrutiny should attach to procurement decisions, contract terms, and oversight mechanisms. But collapsing "handles release of information for Ohio Medicaid providers" into an implication of fraud access conflates the compliance infrastructure with the political story being told around it.
The more uncomfortable question for Ohio Medicaid isn't who moves the records, but whether the audit demand cycles driving all those ROI requests are themselves surfacing the fraud patterns or just generating paper that buries them.
https://www.onhealthcare.tech/p/transforming-release-of-information?utm_source=x&utm_medium=reply&utm_content=2054721052517814672&utm_campaign=transforming-release-of-information
Yes to Katy.
If I was a manufacturer I would contract directly with the employer.
The health plan when self funded has a fiduciary obligation to plan assets…
Nice work @KatyTalento
Contracting directly sounds clean, but the mechanism that makes it legally unavoidable is already in motion before most employers realize it.
The Johnson & Johnson case documented a 90-day supply of a generic MS drug costing the plan over $10,000 when the identical drug was available for $40 cash at a retail pharmacy. That gap is no longer just a negotiating grievance. Under ERISA's prudent expert standard, plan fiduciaries, meaning the HR executives and CFOs sitting on benefits committees, face personal liability for exactly that kind of pricing failure. And unlike the tobacco settlements that required state attorneys general to move, ERISA cases can be brought by individual plan participants. Sixty million people covered by self-insured employers are each a potential plaintiff, independently.
Direct manufacturer contracting is one destination this litigation pressure points toward. But employers also need to understand they are simultaneously exposed as defendants in employee suits and potential plaintiffs against their own TPAs and PBMs. Most haven't documented the fiduciary diligence that would survive discovery.
https://www.onhealthcare.tech/p/the-coming-storm-why-erisa-fiduciary?utm_source=x&utm_medium=reply&utm_content=2055780509146337524&utm_campaign=the-coming-storm-why-erisa-fiduciary