Connect with us

Artificial Intelligence

US health departments to pilot OpenAI and Anthropic AI tools under new PULSE program

Published

on

OpenAI and Anthropic AI

Why public health agencies are turning to generative AI

A new initiative called PULSE will let 10 US public health jurisdictions trial generative AI tools from OpenAI and Anthropic. The goal? Figure out what works — and what doesn’t — before the technology spreads further.

The program, formally named the Public Health Use Case and Learning Scaling Engine, is backed by the Coalition for Health AI (CHAI), Accenture, and the two AI companies. It will run across state, local, tribal, and territorial health agencies.

OpenAI and Anthropic have each donated 10 enterprise licenses, giving up to 2,000 public health practitioners access to their commercial AI products. Accenture will handle participant onboarding and help build playbooks from the trial results.

“Every major technological transformation succeeds or fails based on trust, governance and execution,” said Dr. David Lakey, former Texas health commissioner, in a statement. “PULSE will support agencies in this endeavour, and is specifically designed for practical implementation.”

Five focus areas for the pilot

CHAI’s leadership council will pick the participating jurisdictions. Practitioners will then be grouped into communities tackling five specific use cases:

  • Biosurveillance and drug-wave prediction — spotting disease outbreaks and tracking illicit drug trends.
  • Social determinants of health (SDoH) mapping — using AI to identify how housing, income, and environment affect community health.
  • Operations and community-feedback analysis — automating the review of public comments and internal workflows.
  • Public communications and multilingual translation — generating health messages in multiple languages.
  • Automated clinical-data retrieval and FHIR query engine — pulling electronic health records using the FHIR standard.

Notably, CHAI hasn’t specified which OpenAI or Anthropic products will be used, nor the model versions or configurations. The announcement also leaves unclear how the two providers will be assigned across the pilots.

What about FHIR and human oversight?

FHIR — an HL7 standard for exchanging healthcare data electronically — features in the clinical-data retrieval use case. But the announcement doesn’t define exactly how generative AI fits into that workflow. Will the models write queries, fetch records, summarize results, or do all three?

It also doesn’t say whether staff will check for incorrect queries, incomplete retrievals, or unsupported summaries before using the information. That’s a critical gap, especially for applications that could involve demographic, geographic, clinical, or population-health data.

CHAI hasn’t disclosed whether the pilots will use identifiable records, de-identified information, synthetic data, or aggregated datasets. That distinction matters for compliance with the US Health Insurance Portability and Accountability Act (HIPAA).

HIPAA and data protection: what’s missing

The US Department of Health and Human Services requires organizations covered by HIPAA to protect electronic health information. Its cloud-computing guidance says regulated entities and service providers must meet HIPAA rules when cloud systems create, receive, maintain, or transmit electronic protected health information.

But HIPAA won’t apply to every PULSE participant or workflow — it depends on the agency, the data involved, and the function being performed. The announcement doesn’t set out retention periods, access controls, audit arrangements, or rules for submitting protected health information.

OpenAI says inputs and outputs from its business services — including ChatGPT Enterprise and its API — are not used to train or improve its models by default. Anthropic makes a similar claim. However, those policies don’t define how the PULSE deployments will be configured in practice.

“We believe AI should be useful, safe and accessible to the people tackling society’s most important challenges,” said Felipe Millon, OpenAI’s head of government go-to-market. He added that the donated licenses were designed to help public health organizations evaluate the tools through a structured process.

Governance and evaluation remain vague

The pilots are scheduled to begin in autumn 2026. CHAI expects to release the resulting playbooks in 2027, which other public health agencies can use as reference material.

But CHAI hasn’t published the measures it will use to assess the pilots. It hasn’t explained whether each use case will be evaluated under separate technical, operational, privacy, and safety criteria. The announcement also doesn’t detail how model outputs will be reviewed — whether staff must approve generated public communications, verify translations, validate retrieved clinical information, or check biosurveillance outputs before use.

The US National Institute of Standards and Technology (NIST) recommends identifying which AI functions need human oversight and training users to understand system performance and limitations. Its generative AI guidance also covers testing, validation, monitoring, documentation, privacy, and management oversight.

“Public health teams are being asked to do more with less, and AI can help — as long as it’s brought in with care and the right guardrails,” said Elizabeth Kelly, Anthropic’s head of beneficial deployments. She said PULSE would let practitioners test the tools in their own environments with privacy, governance, and responsible-use measures built in from the start.

Who can participate — and what’s still unknown

Eligible participants include state and territorial health departments, county and municipal agencies, tribal authorities, Indian health organizations, and large city health departments. But CHAI hasn’t specified minimum staffing, infrastructure, interoperability, or cybersecurity requirements for participating jurisdictions.

Data from the National Association of County and City Health Officials, cited by CHAI, shows nearly 40% of local health departments aren’t using AI at all. The coalition said some departments are interested in revising workflows and improving operational efficiency.

PULSE plans to convert findings from 10 jurisdictions into guidance for wider use. Yet the announcement doesn’t explain how the playbooks will account for differences in agency size, technical systems, legal responsibilities, staffing, or procurement arrangements.

It also doesn’t say whether outputs from biosurveillance, drug-wave prediction, or clinical-data retrieval will be used only for testing, presented to staff for review, or incorporated into operational workflows.

“We know AI is going to reshape how we deliver public health — the question is whether we do it thoughtfully or not,” said Dr. Ashish Jha, a former White House COVID-19 response coordinator. He said the program would test which applications work and document the findings for other agencies.

Broader context: CHAI’s governance work

PULSE is part of CHAI’s larger effort on governance standards for healthcare AI. In May, the organization announced plans to develop guidance covering eight governance areas through workshops and working groups involving more than 150 healthcare AI representatives. It has since started publishing playbooks on organizational AI policies, governance structures, and internal resources.

Separately, CHAI has worked with the Joint Commission on governance playbooks aligned with its voluntary Responsible Use of AI in Healthcare certification. The PULSE announcement doesn’t state that participating public health agencies will be assessed under that certification.

Dr. Brian Anderson, chief executive of CHAI, said public health agencies entered the COVID-19 pandemic after years of limited investment in technology. He said PULSE was intended to give agencies practical experience with AI before wider implementation.

For more on how AI is being applied in healthcare settings, read our coverage of Bunkerhill’s $55M raise for agentic AI across health systems. And if you’re interested in the broader AI landscape, check out our analysis of AI and big data trends in healthcare.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Artificial Intelligence

Moonshot’s Kimi K3 Could Match Anthropic’s Best — And It’s Open Source

Published

on

Moonshot Kimi K3

The Next Leap in Open-Weight AI

Chinese AI lab Moonshot AI is about to release a model that could fundamentally change how enterprises think about paying for frontier artificial intelligence. According to a report from the Financial Times, the upcoming Kimi K3 is expected to perform on par with — or even surpass — Anthropic’s Opus 4.8. That’s a bold claim, but one backed by the lab’s recent track record.

The Kimi K2 models already turned heads in the open-source community. They scored high on standard benchmarks and showed capabilities that weren’t far behind the latest proprietary systems. K3, insiders say, takes that momentum further. It’s designed to close the gap with closed-source giants like OpenAI and Anthropic.

What Makes Kimi K3 Different

Size matters here. The Kimi K3 will reportedly be the largest open-weight AI model ever released from China, with a parameter count landing somewhere between 2 trillion and 3 trillion. For context, that dwarfs many of the most capable models on the market today. And it won’t stay behind closed doors for long — the FT report says it will be released “in the coming days.”

That timeline is aggressive. It suggests Moonshot is racing to capitalize on a moment when enterprises are rethinking their AI budgets. Why pay a premium for proprietary models when an open-weight alternative can do the same job for a fraction of the cost?

A Valuation That Reflects the Ambition

Moonshot is also reportedly raising fresh capital at a valuation of $31.5 billion. That’s a significant jump from the $20 billion valuation it commanded back in May, when it raised $2 billion. Investors are clearly betting that open-weight models will carve out a major slice of the AI market — and that Moonshot will be the one delivering them.

The Enterprise Shift Toward Open-Source AI

The timing couldn’t be better for Moonshot. A growing number of business leaders are questioning whether it’s worth paying for expensive, closed-source models from labs like OpenAI and Anthropic. The fear? That these companies will somehow extract and use the data clients submit through products like ChatGPT and Claude.

That worry isn’t theoretical. It’s driving real decisions. Executives are now actively pitching their own in-house models as safer alternatives. Others are telling companies to take cheaper open-source models — from labs like DeepSeek, Z.ai, or Moonshot — and fine-tune them for specific use cases.

The argument is gaining real traction. Especially as Chinese open-source models continue to close the performance gap with their more expensive, frontier counterparts.

How Kimi K3 Could Reshape the Market

If the Kimi K3 delivers on expectations, it could accelerate a trend that’s already underway: the commoditization of large language models. When a 3-trillion-parameter open-weight model can match Anthropic’s Opus 4.8, the premium for proprietary systems becomes harder to justify.

This isn’t just about benchmarks. It’s about control. Companies that adopt open-weight models retain ownership of their data and can customize the model to their needs. They aren’t locked into a vendor’s roadmap or pricing changes.

Of course, there are trade-offs. Running a model the size of Kimi K3 requires serious infrastructure. Not every organization has the compute to handle it. But for those that do, the economics are compelling.

The Bigger Picture: China’s AI Ambitions

Moonshot’s rise is part of a broader story. Chinese AI labs have been quietly catching up — and in some areas, leapfrogging — their Western counterparts. The release of models like DeepSeek’s R1 and now Moonshot’s Kimi K3 shows that open-weight development in China is no longer a sideshow. It’s a central force.

The political implications are real, too. As export controls tighten on advanced chips, Chinese labs have had to innovate on efficiency and architecture. The fact that Kimi K3 can compete with top-tier Western models despite these constraints says a lot about the ingenuity at play.

For enterprises evaluating their next AI move, the message is clear: the days of assuming closed-source is inherently better are ending. The Kimi K3 might just be the model that proves it.

Continue Reading

Artificial Intelligence

Claude Cowork can now keep working even after you close your laptop — here’s what that means

Published

on

Claude Cowork

Your laptop is no longer the anchor

Until today, using Anthropic‘s Claude Cowork meant keeping your desktop awake and running. Shut the lid? The task died. That limitation is gone now.

Anthropic is rolling out Cowork to web and mobile starting today. Max plan subscribers get first dibs on the beta over the next few days. Other plans will follow in weeks, not months. The core promise? You can start a job on your desk machine, walk away, and check progress later from your phone.

How Claude Cowork actually works

If you haven’t tried it, the concept is straightforward. You hand Claude a task — sorting files, scanning your inbox, updating a calendar, whatever — and it churns through your connected tools until it’s done. The catch, until now, was that your laptop had to stay on and awake. That made long-running jobs impractical.

The new update shifts those tasks to the cloud. Scheduled work runs remotely. So you can close your laptop, take a nap, and Claude keeps going. When it hits a fork in the road — a decision requiring your input — it pings you. Nothing gets sent or executed without your approval. No rogue agents here.

Who’s actually using this thing?

Anthropic also dropped usage numbers alongside the expansion. And they surprised me. I assumed Claude Cowork was primarily a developer tool — something for coders and programmers. Turns out I was wrong. The company says over 90% of Cowork usage falls into everyday business operations and content creation.

That’s a telling stat. It suggests AI agents are crossing over from niche technical tools to something broader. People drowning in spreadsheets, repetitive emails, or presentation decks are the real users. If you’ve ever wanted to hand off the boring parts of your job, this update makes that more practical.

What changes with web and mobile access

The expansion means two things. First, you’re no longer tethered to a single machine. You can kick off a task from your desktop, then check in from your phone while commuting or grabbing coffee. Second, cloud-based scheduling means tasks don’t pause when you step away.

This matters for anyone who has tried running a long data cleanup or report generation on a laptop. The old model forced you to keep the machine awake, which is annoying and wasteful. Now, the work lives in Anthropic’s cloud infrastructure. Your laptop is just the launchpad.

What you still can’t do

Claude still won’t act autonomously on sensitive actions. Every output that touches the outside world — sending an email, posting a file, updating a shared calendar — requires your explicit go-ahead. Anthropic is clearly cautious about trust and safety. That’s a good thing, even if it means you can’t fully “set and forget” everything.

What this means for the AI agent space

The usage numbers tell a story. Over 90% non-developer usage is a strong signal that AI agents are finding a real product-market fit outside of engineering teams. Business operations folks, marketers, writers, and project managers are the ones leaning in.

I expect this expansion to accelerate that trend. Once people realize they don’t need to keep a laptop running to use Anthropic’s AI agent, adoption should climb. The friction point was obvious: nobody wants to babysit a laptop for a background task. Removing that barrier makes the tool dramatically more useful.

If you are a Max subscriber, you can test the beta now. For everyone else, the wait is weeks, not months. In the meantime, start thinking about which tedious tasks you’d hand off to an AI that never sleeps.

Continue Reading

Artificial Intelligence

The AI agent evaluation gap: Enterprises trust their tests less than they trust their agents

Published

on

AI agent evaluation gap

Half of enterprises shipped a failing agent

That number should stop anyone building AI agents for customers cold. According to new VentureBeat Pulse Research surveying 157 enterprise organizations, 50% have deployed an agent or LLM feature that passed internal evaluations — and then caused a customer-facing failure in production. A quarter have seen it happen more than once.

The finding lands like a punch. It means the standard pre-deployment gauntlet — unit tests, red-teaming, automated evals — is letting bad agents through. The test says go. The agent breaks. The customer pays.

Only 36% of organizations report no such failure. The rest either don’t run pre-deployment evaluations at all (8%) or don’t track root causes closely enough to know (6%).

Trust in automated evaluation is almost nonexistent

Ask enterprise leaders how much they trust automated evaluation today, and the answer is brutal: just 5% say they fully trust it. That leaves 95% with a specific complaint holding them back.

The top grievance, cited by 29% of respondents, is the one that explains the failure rate: evaluations align poorly with real-world outcomes. A passing score in the test environment doesn’t predict what happens when real users, real data, and real edge cases show up.

Bias and inconsistency (21%) come next, followed by lack of explainability (18%) — organizations can’t always understand why an evaluation reached its verdict. Another 17% cite data leakage or privacy concerns in the evaluation process itself.

So the tests meant to certify agents are, broadly speaking, not trusted to certify them. That makes what comes next all the more surprising.

Autonomy is accelerating — despite the trust gap

Here’s the paradox at the heart of the research. Even though almost no one fully trusts automated evaluation, two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within twelve months (33%).

Only 22% rule it out for the foreseeable future.

The direction is clear: enterprises are moving to let evaluations gate production autonomously, removing the human check, at the same moment they say those evaluations don’t reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it.

Notably, this isn’t just a startup phenomenon. Larger enterprises (2,500+ employees) are slightly more likely than smaller ones to be on the zero-human-review path (70% versus 64%) and slightly more likely to have shipped a failing agent (54% versus 48%). The assumption that big, regulated organizations hold the human in the loop longest is, in this sample, backwards.

The evaluation stack is fragmented — and provider-led

Ask which agent reliability or evaluation platform enterprises primarily use, and the market has no clear leader. Provider-native tooling leads: OpenAI‘s native evals and traces (17%) and Anthropic‘s Claude Console evals (13%) together outweigh any independent platform.

But they’re tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all.

The specialist vendors — DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize — are scattered across single to low double digits. Another 11% have built their own. No independent platform has yet become the category standard.

Production monitoring mostly watches uptime, not correctness

Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up, how fast, at what cost, any errors. Or it can watch whether the agent’s output is correct — automated checks on each answer’s content.

The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring. The request completes. The response is fast. No error is thrown. Everything reads healthy.

The split is stark: 51% of organizations monitor only whether the agent is functioning, while just 23% monitor whether its answers are right. Roughly three-quarters run no automated, real-time evaluation of output correctness in production. They’re taking correctness on faith.

What drives tool selection — and what’s next

Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%). Breadth of observability (13%) and vendor roadmap (4%) matter far less.

On what success looks like, more than a third (36%) name evaluation consistency — getting the same verdict on the same behavior every time. That’s well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%).

The emphasis on consistency is telling: before enterprises can trust an evaluation’s verdict, they need it to be stable — the very property whose absence (bias and inconsistency) ranked among the top trust limitations.

A tooling reshuffle is coming

The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months. 31% plan to do so within the next quarter.

The consideration set points where current usage is thinnest: DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI’s native evals (13%) and Braintrust (9%). The open-source specialists are drawing more interest than their present footprint suggests.

Given that so many enterprises today rely on provider-native tools or nothing at all, this is less a defection than a first real wave of tooling adoption — the moment the evaluation layer starts to consolidate.

The bottom line: An evaluation gap that autonomy will widen

Organizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer. Almost none fully trust automated evaluation, chiefly because it doesn’t match real-world outcomes. Most watch production for uptime and cost rather than for whether the agent’s answers are right.

Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone.

The vendor market is early and unsettled. Encouragingly, the next dollar is going to observability and — pointedly — human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market. But the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust.

The evaluation gap is not a coverage problem that more tests alone will close. It is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy — or whether the false-confidence failures move from customer incidents into changes that deploy themselves.

Continue Reading

Trending