Connect with us

Artificial Intelligence

Enterprise AI Has a Trust Problem, Not a Retrieval Problem — And the Fix Is Still Under Construction

Published

on

AI context gap

The Numbers Behind the AI Hallucination That Isn’t

A new VentureBeat Pulse Research survey of 101 enterprises with more than 100 employees delivers a sobering finding: 57% of organizations have already watched an AI agent produce a confident, wrong answer — and traced the error back to missing or inconsistent business context. More than half of those saw it more than once.

This isn’t the classic hallucination problem, where a model makes up facts out of thin air. It’s worse. The agent sounds authoritative. It cites documents, metrics, or definitions that are real — but stale, incomplete, or contradictory. The AI context gap is the distance between how confidently an agent answers and how reliable the foundation beneath it actually is.

And right now, that foundation is being built faster than it can be trusted.

RAG Is the Default — and the Main Failure Surface

Retrieval-augmented generation (RAG) has quietly become the backbone of enterprise context. For 38% of organizations, RAG over documents or a vector index is the primary way AI agents understand the business. That’s nearly double the share of the next most common approach, a governed semantic layer or ontology (21%).

The concentration matters. Because so much enterprise context flows through retrieval, the quality of that retrieval is the quality of the answer. Thin retrieval isn’t an edge case — it’s the main failure surface. When 57% of enterprises have already seen agents go confidently wrong, and most of those have seen it more than once, the retrieval pipeline is the obvious culprit.

Notably, fine-tuning — once the darling of enterprise AI customization — has all but vanished from the conversation. In a separate VentureBeat survey wave (April–May, n=136), fine-tuning ranked dead last among six factors in model selection at 5%. Context injection at run time is how enterprises make agents knowledgeable. The question is whether that context can be trusted.

Provider-Native Retrieval Has Quietly Won — For Now

One of the survey’s most surprising findings: the dedicated vector database is no longer the center of the RAG universe. OpenAI‘s file search (40%) and Google‘s Vertex AI Search (38%) already lead every purpose-built vector database in production usage.

Among the specialists, Elasticsearch/OpenSearch (20%) and pgvector (12%) — tools enterprises already run for other reasons — beat the pure-play vector databases like Weaviate, Qdrant, Pinecone, and Milvus, each sitting in single digits or low double digits. The category that coined the term “vector database” is being absorbed by the platforms enterprises already buy from.

Yet here’s the tension: a plurality of enterprises (36%) say they intend to keep best-of-breed standalone tools rather than consolidate onto a provider’s native context stack. Only 21% plan to fully consolidate. The gap between what enterprises run and what they say they want is the strategic question of the category. They’re adopting bundled retrieval for convenience while insisting they want independence.

Hybrid Retrieval Is the Consensus — But Uncertainty Runs Deep

Vector-only retrieval is already viewed as insufficient. A third of enterprises (34%) expect hybrid retrieval — embeddings combined with reranking and access controls — to dominate their production systems by the end of 2026. That’s three times the share who expect pure vector search to prevail.

The second-largest answer? Uncertainty. 17% simply don’t know, and 14% expect to move beyond a dedicated vector layer entirely toward tool-first or long-context retrieval. The consensus isn’t a single tool — it’s a layered pipeline, and that pipeline isn’t fully formed yet. The access controls that hybrid retrieval promises are the very controls whose absence produces the confident-but-wrong failures.

For more on how to improve the quality of your AI outputs, see our guide on improving RAG accuracy with better data preparation.

The Governed Semantic Layer: Under Construction, Not Yet in Production

The industry’s answer to the AI context gap is a governed semantic layer — a shared, consistent definition layer that gives agents and business intelligence tools a common understanding of metrics, terms, and data sources. Think of it as a single source of truth for agent context.

Well over half of enterprises (58%) either run a governed semantic layer in production (25%) or are piloting and building one (34%). Another 17% are actively evaluating. That means three-quarters of enterprises are engaged with the idea in some form.

But here’s the catch: more are building than have shipped. For most organizations, the governed layer that would prevent inconsistent context is still a work in progress. The survey catches this wave mid-construction — ambition well ahead of production reality. The fix is being built, but agents are already running on the old, unreliable foundation.

How Enterprises Buy and Monitor Retrieval Systems

Enterprises choose retrieval systems on operability, not accuracy. Ease of data ingestion (36%), latency and performance (32%), and operational simplicity (29%) lead the selection criteria — ahead of retrieval accuracy and access control (23% each), the two factors most directly tied to the failures.

Once systems are running, the emphasis shifts toward trust. The most-tracked metrics are response correctness (42%) and security and access control (38%), ahead of latency (28%), operational stability (27%), and answer relevance (23%). Enterprises buy for how easily a system runs and watch it for whether it can be trusted.

Satisfaction with current systems is moderately positive — averaging 4.0 on a five-point scale — but not enthusiastic. Ease of implementation and value for money both hover around 3.9. The message: current tools are workable, but nobody is thrilled.

A Provider Shuffle Is Coming

The retrieval stack is not settled. While 43% of enterprises have no plans to change, a small majority (57%) intend to switch or add a provider within twelve months. A quarter plan to move within the next quarter.

The consideration set reveals an interesting dynamic. Provider-native retrieval still leads what enterprises are evaluating (OpenAI 22%, Vertex AI Search 21%), but the open-source vector specialists punch above their current footprint. Qdrant (14%) and Milvus (13%) draw more switching interest than their present usage (10% and 6%) would suggest.

Read alongside the best-of-breed preference, the picture is a market in flux: enterprises run provider-native today, are evaluating a broader field, and say they want to keep their options open. The reshuffle ahead will test whether best-of-breed intent survives contact with the convenience of the bundle.

The Bottom Line: More Retrieval Alone Won’t Close the Gap

The AI context gap is not a volume problem. Throwing more documents or bigger indexes at it won’t solve it. The problem is governed, consistent, access-aware context — and that requires a semantic layer, hybrid retrieval with reranking and access controls, and a shift from buying for operability to buying for trust.

Right now, agents are running ahead of the infrastructure that feeds them. The context layer is the next contested tier of the AI stack. The open question for later survey waves is whether enterprises finish building that layer before the confident-but-wrong failures move from the lab into decisions that matter.

For a deeper look at how to build trust in AI systems, read our analysis on enterprise AI governance best practices.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Artificial Intelligence

OpenAI drops GPT-5.6 Sol, Terra, and Luna to the public this week

Published

on

GPT-5.6 Sol Terra Luna

After weeks of restricted preview, OpenAI is flipping the switch

On Thursday, July 9, OpenAI will finally make its GPT-5.6 Sol, Terra, and Luna models available to everyone. The company confirmed the date in a post on X, ending a rollout that was anything but routine.

If you’ve been watching from the sidelines since the limited preview kicked off in late June, the wait is almost over. But the delay wasn’t a technical hiccup — it was political.

Why the government put the brakes on GPT-5.6

When OpenAI first unveiled the GPT-5.6 family on June 26, access was locked down to about 20 trusted partners. The reason? The US government asked for time to review the models before a wider release. With preview access now expanding globally and a full public launch set for Thursday, that review appears to be finished.

OpenAI isn’t alone in facing this kind of scrutiny. Earlier this year, Anthropic had to suspend access to Claude’s Fable and Mythos models after the US Commerce Department raised concerns under export controls. Anthropic restored those models on July 1 after a similar review.

Three models, three purposes

Instead of one monolithic release, OpenAI split GPT-5.6 into a trio. Here’s how they break down:

  • Sol — the flagship. Built for advanced coding and cybersecurity work. It includes new Max and Ultra reasoning modes for complex, multi-step tasks.
  • Terra — the balanced option. Meant for everyday workflows where you need solid performance without burning through your budget.
  • Luna — the speedster. The fastest and most affordable variant, aimed at straightforward queries and high-volume use.

The naming scheme is deliberate. OpenAI wants developers to pick based on intelligence, speed, or cost — no guesswork required.

What Sol’s Max and Ultra reasoning modes actually do

Sol isn’t just bigger. It introduces two new reasoning modes: Max and Ultra. Max handles long-running agentic tasks — think code generation across hundreds of files or multi-hour cybersecurity audits. Ultra goes further, applying deeper inference chains for problems that require sustained logical consistency.

OpenAI says the entire GPT-5.6 family brings improvements to reasoning, coding, and long-running tasks. But Sol is the only one that gets the Max and Ultra treatment.

What this means for developers and everyday users

For developers, the split means you can finally stop overpaying for a model that’s too powerful for simple tasks. Need a quick API call? Luna handles it cheaply. Building a security scanning tool? Sol with Ultra reasoning is your pick.

For regular users, the headline is simpler: GPT-5.6 is faster and smarter, and you can try it starting Thursday. If you want to compare it to existing options, OpenAI’s model comparison tool is a good place to start.

The broader lesson from this rollout is clear. Big AI releases are no longer just product launches — they’re diplomatic events. Governments want a look before the world gets one. For now, the review process is done, and GPT-5.6 Sol, Terra, and Luna are cleared for takeoff.

Continue Reading

Artificial Intelligence

US health departments to pilot OpenAI and Anthropic AI tools under new PULSE program

Published

on

OpenAI and Anthropic AI

Why public health agencies are turning to generative AI

A new initiative called PULSE will let 10 US public health jurisdictions trial generative AI tools from OpenAI and Anthropic. The goal? Figure out what works — and what doesn’t — before the technology spreads further.

The program, formally named the Public Health Use Case and Learning Scaling Engine, is backed by the Coalition for Health AI (CHAI), Accenture, and the two AI companies. It will run across state, local, tribal, and territorial health agencies.

OpenAI and Anthropic have each donated 10 enterprise licenses, giving up to 2,000 public health practitioners access to their commercial AI products. Accenture will handle participant onboarding and help build playbooks from the trial results.

“Every major technological transformation succeeds or fails based on trust, governance and execution,” said Dr. David Lakey, former Texas health commissioner, in a statement. “PULSE will support agencies in this endeavour, and is specifically designed for practical implementation.”

Five focus areas for the pilot

CHAI’s leadership council will pick the participating jurisdictions. Practitioners will then be grouped into communities tackling five specific use cases:

  • Biosurveillance and drug-wave prediction — spotting disease outbreaks and tracking illicit drug trends.
  • Social determinants of health (SDoH) mapping — using AI to identify how housing, income, and environment affect community health.
  • Operations and community-feedback analysis — automating the review of public comments and internal workflows.
  • Public communications and multilingual translation — generating health messages in multiple languages.
  • Automated clinical-data retrieval and FHIR query engine — pulling electronic health records using the FHIR standard.

Notably, CHAI hasn’t specified which OpenAI or Anthropic products will be used, nor the model versions or configurations. The announcement also leaves unclear how the two providers will be assigned across the pilots.

What about FHIR and human oversight?

FHIR — an HL7 standard for exchanging healthcare data electronically — features in the clinical-data retrieval use case. But the announcement doesn’t define exactly how generative AI fits into that workflow. Will the models write queries, fetch records, summarize results, or do all three?

It also doesn’t say whether staff will check for incorrect queries, incomplete retrievals, or unsupported summaries before using the information. That’s a critical gap, especially for applications that could involve demographic, geographic, clinical, or population-health data.

CHAI hasn’t disclosed whether the pilots will use identifiable records, de-identified information, synthetic data, or aggregated datasets. That distinction matters for compliance with the US Health Insurance Portability and Accountability Act (HIPAA).

HIPAA and data protection: what’s missing

The US Department of Health and Human Services requires organizations covered by HIPAA to protect electronic health information. Its cloud-computing guidance says regulated entities and service providers must meet HIPAA rules when cloud systems create, receive, maintain, or transmit electronic protected health information.

But HIPAA won’t apply to every PULSE participant or workflow — it depends on the agency, the data involved, and the function being performed. The announcement doesn’t set out retention periods, access controls, audit arrangements, or rules for submitting protected health information.

OpenAI says inputs and outputs from its business services — including ChatGPT Enterprise and its API — are not used to train or improve its models by default. Anthropic makes a similar claim. However, those policies don’t define how the PULSE deployments will be configured in practice.

“We believe AI should be useful, safe and accessible to the people tackling society’s most important challenges,” said Felipe Millon, OpenAI’s head of government go-to-market. He added that the donated licenses were designed to help public health organizations evaluate the tools through a structured process.

Governance and evaluation remain vague

The pilots are scheduled to begin in autumn 2026. CHAI expects to release the resulting playbooks in 2027, which other public health agencies can use as reference material.

But CHAI hasn’t published the measures it will use to assess the pilots. It hasn’t explained whether each use case will be evaluated under separate technical, operational, privacy, and safety criteria. The announcement also doesn’t detail how model outputs will be reviewed — whether staff must approve generated public communications, verify translations, validate retrieved clinical information, or check biosurveillance outputs before use.

The US National Institute of Standards and Technology (NIST) recommends identifying which AI functions need human oversight and training users to understand system performance and limitations. Its generative AI guidance also covers testing, validation, monitoring, documentation, privacy, and management oversight.

“Public health teams are being asked to do more with less, and AI can help — as long as it’s brought in with care and the right guardrails,” said Elizabeth Kelly, Anthropic’s head of beneficial deployments. She said PULSE would let practitioners test the tools in their own environments with privacy, governance, and responsible-use measures built in from the start.

Who can participate — and what’s still unknown

Eligible participants include state and territorial health departments, county and municipal agencies, tribal authorities, Indian health organizations, and large city health departments. But CHAI hasn’t specified minimum staffing, infrastructure, interoperability, or cybersecurity requirements for participating jurisdictions.

Data from the National Association of County and City Health Officials, cited by CHAI, shows nearly 40% of local health departments aren’t using AI at all. The coalition said some departments are interested in revising workflows and improving operational efficiency.

PULSE plans to convert findings from 10 jurisdictions into guidance for wider use. Yet the announcement doesn’t explain how the playbooks will account for differences in agency size, technical systems, legal responsibilities, staffing, or procurement arrangements.

It also doesn’t say whether outputs from biosurveillance, drug-wave prediction, or clinical-data retrieval will be used only for testing, presented to staff for review, or incorporated into operational workflows.

“We know AI is going to reshape how we deliver public health — the question is whether we do it thoughtfully or not,” said Dr. Ashish Jha, a former White House COVID-19 response coordinator. He said the program would test which applications work and document the findings for other agencies.

Broader context: CHAI’s governance work

PULSE is part of CHAI’s larger effort on governance standards for healthcare AI. In May, the organization announced plans to develop guidance covering eight governance areas through workshops and working groups involving more than 150 healthcare AI representatives. It has since started publishing playbooks on organizational AI policies, governance structures, and internal resources.

Separately, CHAI has worked with the Joint Commission on governance playbooks aligned with its voluntary Responsible Use of AI in Healthcare certification. The PULSE announcement doesn’t state that participating public health agencies will be assessed under that certification.

Dr. Brian Anderson, chief executive of CHAI, said public health agencies entered the COVID-19 pandemic after years of limited investment in technology. He said PULSE was intended to give agencies practical experience with AI before wider implementation.

For more on how AI is being applied in healthcare settings, read our coverage of Bunkerhill’s $55M raise for agentic AI across health systems. And if you’re interested in the broader AI landscape, check out our analysis of AI and big data trends in healthcare.

Continue Reading

Artificial Intelligence

Moonshot’s Kimi K3 Could Match Anthropic’s Best — And It’s Open Source

Published

on

Moonshot Kimi K3

The Next Leap in Open-Weight AI

Chinese AI lab Moonshot AI is about to release a model that could fundamentally change how enterprises think about paying for frontier artificial intelligence. According to a report from the Financial Times, the upcoming Kimi K3 is expected to perform on par with — or even surpass — Anthropic’s Opus 4.8. That’s a bold claim, but one backed by the lab’s recent track record.

The Kimi K2 models already turned heads in the open-source community. They scored high on standard benchmarks and showed capabilities that weren’t far behind the latest proprietary systems. K3, insiders say, takes that momentum further. It’s designed to close the gap with closed-source giants like OpenAI and Anthropic.

What Makes Kimi K3 Different

Size matters here. The Kimi K3 will reportedly be the largest open-weight AI model ever released from China, with a parameter count landing somewhere between 2 trillion and 3 trillion. For context, that dwarfs many of the most capable models on the market today. And it won’t stay behind closed doors for long — the FT report says it will be released “in the coming days.”

That timeline is aggressive. It suggests Moonshot is racing to capitalize on a moment when enterprises are rethinking their AI budgets. Why pay a premium for proprietary models when an open-weight alternative can do the same job for a fraction of the cost?

A Valuation That Reflects the Ambition

Moonshot is also reportedly raising fresh capital at a valuation of $31.5 billion. That’s a significant jump from the $20 billion valuation it commanded back in May, when it raised $2 billion. Investors are clearly betting that open-weight models will carve out a major slice of the AI market — and that Moonshot will be the one delivering them.

The Enterprise Shift Toward Open-Source AI

The timing couldn’t be better for Moonshot. A growing number of business leaders are questioning whether it’s worth paying for expensive, closed-source models from labs like OpenAI and Anthropic. The fear? That these companies will somehow extract and use the data clients submit through products like ChatGPT and Claude.

That worry isn’t theoretical. It’s driving real decisions. Executives are now actively pitching their own in-house models as safer alternatives. Others are telling companies to take cheaper open-source models — from labs like DeepSeek, Z.ai, or Moonshot — and fine-tune them for specific use cases.

The argument is gaining real traction. Especially as Chinese open-source models continue to close the performance gap with their more expensive, frontier counterparts.

How Kimi K3 Could Reshape the Market

If the Kimi K3 delivers on expectations, it could accelerate a trend that’s already underway: the commoditization of large language models. When a 3-trillion-parameter open-weight model can match Anthropic’s Opus 4.8, the premium for proprietary systems becomes harder to justify.

This isn’t just about benchmarks. It’s about control. Companies that adopt open-weight models retain ownership of their data and can customize the model to their needs. They aren’t locked into a vendor’s roadmap or pricing changes.

Of course, there are trade-offs. Running a model the size of Kimi K3 requires serious infrastructure. Not every organization has the compute to handle it. But for those that do, the economics are compelling.

The Bigger Picture: China’s AI Ambitions

Moonshot’s rise is part of a broader story. Chinese AI labs have been quietly catching up — and in some areas, leapfrogging — their Western counterparts. The release of models like DeepSeek’s R1 and now Moonshot’s Kimi K3 shows that open-weight development in China is no longer a sideshow. It’s a central force.

The political implications are real, too. As export controls tighten on advanced chips, Chinese labs have had to innovate on efficiency and architecture. The fact that Kimi K3 can compete with top-tier Western models despite these constraints says a lot about the ingenuity at play.

For enterprises evaluating their next AI move, the message is clear: the days of assuming closed-source is inherently better are ending. The Kimi K3 might just be the model that proves it.

Continue Reading

Trending