Connect with us

Artificial Intelligence

Apple’s new Home AI features come with a hidden price tag — and it stings

Published

on

Apple Home AI features

Apple’s new Home AI features are here — but only if you’re willing to pay

At WWDC 2026, Apple unveiled a suite of new AI-powered features for its Apple Home app. Auto-updating notifications. Smarter camera search. Automatic tracking and stitching of multiple video feeds for a single event. Higher-resolution recordings. They sounded like genuine quality-of-life upgrades.

But there was a catch. Apple didn’t say which iCloud+ plans would get access. Now, buried in the release notes of macOS Golden Gate beta 3, we have the answer. And it’s not good news for anyone on a budget.

The hidden cost of Home AI features

Apple’s iCloud+ pricing tiers range from cheap to steep:

  • 50GB: $0.99/month
  • 200GB: $2.99/month
  • 2TB: $9.99/month
  • 6TB: $29.99/month
  • 12TB: $59.99/month

I wasn’t expecting the 50GB plan to get the new AI features. That plan only supports one camera via HomeKit Secure Video. But the 200GB tier felt like a safe bet. It supports up to five cameras — exactly the kind of multi-camera setup that benefits most from AI summaries and cross-camera search.

Wrong. Apple has locked the Apple Home AI features behind the 2TB iCloud+ plan and above. That means you have to pay at least $9.99/month to use them. No exceptions.

Why does 2TB feel like the wrong cutoff?

Let’s look at how HomeKit Secure Video works today. The 50GB plan gets you one camera. The 200GB plan supports up to five. The 2TB plan removes the camera limit entirely.

So here’s the logic gap: a user with a 200GB plan and five cameras already has the hardware setup that would benefit most from AI features. They have multiple angles, multiple events to track, and a real need for smart summaries. Yet Apple is telling them to pay triple — jumping from $2.99 to $9.99 — just to unlock software that their existing setup is perfectly suited for.

A cash grab or a cost shift?

This feels like a deliberate push. Apple’s AI infrastructure is expensive to run. Processing video on-device is one thing, but cloud-based AI features — especially those that stitch multiple feeds and generate summaries — require serious server power. By restricting these features to the 2TB tier, Apple is effectively asking users to subsidize its rising AI costs.

Is that unfair? A little. HomeKit Secure Video already requires a paid plan. Adding AI on top of that and demanding a higher tier feels like double-dipping. Especially when the 200GB plan already supports multi-camera setups that would benefit most.

What you actually get with the Home AI features

For those willing to pay, here’s what the new features include:

  • Auto-updating notifications: Alerts that evolve as events unfold, instead of static pings.
  • Smarter camera search: Search by object, person, or activity across all your cameras.
  • Automatic video stitching: Multiple camera feeds for a single event are merged into one timeline.
  • Higher-resolution recordings: Better clarity for playback and analysis.

These are genuinely useful — especially for people with multiple cameras covering a large property. But the price of admission is steep.

Bottom line: The 200GB tier deserved better

Apple could have made the Apple Home AI features available starting at the 200GB tier. That would have aligned with the existing camera limits and given users a clear upgrade path. Instead, they’ve forced a jump from $2.99 to $9.99 — a 234% increase — for features that many 200GB users would actually use.

It’s a smart business move for Apple, no doubt. But for users, it feels like a cash grab. If you’re running a five-camera HomeKit setup on a 200GB plan, you now have a tough choice: pay more or go without.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Artificial Intelligence

Claude Cowork can now keep working even after you close your laptop — here’s what that means

Published

on

Claude Cowork

Your laptop is no longer the anchor

Until today, using Anthropic‘s Claude Cowork meant keeping your desktop awake and running. Shut the lid? The task died. That limitation is gone now.

Anthropic is rolling out Cowork to web and mobile starting today. Max plan subscribers get first dibs on the beta over the next few days. Other plans will follow in weeks, not months. The core promise? You can start a job on your desk machine, walk away, and check progress later from your phone.

How Claude Cowork actually works

If you haven’t tried it, the concept is straightforward. You hand Claude a task — sorting files, scanning your inbox, updating a calendar, whatever — and it churns through your connected tools until it’s done. The catch, until now, was that your laptop had to stay on and awake. That made long-running jobs impractical.

The new update shifts those tasks to the cloud. Scheduled work runs remotely. So you can close your laptop, take a nap, and Claude keeps going. When it hits a fork in the road — a decision requiring your input — it pings you. Nothing gets sent or executed without your approval. No rogue agents here.

Who’s actually using this thing?

Anthropic also dropped usage numbers alongside the expansion. And they surprised me. I assumed Claude Cowork was primarily a developer tool — something for coders and programmers. Turns out I was wrong. The company says over 90% of Cowork usage falls into everyday business operations and content creation.

That’s a telling stat. It suggests AI agents are crossing over from niche technical tools to something broader. People drowning in spreadsheets, repetitive emails, or presentation decks are the real users. If you’ve ever wanted to hand off the boring parts of your job, this update makes that more practical.

What changes with web and mobile access

The expansion means two things. First, you’re no longer tethered to a single machine. You can kick off a task from your desktop, then check in from your phone while commuting or grabbing coffee. Second, cloud-based scheduling means tasks don’t pause when you step away.

This matters for anyone who has tried running a long data cleanup or report generation on a laptop. The old model forced you to keep the machine awake, which is annoying and wasteful. Now, the work lives in Anthropic’s cloud infrastructure. Your laptop is just the launchpad.

What you still can’t do

Claude still won’t act autonomously on sensitive actions. Every output that touches the outside world — sending an email, posting a file, updating a shared calendar — requires your explicit go-ahead. Anthropic is clearly cautious about trust and safety. That’s a good thing, even if it means you can’t fully “set and forget” everything.

What this means for the AI agent space

The usage numbers tell a story. Over 90% non-developer usage is a strong signal that AI agents are finding a real product-market fit outside of engineering teams. Business operations folks, marketers, writers, and project managers are the ones leaning in.

I expect this expansion to accelerate that trend. Once people realize they don’t need to keep a laptop running to use Anthropic’s AI agent, adoption should climb. The friction point was obvious: nobody wants to babysit a laptop for a background task. Removing that barrier makes the tool dramatically more useful.

If you are a Max subscriber, you can test the beta now. For everyone else, the wait is weeks, not months. In the meantime, start thinking about which tedious tasks you’d hand off to an AI that never sleeps.

Continue Reading

Artificial Intelligence

The AI agent evaluation gap: Enterprises trust their tests less than they trust their agents

Published

on

AI agent evaluation gap

Half of enterprises shipped a failing agent

That number should stop anyone building AI agents for customers cold. According to new VentureBeat Pulse Research surveying 157 enterprise organizations, 50% have deployed an agent or LLM feature that passed internal evaluations — and then caused a customer-facing failure in production. A quarter have seen it happen more than once.

The finding lands like a punch. It means the standard pre-deployment gauntlet — unit tests, red-teaming, automated evals — is letting bad agents through. The test says go. The agent breaks. The customer pays.

Only 36% of organizations report no such failure. The rest either don’t run pre-deployment evaluations at all (8%) or don’t track root causes closely enough to know (6%).

Trust in automated evaluation is almost nonexistent

Ask enterprise leaders how much they trust automated evaluation today, and the answer is brutal: just 5% say they fully trust it. That leaves 95% with a specific complaint holding them back.

The top grievance, cited by 29% of respondents, is the one that explains the failure rate: evaluations align poorly with real-world outcomes. A passing score in the test environment doesn’t predict what happens when real users, real data, and real edge cases show up.

Bias and inconsistency (21%) come next, followed by lack of explainability (18%) — organizations can’t always understand why an evaluation reached its verdict. Another 17% cite data leakage or privacy concerns in the evaluation process itself.

So the tests meant to certify agents are, broadly speaking, not trusted to certify them. That makes what comes next all the more surprising.

Autonomy is accelerating — despite the trust gap

Here’s the paradox at the heart of the research. Even though almost no one fully trusts automated evaluation, two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within twelve months (33%).

Only 22% rule it out for the foreseeable future.

The direction is clear: enterprises are moving to let evaluations gate production autonomously, removing the human check, at the same moment they say those evaluations don’t reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it.

Notably, this isn’t just a startup phenomenon. Larger enterprises (2,500+ employees) are slightly more likely than smaller ones to be on the zero-human-review path (70% versus 64%) and slightly more likely to have shipped a failing agent (54% versus 48%). The assumption that big, regulated organizations hold the human in the loop longest is, in this sample, backwards.

The evaluation stack is fragmented — and provider-led

Ask which agent reliability or evaluation platform enterprises primarily use, and the market has no clear leader. Provider-native tooling leads: OpenAI‘s native evals and traces (17%) and Anthropic‘s Claude Console evals (13%) together outweigh any independent platform.

But they’re tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all.

The specialist vendors — DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize — are scattered across single to low double digits. Another 11% have built their own. No independent platform has yet become the category standard.

Production monitoring mostly watches uptime, not correctness

Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up, how fast, at what cost, any errors. Or it can watch whether the agent’s output is correct — automated checks on each answer’s content.

The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring. The request completes. The response is fast. No error is thrown. Everything reads healthy.

The split is stark: 51% of organizations monitor only whether the agent is functioning, while just 23% monitor whether its answers are right. Roughly three-quarters run no automated, real-time evaluation of output correctness in production. They’re taking correctness on faith.

What drives tool selection — and what’s next

Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%). Breadth of observability (13%) and vendor roadmap (4%) matter far less.

On what success looks like, more than a third (36%) name evaluation consistency — getting the same verdict on the same behavior every time. That’s well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%).

The emphasis on consistency is telling: before enterprises can trust an evaluation’s verdict, they need it to be stable — the very property whose absence (bias and inconsistency) ranked among the top trust limitations.

A tooling reshuffle is coming

The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months. 31% plan to do so within the next quarter.

The consideration set points where current usage is thinnest: DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI’s native evals (13%) and Braintrust (9%). The open-source specialists are drawing more interest than their present footprint suggests.

Given that so many enterprises today rely on provider-native tools or nothing at all, this is less a defection than a first real wave of tooling adoption — the moment the evaluation layer starts to consolidate.

The bottom line: An evaluation gap that autonomy will widen

Organizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer. Almost none fully trust automated evaluation, chiefly because it doesn’t match real-world outcomes. Most watch production for uptime and cost rather than for whether the agent’s answers are right.

Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone.

The vendor market is early and unsettled. Encouragingly, the next dollar is going to observability and — pointedly — human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market. But the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust.

The evaluation gap is not a coverage problem that more tests alone will close. It is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy — or whether the false-confidence failures move from customer incidents into changes that deploy themselves.

Continue Reading

Artificial Intelligence

Microsoft pushed Copilot everywhere, but barely anyone bought it, and even fewer use it: Report

Published

on

Copilot adoption

Microsoft’s big AI bet isn’t paying off — yet

Over the past two years, Microsoft has shoved its Copilot AI assistant into almost every corner of its ecosystem. It’s in Windows 11. It’s in Edge. It’s baked into Word, Excel, and Teams. New laptops ship with a dedicated Copilot key on the keyboard. The company wanted AI to become as routine as checking email.

But the numbers tell a different story. A very different story.

Microsoft recently disclosed that Copilot 365 has more than 20 million paid seats. That sounds big — until you realize Microsoft 365 has over 450 million paid commercial users. That’s less than 4.5% of the customer base. And the real usage figures are even worse.

Most paid Copilot seats sit idle

Paying for Copilot doesn’t mean people actually use it. According to enterprise surveys cited in a new report, only 20% to 30% of licensed Copilot seats see weekly engagement. Do the math: that leaves roughly 4 million to 6 million weekly active users. That’s about 1% of Microsoft 365’s total commercial audience.

One percent.

These numbers cover the paid Microsoft 365 Copilot product — the one that works across emails, meetings, and internal documents. They don’t include the free consumer chatbot or the basic Copilot Chat that comes with eligible subscriptions. So the paid product, the one Microsoft charges extra for, is barely being touched.

Why companies buy seats nobody uses

Here’s the pattern: an enterprise buys thousands of Copilot licenses during a company-wide AI rollout. Leadership gets excited. Then employees mostly ignore it. The floating Copilot button becomes digital wallpaper.

Microsoft knows this. The company has started letting Office users hide that floating Copilot button. Some organizations can even uninstall the Windows Copilot app entirely. After a wave of user backlash — sometimes called “Microslop” online — Microsoft also pulled back Copilot branding from some inbox apps.

It’s a quiet admission that the forced integration strategy has limits.

Microsoft raised prices anyway

Despite the lukewarm reception, Microsoft didn’t hesitate to hike prices. Earlier this month, the US monthly price for Business Basic jumped from $6 to $7. Business Standard went from $12.50 to $14. Several enterprise and frontline plans increased by 5% to 33%.

And if you want the full Copilot experience bundled in, you’re looking at subscriptions of $23.50 per user per month for Business Standard with Copilot, or $32 per user per month for Premium with Copilot.

That’s a steep ask for a tool most employees aren’t opening.

What this means for enterprise AI adoption

The Copilot adoption numbers highlight a hard truth about enterprise AI: buying licenses is easy. Changing how people work is not.

Microsoft’s strategy of flooding every product with AI prompts assumes users will naturally discover and adopt the assistant. But the data suggests most office workers either don’t see the value, find the feature intrusive, or simply forget it exists. The Copilot key on new laptops may be the most visible symbol of this mismatch — a button dedicated to a service most people don’t use.

For IT leaders, the lesson is clear. Rolling out AI tools requires more than just flipping a switch. It demands training, workflow integration, and a clear answer to the question every employee asks: “What does this do for me?”

Microsoft hasn’t answered that question convincingly enough. And until it does, Copilot will remain what it is today: a feature that’s everywhere, but used almost nowhere.

Continue Reading

Trending