Connect with us

Artificial Intelligence

Google’s Gemini 3.5 Pro is late. Here’s why that matters more than you think

Published

on

Gemini 3.5 Pro delay

Google’s next big AI model is slipping

Google helped ignite the modern AI race, but staying on top has proven far harder than joining it. A new Bloomberg report claims the company has fallen months behind its internal schedule for launching Gemini 3.5 Pro, its next flagship model. The culprit? Coding.

Engineers are reportedly still wrestling with the model’s ability to write and debug software — a weakness that has become impossible to ignore. The delay isn’t just about polishing a chatbot. It exposes a deeper problem: Google’s sprawling structure, competing divisions, and stricter safety reviews are slowing its response to rivals who move at breakneck speed.

While OpenAI, Anthropic, and Meta keep shipping increasingly capable models, Google appears stuck balancing innovation with the trust it has built across products used by billions. That balancing act is getting harder by the week.

Coding remains Gemini’s biggest weakness

Bloomberg’s report, citing multiple current and former Google employees, says Gemini 3.5 Pro has been delayed because the company hasn’t hit its internal targets for coding performance. Google even refreshed the model’s training data late last month to boost its coding chops. The results? Reportedly underwhelming.

This matters because writing code has become the clearest benchmark separating today’s leading AI models. OpenAI, Anthropic, and Meta have all poured resources into developer-focused systems that can write, debug, and reason through complex software projects. According to the report, both OpenAI and Meta currently outperform Google’s available models in this arena.

Google, for its part, insists progress is happening. In a statement cited by Bloomberg, the company said it is testing Gemini 3.5 Pro, an upgraded Flash model, and other AI systems with partners. It also noted ongoing discussions with the US government around testing standards and AI safety.

The timing is awkward. Many observers expected Gemini 3.5 Pro to debut at Google I/O earlier this year. Instead, Google offered incremental Gemini updates while competitors kept shipping frontier models.

Google’s scale is a double-edged sword

Unlike most AI startups, Google isn’t building models in isolation. Every major Gemini release eventually needs to work across Search, YouTube, Maps, Android, Workspace, Cloud, and dozens of other products. That scale gives Google massive advantages — including unmatched access to real-world data. But it also creates layers of internal coordination that can slow decision-making to a crawl.

Employees describe competing priorities across DeepMind, Google Cloud, Android, and other teams. Overlapping AI coding efforts make it harder to maintain a unified strategy. Former employees also mentioned internal disagreements over AI-generated code, plus earlier restrictions on using Gemini for software development, which limited experimentation during the tech’s early rollout.

Policy shifts and internal friction

Google says those policies have evolved. The company claims roughly 75 percent of its production code is now generated using AI. Internal coding tools are being consolidated under a common platform called Google Antigravity. Engineers are now expected to use AI for coding, although some still face computing capacity constraints due to intense internal demand for GPU resources.

But frustration is simmering. Some researchers have reportedly left for competitors like Anthropic. Customers are split on Gemini 3.5 Flash, too. Companies like Figma praise its balance of speed and quality. Others, including education platform Platzi, say it sits in an awkward middle ground — higher costs than previous Flash models without matching the reasoning power of premium rivals.

The real question for Google

The bigger picture is that Google’s AI challenge is no longer about proving it can build frontier models. Few doubt that it can. The real question is whether a company of Google’s size can ship those models quickly enough in an industry where competitors now measure progress in weeks, not months.

For now, the Gemini 3.5 Pro delay is a telling sign. Google’s next move will reveal whether it can adapt its internal machinery to the pace of the AI race — or whether it gets left behind in the very revolution it started.

Want to stay updated on Google’s AI moves and the broader AI model competition? Keep an eye on this space. Also check out our breakdown of Google Antigravity coding tools and how they fit into the bigger picture of developer-focused AI systems.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Artificial Intelligence

Cognition, the startup behind AI coder Devin, is reportedly raising again at a $40B valuation

Published

on

AI coding startup

Another round, another massive jump

Three months. That’s how long it’s been since Cognition last announced a funding round. Now, according to Bloomberg, the AI coding startup is already talking to investors again — and the valuation target is eye-watering.

The company, best known for its AI coding agent Devin, is reportedly seeking a round that would value it at no less than $40 billion. That’s a hefty leap from the $26 billion valuation it landed in May, when it raised $1 billion in a round led by some of the biggest names in tech investing.

What’s driving the jump? A simple number: $1 billion in annualized revenue run rate. Sources familiar with the talks told Bloomberg that hitting that milestone is the key to securing the $40 billion price tag.

Devin’s revenue is climbing fast

When Cognition announced its May raise, CEO Scott Wu told TechCrunch the company had already reached a $492 million annualized revenue run rate. He also shared a figure that explains why investors are so keen: enterprises were growing their usage of Devin by 50% month-over-month for six straight months.

That kind of growth doesn’t stay quiet for long. If the current trajectory holds, hitting $1 billion in annualized revenue run rate within a few months is not out of the question.

Who’s using Devin?

Cognition says its customers include some serious names: Mercedes-Benz, NASA, and Goldman Sachs. That’s a mix of automotive, aerospace, and finance — hardly a narrow niche.

Not a human replacement, says Wu

Despite the hype around AI agents taking over coding jobs, Wu has been careful to frame Devin differently. He told TechCrunch the tool isn’t being sold as a replacement for human programmers. Instead, it’s often assigned the long-tail grunt work that many developers dislike — things like updating legacy software or migrating applications from one platform to another.

That positioning seems to be working. It’s a lot easier for an enterprise to say yes to a tool that handles the boring stuff than one that threatens to put its engineering team out of work.

What this means for the AI coding market

If Cognition closes this round at $40 billion, it would cement its place among the most valuable AI startups in the world. It also signals that investors are still willing to pay premium prices for AI coding agents, even as the broader market for AI funding shows signs of cooling.

The bigger question is whether Devin can keep up the momentum. A 50% monthly growth rate is impressive, but it also gets harder to sustain as the base gets bigger. And competitors are circling — from established players to newer entrants in the AI coding assistant space.

For now, though, the numbers are on Cognition’s side. And if the company hits that $1 billion run rate, the $40 billion valuation might not even be the peak.

Continue Reading

Artificial Intelligence

China’s Kimi K3 is an open-weight giant that’s nipping at OpenAI’s heels

Published

on

Kimi K3 open-weight

China’s new heavyweight challenger

Moonshot AI just dropped Kimi K3, a 2.8-trillion-parameter model that’s squarely aimed at the top of the AI food chain. It’s built for coding, research, reasoning, and vision tasks, and the company isn’t shy about comparing it to the best closed models from OpenAI and Anthropic.

The catch? Moonshot itself admits K3 still trails Claude Fable 5 and GPT 5.6 Sol overall. But the benchmark numbers tell a more interesting story.

How close is Kimi K3 to the best closed models?

On several key tests, K3 actually edges out both rivals. It scored 77.8 on Program Bench, narrowly beating Fable 5’s 76.8 and GPT 5.6 Sol’s 77.6. It also led BrowseComp with 91.2 and SWE Marathon with 42.0, ahead of both.

Other results show the gap hasn’t closed entirely. K3 scored 67.5 on DeepSWE, versus 70.0 for Fable 5 and 73.0 for GPT 5.6 Sol. Moonshot also says the overall user experience still lags behind the proprietary pair.

There’s an important caveat to all this. Moonshot notes that Fable 5 results may include fallbacks to another model, and GPT 5.6 Sol results may include cyberguards that restrict certain responses. In other words, the comparison is useful — but not perfectly even.

Open weights put pressure on OpenAI and Anthropic

The most significant part of K3 might not be its performance at all. Moonshot plans to release the full model weights by July 27, meaning you can download and run K3 locally — provided you have the hardware to handle a model this size. Developers can also modify and fine-tune it for specific tasks.

That open-weight approach is a direct challenge to the closed strategies of OpenAI and Anthropic. When a near-frontier model is freely available, it gets harder to justify the premium prices those companies charge.

Pricing that undercuts the U.S. giants

K3 isn’t the cheapest Chinese model on the market, but it’s still significantly more affordable than its U.S. counterparts. Its API pricing is $0.30 per million cached input tokens, $3 for uncached input, and $15 for output. That puts it closer to Anthropic’s mid-range offerings than to the rock-bottom prices of earlier Chinese releases.

Chinese AI models have long been cheaper than American ones, and that gap is already pushing some U.S. startups to adopt them to cut costs. Kimi K3’s near-frontier performance at a lower price only adds to that pressure.

What this means for the AI race

The release of Kimi K3 signals that the competitive landscape is shifting. Open-weight models are no longer a step behind on every metric — they’re trading blows with the best closed systems on specific benchmarks. For developers who want control and cost savings, that’s a compelling combination.

It also raises a question: how long can OpenAI and Anthropic maintain their premium pricing when a model like K3 is close enough on performance and free to download? The answer may determine the next phase of the AI market.

If you’re weighing your options, it’s worth keeping an eye on how K3 performs in real-world use, not just on benchmarks. And for those who prefer the closed-model route, the pressure from open-weight rivals could lead to better pricing or features down the line.

Continue Reading

Artificial Intelligence

OpenAI pushes ChatGPT into patient health records

Published

on

ChatGPT Health feature

OpenAI’s ChatGPT Health feature goes live

OpenAI has switched on a new Health feature inside ChatGPT that lets users link Apple Health data and medical records directly to the chatbot. Anyone logged in, aged 18 or over, can access it now on web and iOS, across the Free, Go, Plus, and Pro tiers.

The integration pulls in medications, lab results, recent visits, sleep data, and activity logs. Once synced, ChatGPT can use that context in any conversation, not just a dedicated health section. That’s a deliberate design shift, and it stems from something surprising the company found during early testing.

Why OpenAI redesigned the health experience

Earlier, OpenAI ran a pilot where users had to open a separate health area to get responses grounded in their own data. The result? More than 70 percent of health-related conversations happened elsewhere—smack in the middle of meal planning or an unrelated symptom query, not inside the dedicated space.

That insight drove the redesign. Instead of forcing users into a specific mode, ChatGPT now draws on connected health information across any conversation, provided the user has granted permission. A person planning a dinner out might get a restaurant suggestion that accounts for a logged dietary restriction. Someone asking about weekend plans might get activity suggestions adjusted for a recent injury noted in their synced records.

The Health tab in the sidebar still exists, but its role has shifted to being a management hub: connecting accounts, reviewing synced data and trends, browsing suggested prompts, and returning to past health-related chats.

What early testers of the Health feature in ChatGPT report

OpenAI published numerous accounts from its early access group, and the details are worth weighing against the company’s own framing of the tool as support rather than diagnosis.

Blake, a technical program manager, said: “The most useful part has been turning scattered medical history into something I can actually understand and explain. I have multiple overlapping issues and ChatGPT helped connect those pieces into a clear timeline, explain the medical terms in plain English, and create summaries I could share with a physical therapist or trainer.”

On the shift from disconnected records to a usable pattern, Blake added: “Instead of just seeing disconnected diagnoses, imaging results, and surgery notes, I could understand the bigger pattern. It made the information more usable and gave me better language to advocate for myself with providers and trainers.”

Reweti, a portfolio manager, pointed to longitudinal analysis as the differentiator over a standard search or a one-off doctor visit: “What I want is to infer patterns that aren’t obvious and make connections I wouldn’t have made on my own. ChatGPT can access my existing labs because I’ve connected everything, and it can look over time—that’s the big advantage. It’s like having a research analyst. It allows me to be more proactive and own more of my health journey.”

However, not every account was frictionless, and one is worth flagging given the stakes involved in surfacing clinical data through a chatbot interface. Shannon, a nurse, described finding an unexpected entry in her own chart through the tool.

“Using Health has actually reinforced something I’ve believed for a while: one of AI’s greatest strengths isn’t replacing healthcare professionals, but helping patients better understand and navigate their own health information,” explained Shannon. “Discovering an unexpected chart entry through Health really highlighted that for me. It wasn’t AI creating a problem. It helped me identify something I can now appropriately follow up on with my healthcare providers.”

That distinction—between a tool surfacing something for human follow-up versus a tool making a clinical call—is the line OpenAI needs the product to hold as usage scales.

Other testers focused less on clinical nuance and more on the practical grind of manual data wrangling the feature is meant to replace. Daniel, a consultant, connected multiple sources and found the combined view more useful than isolated chat sessions.

“It’s been great for coordinating labs and translating them into language I understand. Connecting Apple Health and MyChart makes the insights more grounded in what is happening across my life outside of just chat interactions,” said Daniel.

Kathleen, a small business owner who’d lost a decade-long fitness habit to work pressure, described a lower-stakes but still concrete use case, with the model adjusting suggestions based on activity gaps it could see directly.

“I’ll say, ‘I need something to help me move today,’ and Health can see I haven’t worked out in the last seven days and suggest starting slowly—with a walk or some stretching. It’s helping me make things manageable and get back into it in a reasonable way,” explained Kathleen.

Carlton, an operations manager, had previously resorted to exporting spreadsheets from Apple Health and uploading them manually before every chat, a workaround the new integration is designed to eliminate.

“Prior to Health, I was exporting massive spreadsheets from Apple Health and importing them into ChatGPT. Now, it feels a lot more streamlined. Being able to see years and years of my fitness journey in Health helped me see things in a different light—a bigger picture,” said Carlton.

OpenAI puts weekly health-related ChatGPT queries at north of 300 million people, covering everything from decoding a lab result to prepping for a doctor’s visit. That figure, if accurate, puts ChatGPT in a position most digital health platforms would need years and considerable marketing spend to reach.

The company is explicitly positioning Health as a support tool rather than a diagnostic one, and telling users to confirm anything important with their actual healthcare provider.

OpenAI’s AI model performance claims and how they were tested

OpenAI attributes the feature’s viability partly to newer models: GPT-5.5 Instant, available to Free users, and GPT-5.6 Sol, reserved for paid tiers.

The company says GPT-5.5 Instant made gains in recognising when urgent care might be needed and in explaining uncertainty, and that on its toughest health evaluations it performed comparably to OpenAI’s frontier Thinking models at the time. GPT-5.6 Sol is described as the company’s strongest health model so far, built for reasoning across multiple data points such as lab trends over time.

To validate these claims, OpenAI says it worked with hundreds of physicians to build health scenarios and rubrics scoring responses on accuracy, safety, communication, context awareness, completeness, and appropriate escalation to professional care. The company reports that every GPT-5.6 model outperformed GPT-5.5 on HealthBench Professional, an internal evaluation built for this purpose.

A chart included in OpenAI’s announcement shows GPT-5.6 Sol scoring higher than GPT-5.5 Instant and GPT-4o across categories including accuracy, communication, completeness, and following instructions, with physician-written responses used as a comparison baseline.

OpenAI does say physicians tested the live Health product before release specifically to assess real-world performance and safety with connected data, which is a step beyond benchmark scoring alone, though the company hasn’t published the methodology or results of that testing in detail.

Health data handling and the permission architecture

Connected medical records and Apple Health data—along with any conversations that draw on them—are excluded from foundation model training and ad targeting, according to the company, regardless of a user’s broader ChatGPT training settings. Conversations that don’t touch Health data still follow whatever training preference a user has set separately.

Access is permission-gated by default. ChatGPT asks before using connected health data to personalise a response, though users can switch to “always allow” and turn off the prompts entirely. That setting lives in Settings > Plugins > Health and can be reversed at any time. Disconnecting a data source triggers deletion from OpenAI’s systems within 30 days, though anything already surfaced in existing chat history sticks around until the user deletes those conversations manually.

Memory creation is scoped narrowly, too. Memories can be created from health conversations but not directly from the raw connected records or Apple Health data itself. Users wanting to avoid memory creation altogether can use Temporary Chat or disable memory in settings.

OpenAI also flags a specific edge case: actions that could expose Health data through other connected plugins, such as sending a training plan built from Apple Health metrics to a running partner. The company says additional checks apply before such actions execute, and that for sensitive cases ChatGPT may ask for explicit confirmation. OpenAI states it runs red teaming exercises targeting these scenarios, though no findings or failure rates from that testing have been made public.

What happens to accuracy at the edges

The practical friction point sits with data quality. A medication can stay listed in a patient’s synced record long after they’ve stopped taking it, and OpenAI uses that exact scenario to illustrate why synced data isn’t automatically current. Its guidance to users is blunt: flag changes to ChatGPT directly and check anything important against the original source rather than trusting the sync to stay accurate on its own.

Wearable and fitness app data carries its own gaps, since availability depends on what each third-party app chooses to share through Apple Health, and OpenAI notes some proprietary scores from fitness apps may not transfer at all.

The relevant question isn’t whether ChatGPT can summarise a lab result correctly in a demonstration, it’s whether the permission model, data deletion timelines, and escalation logic hold up when a user’s synced records are three months stale and the model is asked to reason across contradictory inputs.

OpenAI’s physician testing addresses part of that concern; it doesn’t close the distance between a controlled evaluation and a user managing multiple chronic conditions with incomplete data syncing from four different apps.

For those wishing to give the feature a try, it’s live now for eligible US users through the sidebar Health menu.

See also: OpenAI Presence sells enterprise AI agents with engineers attached

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.

AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.

Continue Reading

Trending