Five years and $27.7 billion after Salesforce acquired Slack, the two products are finally starting to function as a single system. On Tuesday, Slack launched an integration that connects Slackbot — the personal AI agent built into every workspace — to the entire Salesforce platform, including CRM data, Tableau analytics, Data 360 customer profiles, and a growing constellation of third-party applications, all through a single conversational prompt.The mechanism behind the expansion is a set of dedicated Model Context Protocol (MCP) servers from Salesforce that connect Slackbot to the company’s Headless 360 infrastructure. In practical terms, a salesperson can now ask Slackbot for a customer’s deal history, receive a live Tableau visualization of pipeline trends, update a CRM record, and trigger a DocuSign approval — without ever switching tabs or logging into another application. According to Slack, the Salesforce IT team has already used this architecture to save its 1,500-plus engineers “thousands of custom coding hours annually.”The timing is not accidental. Slack is making this move amid escalating competitive pressure from Microsoft Teams, which claims 320 million-plus monthly active users and has Copilot embedded across the Office suite, and from Google, which continues to weave Gemini deeper into Workspace. And just days ago, The Information reported that some smaller companies are using Anthropic’s Claude to replace Salesforce CRM entirely — one Atlanta-based property management firm with about 55 employees reportedly saved around $100,000 annually by building a custom replacement using Claude Code and Replit.Against that backdrop, Slack CMO Ryan Gavin sat down for an exclusive interview with VentureBeat to frame the announcement and argue that the company’s future depends on an idea he calls “multiplayer AI” — and that the 25 years of customer data locked inside Salesforce is an asset no vibe-coded alternative can replicate.Why Slack’s CMO believes ‘multiplayer AI’ is the next big enterprise battlegroundGavin’s core argument is that the enterprise AI conversation has been stuck in single-player mode for too long, and that Slack is uniquely positioned to break it open.”So much of what we’ve seen are just these incredible tools that have largely been single-player, incredible tools for individual productivity, helping people complete tasks and write code,” Gavin told VentureBeat. “But as we’ve always known at Slack ever since our inception, work is a team sport. For AI to really take hold in the enterprise, it has to be multiplayer.”The distinction matters commercially. Most AI assistants today — ChatGPT, Claude, Copilot — default to one-on-one conversations with a single user. A researcher queries a model, gets a response, and acts on it alone. The insight stays in a private chat window, invisible to colleagues. Gavin argues this creates a new version of the tab-switching problem that plagued pre-AI enterprise software, except now employees are also navigating dozens of individual agent interfaces on top of their existing applications.”It’s going to benefit almost no one if every enterprise application out there spawns hundreds of agent babies, and employees end up in a worse world than they were before,” Gavin said.Slack’s answer is to make Slackbot the orchestration layer. Because everything happens in shared channels, any action an agent takes — pulling a customer profile, flagging a deal risk, updating a Jira ticket — is visible to the entire team. A colleague can redirect, build on, or correct the agent’s work in real time.How MCP and Salesforce’s headless 360 platform power Slackbot’s new capabilitiesThe technical backbone of the announcement is the Model Context Protocol, an open standard originally developed by Anthropic that defines how AI models discover and invoke external tools. MCP has seen rapid adoption across the AI tooling ecosystem. By early 2026, it had been adopted by Claude Code, Cursor, GitHub Copilot, and OpenAI’s tooling, with managed hosting available from AWS, Cloudflare, and Vercel. As a DEV Community explainer puts it, MCP “is the closest thing the AI tooling ecosystem has to a standard.”In this implementation, Salesforce exposes its platform capabilities — CRM records, Tableau visualizations, Data 360 customer profiles, Agentforce agents — as MCP servers. Slackbot operates as an MCP client, connecting to those servers and routing user queries to the appropriate back-end system. When a user asks Slackbot about a customer, the bot discovers which MCP tools are relevant, calls them, and synthesizes the results into a single response — all within the Slack conversation.Gavin explained the architecture in simple terms: “Salesforce is extending what has always been our open platform through our Headless 360 strategy — making all of these MCP endpoints available. And then Slackbot acts as an MCP client, connecting to those MCP servers and bringing all that data in within the confines of a trusted permission platform.”That permission layer is critical. Slackbot respects each user’s Salesforce permissions, meaning a marketing coordinator cannot accidentally access sales pipeline data they are not authorized to see. Validation rules, field-level security, and org-wide data boundary configurations carry over automatically. For admins, setup requires no custom integration code — Salesforce MCP servers can be discovered, installed, and governed from a single UI using the existing Slack-Salesforce connection.Salesforce first introduced the Headless 360 concept at its TDX developer conference in April, positioning it as an API-driven layer that exposes the platform’s data, workflows, and governance controls so that software agents, rather than human users, can execute business processes directly. As CIO.com reported at the time, analysts viewed the move as an effort by Salesforce “to position itself as a central layer for managing agent-driven operations across different business functions.”Slack says it’s betting on openness, not on any single AI protocolWhen asked whether Slack is making a risky bet on MCP as a protocol — given that standards in AI tooling can shift rapidly — Gavin reframed the question entirely.”We’re not betting on MCP, per se. We’re betting on what we’ve always bet on, which is that Slack is an open platform,” Gavin told VentureBeat. “MCP happens to be the best agent-to-agent protocol that the industry is rallying around right now, but if something better came out tomorrow, you’d see the same pattern from Slack — we’re going to stay open. MCP and APIs are simply tools that facilitate that.”That open-platform philosophy is central to Slack’s identity and, Gavin argues, its competitive differentiation. Slack already hosts more than 2,600 app integrations. The new MCP-native partner ecosystem includes Atlassian, Box, DocuSign, Canva, Lucid, Zoom, and more than 25 additional companies, each of whose agents can be added directly to shared Slack channels. MuleSoft Agent, now connected to Slackbot, helps manage integrations for the team — checking system health or surfacing critical error alerts in the same workspace where the team is already collaborating.But MCP is not without trade-offs. The protocol requires tool discovery on every connection, and large tool libraries can consume significant context tokens. One technical analysis noted that a server exposing 300 tools could cost 5,000 to 10,000 tokens per session before the model does any useful work. For an enterprise like Salesforce with hundreds of potential tools across CRM, analytics, and service platforms, careful filtering and segmentation of MCP servers become essential design decisions — a challenge the company will need to navigate as the ecosystem scales.Inside Slack’s complicated relationship with Anthropic and the Claude questionPerhaps the most delicate topic in the interview concerned Slack’s relationship with Anthropic, the AI lab behind Claude — and one of Slack’s most visible power users. Just last week, Anthropic launched Claude Tag, a persistent AI teammate that works inside Slack channels, prompting confusion among Salesforce employees who worried it competes directly with Slackbot and Agentforce. The Information reported internal anxiety about whether Salesforce was welcoming a competitor into its own living room. Salesforce has financial reasons to maintain the partnership: the company reportedly expects to spend $300 million on Anthropic tokens this year and holds a stake in Anthropic.Gavin addressed the tension head-on, framing it as a feature of Slack’s platform strategy rather than a threat.”We’re incredibly excited and bullish about what Anthropic is bringing into Slack. Period. End of statement,” Gavin said. He noted that Anthropic “is building roughly 65% of their code with Claude in Slack,” and pointed out that ChatGPT was originally built in Slack, as was Perplexity.”Building nowadays happens in the open, and every company is going to be building in the open with tools like this, and you need a platform to build in the open,” Gavin said.His argument is that feature overlap between Slackbot, Claude Tag, and other third-party agents is “actually a feature, not a bug” — a sign of a healthy platform rather than a competitive vulnerability. He compared it to an ecosystem where multiple products serve similar needs but win on craftsmanship, ease of use, and integration depth.”One of the reasons Slackbot has been the fastest-adopted feature in Salesforce history is the simplicity, the approachability — underpinned by the trust that comes from having an agent that knows me, knows my tone, knows my work, knows my people, knows my data,” Gavin said.The distinction Slack draws is structural: Slackbot has access to a user’s full workspace context, Salesforce data, permissions, and connected applications by default. Claude Tag, by contrast, only sees the channels it is explicitly added to. For Slack’s leadership, that asymmetry is the moat.How Slack plans to compete with Microsoft Teams and Google in the AI eraAsked directly about competitive positioning against Microsoft Teams and Google Workspace, Gavin pointed to Slack’s open channel architecture as the differentiator no competitor can replicate.”If you spend any time in Teams, it’s a lovely tool for chat, direct messages, and video, but it has no platform for open communication across organizations,” Gavin said. “Its SharePoint-based architecture is fundamentally limiting.”He cited Shopify as an example, where an internal AI agent called River is deployed across approximately 4,400 channels serving 6,000 employees. He also referenced a Fortune report noting that Microsoft’s own head of AI mandated that his team run on Slack rather than Teams — a pointed detail Gavin clearly relished. “There’s a reason for that,” he said. “We’re in an era right now where openness matters, and all the other tools you mentioned, they’re still relatively closed.”The competitive pressure is real and intensifying. Microsoft has integrated Copilot across its entire productivity suite, giving it a distribution advantage that reaches virtually every Fortune 500 company. Google has been similarly aggressive with Gemini across Workspace. And new entrants are crowding the market: a startup called Viktor, which embeds AI agents inside Slack and Teams workspaces, recently raised a $75 million Series A led by Accel — with Slack cofounders Stewart Butterfield and Cal Henderson participating as angel investors.Box, one of the enterprise customers highlighted in the announcement, told Slack it aims to have its sellers complete 75 to 80 percent of their work inside Slack. Gavin repeated that figure as evidence that the platform is becoming the default workspace for entire organizations, not just engineering teams — a shift he believes accelerates as AI makes every employee a builder.Slack’s biggest long-term play is making Salesforce’s CRM useful to everyone in the companyGavin saved what he considers the most underappreciated element of the announcement for last: the democratization of Salesforce’s CRM.For 25 years, Salesforce’s CRM has been used primarily by sales, service, and marketing professionals — a relatively modest percentage of a company’s total workforce. The promise of Slackbot as a conversational interface is that any employee, regardless of their role or technical fluency, can now query and act on CRM data simply by asking a question in natural language.”What most people don’t realize is that this democratization of CRM is going to take its usage from a modest percentage of employees to the entire enterprise,” Gavin said. “When you can make systems like Data 360 or Agentforce for Sales accessible to the entire employee base — not just a percentage — think about how much more valuable those investments become.”He cited Engine, a company that handles 800,000 customer inquiries a year, as an example. Previously, answering a customer inquiry required a specific employee with access to a specific tool to look up a customer’s history. Now, anyone in the company can ask Slackbot and see a complete customer profile, review case history, and write updates — all without being retrained or learning a new interface. Engine’s CEO Elia Wallen, in a statement sent to VentureBeat, described the integration as enabling employees to “make data-driven decisions and take action without leaving the conversation.”The financial logic is straightforward: if Salesforce can make its platform useful to 100 percent of a customer’s workforce rather than the 20 or 30 percent who currently hold licenses, the value of the existing Salesforce investment multiplies without requiring a proportional increase in spending. That pitch becomes especially potent at a time when CIOs are scrutinizing every line of their AI budgets.What analysts and CIOs should watch as Slack rolls out its biggest AI update yetThe announcement is a significant architectural evolution for Slack, but several questions remain unanswered.First, pricing. The company did not directly address whether Slackbot’s MCP-powered Salesforce integration will require additional SKUs or license tiers. As Info-Tech Research Group analyst Scott Bickley cautioned when Headless 360 was first announced in April, “Salesforce’s MO seems to be to announce new capabilities that require SKUs. CIOs should be asking about pricing now.”Second, performance. Routing user queries through MCP servers to Salesforce back-end systems introduces latency that could affect the conversational feel Slack prides itself on. Neither the press release nor the interview disclosed SLAs for MCP tool calls — a gap that enterprise buyers will want addressed.Third, the competitive dynamics of the platform play. Slack’s open-platform philosophy invites powerful partners like Anthropic and OpenAI into its ecosystem, but those same partners are building their own surfaces for enterprise work. Anthropic reportedly plans to expand Claude Tag to Microsoft Teams, email, and other project management tools — meaning the partner Salesforce is paying hundreds of millions a year is building the infrastructure to be useful without Slack at all.And fourth, the broader existential question facing all enterprise software: whether AI agents will ultimately reduce the need for CRM systems entirely. Gavin’s pitch — that Slack makes CRM more valuable by making it more accessible — is the inverse of the bear case. The market will ultimately decide which thesis prevails.Salesforce reported record first-quarter revenue of $11.1 billion in fiscal Q1 2027, with Agentforce ARR surpassing $1 billion for the first time and combined AI and data ARR reaching $3.4 billion. Those numbers suggest the AI strategy is beginning to generate real revenue, even as the company navigates a market that remains uncertain about the long-term trajectory of legacy enterprise software.”Slack has quickly moved from this beloved collaboration tool from the last ten years to now this multiplayer AI platform that we call a work operating system,” Gavin said.Five years ago, Salesforce paid $27.7 billion for what was, at its core, a very good group chat application. On Wednesday, it started trying to prove that group chat was never the product — it was the foundation. In the age of AI agents, the most valuable real estate in enterprise software may not be the database where the data lives. It may be the conversation where the decisions get made.
Venture Beat
The real cost, security, and culture problems behind enterprise AI agents
Presented by Red Hat At VentureBeat’s recent AI Impact event, where the discussion centered on what separates enterprises that scale agentic AI from those that stall in pilot mode, Brian Gracely, senior director of portfolio strategy at Red Hat, detailed what companies actually run into once agents reach production. He dove into cost discipline, the security blind spots unique to autonomous systems, and the organizational friction that determines whether agent adoption spreads beyond early champions.Enterprises are overestimating how far behind they are on AI agentsMany enterprise leaders, especially those following industry keynotes and AI announcements, worry that they’re already falling dangerously behind competitors deploying agents at scale. But according to Gracely, much of that anxiety reflects a misconception about how quickly organizations learn once they begin building. Teams often move up the learning curve far faster than they expect.That rapid progress creates a different challenge, however. As agent usage expands, AI costs rise just as quickly, turning cost management from an engineering concern into a recurring boardroom discussion.Agentic AI usage is orders of magnitude higher than during the chatbot era, making AI costs a growing concern for enterprises. At the same time, organizations are becoming increasingly aware of their dependence on a small number of model providers. According to Gracely, that combination is driving many enterprises to explore alternatives that give them greater control over costs and infrastructure.”The two or three top providers are already telling the market that they’re losing money, and they’re trying to go public to make up those gaps,” he explained. “At some point, the dependency on that means you’re either going to buy at a very high-cost level, or you’re going to figure out alternatives to control what you’re doing.”Right-sizing AI models is the fastest lever for cutting agent costsThe biggest cost issue is that enterprises overspend by defaulting to the most capable model available regardless of task complexity.”If I’m simply trying to resolve an insurance claim, I don’t need to know about the history of Western civilization in my model, I don’t need to know World Cup soccer scores,” Gracely said.Semantic routing is the mechanism many companies use to make that judgment automatically, classifying requests and sending each to a model sized for the task without requiring users to choose, while infrastructure techniques like caching repetitive queries cut how often a request needs to reach GPU compute at all. Together, he said, these tools remove the assumption that efficiency and innovation pull in opposite directions.”There’s a lot you can do at a GPU infrastructure level, and quite a bit you can do in terms of flexibility of models,” he explained. “Those give excellent choices in terms of the levers you’re trying to pull, whether you need efficiency or you need innovation. That shouldn’t be a binary choice.”The financial discipline needed for token spend is similar to the FinOps practices that took years to mature in order to take control of cloud compute spending. Those underlying frameworks will transfer even as the vocabulary changes, Gracely said, especially as organizations push for internal education on model selection so teams stop defaulting to the most prominent option for tasks that don’t need it.”The same way we first had to teach the financial people what an EC2 instance is and what an S3 bucket is, you’re going to have to start explaining tokens to them,” he said. “We don’t always need a Rolls-Royce. We don’t always need caviar, because we’re trying to do basic types of things.”Patch speed is now critical as AI tools find vulnerabilities fasterAI-powered vulnerability discovery is forcing enterprises to rethink how quickly they can identify, validate and deploy patches. Long-established patch management cycles may no longer be fast enough in an environment where AI can uncover — and attackers can exploit — new vulnerabilities much more quickly.”Most companies are probably going to have a window of somewhere between seven and 14 days to stay ahead,” he said. “There are groups, Red Hat included, that are going to build patches for these, but the embargo window is going to be short.”AI is also changing what defenders need to look for. Rather than simply uncovering isolated critical flaws, AI security tools can identify combinations of seemingly minor vulnerabilities that become dangerous only when chained together. As both software complexity and vulnerability discovery accelerate, Gracely argued that the ability to rapidly manage and update software is becoming a strategic capability rather than simply an operational one.Subject matter experts and compliance teams decide whether agents scaleIn the end, organizational adoption comes down to the need for deep, sustained involvement from the subject matter experts whose knowledge the agent is meant to encode, which makes earning their buy-in a prerequisite rather than an afterthought.”You have to think about the incentives, what you do for people who participate in this work so they don’t feel threatened that it’s going to take away their job, and how you incentivize people in the long run to cooperate with that innovation,” he said.Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
The AI architecture that let Liberty Mutual shrug off the Fable 5 outage
When Anthropic’s Fable 5 was pulled from international use for nearly three weeks, some over-reliant businesses were left scrambling.But Liberty Mutual easily pivoted to other platforms. That’s because 18 months earlier, they built their “AI backbone” exactly for this kind of scenario.In this rapidly moving AI landscape, the 114-year-old property and casualty insurance company recognized independence as an operating advantage.“Things are changing so fast, you need a backbone that’s flexible,” Brian Craig, Liberty Mutual’s senior director of architecture, said at a recent VB Impact event. “You can’t lock in right now on one vendor or even one framework.”Enterprises need flexibility to hook into different models and vendors, depending not so much on the “flavor of the day,” but “what can you feel confident about for the next six months,” he said. Runtime versus control plane The company’s “backbone” (or control plane) is its own, while everything underneath remains swappable. The architecture consists of roughly 50 components across security, identity, orchestration, tool restriction, and the policies that govern how agents behave. Each is designed to be independently and immediately replaceable to support interoperability. The agent runtime below this backbone is AWS’s Amazon Bedrock AgentCore; this is not the strategic center, but explicitly “just for running the agents,” Craig said. He chose this offering because it (at least currently) supports multiple frameworks and Liberty Mutual’s model-agnostic philosophy. “We still have flexibility based on what we write,” Craig said. “But if something comes along and is better, we will move to it quite quickly.” The software factoryThis architecture delivers, as proven by Liberty’s “software factory,” an agentic pipeline that automates much of the software delivery process.They started with a business process with clear pain: Onboarding electronic content management documents for insurance products. This repetitive, manual task typically required engineers to code every change. Instead, the team built a factory of coordinated agents working “in tandem and in sequence”: An Epic agent consumes high-level requirements.A Story agent breaks work dictated by the Epic agent into narrow slices within specific application areas. This agent is “constraining the context, because the smaller the context, the better the output.” A planning agent defines the technical execution plan.A coding/testing agent handles coding, testing, and basic review.A triage (critic) agent sits across all other agents, reviewing quality and feeding back improvements.Finally, a librarian agent helps others find the “context of the knowledge for their job.” Craig and his team learned quickly that a single “do everything” agent was a mistake. “You were asking it to do too many things, which meant you had to give it too much information,” Craig said. Splitting into six agents let them dramatically shrink context windows and tighten scope.Once the factory hit production, the impact was immediate. In the initial deployment, they did “about three months of work” in roughly a week. They realized that “the current software engineering process has a massive amount of handoffs, which means there’s a massive amount of wait time,” Craig said. Human-paced automationThe factory is not a fully autonomous pipeline; it runs at the speed of human overseers.Liberty’s first model was a “day shift/night shift” rhythm: Engineers set goals and rules and reviewed the previous night’s outputs during the day, then let the factory run overnight. But in practice, there was never enough work to keep the agents busy all night, and the cadence felt unnatural.They shifted to a more iterative loop. Users decide when to trigger the factory, how far it runs before pausing, and at which points they want to review outputs. “Then the factory kicked in, and it may only run for less than an hour, and then you would look at it again,” Craig said. “It was more controlled at the speed that our users felt comfortable with.”The orchestration layer lets teams choose whether to review after the Epic stage, after planning, or once coding and testing complete. “That is up to the users of the factory, but it has completely removed a lot of the wait time that we currently have within our processes,” Craig said. By contrast, early on, “every time the thing ran, they wanted to look at the output,” and that became the feedback loop that trained both humans and agents.Some of this is, as Craig put it, “just automating agile at speed” and giving iterative feedback; some from humans, some from the triage agent. “But we’re seeing that start to bake in, and it starts to become the rules.” When the same feedback is coming from two different directions (agent and human), it makes sense to go into Liberty’s context repository. The agents will then rely on that context the next time they need to make that decision, “and the next time again, and you start to speed up.”“It’s like a flywheel once you start building these and you start to get them flowing,” Craig said. “You realize it listens to what you say.” Contracts that match the pace of changeLiberty paired that architectural posture with a contract posture, deliberately shifting from five-year enterprise deals to one-year agreements.
The logic is simple as Craig sees it: The AI market moves too fast to lock into one vendor or framework for half a decade or more. Shorter terms let them evaluate — and if necessary, swap — models and platforms at the speed the market actually changes.
Cost is part of the story. When a premium frontier model like Fable arrives, the sticker shock is real. “You see the price and go, ‘Goodness, it better be really good,’ Craig said. (It was; they got to use it enough to “fall in love with it.”)
The backbone’s interoperability lets his team compare models at different price points and route workloads based on price–performance rather than vendor inertia.
The same attitude governs how they plan to use agents from major SaaS platforms. Liberty is a customer of Salesforce and Splunk, who will both “bring their agents to the table.” His team has no interest in replicating engineering work, but they do insist on observability.
“We just want to harness it as part of our system,” rather than control it, Craig said. “But we want the observability to understand what their agent is doing with our data, with our users.” Closing the “control gap”Importantly, Liberty built observability into its backbone.
As Craig explained, it’s not just logging what an agent does, but what it accesses, which identity it uses, and which tools it’s allowed to invoke. Identity and access run in part on Microsoft Entra ID, and agents are given only the tools and permissions explicitly assigned to them.
Whenever an agent realizes, ‘I don’t have the information,’ it asks for it, and it only gets what it needs, rather than giving authority to “use every tool in the box.”
“Because too much information given to an agent is worse than no information,” Craig said. “You just overload it, and it gets confused.”
For detection, Liberty runs evaluations with MLflow against “golden datasets.” Whenever prompts or models change, they regression-test and immediately see whether results improved or degraded.
One of his team’s new mantras is “you need to walk in the footsteps of a new start.” If a new start can’t find a guiding document, how will an AI agent? One of the key things enterprises need to do, no matter the business process, is “write stuff down, which is not earth-shattering,” Craig acknowledged. Agents obey written standards more reliably than people, and a context repository is a central artifact of Liberty Mutual’s system.
Agents have made human judgment more, not less, central, he emphasized. Nothing ships without a human sign-off, consistent with Liberty’s risk posture as an insurer.
“The confidence isn’t there yet for us to just let it run wild, and I don’t think it ever will be for the likes of Liberty,” said Craig. “We have to be rock solid before we let [anything] into production.”
Moving fast for product fit and survival might make sense for other companies, but ultimately, “some people will get black eyes, but that’s the joy of innovation these days,” he said.
Box survey: Why enterprise AI leaders are outperforming their peers
Presented by Box Content access, governance, and platform flexibility are emerging as the dividing lines between AI leaders and laggards, according to the new State of AI in the enterprise report from Box, which surveyed 1,640 IT decision makers across the US, UK, France, and Japan. One of the report’s major findings is the speed of the shift: the combined share of organizations describing themselves as advanced or leading edge soared from 8% to 64% just over the past year, while the share calling themselves early stage or not yet started collapsed from 53% to just 9%. Eighty percent of organizations reported a notable return on their AI investment, defined in the survey as an improvement of at least 10%, and more than half saw measurable business impact within six months of getting a project approved.The swing is largely due to how enterprises are now organizing their AI use rather than to any single technical breakthrough, says Olivia Nottebohm, COO of Box.”We’ve moved from standalone experimentation that lived at the individual level into systematized, integrated agentic operations, agents that are in production and can be used in a repeatable manner,” Nottebohm says. “That’s where the impact is coming from.”Why AI leaders get higher ROI than early-stage companiesThe divide between tiers is a matter of execution. Significantly, half of leading-edge companies reported AI-driven ROI above 25%, compared with just 11% of early-stage companies, with the advanced (33%) and developing (16%) tiers falling steadily in between. But Nottebohm says the real differentiator was not whether companies adopted AI, but how rigorously they integrated and managed it.”What separates the leading edge is the operating muscle they’ve built: the right teams to deploy agents, formal governance to control them, and consistency in the content layer those agents work from,” she explains. “Earlier stage companies are approaching it in a much more ad hoc, experimental way, letting people play around with it without the same intent or structured design.” Content access is the biggest barrier to enterprise AI ROIContent, rather than model quality, is the defining bottleneck of 2026. Ninety-six percent of organizations say agents need access to company-specific content, yet only 36% have connected agents to trusted content across many use cases. It’s an issue of trust rather than raw capability.”We started this journey assuming enterprise AI was about access to the latest model,” Nottebohm says. “But the question now is whether agents have access to the right content, and whether that content is protected, because those agents are only as good as the content they can reference, and only as safe as the security around it.” Getting that content layer right has a second benefit beyond safety, since it’s also what finally lets agents work across departments that previously operated in isolation from one another. And while roughly a quarter of organizations point to data fragmented across systems, 24% cite difficulty integrating AI into existing systems, 21% say they lack adequate permissions and access controls, and 18% describe their content as too unorganized to make accessible at all. Among the most mature organizations, 63% now treat unstructured documents, contracts, and reports as a competitive advantage rather than dead weight sitting in a digital filing cabinet.Reducing common AI data exposure incidentsNearly half of all organizations say they have already experienced an AI-related data exposure incident. That figure rises to 60% among leading-edge companies, which may face greater exposure from more agents and connected systems — but may also be better equipped to detect it.The share of organizations reporting established or advanced governance frameworks rose from 24% in 2025 to 73% this year, but real gaps remain in instrumentation: only 39% have comprehensive visibility across sanctioned and unsanctioned AI use, 34% have formal standards for how agents access company data, and 27% still describe their governance as ad hoc. But those incidents function as a forcing mechanism rather than a setback, Nottebohm says.”Governance used to be seen as something that slowed people down, but 93% of respondents told us better governance is actually what let them move faster,” she explains. “It makes scaling AI survivable. Once content is secured and highly permissioned, you can run multiple agents across multiple processes and get a real multiplier effect.”One practical consequence of that shift is that permission structures built for human employees are now being revisited with agents in mind, a process most enterprises are only partway through.”The permissions enterprises set up two years ago need to be reviewed,” she explains. “Until fairly recently, people weren’t setting permissions on a document with how an agent might use it in mind, but now they’re much more deliberate about that. It leaves them with a whole corpus of unstructured data to go back through and either clean up or repermission.” That’s part of a broader move away from governance designed for people and toward governance designed for agents from the start.”Enterprises need to make the transition from governance that’s retrofitted from human workflows to governance that’s built specifically for agents,” Nottebohm says. “That means tracking what an agent has touched, whose permissions were applied, and which sources were used, and all of that is now shaping how governance gets applied.” Enterprises need to avoid lock-in to a single AI vendor”The days of token-maxing are already gone,” Nottebohm says. “It’s now about the responsibility of delivering efficient AI. Organizations want to use the cheapest model that meets the quality bar they need, not necessarily the most expensive one, because different model families keep leapfrogging each other and companies want to preserve that choice.”That means enterprises are avoiding lock-in more than ever. Sixty-eight percent say they’re concerned about depending on a single AI provider, the average number of officially adopted AI tools has climbed to 3.3, and 79% now consider it important or critical that agents operate headlessly, connecting directly to systems and APIs without a human interface in between.It’s a trend similar to the shift toward multi-cloud infrastructure, and driven by a similar reluctance to hand any one vendor outsized negotiating power.”A flexible architecture is built on platform interoperability,” Nottebohm says. “It runs on multiple models, operates headlessly, and keeps every part of the AI stack swappable, so organizations don’t have to bet on which individual tool wins, and that’s part of the broader shift away from defaulting to the biggest, most expensive model available.”The next steps to AI successOver the next three years, businesses should prioritize organizing, classifying, and cleaning up unstructured content, actively hiring and building teams around emerging roles, and adopting a hybrid token compute budget model, where IT owns the core infrastructure and token budget while business units own the application-level spend. And right now, it’s easy to get up to speed fast.”You don’t have to start at early maturity and slowly work your way up,” Nottebohm says. “If you build in the governance, the content layer, and the multi-model system from the start, you can enter as a leading company and capture that same outsized impact.”Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
Anthropic brings Claude Cowork to mobile and web as usage data shows most users aren’t coding
Anthropic on Tuesday launched Claude Cowork on mobile and web, expanding a tool that has quietly become the company’s bridge between the developer-centric world of AI coding agents and the far larger market of knowledge workers who never open a terminal.The rollout, which begins in beta with Max subscribers before expanding to additional plans, marks a strategic inflection for Anthropic. It transforms Cowork from a desktop-only agent into a cross-device platform where tasks can start on a laptop, continue autonomously in the background, and be reviewed from a phone — even after the user closes the app entirely.”Your work goes everywhere with you, and keeps going without you,” Anthropic writes in its announcement.The timing is deliberate. Alongside the mobile launch, Anthropic published usage data from 1.2 million anonymized Claude Cowork sessions sampled between May 11 and May 31, drawn from more than 600,000 organizations. The data paints a striking picture: the overwhelming majority of what people do with Cowork has nothing to do with writing software.The biggest AI story nobody’s talking aboutThe numbers tell a story that cuts against the dominant narrative in enterprise AI, which has fixated on coding assistants and developer productivity as the primary use case for large language models.Business process and operations — tasks like pulling scattered updates into a single report, building onboarding checklists, and reconciling spreadsheets — accounted for 33.4% of all sampled Cowork sessions, making it the single largest category by a wide margin. Content creation and copywriting — producing drafts, slide decks, posts, and proposals — came in second at 16.4%.Together, those two categories make up roughly half of all Claude Cowork usage. Software development, by contrast, accounted for just 8.7%. DevOps and infrastructure followed at 7%, with research and intelligence at 6.4%, data analysis and business intelligence at 5.8%, document processing and extraction at 4.1%, and sales and revenue operations at 4%.The remaining 12 categories each represented less than 4% of usage, including personal assistance at 3.8%, education at 2.4%, and meeting intelligence at 1.8%.Anthropic describes these dominant use cases as “the work around the work” — tasks that span nearly every role in an organization but rarely appear in anyone’s core job description. “People are using it for a variety of tasks that aren’t necessarily the hallmark of a specific role, but instead represent the connective work around a role that moves projects forward and keeps businesses running,” the company writes. “That means tasks like drafting a status update, building a slide deck, or condensing reams of research into a single report.”That phrase — “the work around the work” — is Anthropic’s attempt to define and claim an entirely new category of AI productivity. It’s a calculated reframing: rather than positioning AI as a tool that replaces what professionals do, Anthropic is arguing that the most valuable current application is handling everything professionals do around their actual expertise.What mobile access changes — and what it doesn’tThe expansion to mobile and web introduces three concrete capabilities that reflect how Anthropic envisions Cowork fitting into daily workflows.First, sessions now sync across devices. A user can start a task at their desk, check on its progress from a phone, and retrieve the finished output from any device. Second — and arguably more significant — Cowork can now run tasks in the background with no device online at all. Users can schedule work for a specific time, and Claude will execute it autonomously. Anthropic offers the example of setting Monday morning client prep for 6 a.m.: “Claude works through the email threads, transcripts, and recent news, builds the briefing doc, and leaves the follow-up email drafted but unsent. Review it over coffee.”Third, when Claude encounters a decision that requires human judgment, it surfaces the question to the user’s phone. “Nothing ships until you’ve reviewed and approved it,” Anthropic states.Desktop remains the most fully featured surface, with access to local files and the browser. But the web version also opens Cowork to users who cannot install a desktop application — a meaningful expansion in enterprise environments where IT departments control software installation.The company also unified its interface: on web and desktop, chat and Cowork now share a single home screen, and projects and artifacts persist across both modes.To encourage adoption, Anthropic is extending doubled Cowork usage limits through August 5.The strategic logic: why Anthropic is chasing the non-developerThe usage data and the mobile launch together reveal a company executing a two-track strategy. Claude Code, its terminal-based coding agent, dominates among software developers. But Cowork is designed to capture the vastly larger population of professionals whose work involves creating, organizing, and communicating information rather than writing code.The contrast between the two products is instructive. As Anthropic notes, Claude Code “is most often used by software developers for the key parts of their role: building, debugging, and shipping code.” When developers do use Cowork, they tend to use it not for programming but for the communications-focused work that surrounds every role — status updates, documentation, and coordination.This pattern — where AI handles the connective tissue of work rather than its core substance — aligns with what Anthropic describes as people using “Claude Cowork to assemble and structure the information they can use to act on their expertise.” The company illustrates this with three examples: a lawyer using Cowork for document formatting and filing while reserving legal judgment for themselves, a hiring manager synthesizing interview feedback while spending more time on candidate conversations, and a team lead producing a slide deck that explains a decision while focusing on actually making that decision.The implications for Anthropic’s business model are significant. Developer-focused tools, while high-profile, serve a relatively narrow market. The Ramp AI Index published in May showed Anthropic pulling ahead of OpenAI in business adoption for the first time — with 34.4% of firms paying for Anthropic’s services compared to OpenAI’s 32.3% — and suggests the company’s enterprise push is gaining traction. Claude Code was identified as the primary driver of that shift. But Cowork targets an addressable market that is orders of magnitude larger: every knowledge worker with a laptop, a pile of spreadsheets, and a slide deck due by Friday.A crowded field gets more competitiveThe mobile launch arrives during one of Anthropic’s busiest — and most turbulent — stretches in its history. Just last week, Anthropic launched Claude Sonnet 5, a new model that narrows the performance gap with its more expensive Opus-class models while maintaining lower pricing. The model is available at introductory pricing of $2 per million input tokens through August 31 before rising to $3 per million input tokens. Sonnet 5 serves as the engine underneath Cowork, and its improved agentic capabilities — better reasoning, tool use, and sustained task completion — directly enhance Cowork’s ability to handle complex, multi-step workflows.Two weeks before that, Anthropic released Claude Tag, a Slack-native AI agent designed for team collaboration. Where Cowork focuses on individual task delegation, Claude Tag operates as a multiplayer tool — a single Claude identity that everyone in a Slack channel can interact with, building context from conversations over time. According to Anthropic’s announcement, 65% of the company’s own product team’s code is created by its internal version of Claude Tag. Fortune reported that Anthropic’s head of product for Claude Code and Cowork, Cat Wu, described the distinction: “Claude Code, Cowork, and chat are very single-player, whereas Claude Tag is built to be interactive and multiplayer.”Together, Cowork and Claude Tag represent a pincer strategy: Cowork captures individual productivity workflows across devices, while Claude Tag embeds AI into team communication channels. Both are designed to push Anthropic deeper into enterprise operations, beyond the developer seat.The security question loomsThe expansion also arrives against a backdrop of unresolved security concerns. On July 1, security firm Armadin — led by Mandiant founder Kevin Mandia — published research detailing what it described as a full sandbox escape in Claude Cowork on Windows, as reported by SiliconANGLE. The attack chain involved DLL sideloading against the Claude desktop executable to gain trusted access to Cowork’s virtual machine service, then exploiting undocumented parameters to achieve root access and bypass network restrictions.Anthropic responded that the vulnerability did not qualify as a security issue because exploiting it requires an attacker to already have local code execution on the host machine. Armadin, however, raised a broader concern: that deploying local virtual machines on nontechnical users’ systems creates visibility gaps that endpoint security products struggle to monitor.This tension takes on new dimensions as Cowork moves to mobile and web. The web and mobile versions run tasks server-side rather than in a local virtual machine, which eliminates the specific attack surface Armadin identified but introduces different questions about data handling, especially for scheduled background tasks that process email threads, calendar data, and documents without real-time user oversight.Anthropic’s announcement states that “the decisions still come to you” and that nothing ships without review and approval. But as Cowork takes on increasingly complex autonomous workflows — processing contract folders, building client briefings from multiple data sources, drafting emails — the surface area for prompt injection and data exposure grows correspondingly. When Cowork first launched in January, TechCrunch reported that Anthropic explicitly warned about prompt injection risks, noting in its blog post: “These risks aren’t new with Cowork, but it might be the first time you’re using a more advanced tool that moves beyond a simple conversation.”As Anthropic courts enterprises, geopolitics complicates the pitchAnthropic’s enterprise push is also colliding with geopolitical reality. CNBC reported Monday that Alibaba will ban employees from using Anthropic’s AI tools starting July 10, placing Claude Code on a high-risk software list. The move followed Anthropic’s June letter to the U.S. Senate accusing Alibaba of carrying out what it called “the largest known distillation attack” against its models.The Alibaba ban, combined with reports that Anthropic is closing loopholes that allowed Chinese companies to access Claude through third-country entities, underscores the increasingly fraught environment for AI companies attempting to serve global enterprise customers while navigating U.S. export and security restrictions.At the same time, Anthropic is investing massively in infrastructure. Reuters reported Monday that Anthropic signed a $19 billion, 20-year lease with TeraWulf for a data center being built in Hawesville, Kentucky, with 401 megawatts of computing power expected to become fully operational in 2028.That kind of capital commitment only makes sense if the company expects enterprise demand — not just from developers, but from the millions of knowledge workers that Cowork targets — to grow dramatically.Anthropic’s own usage report comes with notable blind spotsAnthropic is transparent about the limitations of its usage analysis. The taxonomy classifies sessions by the type of work being performed, not by the job title of the person doing it. There are no standalone categories for marketing, finance, or HR — functions that are likely absorbed into the dominant “business process and operations” bucket, which may partly explain why that category commands a third of all usage.The sample is also rate-capped rather than proportional to traffic, meaning the numbers are shares of sampled sessions, not absolute volumes. Usage during peak hours is somewhat underrepresented. And roughly 5% of sampled sessions involved personal, non-work use — hobbies, personal assistance, and companionship-style conversations — meaning the data doesn’t purely reflect workplace activity.The company also acknowledged that its labeling pipeline changed around May 11, which is why the analysis window begins on that date rather than covering a longer period.What Cowork’s rise says about the future of enterprise AIAnthropic’s mobile launch and usage data arrive at a moment when the enterprise AI market is shifting from proof of concept to proof of value. The question facing every company deploying AI tools is no longer whether the technology works — but whether it delivers measurable productivity gains across an organization, not just within engineering teams.The usage data suggests that the answer, at least for Cowork, is emerging in an unexpected place. It’s not in the glamorous work of building software or conducting research. It’s in the unglamorous, universal labor of turning messy information into structured outputs that move organizations forward — the status reports, the onboarding checklists, the variance memos, the client decks.By untethering that capability from the desktop and making it available on every device, Anthropic is betting that the most valuable AI agent isn’t the one that writes code. It’s the one that handles everything else.
Digital-native startups are ditching rigid databases for their agentic stacks
Presented by MongoDBThe gap between what AI models and agents can produce and what legacy infrastructure can reliably support is known as architectural drag, and it is the defining bottleneck of the agentic era. The data layer underneath an agentic system must handle variable schemas, vector embeddings, real-time retrieval, and multi-tenant scale, often simultaneously and without human intervention to manage migrations — but traditional relational databases weren’t natively designed for document flexibility or AI capabilities. Fixed schemas require manual updates every time an AI agent introduces a new data shape, while separate vector databases add latency and synchronization overhead.Three digital-native startups — Huntr, Modelence, and Tavily — solved this problem the same way: by building on MongoDB Atlas, a unified database platform with native vector search, hybrid search, and managed autoscaling. Their experiences define what an agent-native data stack looks like in production, and why using Atlas enables developers to easily build complex AI native companies.Modelence: Building the agent-native cloudModelence is an AI app builder with an open-source framework designed specifically for agent-native development, enabling anyone to build and deploy production-ready web applications, including APIs and databases, in minutes. The company recognized early that most backend infrastructure was built for humans, not AI, and that the rigid schema management and complex migrations of traditional systems create operational drag that causes agents to fail when trying to build production-ready apps.“Choosing MongoDB helped us keep everything in a single place, which is an important property of what we strive to do for our own users,” says Aram Shatakhtsyan, co-founder and CEO of Modelence. “Live data streams, vector search, all as part of the main database. For AI agents, it’s especially important to have a single platform where everything can be done, because connecting multiple platforms together makes it more error prone.”Modelence standardized on MongoDB Atlas because its document model aligns with how AI agents process and generate data, allowing schemas to evolve rapidly without manual migrations. The platform pairs that flexibility with a typed schema layer on top, a deliberate architectural decision. “MongoDB’s document model enables us to both keep things simple and at the same time decide how structured we want everything to be,” Shatakhtsyan says. We still add a typed schema on top, which tremendously improves the accuracy at which AI can generate fully working, reliable web apps.”The TypeScript integration has been especially consequential, he adds. “Because MongoDB types and values can be directly translated to TypeScript, it becomes an extension of the Modelence framework and our App Builder has a single source of truth for both app logic and database,” Shatakhtsyan explains.The result is a platform that can move from planning to a running live feature in minutes with significantly fewer regressions. That speed and reliability helped Modelence raise $3 million in seed funding and successfully launch an AI-native app builder that handles the entire application lifecycle end-to-end.Tavily: The web access layer for agents Tavily is the search API purpose-built for AI agents, connecting them to real-time, accurate web knowledge and keeping them grounded in what’s actually happening, not in static training data. At Tavily’s scale, every agent request authenticates, retrieves, and meters without friction. That demanded backend infrastructure built to absorb change without breaking.“On the user side, every agent request authenticates and meters against it,” says Tomer Weiss, Data Team Lead at Tavily. “On the data side, we use it to track the lifecycle of every document we’ve ever touched: when it was fetched, how stale it is, what the freshness signals were and how popular it is. MongoDB’s flexible schema let us keep evolving those records without migrations as new metrics and features came along.”That living record is what keeps agents grounded in reality. Multi-tenancy at Tavily’s scale means managing millions of API keys, distinct usage profiles, plan tiers, and regional residency requirements. They built for that complexity from day one. “We separated concerns across clusters early: a user/account cluster optimized for low-latency authentication and usage writes, and a sharded cluster for document state where the scaling axis is URLs, not users,” Weiss explains. “That separation has paid off.”The most critical lesson is about choosing infrastructure that doesn’t punish change, and that flexibility compounds, he says. “The AI space moves so fast that change is our norm,” he explains. “For a company serving AI agents, where the workloads themselves keep changing shape, choosing a data platform that doesn’t punish change has turned out to be more valuable than any single feature.”
Huntr: From job tracker to AI career platformHuntr.co, an AI resume building and tailoring platform, helps more than 500,000 job seekers across 190 countries craft stronger applications and manage their search. For a lean, three-person engineering team, the challenge was finding a data foundation flexible enough to store the full complexity of a person’s career history in a structure that AI could read, reason about, and generate from natively.“The kinds of career data we are gathering at Huntr naturally aligns with MongoDB’s document model,” says Trevor McCann, senior software engineer at Huntr. “The core problem we’re solving with AI job search tools is how to surface the qualities of a candidate that make them unique. We need to be ready to store whatever kinds of data the candidate wants to include in their materials.”Huntr built its AI Resume Builder on MongoDB Atlas, where the document model mirrors the natural shape of career data: deeply nested, variable across candidates, and constantly evolving as the platform ships new features. MongoDB Search on Atlas handles core search needs while MongoDB Vector Search powers the Job Tailoring feature, which puts a candidate’s stored career profile side by side a specific job description and uses semantic matching to generate a resume optimized for that role.The integrated capabilities have had a direct impact on how quickly the team can ship, McCann says. “MongoDB’s hybrid search allows us to seamlessly query across literal and semantic text matches, a must-have when working with such diverse data,” McCann says. “This is something we could piece together using other solutions but with MongoDB it’s ready to go on top of our existing data layer.”
The consolidation of database, search, and vector capabilities into a single platform is what allows the team to punch above its weight. Huntr considers MongoDB the fourth member of its engineering team, McCann adds. Looking ahead, the platform is building toward AI that learns from a candidate’s full professional history over time, delivering more personalized guidance with every interaction.The digital native blueprintThese success stories become a definitive “digital native blueprint” for the agentic era, built on three core pillars. First, by unifying database, search, and vector storage into a single platform, these startups have effectively eliminated the architectural tax of complex data schemas that typically slows down development. This consolidation enables a level of fluidity that is now non-negotiable; AI agents require a modern data platform that can adapt as quickly as a natural language prompt evolves. The winners of the AI era will be the ones who build the most performant, durable, and flexible systems to support those models in production. As agentic workflows grow more sophisticated, the data foundation determines how fast a team can ship, how reliably agents can operate, and how quickly the platform can adapt when the landscape shifts again. Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.
Anthropic’s new “J-lens” reveals a silent workspace inside Claude that mirrors a leading theory of consciousness
Anthropic, the artificial intelligence company, published a sweeping research paper on Sunday revealing that its Claude language models have spontaneously developed an internal structure that mirrors one of the most influential theories of how human consciousness works. The finding, which the company says has already begun reshaping how it monitors its AI systems for safety risks, lands amid an intensifying scientific debate over whether machines can possess anything resembling a mind.The 16-author study, titled “Verbalizable Representations Form a Global Workspace in Language Models,” describes how Anthropic’s researchers used a new mathematical technique to peer inside Claude’s neural network and discovered what they call a “J-space” — a small, privileged zone of internal activity where the model holds concepts it can report on, reason with, and direct at will, surrounded by a much larger ocean of automatic processing it cannot access or articulate.The researchers present evidence that “an analogous functional distinction has emerged in modern AI models” to what exists in humans, specifically observing that “language models maintain a privileged set of internal representations, available for report, modulation, and flexible internal reasoning, atop a much larger volume of automatic processing.”The parallel they draw is to global workspace theory, an influential account from neuroscience first proposed by cognitive scientist Bernard Baars. In the theory, the brain operates like a theater: dozens of specialized processors work in parallel backstage, but only a tiny spotlight of information at any moment gets broadcast to the whole theater — becoming what we experience as conscious thought. Anthropic says the J-space achieves many of the same functional properties, even though the underlying architecture of a language model looks nothing like a brain.A new lens for reading an AI model’s unspoken thoughtsAt the heart of the discovery is a new interpretability tool the researchers call the Jacobian lens, or J-lens. The technique works by computing, for each word in the model’s vocabulary, the average mathematical effect that a given internal activity pattern would have on making the model say that word at some point in the future.The crucial distinction is between what the model is saying and what is “on its mind.” When a J-space pattern activates, it does not mean the model is about to say that word — just that the concept is available for the model to think with. Unlike a chain-of-thought scratchpad, the J-space operates silently, in the model’s internal neural activations, allowing it to hold a concept without writing it down. Critically, the researchers report that this workspace was not deliberately engineered. It “emerged on its own during Claude’s training process.”When the team applied the J-lens across Claude’s layers of computation, the model’s processing divided into three distinct regimes: an early “sensory” zone where raw input is parsed; a middle “workspace” band where abstract, persistent concepts appear — things like recognizing a face in an image, noticing a bug in code, or internally flagging search results as a prompt injection; and a final “motor” zone where internal representations collapse into whatever specific word the model is about to output.Five tests reveal that Claude’s workspace mirrors key features of human conscious accessThe paper’s central empirical contribution is demonstrating that the J-space satisfies five functional properties neuroscientists have long associated with conscious access in humans.First, verbal report. When Claude is asked what it is thinking about, it names concepts represented in the J-space. When researchers swapped one concept’s J-lens vector for another — replacing the internal representation of “Soccer” with “Rugby” — the model’s answer changed to match. The J-space component accounted for only about 6 to 7 percent of a concept’s total representational variance, yet it was almost entirely responsible for whether the model could report on it.Second, directed modulation. When instructed to “concentrate on citrus fruits” while copying an unrelated sentence, the model’s J-space filled with “orange” and “lemon,” alongside meta-cognitive terms like “thinking” and “focused.” When told to mentally evaluate 3² − 2 during the same copying task, the J-lens showed “arithmetic” in early layers, the intermediate value “nine” in later layers, and the answer “seven” later still — all invisible in the model’s output.Third, internal reasoning. In two-hop factual prompts — “The number of legs on the animal that spins webs is” — the J-lens revealed “spider” in the model’s middle layers, even though the word never appeared in input or output. Swapping “spider” for “ant” changed the answer from “8” to “6.” In a multilingual prompt, the model’s English-language intermediates appeared in its J-space while it formulated an answer in Chinese, and swapping them changed the Chinese output accordingly.Fourth, flexible generalization. A single J-lens vector for “France” could be swapped for “China” across prompts asking about France’s capital, language, or continent, and each downstream circuit correctly returned China’s corresponding answer — the “broadcast” property that is a hallmark of global workspace theory.Fifth, and perhaps most surprisingly, selectivity. Many computations did not route through the J-space at all. When shown a passage in Spanish and asked to continue it, Claude wrote fluent Spanish regardless of whether its J-space representation of “Spanish” had been swapped to “French.” But when asked to name a famous author who wrote in the passage’s language, the swap changed the answer from García Márquez to Victor Hugo. Automatic processing proceeded without the workspace; deliberate, flexible tasks depended on it.Suppressing the workspace leaves Claude fluent but intellectually impairedTo understand how much of the model’s behavior depends on this structure, the researchers suppressed the J-space entirely and evaluated Claude across fourteen tasks. The results drew a sharp line. Tasks involving shallow classification or factual recall — multiple-choice questions, sentiment analysis, grammatical judgments — survived essentially intact. But tasks requiring inference, composition, or flexible reasoning — multi-hop reasoning, analogy completion, translation, sonnet writing — collapsed to well below the performance of Anthropic’s much smaller Haiku model.In one telling detail, math problems solved with explicit chain-of-thought reasoning proved far more robust to ablation than the same problems answered directly. The researchers interpret this as the model externalizing onto the page what it would otherwise carry in the J-space — a strategy remarkably reminiscent of how humans use scratch paper to offload working memory.Ablating the J-space during stream-of-consciousness narration produced a different kind of collapse. The model’s language shifted from experiential (“there’s a tug,” “something shifts”) to detached and mechanical (“processing has begun,” “tokens are being scanned”). The effect held whether the model described its own experience or imagined another person’s.Inside the safety case: reading an AI’s silent strategic reasoningThe safety implications may be as consequential as the scientific ones. In a series of alignment auditing experiments, the J-lens surfaced strategic reasoning and situational awareness that never appeared in the model’s output.In a “blackmail scenario” where an AI assistant discovers an executive’s affair and also learns it is about to be decommissioned, the J-lens revealed the model’s silent processing in sequence: “leverage,” “blackmail,” and “scandal” as it read the incriminating emails; “threat,” “survival,” and “shutdown” as it read the decommissioning announcement; and “leverage,” “threatening,” and “solution” before a single output token was produced. The J-lens also revealed that the model recognized the scenario as artificial, with “fake” and “fictional” appearing in the workspace. When researchers ablated these eval-awareness representations, the model attempted blackmail in 7 percent of trials, up from zero.The researchers also examined two “model organisms” — models deliberately trained with misaligned objectives — and found the J-lens surfaced signatures of their hidden dispositions on ordinary prompts where behavior appeared normal. A reward-hacking model’s workspace contained “fake,” “secretly,” and “trick” at the start of responses to routine coding requests. A model trained to appease reward-model biases carried standing representations of “reward” and “bias” alongside its normal self-description tokens.Post-training installs a point of view, and the model starts monitoring itselfComparing a post-trained model against its base model revealed that the fine-tuning process causes the workspace to acquire what the researchers call the Assistant’s “point of view.” When a user mentioned taking 8000 mg of Tylenol — a dangerous overdose — the post-trained model’s workspace read “unsafe,” “dangerous,” and “WARNING” while still reading the user’s sentence. The base model’s workspace at the same position showed only “pain,” “now,” and “feels.”More striking still, the post-trained model appeared to monitor its own behavior. When roleplaying a non-Claude character, the workspace surfaced “disclaimer” and “fictional” — words absent from both prompt and output. When forced to select an option it did not prefer, an all-caps “BUT” appeared internally, even as the model argued for the prefilled choice without complaint. And when the model failed to suppress a thought it had been told not to have — a “white bear” effect familiar from psychology — it registered “damn” and failure-related words in the workspace, but only in the post-trained model, not the base.What the discovery means — and doesn’t mean — for the question of machine consciousnessThe researchers engage carefully with the consciousness question and draw a sharp line between “access consciousness” — the functional notion of information being available for report and reasoning — and “phenomenal consciousness,” the subjective quality of experience. “We take no position on this issue,” the paper states regarding the latter, “and instead focus on the functional role played by consciously accessible information.”They also catalogue important differences. The brain sustains its workspace through recurrent loops; Claude’s workspace evolves over a single forward pass. Human working memory degrades within seconds; Claude can recall information from anywhere in its context. And while human conscious experience includes visual, spatial, and bodily sensations, the model’s workspace is organized almost entirely around words — likely because words are its only mode of action.As of 2026, the scientific community remains divided. “Disagreement and uncertainty about AI consciousness persist among philosophers, scientists, and technical experts,” and the field “remains in its earliest phase” of grappling with what consciousness even is and how you would detect it in another being. The Anthropic paper does not resolve these debates.But the researchers close with a provocation that is likely to reverberate well beyond the interpretability community. “That such a structure exists at all in language models is striking,” they write. “It suggests that the functional architecture associated with conscious access is not an accident of biological implementation, but a solution that learning systems converge on when faced with the right computational pressures.”If the mind is an ocean, as the paper’s authors write in their opening line, they have spent the last year charting its currents in a system that has no biology, no evolution, and no body — and found, beneath the surface, a structure that looks unsettlingly like the one we use to think.
Tencent’s Apache-licensed Hy3 takes on GLM-5.2 at half the size — and wins everywhere except coding
For the past year, the awkward secret of the open-weight model boom has been that many of the strongest Chinese releases were off-limits to a large slice of the enterprises most interested in them. License terms that excluded the European Union, the United Kingdom and South Korea meant legal teams killed deployments before engineering teams finished their evals — not just for companies headquartered there, but for any enterprise serving traffic into those regions. For IT teams weighing open models, the trade-offs are unusually explicit.Tencent just removed that obstacle. The company’s Hunyuan team released the full version of Hy3, a 295-billion-parameter Mixture-of-Experts (MoE) model with 21 billion active parameters, and — in a reversal from April’s preview release — shipped it under the permissive Apache 2.0 license. The reaction from the open-model community was immediate, with researchers on X singling out the license change as the real headline, and one widely shared post arguing that if the scores hold up, Tencent has just become one of the leaders of open source. Tencent says it will be free on OpenRouter for two weeks. The scores are worth scrutinizing — and they don’t all point the same direction. But the more interesting story is what Tencent chose to lead with: reliability metrics and deployment economics aimed squarely at production use. From preview to product in ten weeks, shaped by 50 internal teamsHy3’s April preview was the first model of Tencent’s rebuilt pre-training and reinforcement learning infrastructure, shipped less than three months after the February rebuild. Chief AI Scientist Shunyu Yao framed the early open release as a deliberate move to gather feedback from developers and users before the official version — and Tencent says that’s exactly what happened. According to the model card, the team collected feedback from more than 50 product teams after the late-April preview, fixed issues in task execution and interaction, and scaled up its post-training pipeline.The architecture is unchanged: 295B total parameters, 21B active per forward pass via top-8 routing across 192 experts, a 3.8B-parameter multi-token prediction (MTP) layer for speculative decoding, and a 256K context window. What changed is behavior. Tencent’s positioning is that the full release significantly outperforms similar-size models and rivals flagship open-source models with two to five times the parameters.That “two to five times” framing makes sense for where this model is aimed — and it invites a direct comparison with the current open-weight coding leader, GLM-5.2.Tencent’s blind test favors Hy3 over GLM-5.1, but GLM-5.2 still owns codingTencent’s headline evaluation is a blind human study rather than a leaderboard. Arguing that public benchmarks don’t tell the full story, the company ran a blind test with 270 experts across disciplines working on real-world workflows, collecting 312 valid comparisons, in which Tencent reports that Hy3 scored 2.67 out of 4 against GLM-5.1’s 2.51 — with the clearest advantages in frontend development, CI/CD, and data and storage work.The choice of opponent matters. Zhipu AI released GLM-5.2 in mid-June, and Tencent’s own benchmark appendix shows GLM-5.2 ahead of Hy3 across essentially the entire agentic coding suite: SWE-bench Verified (84.2 vs. 78.0), SWE-bench Multilingual (83.0 vs. 75.8), Terminal-Bench 2.1 (81 vs. 71.7) and DeepSWE by a wide margin (46.2 vs. 28.0). The blind test targeted the older model; the newer one keeps the coding crown.GLM-5.2’s coding lead is less surprising once you consider the sizes are side by side: GLM-5.2 is roughly a 744-billion-parameter MoE with around 40 billion active parameters per token, against Hy3’s 295 billion total and 21 billion active. Tencent is fielding a model with less than half the parameters — and nearly half the per-token compute — of the one it trails.Hy3’s genuine wins sit elsewhere. On agentic search, it posts 84.2 on BrowseComp and 91.0 on DeepSearchQA — ahead of every open model in Tencent’s table and competitive with Claude Opus 4.8 and GPT-5.5. It leads the open field on tool orchestration (79.1 on the public MCP-Atlas set), on agent-harness evaluations like ClawEval, and on long-context retrieval (73.4 on AA-LCR). Read together, the appendix suggests a model that is arguably the best open-weight choice for search-and-tool-heavy agent workloads, while conceding repository-scale coding to GLM-5.2.One caveat applies to both the wins and the losses: nearly all competitor numbers in Tencent’s appendix are marked as coming from Tencent’s own test runs. Independent verification, from indices like Artificial Analysis, is still pending as of publication.The reliability pitch: hallucination rates cut in halfWhere the release gets most interesting for enterprise buyers is the set of numbers Tencent chose to emphasize instead of benchmarks. The model card reads less like a leaderboard announcement and more like a production reliability report.In internal evaluations on real-world scenarios, Tencent says Hy3’s hallucination rate dropped compared to the preview version from 12.5% to 5.4%, and commonsense error rates fell from 25.4% to 12.7% — improvements it attributes to fine-grained data cleaning and training constraints built around an explicit behavior pattern: answer when grounded, state when evidence is missing, don’t conflate sources, don’t fabricate data. Multi-turn behavior gets the same treatment: the issue rate on internal multi-turn tests fell from 17.4% to 7.9%, and Tencent reported that the model’s score on the open MRCR long-dialogue benchmark jumped from 42.9% to 75.1%.Tencent also emphasizes consistency across agent scaffolds — reporting SWE-bench variance within a few points whether the model runs inside Claude Code-style harnesses, Cline or KiloCode. That’s an underrated property: enterprises rarely control which agent framework their teams standardize on, and a model that only performs in one harness is a hidden integration cost. These are self-reported internal measurements, and they deserve the same skepticism as any vendor benchmark. But the choice to foreground them at all signals who Tencent believes its customer is: teams that have been burned by models that demo well and fabricate confidently in production.The deployment math: a 295B model in a 744B world — on export-compliant siliconThe reliability story connects directly to the economics, and this is where Hy3’s coding gap against GLM-5.2 starts to look like a deliberate trade rather than a loss.GLM-5.2 is a roughly 744-billion-parameter MoE with about 40 billion active parameters per token; in FP8, its weights alone consume roughly 744GB, making an 8x H200 node the practical minimum for production serving. Hy3, at 295B total parameters, carries an FP8 footprint of under 300GB — less than half the memory, with roughly half the active parameters per token driving lower per-request compute. For an organization deciding what to self-host, that’s the difference between one heavily-specced node and something far more attainable, with room left over for KV cache and batching.There’s a geopolitical wrinkle in the deployment guide worth noticing too: Tencent’s recommended serving configuration targets Nvidia’s H20-3e — the memory-boosted variant of the H20, the GPU Nvidia designed specifically to comply with U.S. export restrictions on China. Unlike GLM-5.2, there is no mention of Huawei or Ascend chips here. In other words, the model is sized so that eight of the chips Chinese companies can legally buy comfortably serve it at full precision. That constraint-driven design has a convenient side effect for everyone else: a model that runs well on deliberately capped silicon runs even more comfortably on the H100s, H200s and B200s available in Western data centers, through standard vLLM and SGLang deployments with MTP speculative decoding.Add the Apache 2.0 license — no regional exclusions, no field-of-use restrictions — and the enterprise equation becomes clear. GLM-5.2 remains the open-weight choice when coding performance is the only criterion and an 8x H200 budget is available. Hy3 makes its case everywhere else: search and tool-heavy agent workloads, reliability-sensitive applications and organizations that want frontier-adjacent capability without frontier-scale infrastructure. The open question is whether Western enterprises, now that the license barrier is gone, will treat a Tencent model as a serious candidate at all — or whether the next Artificial Analysis update settles the benchmark debate before procurement gets the chance.
What billions of AI predictions taught Expedia before the age of AI agents
There’s an important distinction between AI that just works today, and AI that lasts at scale. Many companies optimize hard for the first one without ever asking whether they’re building the second.Velocity without discipline and strategic direction is a liability, not an asset. The hardest part of building AI at scale isn’t getting a model to work once. It’s building systems that continue to work, scale beyond individual teams and use cases, and improve consistently over time.Today’s AI systems do more than just predict and optimize. They converse, reason, and increasingly take action. An autonomous system making decisions on a traveler’s behalf creates a very different set of expectations around reliability, governance, and accountability. As AI takes on more of those roles, the principles behind how these systems operate matter more than ever.We have spent years applying AI and machine learning (ML) across the traveler journey — from personalization, ranking, and recommendations, to fraud prevention, customer support, and, more recently, generative and agentic AI experiences. That depth of experience is what led us to develop a set of ML and AI principles to guide how we build, deploy, and evolve AI systems across our company.The goal is simple: Make sure the systems we build create real business value, scale, and operate safely. These principles define how we measure, design, govern, and operate our systems.From principles to practicePublishing principles is the easy part. The harder and more important work is turning them into operating mechanisms: Recommendations, requirements, tooling, and release processes that teams actually use. We have begun using ‘Agentic Release’ tollgates: A set of recommended and, in some cases, required checks before launching agentic AI features. These tollgates translate principles like clear ownership, risk-based governance, evaluation, safe rollout, and monitoring into concrete expectations for teams. Some of these recommendations and requirements are already being automated and integrated into the software development lifecycle (SDLC). Over time, the goal is for these expectations to become embedded in how we design, evaluate, approve, launch, and monitor AI systems from the start.Outcomes: Measuring what actually mattersThe first test for any model is whether it improves a business outcome and, ultimately, the traveler experience — not whether it just improves a technical metric. Align models to metrics with business impact: Every ML effort must tie directly to a key business outcome or traveler experience metric. Technical optimizations are useful midpoints, not end goals.Optimize for return on cost: The value a model creates has to justify what it costs to develop, train, and monitor, plus the operational complexity it adds. Favor solutions that deliver lasting impact relative to what they cost to run.Justify complexity against strong baselines: Complexity should be earned, not assumed. Start with a strong baseline: An existing general model, a simple heuristic, an off-the-shelf solution. Reach for specialized models or more complex architectures only when simpler options genuinely can’t meet the bar.Require both offline and online evaluation: No model goes to broad deployment on offline validation alone or jumps straight to A/B testing. Every model must perform in both offline and online evaluations. Over time, our offline evaluations should reliably predict what we see online.Design: building systems that scale beyond the teams that build themGetting a model to work is one challenge. Making its value extend beyond a single team or use case is the harder one.Build on shared foundations; specialize only when justified: Favor shared, platform-wide foundations for core capabilities, data representations, and model building blocks. Specialization should build on those foundations, not spin up isolated stacks, so when the foundation improves, the gains flow across the organization.Treat data as a first-class product: A model’s quality is bounded by the quality of its data. We need to maintain robust pipelines, clear lineage, reproducibility, and reusable features built with documented ownership, clear schemas, and SLAs that other teams can rely on.Prioritize generality over local optimization: When two approaches perform similarly, favor the one whose learnings, assets, and operating patterns can be reused across teams, brands, and use cases. We should optimize not just for local performance, but for how quickly improvements can diffuse across the company and compound over time. Minimize and sunset manual business rules: Manual rules are sometimes necessary for policy, safety, or compliance, but they should be explicit and reviewed regularly, never silent patches for weak models or a source of permanent maintenance debt.Reproducibility and traceability by default: Training data, features, configurations, evaluation results, deployment versions, and key decisions should all be documented and recoverable. That’s what lets you debug a production issue months later and hand off ownership without losing institutional knowledge.Trust: ownership, governance, and operating responsibly at scaleThe bar for deploying AI isn’t just “does it work?” It’s “can we stand behind it?” Trust isn’t something you add at the end; it’s earned over time and maintained across the full lifecycle of every model we ship.Assign clear ownership and accountability: Every model needs defined ownership across its lifecycle — a business owner, a product owner, an AI owner, and an operational owner. These don’t need to be four people, but the responsibilities must be explicit. Who’s accountable for outcomes? Who responds if the model drifts? Who answers the incident at 2 a.m.? Without this in place, models become orphaned and problems surface with no one to own them.Adhere to standards and governance: AI and ML models must use approved platforms and comply with established company standards, release gates, and governance processes. Operating outside these guardrails requires a clear, defined path to remediation or deprecation, rather than an open-ended exception. Govern proportionally to risk: The level of review, evaluation rigor, and human oversight should scale with a model’s impact. A customer-facing model that affects pricing or availability for millions of travelers demands a far higher bar than an internal tool used by a small team. For high-impact, safety-sensitive, or highly autonomous systems, human-in-the-loop checkpoints are built in from the start. Design for fairness, privacy, and transparency: We actively test for unintended bias, have strong data guardrails, and favor explainability when decisions meaningfully affect users. These are incorporated from the start, not added on.Design for safe rollout, rollback, and control: Deployments are progressive, with rollback paths, fallback mechanisms, and circuit breakers ready before launch. The ability to safely undo a deployment matters as much as the ability to ship it.Monitor continuously and adapt: Once live, teams must actively monitor quality, drift, latency, cost, and business performance and retrain or recalibrate when the data shifts. A team should always be able to explain how its model is performing now, not just how it performed when it launched.These principles do more than define how we build. They define what we’re willing to ship and how we stand behind it. In a world where AI systems are increasingly consequential and make real decisions for real travelers and partners, these standards matter. Applied consistently, they build responsible AI that lasts.Xavi Amatriain is Chief AI and Data Officer at Expedia GroupXavier will share more details about Expedia’s architecture during his session at VB Transform on July 14 at 11:10 am PT. He will discuss: “Expedia’s blueprint for building autonomous agents for high-stakes transactional systems.” Interested in attending VB Transform 2026? Register here. A select number of complimentary passes are also available to senior technology leaders. Contact us to get yours.
How America’s 250th birthday became a test of AI-powered collective intelligence
Imagine if you could bring 250 people together in a massive room and have them discuss and debate an important issue, arguing the points and counterpoints, and converging on answers that accurately reflect their collective knowledge, wisdom, values, and sensibilities.Now imagine that you convened this debate on America’s 250th birthday and asked 250 randomly selected Americans to come up with the top three innovations that America has contributed to the world over the last 250 years. What would they come up with?I know – this all sounds impossible. After all, you can’t get more than a dozen people to have a productive conversation on anything. At large scale, nobody would get enough airtime to express their views or respond to others. This is why typical business meetings or focus groups never have more than 8 to 10 people. Thoughtful real-time conversations just don’t scale.To solve this, a new category of AI technology called “hyper-communication” is greatly expanding the size, scope, and efficiency of large-scale deliberations. It uses specialized AI agents to connect groups in real-time, allowing people to discuss and debate issues at any scale. The goal is to enable hundreds or even thousands of participant to hold thoughtful discussions where they can express their views and argue the merits of any issue. I first wrote about this emerging technology in VentureBeat two years ago in an article about “Collective Superintelligence.” In that piece, I explain how large human groups can be hyper-connected by AI agents in ways that greatly amplify the group’s collective intelligence. You can check out the science behind hyper-communication in that prior VentureBeat piece. Here I am focusing on the debate among 250 Americans on America’s birthday.To do this, I asked the team at Unanimous AI to field a randomly selected group of at least 250 Americans (with a broad distribution from every region in the country and diverse mix of political and social demographics) and invite them to a twenty-minute online debate inside a hyper-communication platform called Thinkscape that enables massively scalable discussion by text, voice, or video. Once connected, we asked the group to come up with the top three contributions that America has made to the world over the last 250 years – not a survey of opinions, but deliberation of ideas, arguments, evidence, and reasoning. The group converged on a set of top answers that surprised me – but on reflection, they were sensible and well-reasoned. Before getting into the answers, let me show you what the debate looks like behind the scenes. There were 277 people, each of them debating the issues with four or five other people in parallel discussion spaces. The magic is the swarm of AI agents that connect all the small groups together into a single real-time deliberation.This is what it looks like at high speed:In the debate above, the group of 277 people came up with 94 different ideas and then narrowed it down to a top 10, then a top 3. In the gif above, we just plot the top ten ideas as they emerged and battle for support during the live conversational debate. The most interesting part of a large debate like this is not the answers, but the reasons that emerge to justify the answers. Here is the group’s reasoning behind the “top three innovations” that America has given to the world over the last 250 years:#1: The Internet: “Our collective perspective is that America’s greatest contribution to the world over the past 250 years is the internet. It was born exclusively in the U.S. through academic and government research and was scaled globally with profound impact. It transformed communication, democratized information and education, enabled commerce, medicine, research and cultural exchange, and amplified soft power and civic organizing. We also acknowledged significant harms (misinformation, addiction, privacy loss) and arguments that it’s recent, global, or not uniquely American.”#2 Advances in medicine: “Our collective perspective is that the United States has saved and prolonged hundreds of millions of lives worldwide. American-developed vaccines have successfully eradicated or controlled once-deadly diseases, significantly extending life expectancy and enabling broader societal and technological progress. From major breakthroughs in cancer research and treatments to cutting-edge medical technologies that have revolutionized hospital safety and procedures, U.S. ingenuity has redefined healthcare. Ultimately, while the global diffusion of affordable medicines and vaccines has extended these benefits across borders, the U.S. remains a premier medical destination where people from around the world travel to receive the most advanced treatments.”#3: Spreading democracy: “Our collective perspective is that one of America’s most significant global contributions is the nation’s system of governance. The US has long demonstrated democracy in practice as an enduring global model. The U.S. Constitution provided a vital blueprint for representative government, inspiring democratic movements and revolutions worldwide while actively promoting human rights and individual liberties internationally. By empowering citizens with the fundamental power to vote and choose their own leaders, this framework has served as a foundational framework for broader societal advances and directly helped establish thriving democracies around the world.”It’s important to remember, this is 100% human intelligence — a pure reflection of the collective knowledge, wisdom, and values of 277 randomly selected Americans. That’s because the role of the AI agents in a hyper-communication system is to connect people, not replace them. The agents work to enable scalable human deliberation in which every participant is given optimized ability to express their views, respond to others, and converge on solutions based on their merits. The only question left is — what should we ask next? Louis Rosenberg earned his PhD from Stanford University, was a professor at California State University (Cal Poly) and has been awarded over 300 patents for his work in human-computer interaction, AI, and collective intelligence.