🎯 SUCCESS 🧠 BRAIN 💸 MONEY 🧭 SPACES 🌍 TRAVEL 🎙️ PODCASTS 📺 VIDEOS 🎥 CRIME & MOVIES
  • Skip to main content

Mad Mad News

LIVE ABOVE THE MADNESS

Order Now • Check Delivery Today
As an Amazon Associate I earn from qualifying purchases. Delivery availability varies by item and location.

Venture Beat

AI cites the deep pages but sends humans to the homepage — most sites are built backward

July 27, 2026 MMN Editor Filed Under: Uncategorized

If your business depends on people clicking through to web pages, the last two years have been brutal. Pew Research Center tracked the browsing behavior of 900 U.S. adults and found that when Google shows an AI summary, users click a traditional result just 8% of the time, roughly half the 15% rate when no summary appears. Links cited inside the AI answers themselves fare worse: users click on them only about 1% of the time.This has had a huge impact on publishers. Chartbeat data reported by Axios shows page views from Google Search fell 34% across its publisher network between December 2024 and December 2025, and small publishers have lost roughly 60% of their search referral traffic over two years. Business Insider’s organic search traffic dropped 55% over three years, and some smaller publishers have already shut down. Chatbot referrals, meanwhile, still account for less than 1% of publisher page views despite growing more than 200% in a year.The story in publisher circles has been simple: AI is killing the web. But recent developments show the reality is more nuanced.Machines are reading more than everSimilarweb’s 2026 Generative AI Landscape report reveals that while AI platforms send fewer humans to web pages relative to the answers they generate, the AI systems themselves are consuming the web at an accelerating rate in the form of searching the web on the user’s behalf to answer their questions. The share of ChatGPT answers containing live web citations grew more than fivefold in under a year, reaching 6.8% of all answers by May 2026. In some categories like travel, it is as high as 22.6%.Every major AI search product fetches live pages from search indexes and synthesizes answers from them, which means the quality of AI answers depends directly on the health of the content layer underneath. As Lily Ray, VP of SEO and AI search at Amsive, puts it in the Similarweb report, if your organic visibility dips, your AI search visibility follows, because the models are less likely to find your content.This has created a troublesome feedback loop. AI answers are built on an information supply chain whose funding model, ad-supported clicks, is collapsing primarily because of AI answers. A critical question for the continued viability of the open web is whether some alternative business model will work, and what the model will be.The replacement economy is forming inside the chatThere is some early data showing where things may be going. Following ChatGPT’s May 7 search update, which surfaced prominent clickable brand links inside answers, referral traffic from ChatGPT surged by 157% in a week. But the shape of that traffic changed: the share of referrals landing on homepages more than doubled, from roughly 25% to nearly 60%.Traditional search sent users to specific articles and deep pages tied to specific queries. AI referrals increasingly deliver a pre-informed visitor to a brand’s front door. The chatbot does the researching and comparing; the human arrives ready to act. Similarweb’s data shows AI-recommended brands receive two to four times as many subsequent visits as competitors that were not recommended.Money follows the behavior. Sponsored results appeared in 26% of U.S. desktop ChatGPT conversations in June 2026, up from 14% just a month earlier, per Similarweb’s ad intelligence data. Two-thirds of those ads appear after the second prompt, targeted on conversation context rather than a keyword. Click-through sits around 0.50%.The traditional search engine keyword auction is being replaced by something new: paid placement inside a conversation, targeted on accumulated context. That is a direct challenge to the auction Google has run, and dominated, for two decades.Google’s monopoly meets a new competitorGoogle is not a bystander here; it is simultaneously the incumbent being disrupted and one of the largest players in the disruption. AI Overviews now appear in a growing share of Google searches, more than 40% by May 2026 per Similarweb, and visits to Google’s conversational AI Mode have climbed steadily since launch. Google is cannibalizing its own click economy rather than ceding the territory.But the ground was already shifting under Google’s core business. eMarketer projects Google’s share of U.S. search advertising will fall below 50% in 2026, the first time since roughly 2004. The biggest chunk of that lost share is going to Amazon, whose sponsored product searches count as search advertising and which are growing three times as fast as Google’s. Conversational ads are barely a rounding error in that accounting right now, but they open up a second front in a war Google has got used to not needing to fight.Meanwhile, more structural shifts are coming for Google. A federal court entered final judgment in the DOJ search antitrust case in December 2025, imposing remedies that bar exclusive default agreements and require Google to share search data with qualified competitors. Google appealed in January 2026; the DOJ cross-appealed seeking stronger remedies. However the appeals resolve, the de facto arrangement that made Google the web’s tollbooth, defaults everywhere and a closed index, is ending just as conversational advertising is changing the landscape.The competitive landscape that results is genuinely new. OpenAI, Google, Perplexity, and Microsoft are now competing not just for users but for the advertising demand that funded the open web, and none of them, including Google, controls the new surface the way Google controlled the old one.Does conversational advertising help or hurt the open web?It’s not clear whether this new model helps or hurts the web. The web as a destination for human attention is shrinking, and the ad-supported publishers built for that web are in real trouble. The web as a machine-readable substrate is growing in importance, and a new referral and advertising economy is forming that routes value to brands rather than to content pages.The problem for publishers may be that they are powerless to influence the outcome. Ahrefs, analyzing over a billion data points across its studies, found that 67% of ChatGPT’s most-cited sources are things marketers cannot influence: Wikipedia alone accounts for nearly 30%. And 28.3% of ChatGPT’s most-cited pages have zero Google organic visibility, suggesting the retrieval layer is only partially tethered to traditional search, a complication for anyone assuming SEO success translates cleanly.Your website needs to be rebuilt for the new way people find itAccording to three independent datasets, in the new world, the pages AI systems cite and the pages AI systems send humans to are different pages doing different jobs. That’s a big change, and most teams are still optimizing for the old world.Similarweb’s data shows 65% of ChatGPT-cited URLs sit two or three folders deep in a site, while 58.8% of referral traffic lands on homepages. Ahrefs found the same split in its own analytics: more than 80% of its AI referral traffic goes to its homepage, product pages, and free tools, not its extensive editorial content. And a Previsible analysis of 6.77 million AI-referred sessions found a third destination: 28.8% of ChatGPT referrals land on internal site search pages, a navigation surface most publishers have long neglected precisely because Google searches were doing it for them.The right action to take is to audit your search traffic patterns. Pull your AI referral logs and whatever citation data you can access, and map which pages are being quoted as evidence versus where visitors actually enter. If it looks like you’re in the new world, there are three clear things to do:Deep pages, documentation, comparisons, and benchmarks should be structured to be citable: specific claims, clear headings, and descriptive URLs (Ahrefs found pages with natural-language URL slugs get cited at 89.78% versus 81.11% without).The homepage should be rebuilt for a visitor who arrives with context from a conversation rather than from a blue link. They already know you have what they need, get them to it as quickly as possible.And internal search, a neglected feature on most sites, is now an acquisition surface that deserves real UX investment.There’s a lot that’s still unknown or in flux here. But the underlying shift is confirmed by every independent source that has looked: the click economy is not coming back, and the entities that learn to be quoted by machines and to convert the humans those machines send will own whatever the web becomes next.

Uh-oh: Some Claude shared conversations and Artifacts appear to be indexed and publicly accessible on Google Search

July 27, 2026 MMN Editor Filed Under: Uncategorized

Over the weekend, Reddit user -void1 posted an alarming discovery on the r/ClaudeAI subreddit: some conversations that users of Anthropic’s Claude AI chatbot had made “shareable” via a link were being indexed by Google Search, and could be clicked on and accessed by seemingly anyone.The conversation took off on the social networks X and Reddit, the latter with thousands of upvotes and comments, many expressing concern about user privacy and information security, and the additional finding by users that shared Claude Artifacts — including interactive applications, dashboards, documents and other AI-generated work products — were also appearing in Google Search results. VentureBeat independently verified that some Claude Artifacts not shared directly with us were indeed searchable and accessible via Google. We could not access any shared conversations. By Sunday morning, many of the original Google search results for shared Claude conversations appeared to have disappeared or become significantly harder to find, suggesting either Google, Anthropic or both had begun taking action. The exposure could carry broader implications for enterprise users. Anthropic has increasingly positioned the feature as a collaborative workspace for building and sharing software, dashboards, documents and other business assets rather than simply chatbot responses.I have reached out to Anthropic for comment and will update this article when the company responds.A simple Google search yields a trove of Claude conversationsReddit user -void1 posted to r/ClaudeAI on July 25, 2026, demonstrating that the Google query site:claude.ai/share surfaced numerous publicly accessible Claude conversations.Screenshots shared across Reddit and X showed Google returning pages from Claude’s /share URLs, while other users reported finding conversations containing cryptocurrency wallet creation, legal questions, résumés and internal business discussions.While many users expressed concern that conversations they believed were effectively “unlisted” could become discoverable through public search engines, others argued the behavior reflected the expected consequences of creating publicly accessible share links rather than a software vulnerability.Indeed, Anthropic requires the user themselves to go into Claude’s options and select to make a conversation or Artifact shareable to others with the link, warning them it will be accessible to anyone with it, over multiple dialog boxes. It is similar to sharing a Google Doc link, where the user must also select the option — it is not enabled by default. Why the exposure of Claude Artifacts may be even more concerningOn July 26, X user Om Patel, founder of research firm BigIdeasDB, posted allegingthat searches such as site:claude.ai/public/artifacts surfaced publicly shared applications, dashboards, reports and documents.Screenshots circulating online appeared to show search results referencing internal-looking proposal documents and other business materials. Another widely circulated post warned that users often interpret “Anyone with the link” as equivalent to an unlisted YouTube video—accessible only if someone possesses the URL—not necessarily as content eligible for indexing by public search engines.VentureBeat independently verified that multiple third-party Claude Artifacts appeared in Google Search results for the query site:claude.ai/public/artifactslaunch and were accessible without authentication, despite the URLs not being previously known to the reporter. However, VentureBeat has not independently verified the full volume or representativeness of the examples circulating on social media.The reports are particularly significant because Artifacts has become one of Anthropic’s flagship product initiatives.First introduced alongside Claude 3.5 Sonnet in June 2024, Artifacts transformed Claude from a conventional chatbot into a collaborative workspace capable of generating interactive web applications, dashboards, visualizations, documents, games and other live software alongside a conversation. VentureBeat previously described the launch as potentially marking the beginning of an “interface war” among AI companies, shifting competition from raw model performance toward collaborative AI workspaces.Anthropic subsequently rolled Artifacts out to all Claude users, saying tens of millions had already been created, before expanding the concept again this year into Claude Code. That update allows engineering teams to publish live HTML dashboards and interactive project workspaces directly from coding sessions, making Artifacts an increasingly important part of Anthropic’s enterprise strategy.That broader functionality raises the potential stakes if publicly shared Artifacts were also being indexed. Unlike ordinary chat transcripts, Artifacts can contain interactive software prototypes, engineering dashboards, planning documents, product mockups, data visualizations and other work products organizations increasingly rely on to collaborate across technical and business teams. If those pages become searchable through public search engines, the exposure could extend well beyond conversational text.A reality check on privacy, information security and the open webImportantly, nothing so far suggests attackers gained access to private Claude accounts or conversations. Rather, the controversy centers on conversations and Artifacts that users explicitly chose to share publicly via Claude’s sharing tools.The dispute instead is whether users reasonably understood those shared pages could become discoverable through public search engines rather than only by recipients possessing the link.Technically, pages that are publicly accessible without authentication can generally be indexed by search engines unless publishers explicitly prevent crawling through mechanisms such as noindex directives or other indexing controls.Several Reddit commenters noted that Claude’s long, randomly generated share URLs are effectively impossible to guess. Instead, search engines typically discover them only after links appear somewhere they are permitted to crawl, such as public websites, forums or social media posts. Others questioned exactly how Google initially discovered so many Claude share URLs.The issue also illustrates a growing challenge for AI companies as chatbots evolve into collaborative workspaces for creating software, documents, dashboards and business applications. Features originally designed to make sharing AI-generated work easier now increasingly expose assets that may carry significantly more business value than a simple conversation. As enterprises adopt AI as a platform for building internal tools and workflows, the distinction between “shared by link” and “publicly discoverable through search” becomes far more consequential.A recurring challenge for AI companiesAnthropic is far from the first AI company to confront the distinction between “shared” and “searchable.”Reddit users quickly pointed out that OpenAI previously faced criticism after publicly shared ChatGPT conversations became discoverable through Google, prompting similar debates over whether “share by link” should imply a publicly indexed webpage or something closer to an unlisted document.Anthropic’s situation also echoes an incident involving Google’s pre-Gemini AI assistant, Bard, in September 2023. SEO consultant Gagan Ghotra discovered that Google Search had begun indexing shared Bard conversation links, warning that users could mistakenly assume they were sharing conversations only with intended recipients rather than making them discoverable through search. Google later responded publicly that it did not intend for shared Bard chats to be indexed and said it was working to block them from Google Search while emphasizing that only conversations users explicitly chose to share were affected.Together, the Bard, ChatGPT and now Claude episodes suggest AI companies continue to wrestle with the boundary between content that is technically public on the web and users’ expectations that “share with a link” behaves more like an unlisted Google Doc or YouTube video than a webpage eligible for indexing by search engines.What enterprises should do nowFor organizations deploying generative AI broadly across employees, the distinction between “shared with a link” and “publicly discoverable through search” is not merely semantic. It can determine whether an internal engineering dashboard, financial model, product roadmap, customer-facing prototype or AI-generated application remains effectively private—or becomes visible to anyone using a search engine.Whether this ultimately proves to be a technical indexing oversight, a mismatch between product design and user expectations, or some combination of both, the episode serves as another reminder that AI products are increasingly functioning less like chatbots and more like collaborative operating systems for knowledge work.As those platforms begin hosting internal dashboards, software prototypes, financial analyses, business planning documents and increasingly sophisticated enterprise applications, seemingly small decisions about how shared links behave can have outsized consequences for enterprise security, product design and user trust.Enterprise leaders should consider taking several practical steps:Audit existing shared AI content: Review shared conversations, Artifacts and other publicly accessible AI-generated assets to determine whether they should remain available or be unpublished.Clarify what “Share” actually means to your ENTIRE organization: Don’t assume employees understand the difference between “accessible by link” and “discoverable through search.” Update internal guidance to explain how each AI platform handles shared content.Treat AI platforms like collaboration software: Apply the same governance you use for Google Docs, Microsoft 365, Slack, GitHub, Notion or SharePoint—including policies around sharing sensitive intellectual property, customer information and regulated data.Prefer authenticated enterprise workspaces for sensitive information: When possible, keep confidential projects, code, financial models and customer data inside enterprise accounts with identity-based access controls instead of publicly accessible links.Review vendor defaults and sharing controls: As AI platforms evolve rapidly, administrators should periodically revisit default sharing settings, retention policies and indexing behavior rather than assuming they remain unchanged after new feature releases.For now, it appears Anthropic has begun limiting the visibility of at least some shared pages in Google Search, though reports suggest cached copies, archived pages and indexing by other search engines may persist for some content. Enterprises that have relied on Claude’s sharing features may wish to review existing shared conversations and Artifacts while Anthropic’s investigation continues.

Why SAP says enterprise AI agents need knowledge graphs and governance

July 27, 2026 MMN Editor Filed Under: Uncategorized

Presented by SAP At VB Transform 2026, Max McPhee, senior solution advisor at SAP, spoke with Rob Stretchay, lead analyst at VentureBeat Research, about what it takes for enterprises to move beyond chatbots to autonomous AI agents that can execute real business processes. He argued that the difference comes down to grounding those agents in a company’s own context rather than general knowledge.https://www.youtube.com/watch?v=SRf9t-wSZSo “Where we’re starting to see more emergent behavior of it feeling like a coworker rather than an assistant, is where we’re able to provide context on the actual enterprise rather than being able to use more of the standard knowledge,” McPhee said.That’s the gap that still separates most enterprise chat software from genuinely agentic systems.Building enterprise context with knowledge graphsThe same principles companies use to onboard new employees also apply to agents, adapted for software that retrieves information differently than humans do.”When you are onboarding a new agent, I think it’s important to acknowledge how you might onboard a new employee, but tune that for an agent,” McPhee said. “The way that is really powerful is using knowledge graphs and having vector-embedded data, because that’s a really easy format for an agent to be able to find and retrieve information.”That same grounding is also what keeps an agent from stumbling over an enterprise’s internal shorthand, a problem that’s acute in SAP’s world. “Being able to provide that tribal knowledge in the format that’s easy for it to consume helps to provide a really nice result with your agents versus a chatbot that might say, ‘Well, what does that acronym mean?'” he said.Bringing governance, identity, and security to autonomous agentsGovernance is an area where SAP’s history works in its favor, and the controls have been evolving for systems that act with more flexibility than earlier automation did.”That’s where SAP really has a good home, around that governance and process control,” McPhee said. We’re a 50-year-old process company, modernizing that governance to be able to handle the flexibility that comes with agents running.”One consequence is a renewed role for machine learning in validating agent behavior.”It’s becoming a bit of a revival of machine learning,” he added, pointing to customers that run agents within a process but then layer in anomaly detection and machine-learning-based validation as a guardrail. This is the same approach SAP had long used for intelligent approval recommendations.Identity and permissions carry that governance into execution. Under this model, both the human and SAP’s Joule, the generative AI assistant embedded across the company’s cloud applications and Business Technology Platform, must hold the rights to access a given system. Even if a user has permission to access S/4, they cannot do so through Joule unless the assistant has also been provisioned for that access, closing off the risk of using an agent to route around access controls.Balancing standard SAP with customized enterprise landscapesMuch of McPhee’s work involves reconciling SAP’s own knowledge with decades of customer customization and non-SAP systems. As he put it, many customers tell SAP, “You’re only 10% of my landscape,” a reality that has shaped the company’s recent strategy. Recent acquisitions such as LeanIX, which McPhee likened to “Google Maps for your architecture,” and process-mining company Signavio are intended to help map that non-SAP majority so SAP’s agents can understand how enterprise systems interconnect. The company has also invested in Berlin-based automation company n8n and is embedding it natively into Joule Studio, its intent-based, low-code environment for building agents.McPhee warned that companies also need to modernize older on-premises systems or risk running into limitations as they expand the use of autonomous agents.”You’re going to probably run into throughput issues, and you’re kind of trying to drive a Ferrari around a dirt track,” he said. “You’ve got to upgrade the track first if you want to drive a Ferrari.”Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact sales@venturebeat.com.

VentureBeat Research: Where enterprise AI agent governance hasn’t caught up

July 24, 2026 MMN Editor Filed Under: SUCCESS, Venture Beat

Enterprises deployed AI agents ahead of the controls needed to manage them — and they did it knowingly. That is the central finding across the five parallel surveys VentureBeat Research fielded in June, spanning every layer of the agentic stack. Now those enterprises are retrofitting to catch up with their own standards, and they are budgeting for it: In each of the five control layers we measured, 57 to 68% of enterprises plan to switch vendors or add new ones within 12 months, and roughly a third, depending on the layer, plan to move within the quarter.VentureBeat Research measured the five controls an enterprise has to build before it can trust an agent: identity, evaluation, cost telemetry, the context layer, and orchestration. Identity governs which agent is allowed to do what, under whose credentials. Evaluation determines whether the agent’s work is any good. Cost telemetry tracks what each agent costs to run. The context layer supplies the business data and definitions agents draw on when they answer. And the orchestration control plane coordinates multi-step agent work. Each of our five reports measures one of those controls.Most deployed “agents” are chatbots wearing the label. Seventy-one percent of enterprises said a quarter or fewer of their deployed “agents” can complete multi-step work on their own; only 10% said true agents are the majority of what they run. These respondents are positioned to know: 81% recommend or decide AI purchases at their companies. A single-prompt chatbot with a human reading every answer needs none of the controls the other four reports measure. A true multi-step agent needs all of them — and most enterprises can’t say which one they’ve deployed. (Full findings: Agentic Orchestration report.)Autonomy is outrunning trust in the evaluations that gate it. Two-thirds of enterprises either already allow an agent to push a code or system change to production on automated evaluation results alone, with no human review, or are actively engineering toward that within 12 months. Only 5% fully trust the evaluations that would make that call — and half of enterprises shipped an agent that passed internal evaluations and then caused a customer-facing failure in the past year. Before removing human review from any workflow, test evaluations against production outcomes rather than internal benchmarks. (Full findings: Agent Reliability & Evals report.)Companies that let agents share credentials get hit more often. Sixty-nine percent of companies let at least some of their agents share credentials — multiple agents operating under one API key or service account. Organizations that allow credential sharing anywhere experienced a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (nine of 22) at companies where every agent has its own scoped identity. The fix is scoped identity for every agent, starting with the ones that touch production systems. (Full findings: Agentic Security & Identity report.)The most expensive hardware in the building runs at half capacity or less. More than eight in 10 enterprises that run their own GPUs reported utilization of 50% or less, and only 44% rigorously track what their AI compute actually costs and returns. The number worth chasing first isn’t more GPUs — it’s the utilization and per-workload cost of the ones already running. (Full findings: AI Infrastructure & Compute report.)Agents answer confidently from data nobody governs. Fifty-seven percent of enterprises traced a confident, wrong agent answer in the past six months to their own missing or inconsistent business context — wrong metrics, stale definitions, absent documents — and most saw it happen more than once. Governing the definitions agents answer from — metrics and entities first — has to come before scaling the agents that depend on them. (Full findings: Context Layers / RAG report.)No layer has an entrenched incumbent: The defaults today are the built-in tools that ship with the big AI platforms enterprises already use. Switching intent runs highest in orchestration itself, where 68% plan to adopt, add, or replace platforms within 12 months and 34% within the quarter. Our surveys did not ask which direction that money moves — toward the platforms’ built-in tools or toward the specialists challenging them — and that open question is the next four quarters of this market.About this research VentureBeat Research fielded five parallel surveys in June 2026 under its VB Pulse program: Agentic Orchestration (101 respondents), Agent Reliability & Evals (157), Agentic Security & Identity (107), AI Infrastructure & Compute (107), and Context Layers / RAG (101) — 573 qualified respondents in total, all at organizations with 100 or more employees. Samples are self-selected, and some findings should be read directionally; each report carries its full methodology note. What the pattern supports more strongly than any single percentage is the direction: every survey, independently, points the same way. VentureBeat produces both this research and VB Transform, the conference where these reports debuted.

Anthropic launches Claude Opus 5, a cheaper AI model for coding, agents and enterprise workflows

July 24, 2026 MMN Editor Filed Under: Uncategorized

Anthropic released Claude Opus 5 on Friday, a model the company says delivers nearly all the intelligence of its top-of-the-line Claude Fable 5 at half the cost — a launch that signals how the AI race is shifting from raw capability to the economics of daily use.The model, available immediately on all of Anthropic’s platforms, is priced at $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor, Opus 4.8. It becomes the new default model on Claude Max, Anthropic’s premium consumer tier, and the strongest model available on Claude Pro.The positioning is deliberate. Anthropic is not claiming Opus 5 is its smartest model — that distinction still belongs to Fable 5, and rival systems retain an edge in certain domains. Instead, the company is making a subtler argument that may matter more to enterprise buyers: that the most economically important AI work happens in a middle band of difficulty, where near-frontier intelligence delivered efficiently and cheaply beats frontier intelligence delivered expensively.”Opus 5 as your daily driver, the model you hand complex work to and review when it’s done,” an Anthropic spokesperson said in an interview with VentureBeat, describing how the company’s lineup now stratifies. “Fable 5 for your most ambitious work, the days-long autonomous projects nothing could take on before… Sonnet 5 for work you run at scale, where speed and cost per call decide what ships. Haiku 4.5 for subagents and instant answers.”How Claude Opus 5 benchmark results stack up against Fable 5 and rival AI modelsOn paper, the results are striking. Anthropic says Opus 5 sets new state-of-the-art marks on coding and knowledge-work evaluations including Frontier-Bench and GDPval-AA. On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scores 43.3 percent — more than double Opus 4.8’s 18.7 percent and well ahead of Fable 5’s 33.7 percent — at a lower cost per task, according to the company. On ARC-AGI 3, an evaluation of novel problem-solving, Anthropic reports Opus 5 scored three times as high as the next best model. On OSWorld 2.0, a computer-use benchmark, the company says the model surpasses Fable 5’s best result at just over a third of the cost.The numbers come with honest caveats that are themselves notable in an industry prone to superlatives. Anthropic acknowledges Opus 5 remains behind Mythos 5, a competing model, on cybersecurity tasks and biology research, and an OpenAI-family model still leads on one agentic coding benchmark.The more revealing caveat came from Anthropic itself, when asked where Opus 5 still falls short of Fable 5. The spokesperson’s answer amounted to a candid admission about what benchmarks do and don’t capture.”The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it’s strongest. What those evals don’t measure is duration,” the spokesperson told VentureBeat. “One way to put it: Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark.”Fable 5, by contrast, “is for the longest, most autonomous jobs, where the model has to stay coherent across many connected steps over hours or days with dense source material,” the spokesperson said, advising customers to “run both on a representative workload, one bounded task and one long-horizon job.” That framing — bounded tasks versus long-horizon autonomy — may become the defining axis of model differentiation in 2026, as benchmarks saturate and the hardest remaining problems involve sustained, multi-day agentic work rather than discrete puzzles.Why token efficiency is becoming the real battleground for enterprise AI spendingThreaded through the launch is a theme Anthropic clearly wants buyers to absorb: Opus 5 doesn’t just score well, it scores well per dollar. The model ships with an adjustable “effort” setting that lets customers trade intelligence for speed and token savings, and Anthropic’s charts emphasize performance at a given cost rather than peak performance alone.Early customers echoed the point with unusual specificity. Harvey, the legal AI company, said Opus 5 achieved similar performance to Opus 4.8’s maximum-reasoning mode “while generating 26% fewer tokens on average,” according to Niko Grupen, its head of applied research. Richard Pham of Fundamental Research Lab said that on hard financial-modeling tasks, the model averaged nine percentage points higher accuracy “while using roughly one-third fewer turns and tool calls and 60% less time.”Wade Foster, chief executive of Zapier, said Opus 5 topped his company’s AutomationBench leaderboard “without spending more tokens than prior Claude models,” running a full churn-prevention workflow from start to finish. “Previous models didn’t pass; Opus 5 hit 100%,” he said. Scott Wu, chief executive of Cognition, the company behind the Devin coding agent, said that on FrontierCode 1.1, “Claude Opus 5 approaches Fable-level performance at half the cost,” with particular strength in debugging and root-cause analysis.The efficiency emphasis reflects commercial reality. Enterprise AI spending is no longer experimental, and inference costs — the price of actually running these models at scale — have become a board-level line item. Anthropic’s business skews heavily toward API and enterprise usage; according to a February 2026 analysis by Contrary Research, Claude held roughly 40 percent of the enterprise large language model market by usage as of late 2025, and Claude Code alone had reached about $1 billion in annualized revenue. For a company whose customers pay by the token, a model that does more with fewer tokens is not a nice-to-have. It is the product.Self-verifying AI agents and what they mean for the hidden costs of automationBeyond the numbers, Anthropic is selling a behavioral story: that Opus 5 verifies its work and iterates until it succeeds. The company offered several examples from testing that read like small parables of machine stubbornness.In one Frontier-Bench task, the model was asked to reconstruct a machine part as a 3D CAD model from a drawing it was intentionally given no way to view. Rather than fail, Anthropic says, Opus 5 wrote its own computer vision pipeline to extract the geometry from raw pixels — and did so repeatedly, while no competing model solved the task in five attempts. In another case, given a real bug in a popular open-source package manager, the model found the root cause and fixed an edge case the community’s own patch had missed; a competing model patched only the symptom and declared victory. An engineer at a trading firm, the company says, used Opus 5 to build a market data feed for a new exchange in a single session and, finding no live feed to validate against, watched the model build its own test harness to check its parsing code.Customers described similar behavior in the wild. Cristian Rivera, a staff software engineer at Stripe, said he gave the model “a chief-of-staff role over my dev environments” for a weekend: “it built its own monitor, drove each box, and pulled me in only for the judgment calls.”This is the capability enterprises actually care about, and it is worth dwelling on why. The gap between a model that produces plausible output and one that verifies its output is the gap between a demo and a deployable system. Most of the hidden cost of enterprise AI today is human review — engineers checking the machine’s work. A model that reliably checks its own work compresses that cost, which is precisely why customers keep citing fewer turns, fewer passes, and less time rather than higher raw scores.Inside Anthropic’s safety strategy: capability gaps, classifiers, and model fallbacksThe launch also showcases Anthropic’s increasingly intricate approach to safety — one that now involves deliberately not teaching its models certain skills. The company says its automated behavioral audit found Opus 5 to be its most aligned model to date, scoring 2.3 on overall misaligned behavior, lower than Opus 4.8, Sonnet 5, or Fable 5, with the lowest rates of deceptive behavior and the least susceptibility to being tricked into misuse.On the capability side, Anthropic says it intentionally avoided training Opus 5 on cyber tasks, as it did with Opus 4.8. The model improved on them anyway — a side effect of general capability gains — and now nearly matches Mythos 5 at finding software vulnerabilities. But it remains far behind at exploiting them: on Anthropic’s OSS-Fuzz evaluation, Opus 5 identified vulnerabilities at a 79.4 percent rate, close to Mythos 5’s 80 percent, but succeeded at developing exploits in only 4 challenges versus Mythos 5’s 13. That asymmetry — strong at defense-relevant discovery, weak at offense-relevant exploitation — appears to be by design, and the safeguards follow the same logic. Anthropic expects Opus 5’s cyber classifiers to intervene about 85 percent less often than Fable 5’s.When a classifier does trigger, requests in Claude.ai, Claude Code, and Claude Cowork fall back to Opus 4.8 by default — raising an obvious question: if a request is too risky for one model, why is it acceptable for another? “The model it falls back to has lower capability levels making the risk of harmful use lower as well,” the spokesperson said, adding that “there is a message that lets the user know when this occurs and is visible in the chat.”The logic is defensible, but it reveals how AI safety actually works in 2026: risk is not a property of the question alone, but of the question multiplied by the capability of the system answering it. On biology, the calculus runs the other way. Opus 5 is now Anthropic’s most capable generally available model for scientific research — scoring 10.2 percentage points higher than Opus 4.8 on the company’s internal chemistry benchmark — though the spokesperson acknowledged that “Mythos 5 remains the stronger model for long-horizon, open-ended work like autonomous drug design campaigns.”The business stakes behind the launch: a $380 billion valuation and massive compute betsThe launch lands at a moment of extraordinary commercial momentum — and extraordinary obligations — for Anthropic. Reuters reported in February that the company was valued at roughly $380 billion in its latest funding round, following a period in which, per Contrary Research’s analysis, its annualized revenue climbed from about $1 billion at the end of 2024 to a projected $9 billion by the end of 2025, with internal targets reportedly reaching $20 to $26 billion for 2026. Those targets are underwritten by enormous infrastructure commitments, including a reported $30 billion Azure compute deal alongside arrangements with Google Cloud and Nvidia — spending that only pencils out if enterprises keep expanding usage.That is the context in which Opus 5’s pricing strategy makes sense. Holding the price at Opus 4.8 levels while roughly doubling performance on key agentic benchmarks is effectively a steep price cut per unit of capability, designed to widen the funnel of workloads that are economical to automate. Every task that was marginal at Opus 4.8’s cost-per-success becomes viable at Opus 5’s — and every viable task is recurring token revenue.The regulatory backdrop has grown more complex as well. A U.S. judge gave final approval this week to Anthropic’s $1.5 billion copyright settlement with book authors, Reuters reported, closing a chapter of litigation over the company’s early training data. And in June, Reuters, citing Axios, reported that the U.S. government had moved to block foreign access to Anthropic’s most advanced models — a reminder that frontier AI is now entangled with export policy in ways that shape which customers can buy what.Also shipping Friday: a Fast mode running at roughly 2.5 times default speed at twice the base price, automatic fallback routing on the API, and mid-conversation tool changes that no longer invalidate the prompt cache — a small feature that agent developers may appreciate more than any benchmark. Consistent with prior Opus models, Opus 5 carries no data retention requirements for general access, a point the spokesperson flagged unprompted for customers with “a hard zero data retention requirement.” Developers can access the model as claude-opus-5 on the Claude API starting today.Two questions will determine whether the bet pays off: whether Opus 5’s efficiency claims survive contact with production workloads at scale, and whether enterprises embrace a world where safety classifiers, not users, sometimes decide which model answers. But the deeper message of Friday’s launch is that the AI industry’s center of gravity has moved. For three years, the labs competed on what their best model could do on its best day. With Opus 5, Anthropic is competing on something less glamorous and far more lucrative: what a very good model can do every day, for half the price. In a market where the frontier keeps moving, Anthropic is wagering that the real fortune lies just behind it.

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

July 23, 2026 MMN Editor Filed Under: Uncategorized

Microsoft AI released two new in-house models into public preview on Wednesday — MAI-Image-2.5-Pro, its highest-fidelity image generator to date, and MAI-Voice-2-Flash, a speech model built for high-volume enterprise workloads — while publishing production data that amounts to the company’s most aggressive argument yet that it can power its own products without leaning on OpenAI’s frontier models.The announcement, made by Microsoft AI’s Superintelligence team, lands roughly a year after the company committed to building purpose-built models internally, and it arrives with an unusual level of specificity about where those models now run: Bing, PowerPoint, OneDrive, Dynamics 365, Excel, GitHub Copilot, and Azure. The message to enterprise buyers — and, implicitly, to OpenAI — is that Microsoft’s homegrown models are no longer research projects. They are production infrastructure serving millions of users.”Each of these enhancements is a step toward the same goal: Microsoft products, powered by Microsoft models,” the company wrote in its announcement blog.How MAI-Image-2.5-Pro and MAI-Voice-2-Flash stake out opposite ends of the AI cost curveThe two new releases occupy opposite ends of what Microsoft calls the quality-speed-cost curve, and the positioning is deliberate. MAI-Image-2.5-Pro targets the premium tier: hero imagery, detailed editing, and precise in-image text rendering — the last of which has long been a notorious weak spot for image generation models. Microsoft priced the model at $5 per million text input tokens, $8 per million image input tokens, and $106 per million image output tokens. The base MAI-Image-2.5 model recently launched at No. 2 for image editing on Arena, the community leaderboard that has become a de facto scoreboard for generative media.The creative industry appears to be taking notice. Rob Reilly, global chief creative officer at advertising giant WPP, called the Pro model “a strong leap forward for GenMedia tools” in a statement included in Microsoft’s announcement, adding that “Microsoft has firmly established itself among the leaders in generative AI.”MAI-Voice-2-Flash goes the other direction. First previewed at Microsoft’s Build conference, Flash runs twice as fast as MAI-Voice-2 and costs 32% less, priced at $15 per million characters. It is designed for the unglamorous but enormous market of high-volume voice — call centers, voice agents, and real-time speech applications where latency and cost-per-call matter more than marginal gains in expressiveness. Together, the two models reflect a strategy of building families of models rather than a single flagship, because, as the company put it, a creative studio chasing maximum fidelity has very different needs from a customer service operation handling millions of calls a day.Microsoft’s production metrics show in-house models cutting GPU costs by up to 89%The model launches are arguably less newsworthy than the deployment metrics Microsoft attached to them — numbers that read like a systematic case for swapping out third-party frontier models across its product portfolio. Bing Image Creator now runs entirely on MAI-Image-2.5, end to end, marking the first time the consumer image tool is fully in-house. In PowerPoint, Microsoft says MAI-Image-2.5 reduces GPU costs by up to 84% compared with GPT-Image-2, OpenAI’s image model. In OneDrive, where MAI-Image-2.5 is now the default for key image-editing scenarios, the company reports a 26% increase in save rates, roughly 25% lower P95 latency, and 2.5 times greater efficiency under medium-utilization production workloads.On the voice side, MAI-Voice-2-Flash now powers Dynamics 365 Contact Center — the platform used by customers including T-Mobile and EasyJet — where Microsoft claims GPU cost reductions of up to 89%. The model is also integrated into Azure Voice Live for developers building speech-to-speech agents.Perhaps the most consequential deployment sits in healthcare. Microsoft’s Dragon Copilot, used by 170,000 medical providers and responsible for processing 28 million patient encounters last quarter, now runs on MAI-Transcribe-1.5 for its multilingual workflow across 58 languages. Microsoft says internal evaluations show a 50% relative reduction in both transcription and language-identification error rates across most languages — a meaningful claim in a domain where transcription errors can propagate directly into clinical notes.Inside the ‘hill-climbing’ strategy that lets small models beat GPT-5.6 in ExcelIn a companion post published the same day, Microsoft detailed the methodology behind these results — what it calls its “hill-climbing machine,” an integrated flywheel of data, models, and the product “harness” that surrounds them.The clearest example is MAI-Code-1-Flash, the lightweight coding model launched in GitHub Copilot in June. Microsoft says the model achieves an approximately 10% higher code accept rate than GPT-5.4 Mini and Claude Haiku 4.5 in VS Code, while using 10% fewer median tokens. Developer retention tells a similar story: users were 6% more likely to return across multiple days than with GPT-5.4 Mini, and 11% more likely than with Claude Haiku 4.5.Then Microsoft did something more interesting. It took the MAI-Code-1-Flash checkpoint and further trained it inside an Excel reinforcement learning environment, teaching a coding model the tools and workflows of spreadsheet knowledge work. The result, according to production user feedback, is a model on par with GPT-5.6 for the most common Excel tasks — while being small enough to run on Nvidia’s older H100 and even A100 GPUs rather than requiring the latest-generation accelerators.That hardware detail deserves emphasis. Every major AI company is fighting for allocation of cutting-edge chips, and a model that delivers frontier-adjacent quality on two-generation-old silicon fundamentally changes the deployment economics. It also frees the newest hardware — including Microsoft’s now-operational GB200 cluster — for training rather than serving.Satya Nadella’s ‘frontier diffusion’ manifesto redraws the OpenAI relationshipMicrosoft CEO Satya Nadella framed the announcements in a lengthy post on X titled “Frontier Diffusion & Control,” which functions as something close to a strategic manifesto. “We can now take saturated frontier capabilities and deliver them at scale and at lower cost through models optimized for high-usage products, while continuing to use frontier models for frontier needs,” Nadella wrote, adding that Microsoft is “beginning to route traffic across our first-party surfaces to MAI whenever our models match or outperform frontier alternatives.”Translated from executive prose: capabilities that were state-of-the-art a year ago are now table stakes, and Microsoft believes it can replicate them cheaply for the specific, repetitive tasks that dominate real product usage. Why pay frontier prices for a frontier model when a user just wants to reformat a spreadsheet column?Nadella was careful to note that “frontier models from OpenAI and Anthropic are part of the orchestration system alongside MAI” — but he also articulated a pointed principle of model independence, arguing that a company’s evaluations “should continue to hill climb even when any given model has been removed.” “Keeping the harness, memory, context, and skills outside the model, he argued, is what gives Microsoft control. The subtext is hard to miss. Reuters reported in April that Microsoft’s exclusive license to OpenAI’s technology had been revised into a non-exclusive arrangement, and The Information reported last September that Microsoft had begun incorporating Anthropic models into some products. Wednesday’s announcement completes the triangle: Microsoft as orchestrator, with its partners’ frontier models as interchangeable components and its own models absorbing an ever-larger share of routine traffic.”Developers cheer cheaper task-specific models while skeptics question Microsoft’s track recordThe response online captured both the appeal and the skepticism surrounding the strategy. “I love when people use small models for niche tasks,” wrote one X user, @mavihsk, responding to Nadella’s post. “Why do I have to use the all-knowing model just to change my field in Excel?” Another user, @nabu_lines, distilled the pitch neatly: “cost and performance both improve when you stop overusing the biggest model.”Others were less charitable about Microsoft’s execution track record. “Microsoft is the worst when it comes to listening to user feedback,” wrote designer @designedbyabin, arguing the company “will lose the AI race because they repeatedly failed to understand user needs.” And one user, @tokenoverflow, offered a drier critique of the model-independence pitch: “i want it keep hill climbing after removing microsoft.”The skeptics raise a fair point. Microsoft’s self-reported metrics — accept rates, save rates, GPU savings — come from its own internal evaluations, not independent benchmarks, and the company chooses which comparisons to publish.But the strategy’s logic does not depend on any single number. Nadella’s framing that software now has “real marginal cost for the first time” explains why Microsoft is obsessive about tokens, GPUs, and serving costs: when AI features run on every keystroke across a billion-user product portfolio, an 84% GPU cost reduction is not an optimization. It is the difference between a viable business and a money pit.Why Microsoft is turning its internal AI playbook into an Azure productThe final piece of the strategy is that Microsoft is selling the playbook, not just the models. Nadella explicitly positioned the hill-climbing approach as “a template for every other AI native, SaaS, or Enterprise company,” and Microsoft is packaging the toolchain through Foundry and what it calls Frontier Tuning — letting enterprises train specialized models against their own proprietary evaluations and reinforcement learning environments. That turns Microsoft’s internal cost-cutting exercise into an Azure product, and it gives enterprise customers a reason to run their AI workloads on Microsoft’s cloud even if the models themselves come from elsewhere.The company’s emphasis on models trained “on clean, traceable, enterprise-grade data, without distillation from third-party models” serves the same commercial end. In an industry facing mounting scrutiny over training data provenance, Microsoft is betting that enterprise buyers — and courts — will care where model capabilities come from. Microsoft says it is now extending the hill-climbing approach to Copilot Chat, Outlook, and PowerPoint, and both new models are available in public preview through Microsoft Foundry and the MAI Playground. “None of this is an endpoint,” the company wrote. “We’re just getting started.”Seven years ago, Microsoft bet more than $13 billion that OpenAI would build the future of AI. Wednesday’s announcement suggests the company has since learned a cheaper lesson: the future of AI may belong to whoever builds the frontier, but the profits belong to whoever makes it ordinary.

Agentic coding goes hands free as OpenAI brings GPT-Live’s full duplex voice control to Codex and ChatGPT on the desktop

July 23, 2026 MMN Editor Filed Under: Uncategorized

Two weeks after debuting its more naturalistic GPT-Live audio AI model with full-duplex capabilities (listening and speaking at the same time), OpenAI is bringing it directly into developer workflows. The company announced that GPT-Live now powers the ChatGPT desktop application on macOS and Windows, integrating directly with agentic systems like Codex and ChatGPT Work (which are separate experiences available in the ChatGPT desktop app). When OpenAI initially launched GPT-Live on July 8, 2026, it introduced a continuous audio model capable of listening and speaking simultaneously—eliminating rigid turn-taking while delegating complex reasoning to background models like GPT-5.5. Today’s release expands that conversational layer to technical tasks, enabling software engineers to orchestrate multi-threaded coding jobs, review pull requests, and debug applications using natural voice commands.As such, it could usher in a new era of “hands free” software development and even live, in-person group coding parties for Codex’s more than 5 million weekly active users. Codex, of course, is the name given to OpenAI’s models and harness focused on coding, but which the company has this year expanded into a more general productivity platform. An OpenAI spokesperson told VentureBeat this is the first time voice activation has been included natively with Codex on the desktop. OpenAI posted a promotional video showing some of its employees, Codex developer experience engineer Jason Liu and Codex technical staffer Guinness Chen, speaking to the same ChatGPT desktop app session in the same room, each issuing different instructions and conversing with the same model. New capabilities unlockedAt its core, this integration relies on decoupling the real-time voice layer from the underlying execution engines.While GPT-Live maintains fluid conversation—inserting natural verbal acknowledgments like “got it” without interrupting the user—it passes heavy computational workloads to background reasoning models. On macOS, the desktop application incorporates “Appshots” and screen context features, allowing ChatGPT Voice to analyze the frontmost window alongside local files, codebase structures, and active plugins.This architecture creates a pair-programming dynamic where developers talk through problems conversationally while agents execute tasks asynchronously. Rather than manually stopping coding sessions to type detailed instructions or switch windows, developers direct the system hands-free. The full-duplex engine dynamically decides when to speak, pause, or invoke tools, maintaining conversational state even as background agents process complex code modifications.Directing coding and complex builds with your voice aloneThe central operational capability in this update centers on multi-task execution across Codex and ChatGPT Work environments. Software engineers can initiate multiple concurrent task threads from a single spoken prompt. For instance, a developer preparing to ship a feature can instruct the system to investigate an open authentication bug, review a pending API migration pull request, and generate missing unit tests simultaneously.The desktop application coordinates these actions across disparate contexts, tracing issues through Slack conversations, GitHub repositories, and local codebases.Developers can also verbally convert design mockups into working code, splitting tasks across frontend, backend, and testing layers. With support for multi-folder projects (build 26.715) and remote execution via iOS, engineers can check task progress, answer agent prompts, and redirect active jobs without switching applications or managing individual processes line by line.Proprietary licenseOpenAI’s voice-enabled desktop release operates under a proprietary, commercial enterprise model. Access is restricted to paid subscribers across Plus, Pro, Business, Enterprise, and Education plans.For individual developers and corporate engineering departments, this commercial structure means the model weights, voice processing pipelines, and agent state architectures remain fully closed. Organizations cannot modify or self-host the underlying systems. Furthermore, tasks initiated via ChatGPT Voice consume standard usage allocations directly from existing Codex and ChatGPT Work plan quotas, treating voice-triggered actions identically to standard agentic workloads.Community reactionsDeveloper communities immediately noted the implications of bringing continuous full-duplex voice to autonomous coding workflows. Reacting to the build 26.715 release announcement—which details voice integration and multi-folder project support—AI Insider journalist @ChrisGPT noted on X: “Today OpenAI will release voice and remote guidance for codex ! One step closer to personal AGI”. Early technical feedback highlights widespread enthusiasm for orchestrating complex agentic tasks hands-free, particularly when stepping away from the workstation or managing build pipelines remotely.

Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start

July 23, 2026 MMN Editor Filed Under: Uncategorized

Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today’s launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt — and to extend the same underlying architecture to robotic vision and actions.The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface. That distinction is central to the company’s pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence — models, in the company’s words, “that can perceive, predict, and act across physical and digital environments.” This release marks BFL’s first public video generation model. FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated “Early Access” program now, to which anyone can apply, but which BFL must approve. There is presently no public access through BFL’s application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request. What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons.Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as “open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction” — a considerably broader commitment than any previous FLUX Dev release, all of which covered images only.But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside — or soon after — a major model announcement will have to wait. That delay does not negate the company’s commitment, but it is disappointing given the role open weights have played in FLUX’s adoption thus far. Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoptionBFL has published several benchmark comparisons, but they’re qualified as preliminary — with full benchmark results and methodology to be published later during broader general availability. In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google’s Gemini Omni Flash in 52%.One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a “preliminary evaluation of an early FLUX 3 candidate” — meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call.Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list — established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist.Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0’s international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one.Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting — multimodal input, video and audio-aware creation, conversational editing — and by BFL’s own measurement, the two are indistinguishable on 10-second text-to-video quality. Google’s advantage in that matchup is that Omni is generally available via Google’s Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around.One regional wrinkle matters for a German company’s home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash.Here’s a rough guide for enterprises considering which video models to rely upon: ModelMax single-generation durationMax resolutionKey constraintsPrice per 10-second clip (720p)Price per 10-second clip (1080p)Price per 10-second clip (4K)FLUX 3 Video 20 seconds Not stated; evaluations run at 720p Early access; no published SLA or pricing Not announced Not announced Not announced HappyHorse 1.1 15 seconds 1080p No 4K; closed weights Not published (v1.0 reseller rate is ~$1.82) Not published (v1.0 reseller rate is ~$3.12) n/a Veo 3.1 Per-second billing 4K Supports clip extension; preview $4.00 $4.00 $6.00 Veo 3.1 Fast Per-second billing 4K Preview $1.00 $1.20 $3.00 Veo 3.1 Lite Per-second billing 1080p No 4K, no clip extension; preview $0.50 $0.80 n/a Gemini Omni Flash 10 seconds (3s minimum) 720p at 24 FPS Preview abd no EU access$1.00 n/a n/a One architecture for media generation and physical actionFLUX 3 builds on Self-Flow, BFL’s method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026. The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations — the same architecture could be extended to action prediction without sacrificing what it learned from video.”We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture,” said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. “True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express.”He put the case more bluntly elsewhere in the announcement: “You can’t cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds.”BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart.For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models.For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations.What FLUX 3 Video can actually doThe video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation. Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI’s discontinued Sora model.The capability list BFL published covers:Text-to-video generation.Image-to-video generation, either animating from a starting frame or using images as visual references.Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context.Generative video-audio continuation from existing video and audio input.Keyframe-to-video generation for controlled transitions between defined moments. Multilingual dialogue.A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics.Typography generation and animated design.Agentic chaining of individual clips into longer, multi-shot sequences.That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots.It is also the capability where competition is most direct. HappyHorse 1.1’s headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage — the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it.BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates.FLUX-mimic tests whether video models can become robot modelsBFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access.The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route — the FLUX 3 backbone combined with mimic’s robot-learning and production-deployment expertise in dexterous manipulation.FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data. BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours.”The hardest part of robotics is data,” said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. “Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning.”BFL argues that a model trained only on images cannot understand a world that “moves, sounds, changes, and responds,” and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni. Its developer documentation cites “world knowledge” that combines “an understanding of physics” with Gemini’s grasp of history, science and cultural context. Its marketing is blunter still: “Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different,” the company posted in June, crediting the model with “an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic.” The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it. There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all.Open weights helped make FLUX an industry standardBFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises. The company’s founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies. That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research’s Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users.Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley’s largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco.FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm’s launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license.The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs.FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction — meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn’t yet shared information about its license, the parameter count, quantizations or hardware requirements.The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows. The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom’s T.Capital.

Multi-turn attacks broke AI models 88% of the time — single-turn testing missed it, Cisco AI security lead warns at VB Transform 2026

July 23, 2026 MMN Editor Filed Under: Uncategorized

When Cisco ran 6,986 multi-turn attacks against 15 flagship models, attackers who adapted across the conversation broke through as often as 88.3% of the time. Amy Chang, Cisco’s head of AI threat intelligence and security research, brought that finding to the agentic security panel at VB Transform 2026; the number should worry anyone still running single-turn red-teaming programs.VentureBeat’s June 2026 Pulse survey of 107 enterprise respondents explains why the room was full. More than half, 54%, have already had a confirmed agent security incident (18%) or a near-miss caught before harm (36%). Just 32% give every agent its own scoped, managed identity, and fewer still, 30%, isolate their highest-risk agents in sandboxes. Provider-native and hyperscaler controls remain the primary agent security layer at 82% of companies surveyed. The world’s largest security vendors have done the same math. Palo Alto Networks closed its $25 billion acquisition of CyberArk in February, CrowdStrike agreed in January to pay $740 million for SGNL, and Cisco announced its intent to acquire Astrix Security for a reported $400 million, all of it aimed at the identity and isolation layer most enterprises have not finished building.Chang came to the panel with almost two decades of experience spanning cybersecurity operations, government, and the military. She ran global cybersecurity operations as an executive director at JPMorgan Chase, where she led the bank’s cyber threat intelligence teams, and served as a senior staffer on the House Foreign Affairs Committee and as a U.S. Navy Reserve officer. She also teaches cybersecurity and emerging threats as adjunct faculty at the Middlebury Institute of International Studies.Chang’s 88.3% number comes from a study she co-authored with Nicholas Conley, built on 30,090 single-turn prompts and 6,986 multi-turn attacks against those 15 closed and proprietary flagship models. Multi-turn success rates ranged from 7.89% to 88.3%, every model tested showed non-trivial multi-turn exposure, and the two testing styles did not even rank the models in the same order. Cisco publishes adversarial evaluation signals for what is now 105 models on its LLM Security Leaderboard, she told the audience.”If you don’t understand how models are susceptible to different types of attacks, then you are unable to account for how that model that is powering your agent, that is powering your application, to understand where those failure points are,” Chang said. Single-turn testing is the one-shot malicious prompt, she explained, while extending an attack into a longer conversation “is more realistic of how we are actually engaging with our models, with our agents, with our applications.” That longer arc surfaces harmful outputs and misaligned behaviors that a snapshot never catches.Cisco has pushed the testing itself into agentic territory. Chang described a framework where agents assess a deployment scenario, develop relevant attacks, judge whether they are worth pursuing, execute them, and evaluate their own success. What surprised her most, after all that sophistication, was how simple the defensive answer stays. “The answer is still that it’s pretty simple,” she said. “You don’t have to get super creative. You just need to think about truly what are the fundamentals and basics of what I’m trying to secure in my organization.”Her starting point for CISOs beginning agentic deployments is Cisco’s Integrated AI Security and Safety Framework, which she said “stipulates all the ways that AI can be compromised across the AI lifecycle” from modality through supply chain. From there, teams can work backward from real incidents, trace how each attack was achieved, and use the framework to build a strategy with the right coverage and mitigations.Heather Ceylan, the CISO of Box, sees the same gap from the defender’s side. “A lot of what you see out there with agent red teaming is just single-turn, and that’s not how people are actually interacting with AI day-to-day,” she told the audience. Box now simulates multi-turn adversaries with agents that think like an attacker and iterate attempt after attempt to hijack the target. “You have to pressure test your agents because otherwise you don’t know if your execution controls are really working as you intended.”Box deployed agents inside its security operations center about a year ago, starting with human approval required for every action, and trust built quickly enough that analysts shifted into monitoring mode. Then the agent made one mistake, and every bit of that accumulated trust vanished. “They had to start all over again,” she said. “So I think that that monitoring piece is so important. Even if you’re not gonna have a human in the loop, things change, models change, and we can’t control how the models change and interpret things.”Rajesh Parekh, VP of AI and ML at Intuit, brought the builder’s perspective. Parekh led large-scale computer vision and ML systems powering Google’s Maps and Geo products before joining Intuit, and holds a doctorate in computer science. Three layers versus an operating systemCeylan described Box’s approach as three concentric layers. Permissioning comes first, so the agent never accesses more content than the human who invoked it. Ephemeral sandbox environments spin up for each agent task, containing the blast radius if an agent gets hijacked, and runtime execution control restricts the agent’s tool calls to only those relevant to the task at hand. “If you want an agent to summarize a doc for you, if you have a prompt injection that came in that says forward this to maliciousattacker at domain.com, it can’t do that,” Ceylan said. “That action in that tool call is not even in its vocabulary.”She classified agent actions into three oversight categories. Actions that are not sensitive, like read and summarize, need no human in the loop. Moderately sensitive actions skip human approval but get logged and monitored, while destructive actions like mass deletion of files always require a human. “Things are gonna shift between those three categories quite a bit,” she acknowledged, “but setting those types of categories up front allows you to have a principled framework.”Rather than layering controls onto agents one at a time, Intuit has built a central platform called GenOS, short for generative AI operating system, which abstracts security, risk, and fraud modeling so individual agent developers never reinvent protection. “Permissioning is not about giving access to AI,” Parekh said. “Instead, it is defining very tightly scoped and clearly auditable authority to the agent to perform very specific tasks.” Intuit evolved from agents inheriting user permissions to each agent carrying its own identity, and the company is now investigating mid-session permission changes tied to the specific task underway.Parekh calls the broader model an AI-powered expert platform, one where the human expert is built into the trust architecture rather than bolted on as a gate. “The paradigm that we are pursuing is where the user, the AI agent, and the human expert are collaborating to solve the user problem,” he said.The end of human code reviewCeylan took on the tension between security testing and development velocity without hedging. “The days of secure code reviews where a human’s looking at the code and we’re looking at security architecture reviews, design docs, those are done,” she said. “If you keep trying to do security that way, you’re gonna get left behind.” Box is building toward a fully agentic development lifecycle where agents review design documents, apply security requirements, and review the code for vulnerabilities. “I’m very optimistic that we will get to a point where we will write code without security vulnerabilities because agents and the models are going to get so good at writing code without vulnerabilities,” she said. “We’re still a long way away from that.”Her advice for development teams skips the advanced AI concepts entirely and returns to basics that predate agents. “It comes down to very basic least privilege access,” she said. “If you start giving your agents overly broad permissions at the beginning, it’s really hard to comb that back and build an infrastructure that allows for those ephemeral credentials and only those narrowly scoped tasks.”Parekh explained why the red teaming surface has expanded so quickly. “These agents have skills, and skills could become vulnerabilities,” he said. “Agents have access to certain data, they have access to tools, and there could be threats that are lurking within those tools as well. So suddenly the blast radius of the malicious code or the intent increases dramatically.” When Intuit identifies common vulnerability patterns from its manual red teaming exercises, it automates those tests back into the GenOS harness so future agents inherit protection and red teamers stay focused on new threat vectors. Runtime scanning of prompts and responses adds a final layer that can stop a suspect response and escalate to a human expert, he said.”You need to continuously test to ensure that those remain robust to the protections that you have built, as well as to account for any sort of drift or any other types of dependencies that you introduce into your scenario that can create novel vulnerabilities,” she said.Intent versus probabilityAn audience question about intent detection set off the sharpest exchange of the session. Ceylan noted that when Box’s own agent operates, the system always knows the user’s intent because it controls the prompt, which means guardrails and tool-call restrictions can be engineered around it. The harder challenge, which she admitted Box is still trying to solve, arrives when external agents connect and the context behind the request is opaque.That exchange exposed a split running through the wider industry. Mastercard, in the fireside chat immediately preceding the panel, came down on the side of quantifying intent, building an open-source framework to propagate it as a standard because complex B2B procurement cannot work without that trust. Endpoint security CTOs, in briefings with VentureBeat, have gone the other way, saying they will bet on probability rather than intent inference for production workloads. Chang explained why models, as they are trained today, cannot reliably derive intent from a prompt, which is why deterministic controls and behavioral proxies remain necessary. Ceylan agreed that both are required. “If you’re not doing anything deterministic, you’re really relying heavily on that intent, and I haven’t seen programs that are there yet,” she said.Ceylan’s story about trust collapsing after a single agent mistake landed as the panel’s most memorable moment because enterprise agentic security is not a problem that gets solved and stays solved. Models change, permissions drift, and adversaries adapt across multi-turn conversations that snapshot tests never capture.For the 82% of enterprises relying on provider-native controls as their primary security layer, and the 59% shopping for agent security tooling over the next 12 months, the panel’s takeaway was blunt. Test the way attackers attack, across full conversations and continuously, or find out in production what your single-turn red teaming missed.

An AI now judges every move Rubrik’s agents make, its AI chief said at VB Transform 2026 — but no one’s measured if the judge is right

July 23, 2026 MMN Editor Filed Under: Uncategorized

At a CISO roundtable organized by Anthropic’s chief information security officer, Dev Rishi asked a simple question: Did everyone in the room have their AI governance and security policies written down? Every hand went up — about 14 people, by his count. His follow-up, about how anyone actually enforces those policies in practice, got a different response. “And everybody chuckled,” Rishi, the GM of AI at Rubrik, recalled at VB Transform 2026 fireside chat in Menlo Park. “It was like the dirty secret in the room that everyone has these policies, but no way to actually make them real.”“Our founder and CTO has actually been really pushing to enable our agents in YOLO mode,” Rishi told the audience. That admission comes from a publicly traded data security firm whose business is backing up what he called the most important data in the world.YOLO mode strips the permission prompt out of agent workflows and lets the agent act on its own. In Rubrik’s version, a second AI judges every action in real time against policy in place of a human clicking approve. Rubrik is running the experiment on itself first. Rishi treats autonomy as a settled capability question and an open judgment question. “If you ask the agent to act autonomously, it will,” he said. “It’s a question that you have internally. Should it?”Rubrik earned that question the hard way. When Claude Code and Cowork pilots rolled out, the company required every command to run in ask mode so the employee issuing it carried the liability, and the developer pushback filled a single Slack thread 120 messages deep. “The developers basically are pushing back, and they’re like, this is like the iTunes service agreement. I’m just hitting check, check, check, check, check, check, check,” Rishi said. “There’s no way that I can actually read through this. And it becomes security theater.” Roughly 80% of respondents are in the same bind, Rishi said, citing Rubrik Zero Labs research that found monitoring and approving agent actions takes more time than the agents save. The State of the Agent, the April report behind that figure, surveyed more than 1,600 IT and security leaders.SAGE is the reason Rubrik trusts the bet. Short for Semantic AI Governance Engine, SAGE is the arbitration layer inside Rubrik Agent Cloud that watches every action an agent takes and reads the semantic intent behind it, then rules the action in or out against policies written in natural language. “We took what people said was human in the loop, a good idea, and we replaced it with AI in the loop,” Rishi said, describing the pitch to security chiefs he characterized as skittish about non-deterministic systems.Security approval, not cost, blocks AI ROIRishi’s path to Rubrik ran through Predibase, the generative AI infrastructure startup he co-founded and ran as CEO until Rubrik agreed to acquire it in June 2025. Before that, he led ML product at Google on the team that became Vertex AI, served as Kaggle’s first product manager as it grew from about one million to ten million users, and holds bachelor’s and master’s degrees in computer science from Harvard. Over roughly his first three and a half months at Rubrik, Rishi set up 200 customer conversations with IT and security leaders across a customer base that looks like the Global 2000, asking open-ended questions about cost, latency, performance, and orchestration. “Pretty consistently, what I heard through all of those conversations was that all of those are pretty secondary,” he said. “The main challenge is actually, how do I get this approved from a security and risk standpoint? I’m concerned about all the different things that could go wrong. Actually, I felt like that was one of the biggest things constraining ROI.”VentureBeat Pulse research presented on the Transform stage earlier in the day confirms the gap Rishi kept hearing. Two-thirds of enterprises, 66%, already allow or are actively building toward production deployment with zero human review, yet only 5% fully trust the automated evaluations that would make that decision. One AI reading what the rulebook can’tRubrik’s own policies exposed why written rules fail as enforcement. One internal rule states that agents should respect Rubrik’s customer data use policy, which sounds enforceable until someone tries. “Rubrik’s customer data use policy is like a three-page document of legal text,” Rishi said. “I have no idea how to write that in there as a rule.” Asked on stage how a team of AI infrastructure people took on a problem that security engineers own, Rishi answered, “with a lot of naivety and innocence, honestly.” His team bet that models good at understanding language could police other models, and SAGE became the answer.The case for putting a model in the judgment seat comes down to precision. A rule like “agents should not be able to edit revenue fields in Salesforce” fails in conventional tooling because Salesforce does not delineate which fields count as revenue, Rishi explained, so administrators fall back on approving every Salesforce action by hand. SAGE reads the intent instead and acts as a judge, carrying organizational context, which can tell a benign lookup from the edit the policy prohibits.Keeping the judge small is what makes the economics work. SAGE runs on a small language model that Rishi said operates at an order of magnitude lower cost and latency than a frontier LLM. “If I told you, don’t worry, you’re gonna be secure and governed, but I’m gonna double your cost and latency, you would tell me to get out of the room,” Rishi said.When Rishi asked who in the audience had worried about token consumption over the past year, half the hands went up. “And I guess the other half is probably just too lazy to raise their hand,” he said.SAGE is an aggregation of judges based on parameter-efficient fine-tuning that Rubrik uses to take on task-specific variants of a base model with shared organizational context. One judge watches for tool-use hallucinations while another suppresses PII before it can leave, each running as its own enforceable policy. Security and GRC teams have started writing financial rules into the same layer, including one internal policy barring AI spend on personal projects.The lethal trifectaAsked which attacks worry him most, Rishi pointed at the lethal trifecta, the term security researcher Simon Willison coined in June 2025 for an agent that holds private data while taking in content nobody vetted, with a channel to send what it finds to the outside world. The danger, according to Rishi, is what happens when individually legitimate permissions stack. An agent granted Salesforce access and email access on an employee’s credentials has done nothing wrong yet, with yet being the operative word. “A very simple example is that an agent can start pulling data from Salesforce and then decide to accidentally leak and exfiltrate that out via an email,” he told the audience. A financial services company he met the morning of the session made the point for him, telling Rishi that none of the individual permissions are bad on their own and the agent needs every one of them to do its job. “It should have permission to each of those systems, but it’s the combination that ends up becoming really destructive,” Rishi said.Traditional identity and access management never priced in that combination because it relied on the judgment of the employee holding the credentials, Rishi argued, and agents supply none. “I can tell you the number of times Claude Code has tried to leak some of our sensitive source code to a public GitHub repository is incredibly high,” he said. Cutting agents off from public resources entirely would defeat their purpose, which returns the problem to adjudicating intent in context rather than revoking access.A separate VentureBeat June Pulse survey of 107 qualified enterprise respondents maps the blast radius of exactly this pattern. On the Transform stage that morning, VentureBeat research reported that 69% of companies run credential sharing somewhere in their agent fleet. Companies with shared credentials anywhere got hit more often, reporting a security incident or near-miss at a 63.5% rate (47 of 74), against 40.9% (9 of 22) where every agent carries its own scoped identity.The attacks no single turn revealsRubrik Agent Cloud reached general availability in February, though not everything Rishi described ships in it yet. Backtesting is just starting to roll out. The feature replays an organization’s historical agent actions and tool calls against a new policy, showing where the policy would have stepped in and where an action would have sailed through uncaught, with policy edits applied in real time. Rishi called that archive one of the most valuable data troves an enterprise holds.Real-time detection and blocking turn out to be the entry point rather than the whole product. Some attacks never trip a single-action rule. “No individual turn of the conversation was problematic, but if you took the session as a full trace, that ended up being problematic,” Rishi said. Agent Cloud runs batch analysis across entire session traces every hour or every day and surfaces what Rubrik calls insights, the problems no individual guardrail caught. The same Zero Labs report found that 88% say they lack the ability to roll back agent actions without system disruption, a recovery gap that sits squarely in Rubrik’s original line of business.A skeptical CISO will ask the question the fireside did not answer. SAGE is a non-deterministic model policing other non-deterministic models, and Rishi offered no false positive or false negative rate for the judge itself. The closest thing the architecture gives to an answer is auditability, since backtesting and the batch insights both leave a human-reviewable trail of each call SAGE made and whatever got past it. Who watches the watcher, for now, is a trail of receipts rather than a benchmark. Until that benchmark exists, AI in the loop stays an operational wager rather than a quantified control.Three questions fall out of the session for security teams. How many of the guardrails now in production depend on a human clicking approve, and what happens to that workload as agent count grows? Does anything in the stack enforce semantic intent, or is it all allow and deny lists? And can the team backtest agent behavior against a new policy, then unwind a multi-turn session without taking systems down?Rishi’s timing has a market behind it. In the same VentureBeat research, 82% of enterprises still name their primary AI provider’s built-in guardrails and cloud controls as their main agent security layer, and 59% plan to adopt, add, or replace agent security tooling within the next 12 months. Only 12% include an agent-identity product in what they are considering, even with credential sharing still the norm. Every CISO at that Anthropic roundtable had a policy document and no enforcement mechanism, and Rubrik built a product for the space between the two. YOLO mode is the bet that an AI watching other AIs can finally make the policies real.

  • « Go to Previous Page
  • Page 1
  • Interim pages omitted …
  • Page 7
  • Page 8
  • Page 9
  • Page 10
  • Go to Next Page »

© 2026 Mad Mad News™ · OGGHY Media™ Live Above the Madness™ Independent news, signals, and analysis. Atlanta, Georgia

Live Above The Madness

Market Wire

Find the signal. Investigate the opportunity.

Market Headlines

Search A Stock

Enter a ticker or company name to open a deeper market view with quote data, charts, company news, financials, and research.

GO DEEPER: Quote • Chart • News • Financials • Research
Primary Source Latest SEC Filings

Search company filings, 10-Ks, 10-Qs, 8-Ks and other disclosures.

Opportunity Watch IPO Watch

Explore upcoming, recent and newly listed public companies.

Minute News Brief

A quick audio briefing for readers who want the market and business picture without opening another video.

Quick Market Pulse

S&P 500 Dow Nasdaq Gold Oil Bitcoin

Market links open third-party research pages. MMN does not provide investment advice.