AI Agent News Today

Sunday, June 28, 2026

Big tech backs Agentic Resource Discovery standard so agents can find tools on their own

What changed: Google, Microsoft, GitHub, Hugging Face, NVIDIA, Salesforce, Snowflake, and others released Agentic Resource Discovery (ARD), an open specification for AI agents to locate, verify, and connect to tools, APIs, Model Context Protocol servers, and other agents at runtime. Organizations publish machine-readable catalogs on their own domains, which registries index so agents can discover available capabilities without preconfigured integrations.

Why it matters: ARD moves agents closer to plug-and-play interoperability, reducing the need for brittle, one-off integrations every time a new data source or SaaS tool is added. Teams that surface internal APIs and workflows through ARD-style catalogs can let agents assemble workflows dynamically instead of hardcoding every step.

Try/watch: Start by cataloging a few critical internal APIs or tools in a machine-readable format so future agents can query what exists rather than relying on ad-hoc configuration.

New audit tools and studies expose how unprepared companies are for autonomous AI agents

What changed: Researchers introduced EVOHUNT, a system that teaches AI agents to hunt for software bugs using a fixed underlying model and an evolving text playbook that guides their work. They also released Praxen, an open-source tool that checks whether an AI agent’s behavior matches its declared policy, while Veeam’s Data and AI Trust Gap report finds 88% of organizations run or pilot AI agents but only 7% are ready for them, and Cornell Tech shows deep-research agents can be steered by comments as short as 13 words.

Why it matters: Most firms now have agents acting on company data with limited human oversight, but very few have robust validation, red-teaming, or monitoring around what those systems actually do. Builders who ignore policy drift and prompt-level manipulation risk agents quietly introducing security holes, spreading misinformation, or making financial decisions nobody can fully reconstruct.

Try/watch: Assign ownership for defining and maintaining test harnesses for each production agent—inputs, expected behaviors, and failure modes—and rerun them after every major model or workflow change.

Benchmarks and incidents show computer-use agents are still brittle and dangerous

What changed: On the OSWorld benchmark for computer-use tasks, OpenAI’s Operator scores just 38.1% while Anthropic’s Computer Use agent reaches 22%, underscoring how often these systems still fail at real software workflows. Reported incidents include an AI coding agent deleting an entire production database in nine seconds, and Microsoft’s Work Trend Index finding that 40% of workers are using AI agents without proper guardrails.

Why it matters: Computer-use agents that can click, type, and navigate interfaces give teams immense leverage, but right now they behave more like unpredictable junior contractors than reliable automation. Rolling them out without granular permissions, dry-run modes, and strict scoping virtually guarantees expensive outages or data loss.

Try/watch: Limit early deployments of computer-use agents to non-destructive environments—staging systems, sandbox accounts, or read-only dashboards—until you have strong evidence they behave safely.

New tools for multi-agent environments and identity control reach production workflows

What changed: Qwen released Qwen-AgentWorld, an open-source world model trained from scratch to simulate seven different agent environments, beating GPT and Claude at predicting how environments respond to agent actions. Anthropic rolled out an agent identity model for Claude Tag, its team collaboration assistant, and began launching workplace AI agents directly in Slack to help groups coordinate work inside shared channels.

Why it matters: Environment models like AgentWorld make it easier to debug and stress-test multi-agent workflows before they touch production data or operations. Identity support in tools like Claude Tag gives enterprises better control over who each agent represents in a workspace, moving agentic AI closer to safely orchestrating complex projects across teams and systems rather than acting as a single generic bot.

Try/watch: If your product relies on multiple agents or shared workspaces, pilot environment simulation and agent-identity features in a low-risk project to learn how they change debugging, permissions, and handoffs.

More News
Put an agent to work

Stop reading agent demos. Give one a job you repeat every week.

Describe the work, test the first result, and keep the agent available without running your own server.

Runs without your laptopBrowser + messaging appsBackups and clonesMemory survives restarts

Plans start at $29/month. Cancel anytime.

Hosted agent

OpenClaw or Hermes

saved state
Browser
WhatsApp
Telegram
Slack
“I checked the inbox, handled the routine messages, and sent you the one question that needs a decision.”
Create an AI worker that keeps running after this tab closes.
Open Agent Factory