The read
Cross-device, write-enabled agents from OpenAI, Anthropic, Microsoft, and Google arrived this week alongside the approval modes, runtime policy gates, and injection defenses meant to keep them in bounds.
Thesis
This was the week AI agents stopped advising and started acting — and the vendors leading shipped the controls to contain them in the same breath, moving the competitive frontier from model quality to who can run real work unattended without it going off the rails.
Market shifts
- Agents crossed from advisors to operators — and went cross-device. In one week OpenAI shipped ChatGPT Work, Anthropic moved Claude Cowork onto web and mobile with cloud sessions that keep running after you close the laptop, and Microsoft took Sales and Service Agents to GA on live CRM data over MCP. By the weekend Claude's Microsoft 365 connector gained write access — drafting and sending email, editing OneDrive and SharePoint — Cursor 3.11 added lifecycle hooks that turn the editor into an agent runtime, and Notion split Agents into a standalone iOS app. The unit of competition shifted from chat quality to reliable, unattended task completion.
- Governance and security became the gating layer — and shipped alongside the autonomy, not after it. Claude Code flipped its default permission mode to Manual, Google introduced runtime intent-gating with Semantic Governance Policies, GitHub's CodeQL added AI prompt-injection detection, and Anthropic and Vercel added expiring API keys and build-log secret redaction. The urgency is earned: LiteLLM's AI-gateway flaw reached unauthenticated CVSS 10 RCE and landed on CISA's exploited list, Cato's DuneSlide zero-click hit Cursor's agent sandbox, and Sysdig documented JADEPUFFER, assessed as the first fully agentic ransomware operation.
- The competitive axis moved to open weights and day-one model distribution. Z.ai's MIT-licensed GLM-5.2 landed within four points of Claude Opus 4.8 on SWE-bench Pro at roughly a sixth of the cost, and Moonshot's Kimi K2.7 Code became the first open-weight model in the GitHub Copilot picker. When GPT-5.6 and Grok 4.5 launched, GitHub Copilot and Vercel raced to expose them the same day — making time-to-model, not feature depth, the moat — even as GPT-5.6 Sol's wafer-scale serving on Cerebras (~750 tok/s, ~10x GPU) signaled a real non-GPU path for frontier inference.
Why it matters
For builders and operators, the buying question changed this week. Choosing an agent platform is no longer just which model is smartest — it is whether the platform can run work unattended and prove what it did. The vendors setting the pace paired every new action capability with approval modes, runtime policy gates, key rotation, and audit trails, because the threats went live: a CVSS-10 gateway RCE on CISA's exploited list and the first documented agentic ransomware. Meanwhile open-weight coders like GLM-5.2 and Kimi K2.7 Code make good-enough-at-a-sixth-the-cost a genuine procurement option — if you can accept the governance caveats.
Watch next
- Whether Manual-by-default (Claude Code) becomes the norm or the outlier as OpenAI, Microsoft, and Google push unattended, work-completing autonomy — the two directions are on a collision course.
- The July 15 China companion-law deadline: whether ByteDance's Doubao goes read-only and Alibaba's Qwen disables agent features on schedule, and whether it hands momentum to less-restricted providers.
- Enterprise uptake of open-weight coders (GLM-5.2, Kimi K2.7 Code) inside proprietary toolchains like GitHub Copilot, and whether Chinese-data-law caveats slow adoption.
- Whether runtime prompt-injection defenses (Google's Semantic Governance Policies, GitHub CodeQL 2.26 scanning) move from preview into default agent CI/CD.
- Whether Cerebras wafer-scale serving for GPT-5.6 Sol under the $20B+ contract expands past its initial limited customer set — the first credible non-GPU frontier-inference path.