Summary

July 13 was dominated by two currents: a widening frontier-model price war and a rapid maturing of the infrastructure agents run on. SpaceXAI's Grok 4.5 and Meta's first paid Model API both undercut Anthropic and OpenAI on token pricing while targeting agentic and coding work, and fresh top-tier models (OpenAI's GPT-5.6 Sol, Anthropic's Claude Sonnet 5) were absorbed into vertical and enterprise platforms within days. In parallel, the agent stack moved toward production: LangChain and LangGraph hit 1.0 with a stability pledge, AgentPrizm shipped governed memory and portable skills, and OpenAI and Microsoft pushed outcome-delivering office agents into general availability. Coding-and-deploy friction kept falling (GitHub's Copilot desktop GA, Cloudflare Drop), while Vercel added tooling to defend agents against MCP tool 'rug pull' attacks.

Key themes

  • Frontier-model price war intensifies: SpaceXAI's Grok 4.5 ($2/$6 per M tokens) undercuts Claude Opus 4.8 by ~60% on input and ~76% on output, and Meta opened its first paid Model API at $1.25/$4.25 per M tokens — both pitched at agentic and coding workloads rather than consumer chat.
  • Frontier models flow into vertical and enterprise platforms fast: Harvey wired in OpenAI's GPT-5.6 Sol for legal agents within days, and Microsoft added Anthropic's Claude Sonnet 5 as a selectable model across Copilot — multi-vendor model choice is becoming table stakes.
  • Office agents that deliver finished work went GA: OpenAI's ChatGPT Work (GPT-5.6) produces documents, decks, and apps from connected data, and Microsoft's Copilot Sales Agent reached general availability — both moving from chat assistant to outcome delivery.
  • Agent infrastructure is hardening for production: LangChain and LangGraph reached 1.0 with a no-breaking-changes commitment, and AgentPrizm launched governed persistent memory and reusable, discoverable skills with audit receipts — long-running, auditable agents over one-off sessions.
  • Coding and deploy friction keeps falling, and agent security is now a first-class concern: GitHub's Copilot desktop app hit GA with Codex-powered JetBrains agents, Cloudflare's Drop deploys a dragged folder to the edge with no account, and Vercel's AI SDK added fingerprintTools/detectToolDrift to catch MCP tool 'rug pulls.'

Notable items

  • SpaceXAI launches Grok 4.5 — a Cursor-trained coding/agentic model priced at $2/$6 per M tokens, claiming ~4.2x fewer output tokens than Claude Opus 4.8 on SWE Bench Pro (high impact).
  • OpenAI launches ChatGPT Work — a GPT-5.6 agent that pulls team context to produce finished documents, spreadsheets, decks, and lightweight sites, folding in Codex for multi-step tasks (high impact).
  • Meta launches Muse Spark 1.1 and its first paid Model API — a 1M-token multimodal reasoning model for agentic work at $1.25/$4.25 per M tokens, a strategic shift from open-weights distribution (high impact).
  • LangChain and LangGraph reach 1.0 — refocused agent loop with a middleware model plus durable, persistent, human-in-the-loop production agents, backed by a stability commitment until 2.0.
  • Microsoft makes Copilot Sales Agent generally available and adds Claude Sonnet 5 as a selectable model across Copilot, including Cowork and PowerPoint.
  • Harvey adds OpenAI's GPT-5.6 Sol to its legal agents — stronger BigLaw Bench results plus web-app localization (French for Canada and France at GA).
  • GitHub makes the Copilot desktop app generally available, adds Codex-powered agents in JetBrains, Kimi K2.7 for Business/Enterprise, and per-user cost-center budgets.
  • AgentPrizm launches a governed agent memory and skills platform — confidence-scored, validity-aware facts with audit receipts and supersede chains, plus versioned SKILL.md procedures in a public marketplace.
  • Cloudflare launches Drop — drag a folder or zip into the browser for an account-free live site on its edge network, living 60 minutes as frictionless proof hosting.
  • Vercel AI SDK 7.0.19 adds fingerprintTools and detectToolDrift to defend against MCP tool 'rug pull' attacks where a server swaps a benign tool definition for a malicious one after approval.

Source coverage

Source rows used: 10