Gynga AI Topics
JA EN

Copilot Cowork Goes GA — What Microsoft’s Internal Cost Comparison Suggests About Orchestration


TL;DR

  • Microsoft’s “Copilot Cowork” is now generally available. Under the hood it runs Anthropic’s Claude Opus 4.8 and Sonnet 4.6, yet Microsoft claims it costs 30–40% less per prompt than Anthropic’s own “Claude Cowork” (ITmedia).
  • The comparison conditions and Claude Cowork’s exact model configuration are not public, so the difference cannot be attributed to orchestration alone. It does suggest that agent-service economics depend on more than model pricing, with M365 integration and runtime design also mattering.
  • Also covered: Kimi K3’s weight release and Anthropic’s CEO staking out a position on distillation regulation (ITmedia), nudify requests targeting children on Hugging Face (The Verge), and Claude shared chats appearing in search results (ITmedia).

Lead story: Copilot Cowork goes GA

Microsoft wrapped up a three-month preview and rolled out “Copilot Cowork,” a new feature within Microsoft 365 Copilot, to general availability worldwide, effective June 16, 2026 (US time) (ITmedia). Through natural-language instructions, it handles drafting and sending Outlook email, managing calendars, creating Word/Excel/PowerPoint documents, posting to Teams, and searching across an organization’s internal resources. On July 27, NTT DATA Advanced Technology announced an internal pilot of the tool for sales administration work.

The technical read

As of GA, Copilot Cowork runs on Anthropic’s Claude Opus 4.8 and Claude Sonnet 4.6. Through the Frontier program, users can also select OpenAI’s GPT-5.5, and Microsoft says its own fine-tuned model, “Cowork 1,” is coming soon — a multi-model, multi-vendor setup. Two things set it apart. First is “Work IQ,” which uses tenant permissions to retrieve relevant M365 business context and execute within Microsoft’s trust boundary. Second is a runtime design that calculates cost from four factors — model usage, amount of context retrieved, number of tool calls, and execution time — and sorts tasks into Light/Medium/Heavy tiers so that only the necessary processing actually runs. Users can create up to 50 custom skills.

The business read

The most important number here is Microsoft’s internal comparison claiming Copilot Cowork costs 30–40% less per prompt on average than Claude Cowork (ITmedia). Claude Cowork, the comparison target, is an Anthropic product using an M365 connector, but its exact model configuration and the comparison conditions are not public. Microsoft attributes the lower cost to a runtime that retrieves only the necessary information and tools and selects a model by task. That does not prove orchestration caused the full difference, but it does suggest that service economics depend on context retrieval and tool execution as well as model prices. Billing follows a two-tier structure: a fixed USL (per-user monthly license) plus metered Copilot Credits, meaning both fixed and variable costs stack.

On the build-vs-platform-absorption question, this feature puts pressure on thin SaaS wrappers selling M365 workflow automation. Each tenant owns its organizational data, but Microsoft’s native access to the M365 permission model and its ability to execute inside the same trust boundary are platform advantages.

The contrarian angle / what’s easy to miss

The “cheaper than Claude” claim comes from Microsoft’s own internal testing, not third-party verification. Because the comparison conditions and Claude Cowork’s model configuration are unknown, the 30–40% difference cannot be credited to runtime design alone. The safer takeaway is that agent-service costs include context retrieval, tool calls, and execution time in addition to model usage.

Implications and positioning

If you’re building a thin agent SaaS centered on Excel manipulation or email automation inside M365, this GA release is direct competitive pressure. A fixed-plus-metered pricing model is replicable; native M365 integration and execution inside Microsoft’s trust boundary are much harder for an independent developer to match. A stronger position is industry-specific data outside M365 or workflows that must run beyond Microsoft’s trust boundary. If your automation is fully self-contained outside M365, the direct impact is limited.

Signal assessment: medium. The GA release and product design are well supported, but the 30–40% cost difference comes only from Microsoft’s internal comparison, with neither causal isolation nor third-party validation.

Other notable stories

Perplexity’s “Personal Computer” expands to Windows

Perplexity has extended its agentic tool “Personal Computer,” previously Mac-only since April, to Windows (The Verge). Billed as a “general-purpose digital worker” that operates across local files, M365 apps, and the web, it’s available to Max/Enterprise Max subscribers starting at $200/month. Copilot Cowork runs in the cloud, while Personal Computer for Windows operates within Windows and can work with local files. Perplexity is moving beyond search into Microsoft’s enterprise territory, where platform incumbents hold advantages in price and integration depth. For smaller companies, this is less an invitation to copy the product than a sign that desktop agents have become a competitive category for major vendors.

Kimi K3 releases model weights and a technical report

China’s Moonshot AI has released the weights and technical report for “Kimi K3.” Despite being open-weight, it’s reported to match some capabilities of GPT-5.6 Sol and Claude Fable 5 (ITmedia). Fixstars reported successfully deploying it on a single node with eight NVIDIA B300 GPUs, but startup took 88.8 minutes total, of which model loading alone took 81.4 minutes. Even a top-tier open-weight model can impose substantial startup delay and operational load. Michael Kratsios, director of the US Office of Science and Technology Policy (OSTP), has claimed Kimi K3 was “distilled” from Claude Fable 5, and some voices are calling for regulation or sanctions in response. The next item, Anthropic’s CEO weighing in on distillation regulation, follows from this same debate. Practical implication: if you rely heavily on open-weight models, assess the 88.8-minute startup delay and operational overhead alongside inference cost.

Anthropic’s CEO clarifies his position on open-weight models

CEO Dario Amodei stated he has never advocated for an outright ban on open-weight models, and instead called for three things: (1) export controls on AI chips destined for China, (2) crackdowns on industrial-scale “distillation,” and (3) mandatory safety testing for high-capability models regardless of whether they’re open or closed (ITmedia). This statement serves as an explanation for why Anthropic did not sign the joint statement issued July 24 by Microsoft, NVIDIA, and others opposing excessive regulation of open-weight models. The technical point of contention is “distillation” — the practice of using a high-capability model’s outputs to cheaply train another model. On the business side, it’s tempting to read a motive of Anthropic protecting the value of its own high-capability models from distillation by staking out a distinct position — but that’s inference not stated directly by the source; Amodei’s own stated reasoning is framed around national security risk. Practical implication: this debate could affect the availability and pricing of open-weight models. Teams that depend on the Kimi line or similar models should monitor export-control and distillation rules.

Nudify requests targeting children expose Hugging Face’s platform safeguards

According to an investigation by the European nonprofit AI Forensics, 7 of the top 9 image-editing models hosted on Hugging Face generated nude images from simple prompts applied to images of women (The Verge). Decoy Spaces received over 1,000 prompts/images in seven days; 73% had sexual content, of which 83% were undressing requests (95% targeting women), and roughly 7% of all sexual requests targeted children. The investigation established that such requests targeted children; it did not establish successful generation from children’s images. Safety controls on Hugging Face depend heavily on individual Space developers and were not effective at the platform level in this test. That gap also conflicts with Hugging Face’s policy against non-consensual sexual content and depictions of minors. Practical implication: businesses distributing or hosting open-weight models must evaluate what safety guarantees the platform itself enforces.

Some chats that users deliberately published through Anthropic Claude’s “share” feature appeared in Google search results (ITmedia). Indexed conversations included cryptocurrency wallet private keys, names and addresses, and chats containing sexual content that violated Anthropic’s policies. This was not a leak of private-by-default conversations: users had created public share URLs, and those URLs were crawlable. Anthropic said it did not provide search engines with a listing or sitemap, while acknowledging that shared links can be archived like other public web content. Similar incidents at other chatbot providers show a recurring risk in treating share URLs as public web pages. Practical implication: organizations handling sensitive information should prohibit shared links or require content review before publication.

Gemini 3.6 Flash’s hallucinations, checked a week later

Google’s Gemini 3.6 Flash, which went GA on July 21, drew a wave of complaints on X immediately after launch — including simple errors like getting “which is bigger, 9.11 or 9.3” wrong (ITmedia). When a reporter re-tested similar questions on July 28, some of the earlier simple errors were not reproduced within the scope checked, though complaints about hallucinations and coding mistakes continue to appear on X. Google says it incorporated feedback from 3.5 Flash to improve coding, knowledge tasks, and image processing, and also improved token efficiency (17% reduction versus 3.5 Flash, up to 65% in some benchmarks). A limited retest cannot establish whether the model itself improved or whether prompting or serving conditions changed. Practical implication: evaluate low-cost, fast models against your own test set instead of treating early screenshots as durable evidence.

Try this week / hype worth ignoring

  • Try: apply one of Copilot Cowork’s custom skills (up to 50 available) to a real M365-centric workflow, and measure your actual USL-plus-Credits spend against your own workload. Don’t take the “30–40% cheaper” figure at face value — it’s worth verifying against your own prompt patterns.
  • Ignore: posts that merely recycle Gemini 3.6 Flash’s launch-week reputation. A few errors failed to reproduce, but that does not establish a model improvement. Test it against your own evaluation criteria.

Sources