Gynga AI Topics
JA EN

AI Agent Extension Standardization: Claude Is Not Yet on the Supported-Client List


TL;DR

  • Vercel’s proposed “Agent Plugins” is emerging as a common standard that lets Agent Skills (SKILL.md) and MCP configurations be reused across different agents, with OpenAI, Amazon, Cursor, and Microsoft on its technical steering committee. Of the eight supported clients, Claude Code and Claude Cowork are currently not among them.
  • A survey of the Salesforce field found concern over over-reliance on specific individuals (88.1%) and expectations that AI can standardize work quality (86.4%) running ahead of actual practice, while 66.7% of those already using AI said they have no system in place to verify its output — making visible, in numbers, the gap between adopting AI capability and building verification processes.
  • Research on automatically refining skills (SkillProx), research mapping the blind spots in AI risk mitigation tools, and NEC’s “all-AI” organizational restructuring — this week’s common theme is building the infrastructure for those who operate agents.

Lead story: AI agent extension standardization, with Claude absent from the current support list

On August 6 (local time), Vercel announced “Agent Plugins,” a standard format for packaging AI agent extensions. It lets the instruction documents that agents reference — “Agent Skills” (SKILL.md) — and MCP server configurations (mcp.json) be written in a common format that can be reused across multiple agents. Vercel proposed the standard, and its technical steering committee includes OpenAI, Amazon, Cursor, and Microsoft. The eight clients currently supported are VS Code, Cursor, Kiro, Hermes Agent, GitHub Copilot, OpenClaw, ChatGPT, and Codex; Anthropic’s Claude Code and Claude Cowork are not on this list. (ITmedia)

Technical reading

What’s actually happening here is modest but practical. It’s about aligning the previously inconsistent ways each agent product wrote instruction documents and configured MCP servers into a common file format — it doesn’t raise the underlying capability of the models themselves. The supported clients span both IDE-style tools (VS Code, Cursor, GitHub Copilot) and general-purpose agent tools (ChatGPT, Codex), giving the standard broad reach. That said, immediately after the announcement, the article doesn’t make clear whether each client’s support amounts to “full implementation” or merely “reading the format.”

Business reading

While major players like OpenAI, Amazon, Cursor, and Microsoft sit on the technical steering committee, neither Anthropic’s name nor Claude Code/Claude Cowork appears on the list of supported clients. The article does not establish Anthropic’s participation status or the reason for its absence. What it does establish is narrower: Vercel proposed the standard, those four companies joined the technical steering committee, and Claude products are not among the eight supported clients listed at launch. As standardization advances, the Skill format itself becomes a common language anyone can write, and the defensible asset shifts to the “content” of the Skill — the domain knowledge, data, and workflow design it encodes.

Implications and positioning

The parts that can be templated — the Skill format and distribution mechanism — are likely to commoditize rapidly, making investment there a poor bet. Conversely, the content of a Skill that distills a winning pattern for a specific task becomes easier to deploy across multiple agents once it rides on a standardized distribution channel. Codifying your organization’s domain knowledge into Skills now can become an asset that doesn’t depend on which agent you happen to use. On the other hand, if Claude is your primary agent, it is worth watching whether an interoperability gap emerges if this standard becomes mainstream.

Signal assessment: medium. The support list and steering-committee membership are concrete, but the depth of each client’s implementation, the reason Anthropic is absent, and real-world adoption of the standard remain unverified. Confidence that interoperability will become a competitive requirement is therefore moderate.

Salesforce field survey: dependence on individuals meets an AI “verification gap”

In a survey by Copard of 110 practitioners involved in Salesforce development and operations at companies with 300 or more employees, 88.1% said they feel their organization’s Salesforce operations depend on specific individuals. The reasons cited were a shortage of Apex talent (58.8%) and insufficient documentation of development history and change logs (51.5%). In response, 86.4% expect that “AI can standardize the quality of work,” with the most common rationale being that AI “can produce consistent-quality output regardless of the individual practitioner’s skill” (63.2%). A combined 62.8% said they have incorporated generative AI or AI agents into their work, with “development (code generation and configuration)” being the area of most advanced adoption at 68.1%. (ITmedia @IT)

Among those already using AI, the most common challenge cited was “no system or process in place to verify the correctness of AI output” (66.7%), followed by restrictions from security and governance requirements (59.4%) and an uneven distribution of people able to write effective prompts (50.7%). Adoption is running ahead of verification infrastructure. This pattern — “lack of a verification system” as the top-ranked challenge — is likely a general phenomenon and not something confined to Apex specifically. There is clear demand for templating that transfers individualized domain knowledge into AI, and packaging a verification mechanism alongside it would create value a level deeper than a simple automation tool.

Taxonomy of AI risk mitigation tools: tooling clusters in technical and operational layers

A paper published on arXiv in August mapped the capabilities of 21 major open-source LLM evaluation and security tools onto the MIT AI Risk Mitigation and Response Taxonomy (an extended version with 32 subcategories). The researchers used an LLM-based RAG pipeline to extract capabilities from source code and documentation, and reliability across three independent reviewers came in at Fleiss’ Kappa = 0.509 (“moderate agreement”). The analysis found that the tools cluster heavily around technical and operational controls — model evaluation, adversarial testing, runtime guardrails — while controls related to governance, legal/regulatory compliance, and financial markets remain largely untouched. (arXiv)

A Fleiss’ Kappa of 0.509 indicates “moderate agreement,” meaning the reproducibility of LLM-assisted automatic mapping still has room for improvement. Even so, the direction of the imbalance — a thick technical layer and a thin governance/legal/financial-markets layer — corroborates where tool-development effort tends to gravitate. Technical verification tools (evaluation, guardrails) are already heading toward commoditization, making it a tough space for new entrants, while tools handling governance, regulatory compliance, and financial risk controls remain a blank space — which lines up with the “lack of verification systems” need that ranked highest in the Salesforce survey above. This is primary data suggesting the opportunity lies not around technical tools themselves, but one layer up, in controls and audit.

SkillProx: automatically refining skills

A paper titled “SkillProx,” published on arXiv in August, proposes a method for automatically refining, based on diagnostic results, the “skills” — lightweight text artifacts loaded into context without any weight updates — that LLM agents accumulate to adapt to repeated tasks. In the forward stage, it repeatedly performs diagnosis-driven edits and rolls back any that degrade performance; in the backward stage, it decomposes skills into auditable units of knowledge, estimates each unit’s contribution using a leave-one-out approach, and then consolidates, demotes, or removes units through a verification gate. Across multiple backbone LLMs and both in-distribution and out-of-distribution benchmarks, the method reportedly outperformed the best existing gradient-based method by an average of 3.0 percentage points in accuracy. (arXiv)

Read alongside the lead story on Agent Skills standardization, this carries real significance. Even as the “format” of Skills solidifies into an industry standard, research is simultaneously advancing on automatically refining the “content” of Skills. If this kind of self-improvement gets implemented in a product, the operational labor of manually maintaining a Skill library becomes itself a target for automation. For now this remains a benchmark-level improvement (+3.0 points) rather than a production implementation, but it points toward a shift from “writing Skills” to “automating Skill quality management” as the next axis of competition.

NEC’s “all-AI” new division: self-reported numbers, a structure worth watching

On August 10, NEC announced it had established, effective August 1, an internal organization called the “Corporate AI Workforce Division,” in which AI fills every role from department head down to staff. It consists of four tiers — AI Division Head, AI Board (CxO functions), AI Managers, and AI Staff — with final evaluation, decision-making, and governance controls remaining with human executives. In a one-month internal pilot starting in July, AI reportedly handled management analysis, simulation, and risk-signal detection for executive board meetings end to end, cutting the time required for that work to roughly one-seventh. NEC positions itself as the “client zero” for this approach and plans to sell the resulting organizational framework and methodology to other companies going forward. (ITmedia)

The “one-seventh” figure is NEC’s own claim with no third-party verification, and it applies only to a specific management-analysis task over a single month, so it cannot be generalized. What’s worth watching isn’t the underlying technology but the fact that a major systems integrator is trying to commercialize the know-how of running an AI organization itself. If this takes off, the “how to operate AI” expertise that small and midsize companies have built up through their own trial and error will turn into a generic package anyone can buy. The evidentiary strength of the numbers is thin, but selling AI organizational know-how could reduce the scarcity of proprietary in-house operating knowledge — and it warrants close attention.

Worth trying this week / hype worth ignoring

  • Worth trying: Codify your organization’s winning patterns in SKILL.md format now. This gets ahead of the adoption cost if Agent Plugins becomes mainstream.
  • Worth ignoring: Citing NEC’s “one-seventh” figure as a proven, transferable result. It’s a self-reported number from one month and a single task, not a third-party-verified effect.

Sources