The Age of AI Agents: Codex Goes Company-Wide and Robots Build Robots
Today’s news makes it clearer than ever that AI is moving from “generating” to “acting.” Code-generation agents have penetrated even legal departments, AI is completing tasks by operating computer screens, and factories where robots autonomously mass-produce other robots have come online. Infrastructure investment is stacking up in the trillions of dollars with India at the center — this is no longer a story about AI’s future, but about AI in motion right now. Yet Ford’s cautionary tale reminds us that the question of how to combine data with human experience remains very much alive.
OpenAI’s Codex Reaches Legal and HR — First Quantified Look at Agent AI’s Economic Impact
An economic research paper released by OpenAI puts hard numbers behind the workplace penetration of agent AI. As of August 2025, even internally at OpenAI, Codex accounted for less than 10% of token consumption. By May 2026, every department — including non-technical functions like legal and recruiting — was using Codex as a primary AI tool, with a sharp increase in the share of long-duration tasks.
What’s striking is the shift in the unit of work. Where ChatGPT handles one-off question-and-answer exchanges, Codex functions as an agent that autonomously completes tasks spanning minutes to hours. OpenAI describes this as “a transition in the unit of knowledge work from single interactions to delegated, long-horizon tasks,” and says the productivity expansion is now quantitatively confirmed.
Behind Codex’s broad adoption is a combination of improving model capability and a steady cadence of new features. As capabilities grew, the range of applicable tasks expanded and usable non-technical use cases multiplied. The adoption pattern — starting with engineers and quietly becoming a company-wide tool — closely mirrors how Slack and Notion spread. It may not be long before agent AI commoditizes into standard business infrastructure.
Google Bakes “Computer Use” Natively into Gemini 2.5 Flash
Google announced that Gemini 2.5 Flash now ships with Computer Use as a built-in standard tool — the capability for AI to perceive a computer screen and execute mouse clicks and keystrokes. The feature was previously available only as the standalone model “Gemini 2.5 Computer Use”; it has now been integrated into Google’s flagship fast model.
The mechanism is straightforward but powerful. Give the AI a screenshot and a goal, and it returns instructions like “click at these coordinates” or “type this text here.” Developer code executes those instructions and feeds a new screenshot back to the AI — a loop that continues until the task is complete. The AI acts as the “brain” issuing instructions rather than directly driving the browser. Google envisions use cases including continuous software testing and cross-application workflow automation.
Security considerations are addressed alongside the launch. Two optional enterprise safeguards are offered: automatic task termination upon detection of adversarial prompt injection, and mandatory user confirmation before sensitive operations. The feature is currently in preview, and developer documentation explicitly states it should not be used for consequential decisions or irreversible actions. Computer Use has been a differentiating capability for Anthropic’s Claude, but Google’s integration into its flagship model signals that the competition for screen-operating agents is now in earnest.
Source: ITmedia AI+ - Google「Gemini 3.5 Flash」に「Computer Use」を標準搭載
Masayoshi Son Reveals “Robots Mass-Producing Robots” Factory at Shareholder Meeting
“I wasn’t planning to say this today” — at SoftBank Group’s annual shareholder meeting, Chairman and CEO Masayoshi Son made an unplanned disclosure about a robotics factory in which his firm has invested. A facility where “robots are already autonomously mass-producing robots” is operational, he said. “I believe this is probably a world first. We’ll be able to make a formal announcement soon,” Son said, choosing his words carefully while unable to conceal his excitement.
Son frames the AI evolution in two stages. Stage one is “conversational, image, and video generation.” Stage two is “agentic AI” — systems that think for themselves and run around the clock. He argues that CPU performance is the decisive variable in the stage-two competition. Subsidiary Arm Holdings has announced its “Arm AGI CPU” for AI data centers, claiming power consumption half that of competing alternatives. The strategy of differentiating against Nvidia’s GPU dominance through a CPU-plus-efficiency axis is evident.
A factory where robots mass-produce robots is not merely an automation story — it means exponential growth in production capacity. Unlike semiconductor manufacturing, robot manufacturing can itself leverage robots. Once the flywheel of falling production costs and accelerating deployment velocity begins to spin, the impact will ripple far beyond SoftBank’s investment returns to reshape entire industries.
China’s Humanoid Robot Shipments Surge More Than 7x in One Year, Capturing 80% Global Share
Growth data on China’s humanoid robot industry, presented by Nomura Research Institute analyst Li Zhihui, is striking. From 2024 to 2025, annual shipment estimates jumped from roughly 2,800 units to approximately 20,000 — a more than sevenfold increase. The number of manufacturing companies doubled to 200, and global market share exceeded 80%.
What explains the pace? Li identifies the “flywheel effect” as the primary driver: real-world deployment data improves AI models, and improved models generate new data from the field — a virtuous cycle now in motion. Additional tailwinds include the ability to repurpose supply chains (reducers, batteries, and more) built for EV and smartphone manufacturing, as well as strong domestic AI adoption. Smartphone makers like HONOR have entered the space, with the humanoid robot market functioning as a destination for hard-technology capabilities migrating out of a saturated phone market.
The challenges, however, are clearly defined. A 2025 Stanford University study found that even the top-scoring model on “BEHAVIOR-1K” — a benchmark for complex household tasks — completed only about 26% of tasks to an acceptable quality level, with full success on roughly 12%. A wide gap remains between mass-production capability and the maturity of the underlying control AI. Even so, Li’s implicit question is hard to dismiss: can Japan afford to ignore a growth curve that went from 2,800 to 20,000 units in a single year?
Amazon’s India Commitment Reaches $48B — The Shape of the Global AI Infrastructure Race
Amazon announced an additional $13 billion investment in expanding AI and cloud infrastructure in India. The announcement came immediately after CEO Andy Jassy met with Prime Minister Modi, with funds earmarked for AWS data center expansion in Mumbai and Hyderabad. The new commitment brings Amazon’s total pledged investment in India to $48 billion.
Amazon is far from alone. Microsoft has already announced $17.5 billion in AI hub and data center investment through 2029, and Google has committed $15 billion. The common thread is a shared recognition that India is consolidating as the “second front” of next-generation AI infrastructure. A market of 1.4 billion people, an abundant pool of technical talent, and an English-language regulatory environment make it attractive — particularly for Western companies for whom investing in China is difficult, as India increasingly functions as an alternative growth market.
That said, “commitments” and “net-new infrastructure investment” are not the same thing. Large long-term commitments typically include operating expenditures alongside capital spending, meaning not all $48 billion translates directly into new data center construction. Even discounting that, India’s AI infrastructure demand is building on genuine underlying need, and the business opportunities surrounding power, cooling, and network buildout are very real.
Source: TechCrunch AI - Amazon ups India bet with fresh $13B AI infrastructure investment
Ford Reclaims JD Power Top Spot for First Time in 16 Years — What Calling Back Former Engineers Taught Them About Automation’s Limits
Ford has returned to the top of J.D. Power’s Initial Quality Study among mainstream brands — its first such ranking in 16 years. But the backstory is a gritty one involving quality degradation from over-reliance on automation and AI, followed by a corrective course. Ford acknowledged that its automated production and design systems failed to function as intended, forcing the company to call back experienced engineers to address the problems.
The root cause came down to data quality. Ford’s conclusion — “the effectiveness of AI is entirely dependent on the quality of the data it is trained on” — is simple but weighty. In the rush to transition to automated systems, the company had undervalued the tacit knowledge that veteran engineers had accumulated across multiple development cycles. If the nuances captured by humans with years of experience never make it into the training data, the AI cannot learn them.
The failure of judgment in designing how to integrate AI is not a problem unique to manufacturing. Charging ahead with the assumption that “automation equals efficiency” can mean losing the implicit quality-control functions embedded in the existing system alongside the parts you meant to replace. The lesson Ford proved at significant cost applies to AI adoption projects across every industry.
Source: The Verge AI - Ford had to hire back former engineers to fix mistakes made by its automated systems
Claude Managed Agents Adopted by Rakuten — The Race to Own the Fully Managed Agent Execution Layer
Anthropic’s AI agent execution platform “Claude Managed Agents,” currently in beta, has been adopted by Rakuten and several other enterprises. Since its April beta release, Anthropic has shipped features in rapid succession: memory (April), multi-agent orchestration (May), scheduled deployments, and Vault for credential management — reinforcing its identity as an “agent execution infrastructure” well beyond a simple model API.
Building a proprietary agent platform from scratch is expensive work. Controlling the agent loop, implementing the tool execution layer, managing state for long-running tasks, handling security and sandboxing — the cost of building all of this from zero is substantial. Claude Managed Agents delivers it fully managed, enabling companies to focus on what to have the agent do rather than how to make the agent run.
With OpenAI’s Codex already demonstrating company-wide deployment and Google integrating Computer Use into its flagship model, the competition for agent infrastructure is migrating to the platform layer. The debate is shifting from “which model is smartest” to “whose infrastructure will agents run on.” The significance of a large-scale platform company like Rakuten appearing as a reference customer carries more weight than any benchmark score.
The thread connecting today’s stories is the full arrival of an era in which AI moves through the world. Codex deployed across legal departments, Gemini operating computer screens, robots autonomously manufacturing other robots — each is a concrete expression of the same agentic pattern: AI receiving a goal and executing work to complete it. Ford’s story, meanwhile, poses a question that stubbornly refuses to go away: how do you balance the speed of automation against the preservation of human tacit knowledge? As the race for dominance plays out simultaneously across infrastructure, models, and execution platforms, the design judgment of what to automate and what to leave to humans is becoming more important, not less.
Sources
- TechCrunch AI - Amazon ups India bet with fresh $13B AI infrastructure investment
- The Verge AI - Ford had to hire back former engineers to fix mistakes made by its automated systems
- The Verge AI - Facebook’s Creator Studio has been revived as an AI companion app
- OpenAI Blog - How agents are transforming work
- ITmedia AI+ - 男性に美人局容疑で3人逮捕 ChatGPTの示談相場示し脅迫か 警視庁
- ITmedia AI+ - 中国が人型ロボット開発で急成長しているワケ 日本が学ぶべきポイントは? 専門家が解説
- ITmedia AI+ - 「教員を生成AIに置き換える考えはない」東京外大が声明 SNSで拡散した懸念にコメント
- ITmedia AI+ - 白血病など16疾患の診断をAIが支援、日立がAUC0.9以上の新技術
- ITmedia AI+ - シャープブランドのAIサーバ展開も検討 シャープと鴻海、5分野で戦略的協業へ
- ITmedia AI+ - 「AIエージェント基盤の構築は色々大変」 Claude Managed Agentsはどう進化しているのか
- ITmedia AI+ - 「今日言うつもりはなかったが……」 孫正義氏が明かした「ロボット自動量産工場」の実態
- ITmedia AI+ - Google、「Gemini 3.5 Flash」に「Computer Use」を標準搭載──AIが画面を見てブラウザやアプリを操作
- ITmedia AI+ - OpenAI、「GPT-5.5 Instant」をアップデート 会話の文脈維持や箇条書き減など「読みやすさ」を改善
- arXiv - Learning Action Priors for Cross-embodiment Robot Manipulation
- arXiv - RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments
- arXiv - On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity
- arXiv - Real-Time Voice AI Hears but Does Not Listen
- arXiv - Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- arXiv - Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models
- arXiv - A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding
- arXiv - Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
- arXiv - When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance
- arXiv - A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks
- arXiv - The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
- arXiv - When Does Synthetic Data Augmentation Improve Score-Based Imbalanced Classification?
- arXiv - Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining
- arXiv - How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations
- arXiv - AI translation of literary texts is "fine", but readers still prefer human translations