Gynga AI Topics
JA EN

PFN's From-Scratch Domestic AI Strategy: Does the Moat Live in the "Tokenizer"?


TL;DR

  • PFN develops PLaMo from scratch rather than through additional training on existing open models. The aim is accountability and Japanese-language processing cost (stemming from the tokenizer), not a strategy to beat overseas players on general-purpose performance.
  • Anthropic explained that it is implementing a mechanism based on the “SynthID-Text” watermarking technology in Claude as part of its EU AI Act compliance. This is a free regulatory-compliance cost, not a differentiator.
  • Ajinomoto’s official X account posted an AI-edited product image in which the label became distorted. The cause was not the use of AI itself but a lack of pre-publication checks.

Main story: PFN’s From-Scratch Domestic AI Strategy — Does the Moat Live in the Tokenizer?

What happened

Daisuke Tanaka, head of PFN’s LLM development business division, and colleagues explained why the company builds its LLM PLaMo entirely from scratch rather than through additional training on existing open models (ITmedia). The three reasons given are: (1) accountability, since the company controls its own training data and architecture; (2) easier root-cause identification when problems occur; and (3) improved Japanese-language processing performance. At the same time, PFN does not rule out additional training as an option, acknowledging that it has already deployed additional training on open models in fields such as healthcare.

Technical read

Building from scratch lets PFN equip its models with a proprietarily designed tokenizer, which Tanaka says reduces token consumption for Japanese-language processing by roughly 20-30% compared with tokenizers used in overseas models (this is his own account; the article does not cite third-party benchmarks to back it up). Because the company controls the entire training pipeline, it also finds it easier to prepare training data for domain-specific models, such as in finance. He argues that when hallucinations occur, PFN can trace the problem back to specifically where in the dataset it went wrong, contrasting this with additional training on open models, where identifying the root cause can be harder. On the flip side, Tanaka himself acknowledges the weakness: building from scratch is costly and takes longer to develop, and can sometimes struggle to keep pace with the latest overseas open models on performance.

Business read

On unit economics, the key question is the balance between the large upfront investment that building from scratch requires (training cost and time) and the limited cost advantage of a 20-30% reduction in Japanese tokens. That advantage only matters for customers with sufficiently large volumes of Japanese-language processing. On the build-versus-adopt-platform question, Japanese-language tokenizer optimization is a gap that overseas foundation model providers could close through continued improvement, and given the heavy training cost and development time inherent to building from scratch (a weakness Tanaka himself concedes), this edge may narrow over time. As a proprietary asset, PFN’s only real moat is control over the upstream of the model itself — training data and tokenizer design — but the risk of falling behind on performance remains. In terms of market positioning, current customers are concentrated among those committed to domestically built AI or those with on-premises requirements, and penetration among customers who choose purely on performance and price appears still limited. On model dependency risk, PFN notably points to the geopolitics of AI infrastructure, citing the case in June when a U.S. government order temporarily suspended provision of Anthropic’s high-performance models “Fable 5” and “Mythos 5.” This is differentiation through sovereignty risk hedging rather than a performance race, which is a different kind of claim from pure technical superiority.

Contrarian take / blind spots

The simplification that “from-scratch equals superior” is wrong. PFN itself has explicitly stated a hybrid approach — acknowledging both that building from scratch entails cost and that additional training alone is sometimes sufficient — so building from scratch is not a universal solution but should more accurately be understood as a choice suited to specific customers for whom accountability and data sovereignty are requirements.

Implications and positioning

If you’re running Japanese-language-heavy workloads (large-scale document processing, translation, agent operations), it’s worth benchmarking PLaMo 3.0 Prime against leading overseas models on your own workload. That said, avoid making adoption decisions on the grounds that it’s “domestically built” — prioritize actual measured performance and cost data instead. This does offer supporting evidence for Thesis T1 (that the moat lies upstream in the model), but the tokenizer advantage is a time-limited first-mover benefit that could be eroded as foundation model providers continue to improve, so it would be premature to treat it as a permanent moat. Reframing domestic AI’s appeal as “sovereignty risk hedging” rather than “performance” clarifies the axis for investment decisions.

Other notable stories

Anthropic implements invisible watermarking based on “SynthID-Text” in Claude Anthropic explained that, as part of its EU AI Act compliance, it is applying a derivative of Google DeepMind’s open-source watermarking technology “SynthID-Text” to text generated by Claude (The Verge). The mechanism embeds a statistical pattern by using a key and the preceding sequence of words to drive the random number generation involved in choices between words with little difference in meaning (e.g., “cold and overcast” versus “cold and grey”). It’s undetectable to readers, but detectable to anyone holding the key. Technically, this is probabilistic steganography rather than cryptographic proof, so it can be weakened by paraphrasing, translation, or reprocessing through another model. Short text or code generation, which offer fewer low-risk word-choice opportunities, may also allow less information to be embedded. From a business standpoint, as Anthropic itself states there is no added cost and no impact on output quality, this represents a regulatory-compliance baseline rather than a differentiator. Google’s Gemini has had a similar mechanism since 2024, while OpenAI has not disclosed watermarking plans for ChatGPT despite being subject to the same law. So what: if authenticity detection of AI-generated content is a business assumption for you (content moderation, academic-integrity detection, etc.), relying on this watermark alone is risky. You should design around the fact that it’s an optional vendor implementation that can be evaded through paraphrasing; for everyone else, this is regulatory-compliance news that doesn’t require you to change anything.

China’s Unitree releases video of humanoid robot jumping 2m without a running start and running at 45km/h Unitree Robotics released a demo video of a new headless humanoid robot under development. It showed a 2-meter jump with no running start and a top speed of 12.66 meters per second (45.576 km/h), claiming both exceed human world records. Development reportedly began only about three months ago (ITmedia). Technically, this is a single demo video released via the official account, with no information on task generality, payload, repeatability, or number of attempts. Showcasing peak performance follows the same pattern as one-off LLM benchmark stunts and should be viewed as distinct from task-completion rates in real-world operation. From a business standpoint, Unitree is a company that already mass-produces and sells humanoid robots, so this reads more like a showcase for its product pipeline than a research announcement. The numbers race on physical performance also reflects that the market is still at an immature stage where it’s discussed in terms of performance comparisons. So what: this has little direct impact on pure software/AI founders, but if you’re involved in software or integration layers around robotics, it’s worth keeping in mind that the faster hardware performance commoditizes, the more value shifts toward integration, data, and field operations — the same pattern as Thesis T1.

OpenAI funds 14 independent projects exploring AI policy OpenAI announced it will fund 14 independent projects exploring policy ideas, framed around expanding economic opportunity and strengthening societal resilience in the “Intelligence Age” (OpenAI Blog). At this point that single statement is all the information available — no specifics such as project names, funding amounts, or evaluation criteria have been disclosed. So what: this is a PR-style announcement without substance — noise rather than signal. It won’t inform investment decisions until the budget size and selected projects become known. Safe to ignore.

Ajinomoto apologizes after AI-edited image on official X account distorted product label Ajinomoto apologized on August 16 after a recipe image using its “Hondashi” product, posted on August 14 to its official X account “Ajinomoto Park,” was found to have a distorted product label and text. The company explained that it had used AI to adjust the brightness and background of an actual photograph, and published the result with the distorted text and label still intact; it also attached the original, unedited photo (ITmedia). Technically, distortion of fine logos and text is a known weakness class even for AI image edits like background and brightness adjustment, so this isn’t surprising. From a business standpoint, the cause Ajinomoto itself acknowledged was not the use of AI per se but a lack of pre-publication checks, and the article lists similar past incidents at the end (Sakura Color Products, a Nara city council member, science media outlets, etc.). This reads as a structural pattern among Japanese companies — a lack of review processes for AI output — rather than a defect specific to one vendor. So what: if your business handles AI image editing for your own products, you should formalize a checklist of elements that are critical if distorted (logos, labels, legally required labeling, etc.) and keep the original, unedited photos on file. This isn’t a reason to avoid AI editing itself, but without a review step, it leads directly to brand damage.

Worth trying this week / Hype to ignore

Worth trying: If Japanese-language processing cost is a bottleneck in your operations, it’s worth benchmarking PLaMo 3.0 Prime against leading overseas models on your own workload for token count and actual cost. Hype to ignore: The simplification that “from-scratch domestic AI equals beating overseas players on performance”; evaluating Unitree’s peak-performance demo in isolation; OpenAI’s announcement of 14 policy funding projects (while still lacking specifics); and treating Claude’s watermark rollout itself as significant news.

Sources