🗾Stalecollected in 83m

NTT developer analyzes the rapid evolution of AI coding

NTT developer analyzes the rapid evolution of AI coding
PostLinkedIn
🗾Read original on ITmedia AI+ (日本)
#llm-evolution#software-engineeringtsuzumi-2ntttsuzumi

💡Learn how NTT's tsuzumi is evolving to master competitive programming and what it means for AI-driven development.

⚡ 30-Second TL;DR

What Changed

AI coding models have reached competitive programming proficiency in just five years.

Why It Matters

Understanding the trajectory of coding-specific LLMs helps developers anticipate the future of automated software engineering tools.

What To Do Next

Analyze your current coding workflow and identify which repetitive tasks can be offloaded to specialized coding LLMs like tsuzumi.

Who should care:Developers & AI Engineers

Key Points

  • AI coding models have reached competitive programming proficiency in just five years.
  • NTT's tsuzumi model is being optimized for advanced coding tasks.
  • The evolution of LLMs is driven by improved training methodologies and architectural refinements.

🧠 Deep Insight

Background and context from public sources — not the original article. 14 sources cited.

🔑 Enhanced Key Takeaways

  • NTT's tsuzumi LLM is designed to be lightweight, with versions as small as 600 million and 7 billion parameters, significantly reducing power consumption and operational costs compared to larger models like OpenAI's GPT-3, and enabling on-premises deployment for enhanced data security.
  • The tsuzumi model demonstrates superior performance in Japanese language processing, outperforming GPT-3.5 and other domestic LLMs in benchmarks such as Rakuda, while also providing support for English.
  • NTT is advancing an 'AI constellation' strategy, which envisions combining multiple smaller, specialized LLMs rather than relying on a single, massive monolithic model, with tsuzumi serving as a customizable node within this architecture.
  • The latest iteration, tsuzumi 2, released in October 2025, is a purely domestic model developed entirely from scratch by NTT, ensuring full control over training data for trustworthiness and security, and is optimized to run efficiently on a single GPU.
  • NTT has developed a novel lossless vocabulary reduction algorithm, presented in April 2026, which facilitates direct collaboration between heterogeneous LLMs with incompatible vocabularies during inference without compromising accuracy, thereby supporting ensemble models and portable-tuning knowledge transfer.

🛠️ Technical Deep Dive

  • Parameter Sizes: tsuzumi is available in lightweight versions with 600 million (0.6B) and 7 billion (7B) parameters. A medium-sized version with 13 billion (13B) parameters is planned.
  • Architecture: Employs an auto-regressive language modeling approach based on the transformer architecture.
  • Training Environment: Utilizes NTT's IOWN All-Photonics Network to connect GPUs and storage across geographically dispersed data centers, ensuring secure LLM learning with minimal performance degradation.
  • Multimodality: Features planned support for multimodal capabilities, including the comprehension of graphical displays, nuances in voice, facial expressions, photos, and diagrams.
  • Language Support: Primarily optimized for Japanese language processing, where it demonstrates high performance, but also supports English, with future plans for broader multilingual capabilities.
  • Tuning: Incorporates adapter tuning for flexible and lightweight fine-tuning, allowing for customization with industry-specific data and facilitating on-premises deployment.
  • Efficiency: The 7 billion parameter version can perform high-speed inference on a single GPU, while the 600 million version can run on a CPU, significantly reducing tuning and inference costs. tsuzumi 2 also maintains efficient operation on a single GPU.
  • Training Data: The pre-training of tsuzumi involved over 1000 billion tokens from diverse domains. tsuzumi 2 was built from scratch using closely controlled training data to ensure reliability.
  • Tokenizer: Integrates the results of NTT's extensive research in Japanese word segmentation to achieve natural language-like segmentation.
  • AI Constellation Concept: A next-generation AI architecture that combines multiple small, specialized LLMs to address social issues, rather than relying on a single, large, monolithic model.
  • Vocabulary Reduction Algorithm: NTT developed a lossless algorithm that enables heterogeneous LLMs with different token vocabularies to collaborate directly during inference by converting predictions into a shared subset without accuracy loss.

🔮 Future ImplicationsAI analysis grounded in cited sources

The adoption of lightweight, specialized LLMs will accelerate enterprise AI integration.
NTT's tsuzumi demonstrates that high performance can be achieved with smaller models, reducing computational costs and enabling on-premises deployment, which addresses critical concerns for businesses regarding data security and operational expenses.
Competitive programming will evolve into a human-AI collaborative discipline.
As AI models achieve grandmaster-level proficiency in coding, the role of human programmers will shift towards prompt engineering, verification, and oversight, leading to new forms of AI-assisted competitions and skill requirements.
Enhanced interoperability between diverse LLMs will become crucial for advanced AI systems.
NTT's development of a lossless vocabulary reduction algorithm highlights the growing need for different LLMs to collaborate effectively, paving the way for more complex multi-agent AI systems like NTT's 'AI constellation.'

Timeline

2018
NTT began constructing a Japanese BERT model following Google's BERT announcement.
2022-11
OpenAI's ChatGPT was released, significantly increasing global attention on LLMs.
2023-02
NTT officially launched its project dedicated to developing LLMs.
2023-06
Pre-training commenced for NTT's tsuzumi LLM.
2023-11-01
NTT publicly announced its proprietary large language model, 'tsuzumi'.
2024-03-25
Commercial services based on NTT's tsuzumi LLM were officially launched.
2025-10-20
NTT released tsuzumi 2, the next-generation version of its LLM.
2026-04-23
NTT presented a lossless vocabulary reduction algorithm at ICLR 2026, enhancing LLM interoperability.

📎 Sources (14)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. medium.com
  2. rd.ntt
  3. ntt-review.jp
  4. rd.ntt
  5. global.ntt
  6. ntt-review.jp
  7. rd.ntt
  8. group.ntt
  9. ibtimes.com
  10. azure.com
  11. researchgate.net
  12. medium.com
  13. our-ai.org
  14. group.ntt
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: ITmedia AI+ (日本)

This is a summary, not the original. Read the source, or get the weekly briefing.

Weekly AI briefing

One email a week. Unsubscribe anytime.