Open-Source Artificial Intelligence
Overview
2025 marked a watershed moment for artificial intelligence. Open-source large language models matched or surpassed proprietary models across a growing range of benchmarks, while reshaping how AI systems are developed and deployed, how computing resources are priced and consumed, and how commercial and licensing models are structured.
This report provides a systematic review of China’s open-source AI ecosystem in 2025. It covers advances in computing infrastructure and large-scale models, the transition of AI agents from experimentation to large-scale deployment, and the evolution of embodied intelligence from simulation to real-world testing. It also examines the institutionalization of AI ethics, safety, and governance, the growing role of open-source collaboration in global open science, and the expanding adoption of open-source AI across industry verticals.
Overall, China’s open-source AI ecosystem in 2025 can be characterized by three major trends. First, the competitive landscape between open-source and proprietary AI models has been fundamentally reshaped, with open-source models no longer merely playing catch-up. Second, technological innovation has begun to move from an “aesthetics of brute force” toward “fine-grained engineering,” with efficiency and intelligence density emerging as key dimensions of competition. Third, AI governance has undergone a fundamental transformation, evolving from voluntary ethical initiatives toward a binding national regulatory framework, while security, compliance, and sovereignty have become non-negotiable requirements.
Taken together, these developments reveal a distinctive development path for China’s open-source AI community in the context of global AI competition—one characterized by a compliance-first approach, legal and judicial safeguards, a sovereignty-oriented framework, and technological self-reliance.
1. Overview of AI Foundation Models
1.1 Redefining the Landscape: Paradigm Shifts and Strategic Realignment in the 2025 Open-Source AI Ecosystem
2025 marked a watershed moment in the history of artificial intelligence. Open-source large language models (LLMs) matched or surpassed leading proprietary models across an expanding range of benchmarks, while reshaping the economics of AI development and deployment, computing, and commercial and licensing models.
For years, the industry had been preoccupied with the “open source versus closed source” debate. By 2025 and early 2026, however, the debate had moved beyond a simple comparison of technical performance and business models. It had evolved into a deeper strategic question: “Who does openness serve?”
The summer of 2025 marked a historic turning point in the global distribution of large-scale model downloads and the evolution of the open-source AI ecosystem. Data from relevant open-source projects show a significant shift in the balance of downloads toward Chinese models. Among models with 1B+ parameters, Meta led with 23.2% of downloads, followed by Alibaba’s Qwen series at 20%, Mistral at 6.8%, and DeepSeek at 3.8%. More importantly, 15.6% of all downloads came from quantized versions developed and distributed by the broader open-source developer community rather than from the original model releases. This is more than a change in download patterns: it signals a broader shift in where value and influence are created within the open-source model ecosystem—from a small number of major model providers toward a decentralized network of developers. The growing adoption of customized and localized models further underscores the increasingly distributed nature of open-source AI development.
In this new multipolar ecosystem, four major players are pursuing distinct technological approaches and business strategies, each helping to shape the evolution of the industry.
Meta uses open-weight models as a “coordination tool,” expanding the global reach of its platform by helping establish industry standards and toolchain conventions. Llama surpassed the milestone of 1 billion downloads in March 2025, strengthening its position as general-purpose AI infrastructure and a potential foundation for next-generation applications. At the same time, however, the Llama 4 Community License has raised questions about whether the model meets the Open Source Initiative’s definition of open source.
DeepSeek, by contrast, has expanded access to advanced reasoning capabilities that were previously available primarily through proprietary models. Its aggressive cost optimization and innovations in reinforcement learning have narrowed the cost gap associated with advanced reasoning, helping to reset market expectations for the price of high-level reasoning.
Mistral has focused on the needs of highly regulated markets in Europe and elsewhere, offering enterprise customers a value proposition centered on trust and digital sovereignty through its permissive Apache 2.0 licensing model. This positioning has also addressed growing geopolitical concerns over technology restrictions.
Alibaba’s Qwen series, meanwhile, has emerged as a powerful “distribution engine,” combining strong multilingual capabilities with a broad range of model sizes, from 0.5B to 235B parameters. Its models have rapidly expanded across cloud services and edge-computing environments worldwide.
The growing strength of the open-source ecosystem has even prompted OpenAI, which had long relied on a closed-source strategy, to release the GPT-OSS-120B and GPT-OSS-20B models under the Apache 2.0 license in August 2025. This move was not merely a defensive response to growing market pressure; it also represented a reluctant acknowledgment of the reality that open source is becoming a foundational layer for the next generation of AI.
1.1.1 Inflection Points: A 2025 Timeline of Landmark Models and Breakthrough Capabilities
The technological breakthroughs of 2025 were notable for their frequency and breadth, with major advances emerging across multiple fronts. These breakthroughs have not only pushed benchmark performance to new heights but also fundamentally reshaped enterprise AI procurement priorities.
1.1.2 January–February: DeepSeek’s Reasoning Revolution and the Collapse in Inference Costs
In January 2025, DeepSeek released DeepSeek-V3 and its reasoning model, DeepSeek-R1, sending shockwaves through the global technology industry and prompting a reassessment of the economics of AI. The release also led capital markets to reassess the earnings outlook for AI companies, contributing to nearly $1 trillion in swings in technology-stock valuations. DeepSeek’s disruptive impact lies not only in its strong performance, but also in its challenge to the industry’s traditional scaling paradigm—the heavy reliance on scaling up model size, data, and compute to achieve better performance.
Its core technical innovations include the Multi-Head Latent Attention (MLA) mechanism, which substantially improves memory efficiency; a highly optimized Mixture-of-Experts (MoE) architecture; and FP8 mixed-precision training. These techniques help strike a favorable balance between computational cost and communication overhead while allowing the models to extract greater performance from limited hardware resources.
More importantly, DeepSeek-R1 represents a fundamental shift away from the traditional reliance on costly human-annotated data for supervised fine-tuning (SFT). It demonstrates that, through large-scale reinforcement learning (RL) alone, models can develop emergent capabilities such as long-chain reasoning, self-reflection, and self-verification in domains such as mathematics and programming, helping overcome a major bottleneck in the acquisition of high-quality reasoning data.
More importantly, DeepSeek-R1 marks a fundamental shift in the way reasoning models are trained, challenging the traditional reliance on costly human-annotated data for supervised fine-tuning (SFT). It demonstrates that large-scale reinforcement learning (RL) alone can enable models to develop emergent capabilities such as long-chain reasoning, self-reflection, and self-verification in domains such as mathematics and programming, helping overcome a major bottleneck in acquiring high-quality reasoning data.
From a cost-efficiency perspective, DeepSeek-V3 demonstrates a significant competitive advantage. With 671 billion parameters, the model was trained using only 2.788 million H800 GPU hours, with total training costs estimated at approximately $5.5 million. The reinforcement learning phase of DeepSeek-R1 required only an additional $294,000. These figures are orders of magnitude lower than the nearly $100 million training costs incurred by major Silicon Valley AI companies.
DeepSeek further expanded this approach by releasing a series of open-source distilled reasoning models ranging from 1.5B to 70B parameters, built on Llama and Qwen architectures through knowledge distillation. Benchmark evaluations demonstrated that the 14B-parameter distilled model could outperform the larger QwQ-32B model in several mathematical reasoning and coding tasks. This breakthrough showed that the capabilities of large-scale reasoning models could be effectively compressed and transferred to smaller models capable of running on consumer GPUs, opening the door to practical on-device reasoning.
1.1.3 April: The Llama 4 Debut and Scaling to New Extremes
Following initial architectural adjustments and preparations, Meta launched the Llama 4 series in April 2025, raising the bar for open-source LLMs in native multimodal capabilities and long-context understanding. The new series marks a fundamental shift away from the dense architectures of previous generations toward a highly optimized hybrid Mixture-of-Experts (MoE) architecture.
The core model lineup is designed to address a wide range of production workloads. Llama 4 Scout features 109 billion total parameters, with 17 billion activated per token across 16 experts. Its most notable engineering achievement is its 10-million-token context window, which enables the model to process hundreds of legal contracts or very large codebases within a single context.
Llama 4 Maverick features 400 billion total parameters, with 17 billion activated per token across 128 experts, together with a 1-million-token context window. This combination enables strong general-purpose conversational performance and sophisticated multimodal capabilities.
At the high end of the lineup, Llama 4 Behemoth scales to 288 billion active parameters. Early benchmark results showed competitive performance against GPT-4.5 and Claude 3.7 Sonnet across a range of STEM tasks, underscoring Meta's continued expertise in large-scale model development and architectural scaling.
1.1.4 July–August: Proprietary Incumbents Hold the Line as Trillion-Parameter Open-Source Models Arrive
In the third quarter of 2025, the growing strength of open source broke through many of the competitive barriers long held by major proprietary AI companies, while also giving rise to the first homegrown trillion-parameter model.
In July, Moonshot AI released Kimi K2, ushering the open-source ecosystem into the trillion-parameter era. Built on a highly optimized sparse Mixture-of-Experts (MoE) architecture, Kimi K2 activates only about 32 billion parameters per inference despite scaling to 1 trillion total parameters. The breakthrough goes beyond raw parameter scale, demonstrating the engineering feasibility of ultra-long-context processing and laying a strong foundation for handling massive enterprise document workloads and building sophisticated RAG knowledge bases.
At the same time, the growing competitive pressure from open source has forced major proprietary AI companies to reconsider their long-standing strategies. In August, OpenAI, which had historically kept its frontier models closed source, released the open-source GPT-OSS-120B and GPT-OSS-20B reasoning models under the Apache 2.0 license. This marks a fundamental shift in the competitive dynamics of the large-model market: open source is no longer merely a strategy for challengers seeking to close the gap, but an increasingly important infrastructure layer that all major players must engage with. OpenAI's strategic shift can be understood as a defensive response to the growing migration of enterprise customers toward cost-effective private deployments of open-source models.
1.1.5 October–November: Native Multimodality Meets Production-Grade Agentic AI
By the fourth quarter of 2025, benchmark scores alone were no longer sufficient to capture the industry's growing demand for practical AI capabilities. The center of gravity had shifted from optimizing low-level compute efficiency to improving higher-level task execution, while agentic models with multimodal capabilities and tool-use functionality emerged as a major new area of competition.
In October, the release of Qwen3-Omni marked one of the most ambitious advances in open-source multimodal AI. Moving beyond the architectural compromises of earlier approaches, the model natively processes audio, video, and text within a unified architecture rather than relying on separately integrated vision encoders. The result is production-grade multimodal performance for demanding enterprise use cases, including complex financial-document parsing and long-form video understanding, while reducing reliance on fragile multi-stage pipelines.
Launched in early November, Kimi K2 Thinking marked a major step forward by combining advanced reasoning with external tool use. As one of China's first mature open-source agentic models, it can autonomously plan and execute hundreds of API calls within a 256K-token context window. More importantly, its emergence signals a fundamental transformation in the role of open-source models: from passive systems that respond to prompts to autonomous agents capable of executing complex business workflows. This shift lays the technical foundation for the transition from prompt engineering to context engineering explored later in this report.
1.2 The Rise of Domain-Specific SLMs and Unified Multimodal Architectures
Not every enterprise workload in 2025 required the power of a trillion-parameter model. Small language models (SLMs) increasingly came into their own in edge and on-device applications, where high-frequency requests and low-latency response times are critical.
Microsoft's Phi-4 series has emerged as a strong contender in this space. In the first quarter of 2025, Microsoft introduced the 14B text-only Phi-4, the lightweight 3.8B Phi-4-mini, and the multimodal Phi-4-multimodal, which natively processes text, images, and audio.
In the open-source vision-language model (VLM) space, Qwen2.5-VL has emerged as a strong competitor to Janus-Pro. The model can natively process videos longer than one hour, combining high-level video understanding with precise visual grounding. It can identify target objects, return their bounding-box coordinates, and generate structured JSON outputs containing object attributes. These capabilities make Qwen2.5-VL well suited to demanding applications such as financial-document parsing and industrial computer vision.
However, real-world experience has shown that Qwen2.5-VL still faces challenges in sustaining multimodal inference at the million-token scale using its native capabilities alone. As context lengths approach one million tokens, maintaining reliable performance may require additional architectural components, cascaded model designs, or dedicated long-context variants.
1.3 From “Brute-Force Scaling” to “Precision Engineering”: A Generational Leap in Foundation Models
As open-source models proliferated rapidly in 2025, the industry's optimization priorities began to shift from brute-force parameter scaling toward more efficient memory utilization and bandwidth optimization, driven by growing constraints on compute resources and energy consumption.
The Rise of MoE: Hybrid Mixture-of-Experts (MoE) architectures have become increasingly dominant in large-scale models, while dense architectures have largely fallen out of favor at the hundreds-of-billions to trillion-parameter scale. Models such as DeepSeek-V3 and Llama 4 illustrate this architectural shift. The fundamental advantage of MoE is its ability to separate total model capacity from inference-time compute: a model can scale to hundreds of billions or even trillions of parameters while activating only a relatively small subset of them for each inference.
Auxiliary-loss-free load balancing (DeepSeek-V3): DeepSeek-V3 dynamically balances expert utilization without relying on an auxiliary loss, avoiding the potential performance degradation associated with additional loss terms. The approach significantly reduces cross-node communication overhead and enables near-complete overlap between computation and communication.
Multi-Head Latent Attention (MLA) tackles the memory bottleneck of long-context inference: MLA compresses key-value representations into a low-dimensional latent space, substantially reducing the memory footprint of the KV cache. This design can cut inference-time memory usage by roughly an order of magnitude while preserving model performance, enabling a single machine to handle high-concurrency workloads with long context windows.
1.4 The Million-Token Trap: Compute Walls and the “Lost in the Middle” Dilemma
Mainstream models such as Llama 4 Scout with its 10M-token context window, Gemini 2.5 Pro with 2M tokens, and Qwen2.5-1M are pushing the industry into the era of ultra-long context. Yet extending context windows to this scale comes with substantial infrastructure costs, including rapidly increasing memory and compute requirements, as well as persistent challenges in long-context reasoning and information retrieval.
- The Compute and Memory Wall (Infrastructure Bottlenecks)
The standard Transformer attention mechanism has quadratic time complexity, making million-token context windows extremely demanding in terms of computation and memory. Processing 1 million tokens can impose a massive computational burden, while the KV cache alone can require approximately 15 GB of GPU memory for a single user.
Solution: The industry has adopted sequence and context parallelism, achieving 93% parallel efficiency across 128 H100 GPUs.
Key enabling technologies: Zigzag Ring Attention, NTK-aware RoPE positional encoding, and NVFP4 quantization, which can reduce memory requirements by roughly half.
- “Lost in the Middle” and Runaway Verbosity (Model Limitations)
U-Shaped Retrieval Performance: Benchmarks such as HELMET, MMLongBench, and Lara reveal a persistent U-shaped pattern in long-context retrieval performance. Models perform well at the beginning and end of a context, achieving recall rates of 85%–95%, but performance drops sharply to 76%–82% in the middle.
Declining Retrieval Quality Under Noise: As noise increases, key retrieval metrics—including NDCG (Normalized Discounted Cumulative Gain), MAP (Mean Average Precision), and MRR (Mean Reciprocal Rank)—decline consistently, indicating a systematic deterioration in retrieval quality.
Performance Degradation in Long Conversations: Longer responses can further exacerbate these limitations. In multi-turn tasks, models achieve an average performance of only 35.6%, compared with 40.7% for shorter responses, as early incorrect assumptions can propagate through subsequent turns and create strong path dependence.
- The Hard Realities of Enterprise Deployment
The “10-million-token context window” is often more of a marketing ceiling than a practical one. For inputs exceeding 200,000 tokens, such as an entire codebase or a lengthy legal case, processing the full input within a single context window can be extremely expensive, with prefill latency exceeding two minutes, and may increase the risk of hallucinations.
The Practical Approach: For inputs exceeding 200,000 tokens, combining smart slicing with advanced RAG techniques can deliver greater accuracy and stability than relying on a single, extremely long context window.
Multimodal Architecture: Native Fusion vs. External Encoder Integration: By 2025, two major architectural approaches had emerged for multimodal understanding and cross-modal generation, with significant implications for enterprise adoption:
Approach 1: Early Fusion (e.g. Llama 4maverick, Gemma 3)
Architectural Features: Early pre-training architectures mapped text, image, and video tokens into a unified representation and processed them through a shared backbone network.
Key Advantage: Joint visual-text reasoning is the dominant approach for multimodal reasoning.
Benchmark Results: Llama 4 scores 94.4% on DocVQA, 90.0% on ChartQA, and 73.4% on MMMU, covering visual question answering, chart understanding, and broader multimodal reasoning capabilities.
Approach 2: Text-Centric Backbone + External Vision Encoder (e.g., DeepSeek-V3 / R1 / VL Ecosystem)
Architectural Features: The architecture is primarily optimized for logical reasoning and code generation, with compute and VRAM resources focused on these workloads. For multimodal tasks, a lightweight external vision encoder is integrated to encode visual inputs into a compatible feature space.
Core Strength: Strong Logical Reasoning: Built on MLA and DeepSeekMoE, the model achieved a resolution rate of around 49% on SWE-bench, approaching—and in some cases even surpassing—the performance of the closed-source flagship model Claude 3.5 Sonnet.
1.5 Paradigm Shift in Production: Context Engineering Takes Center Stage
As AI evolves from one-off conversations to long-lived enterprise workflows, the way developers build and maintain AI systems changes fundamentally.
- The Demise of Prompt Engineering
Static, complex prompts can lead to context rot as system state accumulates, API call logs grow, and conversation history expands. Without temporal reasoning, the model may struggle to distinguish meaningful state changes from accumulated context.
- The Rise of Context Engineering
The Shift from Prompt Engineering to Context Engineering: The focus is shifting from how to phrase a prompt to how to structure the system and precisely control the information presented to the model.
Dynamic Memory Compaction and Structured Notes: Small, low-cost models (e.g., 8B models) distill long conversation histories and error logs into structured fact tables, reducing context overhead while preserving information relevant to the main model.
Multi-Level Context Pyramid: The architecture avoids brute-force context assembly through a three-layer context hierarchy. The bottom layer provides persistent knowledge and compliance policies. The middle layer maintains dynamic workflow memory and relevant examples. The top layer handles real-time user intent and tool outputs.
Model Context Protocol (MCP): MCP dynamically loads and unloads business interfaces and environment-specific information on demand, keeping irrelevant information out of the model’s context window and maintaining a focused context.
1.6 Context Mapping and Reranking in RAG: Mitigating the “Lost in the Middle” Dilemma
In the face of steep long-context costs, Retrieval-Augmented Generation (RAG) has not become obsolete; rather, it has undergone a profound evolution.
From Chunking to Structured Retrieval: TreeRAG and GraphRAG are emerging as more sophisticated alternatives to brute-force chunking. Instead of relying solely on sliding-window segmentation, these approaches use LLMs to extract entities and relationships from unstructured documents and construct structured knowledge representations. Graph queries can then bridge semantic gaps and support complex reasoning across multiple documents.
Passage Reordering: To leverage the LLM's "Lost-in-the-Middle" U-shaped attention pattern, the highest-ranked relevant segment is placed at the beginning and end of the context, while lower-ranked segments are positioned in the middle before the context is passed to the main model. This plug-and-play strategy aligns the context layout with the model's attention distribution, improving retrieval accuracy with negligible additional computational overhead.
2. AI Infrastructure
2.1 Model Training: Landscape and Current State
After years of development, large language model training has entered a new stage characterized by large-scale deployment, industrialization, and refinement. As the race to scale becomes more rational, training efficiency and model capability density are becoming increasingly important areas of focus. Following the shift from hundreds of billions of parameters in earlier models such as GPT-3 toward much larger models, simply scaling up parameter counts is no longer the only path forward. Instead, the industry is increasingly focused on improving model capability through advances in algorithms, data, and architectural design while keeping computational costs under control.
Another noteworthy development is the evolution of the training paradigm. The traditional “Pre-training → Fine-Tuning” pipeline has evolved into a more comprehensive process involving “Pre-training → Supervised Fine-Tuning → Reinforcement Learning from Human Feedback → Post-Training Alignment.” Alignment has become a key factor in determining whether a model is useful, harmless, and honest.
The refinement of large language models is also reflected in the growing importance of multimodality. As text-only models increasingly approach the limits of purely text-based learning, multimodal pre-training has emerged as a major direction. Current approaches, including those used in models such as GPT-4V, Gemini, and Claude 3.5, jointly train on text, images, audio, video, and potentially sensor data to build increasingly general-purpose models that more closely reflect how humans perceive and interact with the world.
2.1.1 Architecture and Core Fundamentals
The traditional LLM training pipeline typically consists of two main stages: pre-training and fine-tuning.
Model Pre-train can be broadly divided into three categories:
Language Modeling (LM) (e.g., GPT-3): Language modeling is one of the most fundamental pre-training objectives. The model learns by predicting the next token from the preceding context, allowing it to learn statistical patterns in large-scale text corpora.
Denoising Autoencoding (DAE) (e.g., BERT, T5): Parts of the input sequence are masked or otherwise corrupted, and the model is trained to reconstruct the original text. The specific corruption strategy varies across models; for example, BERT uses token masking, while T5 uses span corruption.
Mixture-of-Denoisers (MoD) (e.g., UL2): MoD is a unified framework proposed by Google that combines multiple denoising objectives and training configurations to support diverse downstream tasks.
Fine-Tuning and Downstream Adaptation
Instruction Fine-Tuning (IFT) is a method for training models to follow natural-language instructions. By training on instruction-response pairs, the model learns to generalize across tasks such as question answering and translation.
Representative techniques and subsequent alignment methods include:
LoRA (Low-Rank Adaptation): LoRA is a parameter-efficient fine-tuning (PEFT) method. Instead of updating all pre-trained model weights, it freezes the original weights and introduces trainable low-rank matrices. This substantially reduces the number of trainable parameters and the GPU memory footprint while preserving model performance.
ZeRO (Zero Redundancy Optimizer): ZeRO is a memory optimization technique designed for large-scale distributed model training. It reduces memory redundancy by partitioning model states—including optimizer states, gradients, and model parameters—across devices, helping make the training of very large models more feasible.
Human Alignment refers to the process of steering model behavior toward human values, intentions, and preferences. This is often summarized as producing outputs that are helpful, honest, and harmless (3H).
Representative alignment techniques include:
RLHF (Reinforcement Learning from Human Feedback): RLHF uses human preference data and reinforcement learning algorithms, such as PPO, to steer model behavior toward preferred responses.
DPO (Direct Preference Optimization): DPO is an alignment method that directly optimizes the model using preference pairs without explicitly training a separate reward model or running a conventional reinforcement-learning loop.
2.1.2 The Open-Source AI Ecosystem
- Versatile Distributed Training Frameworks
Production-grade frameworks that support the pretraining and full-parameter fine-tuning of models with hundreds of billions of parameters. Representative open-source projects include:Production-grade frameworks for pre-training and full-parameter fine-tuning of models with 100B+ parameters. Notable open-source projects include:
Megatron-LM (Developed by: NVIDIA | Stars: 14.4k) is designed for highly efficient, large-scale training of Transformer models. Its core architecture combines Tensor Parallelism (TP), Pipeline Parallelism (PP), and Sequence Parallelism (SP), with support for ZeRO-based optimization. It is widely used as a foundation for large-scale model training at major technology companies.
DeepSpeed (Developed by: Microsoft | Stars: 40.9k) is a deep learning optimization library centered on the ZeRO (Zero Redundancy Optimizer) family. It significantly reduces the memory required for model parameters, gradients, and optimizer states, enabling users to train larger models with fewer GPUs. It also provides out-of-the-box integration with PyTorch and Megatron-LM.
Colossal-AI (Developed by: HPC-AI Tech | Stars: 41.3k) is an integrated parallel training system that supports a range of distributed training strategies, including heterogeneous memory management. It provides out-of-the-box support for auto-parallelism and parameter-efficient fine-tuning methods such as LoRA. The framework aims to lower the barrier to entry for large-scale model training.
- Parameter-Efficient Fine-Tuning (PEFT) Frameworks & Libraries
These libraries are designed to adapt pre-trained foundation models to downstream tasks while minimizing computational and memory overhead:
PEFT (Developed by: Hugging Face | Stars: 20.2k) is a comprehensive library for parameter-efficient adaptation, with built-in support for popular methods such as LoRA, Prefix Tuning, P-Tuning, AdaLoRA, and (IA)³. It integrates with Hugging Face's Transformers ecosystem and is widely used for resource-efficient model fine-tuning.
TRL (Transformer Reinforcement Learning) (Developed by: Hugging Face | Stars: 16.5k) provides tools for training and post-training Transformer models using reinforcement learning and preference optimization. It is widely used to implement methods such as RLHF and DPO and natively supports PEFT techniques such as LoRA, making it well suited to post-training alignment and the optimization of model behavior toward human preferences.
2.1.3 Challenges and Future Directions
Training large language models presents several challenges, particularly the high computational cost, the quality and diversity of training data, and issues related to fairness and bias. Computational efficiency remains a major bottleneck for training very large models, driving ongoing research into more efficient distributed training and quantization techniques.
Challenges:
High Resource Requirements: Training cutting-edge models can cost tens or even hundreds of millions of dollars in computing resources. It also requires highly specialized engineering expertise, creating a significant barrier to entry.
Data Bottlenecks and Copyright Concerns: High-quality, clean text data is becoming increasingly scarce. Multimodal data is more abundant, but annotation and cross-modal alignment can be costly. Meanwhile, copyright concerns surrounding training data have become increasingly prominent, creating additional legal and compliance risks.
The Complexity of Alignment: A fundamental challenge is defining what alignment should mean across different contexts. Values and preferences can differ across cultures and groups. There is also often a gap between what models say and what they reliably do, while model hallucinations remain an unresolved challenge.
Limitations of Current Evaluation: How to comprehensively and objectively evaluate LLM capabilities remains an open question, particularly for reasoning, planning, safety, and factuality.
Energy Consumption and Social Responsibility: The substantial carbon footprint of large-scale model training has raised broader concerns about AI sustainability and the environmental impact of AI development.
Future Directions:
Multimodality Becomes Standard: Future foundation models are likely to support multiple modalities, including text, images, audio, and video, while seamlessly switching between them.
From Passive Generation to Active Reasoning and Planning: Models are likely to evolve toward agentic systems with stronger multi-step reasoning and long-term planning capabilities, enabling them to call tools, execute tasks, and interact with the real world.
Continued Architectural Innovation: New architectures beyond Transformers, including state-space models such as Mamba, are being explored to improve long-sequence processing and inference efficiency. Sparse architectures such as Mixture-of-Experts (MoE) are also likely to continue evolving.
Data Synthesis and Self-Evolution: Using LLMs to generate high-quality synthetic training data and enabling models to self-critique and improve may help address the data bottleneck.
Miniaturization, Specialization, and Edge Deployment: Lightweight models optimized for specific use cases are likely to become increasingly common and to be deployed on edge devices such as smartphones, vehicles, and robots, enabling offline operation and low-latency AI inference.
Deeper Integration of Reinforcement Learning and Curriculum Learning: Future training processes may increasingly resemble aspects of human learning, enabling models to progressively acquire complex skills through carefully designed curricula and interactive reinforcement learning.
The development of large language models has evolved from academic research into large-scale engineering practice, with the focus shifting from a “scale-first” mindset toward greater efficiency and intelligence per unit of compute. The growth of the open-source ecosystem has lowered barriers to adoption, but the challenges and costs of frontier model development remain substantial. Looking ahead, competition is likely to increasingly center on advances across multiple dimensions, including multimodal understanding, advanced reasoning, efficient model architectures, and responsible alignment.
2.2 Model Serving: Landscape and Current State
In the past, model development placed greater emphasis on training efficiency, resource utilization, and model optimization, while inference was primarily viewed as a deployment and serving concern. As the cost of training large models continues to rise, inference is now subject to increasingly demanding requirements for latency, throughput, and cost efficiency. This is driving a shift from brute-force scaling toward more fine-grained inference optimization, including attention mechanism improvements, quantization, and model compression. Meanwhile, decoupled architectures, multimodal reasoning, edge computing, and real-time inference are emerging as increasingly important areas of development.
2.2.1 Commercial Inference Services:
Closed-source services: OpenAI API, Anthropic Claude, and Google Gemini offer paid API access.
Hosted open-source models: Platforms such as Hugging Face Inference API, Replicate, and RunPod provide hosted inference services for open-source models.
Hardware acceleration: NVIDIA H100/A100, AMD MI300X, and TPU v4 are widely used to accelerate model inference.
Edge devices: Apple M-series chips with MLX, Intel CPUs with BigDL-LLM, and RISC-V platforms supported by projects such as llama.cpp enable inference on edge devices.
Dedicated accelerators: Specialized chips such as FPGAs and ASICs provide high-performance, low-power inference. Dedicated inference processors, such as Groq's LPU, are also emerging as an alternative to general-purpose GPUs.
Heterogeneous hardware optimization: Inference workloads increasingly require coordinated optimization across GPUs, TPUs, NPUs, and CPUs.
A Unified Abstraction Layer for Cross-Vendor Hardware
Inference Optimization as a New Competitive Frontier: The focus of AI infrastructure is gradually shifting from the traditional "training race" toward a broader competition in inference efficiency, performance, and cost optimization.
The Rise of High-Efficiency Inference Frameworks
Continuous batching: Technologies such as vLLM and TGI significantly improve inference throughput by dynamically batching incoming requests.
Memory optimization: Techniques such as PagedAttention in vLLM and FlashAttention-2 reduce memory usage and improve inference efficiency.
Model compression and quantization: 4-bit and 8-bit quantization methods, such as GPTQ and AWQ, can significantly reduce memory requirements and make models with tens of billions of parameters more practical to run on a single GPU.
MoE inference optimization: Mixture-of-Experts models such as Mixtral 8x7B activate only a subset of experts for each input, reducing the amount of computation required for each inference step.
Training–Inference Co-Optimization: Parameter-efficient fine-tuning (PEFT) techniques can also be integrated into inference workflows. For example, LoRA adapters can be dynamically loaded to adapt a base model to different tasks without modifying the underlying model weights.
2.2.2 Technical Architecture and Core Concepts
Approaches to Inference Optimization
Parameter Level Model Compression: Quantization, Pruning, Knowledge Distillation
Algorithm Level : Parameter Usage Reduction, Maximizing Decoding Tokens
System Level Operator Fusion: Memory Management, Workload Offloading,Parallel Serving
Hardware Level:
Operator Fusion: Combining multiple operations into a single kernel to reduce kernel-launch overhead.
Memory Layout Optimization: Organizing data layouts to better match the characteristics of the target hardware.
Pipeline Parallelism: Executing different model layers in a pipelined manner to improve hardware utilization.
2.2.3 Open-Source Ecosystem
a. High-Performance Inference Serving Frameworks
vLLM(Project: vLLM, GitHub stars: 64.4K): vLLM is a high-performance inference and serving engine that uses PagedAttention to improve throughput and memory efficiency. It supports features such as continuous batching and KV-cache optimization and is compatible with models from the Hugging Face ecosystem.
Use cases: High-concurrency production workloads and distributed inference across multiple GPUs.
Text Generation Inference (TGI)(Project: Hugging Face, GitHub stars: 10.7K): Text Generation Inference is Hugging Face's production-oriented inference serving framework. It supports continuous batching, token streaming, and tensor parallelism, as well as features such as FlashAttention-2 and PEFT adapters.
Use cases: Enterprise-grade model deployment and production API serving.
b. Inference Optimization Engines
TensorRT-LLM (Project: NVIDIA, GitHub stars: 12.3K): TensorRT-LLM is NVIDIA's inference optimization framework for large language models. It supports TensorRT-based quantization, in-flight batching, and multi-GPU and multi-node inference.
Use cases: Performance optimization for LLM inference on NVIDIA GPU infrastructure.
SGLang (Project: sgl-project, GitHub stars: 21K): SGLang is a high-performance serving framework for large language and multimodal models. It is designed to efficiently orchestrate and execute complex LLM workloads, with support for low-latency, high-throughput inference and multi-GPU parallelism.
Use cases: Large-scale deployment of complex prompt workflows, multi-turn conversational services, benchmarking, and load testing.
c. Lightweight Inference Tools
Ollama (Project: Ollama, GitHub stars: 157K): Ollama is a lightweight and extensible framework that simplifies the process of running, managing, and deploying large language models locally.
Use cases: Local application development, rapid prototyping, and resource-constrained environments.
llama.cpp (Project: ggml-org, GitHub stars: 90.8K): llama.cpp is a lightweight C/C++ inference framework that supports CPU inference, GPU acceleration, and low-bit quantization, including 4-bit quantization using the GGUF format.
Use cases: Edge devices, CPU-based inference, and resource-efficient deployments.
d. Inference Server Frameworks
FastChat (Project: lm-sys, itHub stars: 39.3K): FastChat provides OpenAI-compatible APIs for model inference and serving, along with support for multi-model management and a web-based user interface.
ONNX Runtime(Project: Microsoft, GitHub stars: 18.6K): ONNX Runtime provides cross-platform inference for models in the ONNX format and includes optimization tools for large language models, such as attention-layer fusion.
Use cases: Unified deployment across heterogeneous hardware, including CPUs, GPUs, and mobile devices.
OpenVINO(Project: OpenVINO, GitHub stars: 9.3K): OpenVINO is Intel's inference optimization toolkit, supporting CPUs, GPUs, and edge devices while providing optimization pipelines for LLM inference.
Use cases: High-performance inference on Intel hardware.
Selection Guide
For maximum inference throughput: vLLM, TensorRT-LLM, SGLang
For production-grade API serving: TGI, FastChat
For edge and CPU deployment: llama.cpp
For cross-hardware deployment: ONNX Runtime
For rapid prototyping and local development: Ollama, LocalAI
2.2.4 Challenges and Future Directions
Large language model inference currently faces several major challenges:
Memory and KV-Cache Bottlenecks: Larger models and longer context windows place increasing pressure on GPU and system memory. Efficient KV-cache management and memory allocation across GPU memory, system memory, and storage remain challenging.
Latency–Throughput Trade-offs: Batching and scheduling techniques can improve throughput but may increase tail latency. Real-time interactive applications require low and predictable tail latency without sacrificing throughput.
Complexity of Dynamic Workload Scheduling: Input lengths, request priorities, and model routing can vary significantly, making it difficult for automated scaling and scheduling to consistently balance cost and performance.
Multimodal and Heterogeneous Pipelines: Integrating vision, speech, and text introduces heterogeneous operators and more complex scheduling requirements, making production deployment more challenging.
Observability and Risk Management: The unpredictability of model outputs, security risks associated with long-context processing, and growing requirements for auditing, monitoring, and compliance increase the complexity of operating inference systems.
Standardization and Reproducibility: Differences in quantization methods and compilation toolchains can lead to variations in model performance and output quality. Benchmark comparability and cross-platform portability also remain open challenges.
LLM inference is evolving from component-level performance optimization toward full-stack system optimization, and from serving individual models toward dynamically orchestrating heterogeneous models. This evolution is driven by growing demand for AI applications and the ongoing trade-offs among hardware resources, computational cost, and system performance.
In the short term, optimization efforts are likely to focus on memory management, quantization, and tail-latency reduction. In the medium to long term, inference infrastructure is likely to evolve into full-stack platforms designed to address diverse application requirements while balancing cost and privacy.
The key challenge going forward will be to develop inference architectures that strike an effective balance among performance, cost, usability, and flexibility.
2.3 LLMOps:Overview and Current Landscape
LLMOps refers to a set of practices, technologies, and processes for managing the end-to-end lifecycle of large language model applications. It can be viewed as an evolution of MLOps for the generative AI era, with a focus on managing the entire lifecycle of LLM development, deployment, monitoring, and continuous iteration.
The core objectives of LLMOps are:
Standardization: Establishing standardized, end-to-end pipelines from experimentation to production.
Scalability: Supporting the continuous delivery and operation of models with tens of billions of parameters.
Control and Governance: Balancing the competing requirements of security, cost, and performance.
Key Differences from Traditional MLOps:
| Dimension | MLOps | LLMOps |
|---|---|---|
| Model Characteristics | Static prediction (classification/regression) | Dynamic generation (text/multimodal) |
| Data Dependency | Structured data | Unstructured text and instruction data |
| Deployment Challenges | Low-latency requirements | Long-form generation and inference optimization |
| Iteration Frequency | Monthly updates | Fine-tuning on a daily basis |
The current LLM application development landscape remains fragmented. Platforms such as Hugging Face, Weights & Biases (W&B), and MLflow each cover different parts of the development and deployment lifecycle, but a unified LLMOps stack has yet to emerge. Based on the maturity of the overall development workflow, organizations can be broadly categorized into three stages:
Experimental Stage: Notebook-based development and single-GPU fine-tuning using platforms such as Colab and Kaggle.
Engineering Stage: Distributed training combined with production deployment using inference engines such as vLLM, typically seen in startups and smaller AI companies.
Enterprise Stage: Mature operational practices, including compliance monitoring, observability, and A/B testing, as seen in organizations such as OpenAI and Anthropic.
At present, many organizations are in the process of transitioning from the experimental stage to the engineering stage, while only a smaller number have established the infrastructure and operational capabilities associated with the enterprise stage. This transition typically requires significant investment in engineering infrastructure, tooling, and operational processes.
LLMOps provides an agile, iterative approach to improving this development maturity. Rather than following a traditional “Train → Deploy” workflow, teams can adopt a faster iteration loop of “Prompt → Fine-tune → RLHF → Serve”, allowing them to continuously evaluate, refine, and deploy LLM applications.
2.3.1 Open-Source Ecosystem
The open-source LLMOps toolchain is expanding rapidly, with a growing number of projects emerging across several key areas:
| Area | Representative Projects |
|---|---|
| Development Frameworks | Transformers, LitGPT, LangChain, Haystack, Flowise, Dify |
| Prompt Engineering & Management | PromptFlow, DSPy |
| Training Management | DeepSpeed, ColossalAI, Axolotl (fine-tuning toolkit) |
| Deployment & Orchestration | Ollama, NVIDIA Dynamo, OpenLLM |
| Inference Serving | vLLM, TGI, SGLang, TensorRT-LLM, OpenVINO |
| Monitoring & Tracing | EleutherAI LM Evaluation, LangSmith, Langfuse |
| Observability | OpenLLMetry, Helicone |
| Evaluation & Testing | promptfoo, DeepEval, OpenCompass |
| Data Management | LlamaIndex, Dolphin (instruction-data cleaning), Hugging Face Datasets |
| Safety & Compliance | Guardrails, NeMo Guardrails (content filtering), Patrol, LLM Guard |
A typical LLMOps workflow and toolchain might look like the following:
a. Development and Internal Testing
Orchestration: Use LangChain or LlamaIndex to build application logic and orchestrate LLM workflows.
Evaluation: Use promptfoo or DeepEval for unit testing and benchmarking of prompts and workflows.
Prototyping: Use Flowise or Dify to enable non-technical teams to quickly prototype and validate ideas.
b. Pre-Production and Deployment
Deployment: Package applications as APIs using frameworks such as FastAPI, or deploy them directly through Dify.
Knowledge Base: Use vector databases such as Chroma or Weaviate to store and manage vectorized knowledge.
Security Scanning: Integrate LLM Guard to filter and validate model inputs and outputs.
c. Production and Monitoring
Observability: Integrate LangSmith or Langfuse to trace individual requests, including execution paths, latency, cost, and token usage.
Evaluation and Iteration: Collect user feedback and production bad cases, feed them back into promptfoo as new test cases, and continuously refine prompts and workflows.
2.3.2 Challenges and Future Directions
LLM development currently faces several challenges, particularly the length and complexity of development workflows, the lack of unified tooling that can cover the full lifecycle, uncontrolled development and operational costs, and gaps in evaluation frameworks. In many organizations, consistency and operational standards across development and production are still maintained largely through manual processes and organizational controls.
Looking ahead, several key directions are likely to shape the evolution of LLMOps:
Autonomous AI Operations: Models continuously monitor their own workloads and automatically trigger scaling actions, such as Kubernetes HPA-based autoscaling for LLM workloads.
Synthetic Data-Driven Development: LLMs are increasingly used to generate training data, following approaches such as Self-Instruct and its successors.
Cloud-Native LLMOps: Lightweight, cloud-native inference architectures, including WebAssembly-based runtimes such as Fermyon Spin, may enable more efficient deployment.
Built-In Governance and Safety: Governance and safety checks are increasingly integrated directly into AI workflows, enabling real-time policy enforcement and risk screening during inference.
The Rise of AgentOps: As agent-based applications become more prevalent, LLMOps is evolving toward AgentOps, with greater emphasis on multi-agent coordination, tool-call success rates, execution tracing, and debugging of planning trajectories.
LLMOps is transitioning from ad hoc, project-specific workflows toward more standardized and industrialized development practices. Future competition is likely to focus on three areas:
End-to-End Automation: Automating the entire lifecycle, from data annotation and evaluation to automated incident detection and recovery.
Trusted AI Feedback Loops: Establishing closed-loop systems for auditability, monitoring, and alignment with safety and governance requirements.
Cost-Efficient AI at Scale: Making advanced AI capabilities more accessible through smaller models, Mixture-of-Experts (MoE) architectures, and more efficient infrastructure.
Open-source communities have become a major driver of innovation, but enterprise-grade LLMOps still needs to address the fragmentation of the current toolchain and provide more integrated, end-to-end solutions.
3. AI Agent
3.1 AI Agent Landscape and Trends
By 2025, AI agents had evolved from a technological concept into a technology with large-scale adoption and deployment. Their definition has also evolved toward intelligent systems capable of perceiving their environment, making autonomous decisions, and executing tasks. According to IDC's framework, mature AI agents are expected to demonstrate three core capabilities: cognitive generalization, closed-loop action, and evolving memory. These capabilities distinguish AI agents from conventional LLM applications that primarily respond to individual instructions. This shift enables AI agents to autonomously decompose and execute complex tasks, positioning them as an emerging foundation for enterprise digital transformation.
In terms of technological maturity, open-source models have become an important driver of AI agent adoption. The development of open-source models such as DeepSeek has accelerated the commercial deployment of low-cost, locally hosted LLM solutions, significantly lowering the barriers to AI agent deployment while addressing data privacy concerns. According to the cited industry data, 23% of enterprises have adopted local deployment, and this figure is projected to reach 90% by 2028. On the computing infrastructure side, the growing adoption of domestic GPUs, including Huawei Ascend and Cambricon MLU, is accelerating the localization of AI computing infrastructure. Integrated AI computing systems, with an average price of approximately RMB 6.8 million, are also providing infrastructure support for large-scale AI agent deployment.
3.2 AI Agent Market Size and Industry Landscape
The AI agent market is experiencing rapid growth. At the same time, the expansion of the open-source ecosystem is providing strong momentum for the broader adoption of AI agent technologies. Zhu Qigang, Secretary-General of the Shanghai Open Source Information Technology Association, has described the development of China's open-source ecosystem as having two distinct phases: “before DeepSeek” and “after DeepSeek.” According to his assessment, the open-source ecosystem is shifting from an “operations-driven” model toward a “value-driven” model, with developers increasingly contributing in response to practical needs and creating a positive feedback loop in which open-source projects help sustain and improve the open-source ecosystem.
Andrew Aiken, a member of the Technical Oversight Committee of the Linux Foundation's Open Source Foundation for Financial Services (FINOS), has emphasized the importance of transparency in open source for AI development. He argues that greater openness can strengthen community engagement, reduce costs, increase the adoption of AI technologies, and improve trust across the industry.
a. Global Market
According to MarketsandMarkets, the global AI agent market is projected to grow from USD 5.1 billion in 2024 to USD 47.1 billion by 2030, representing a compound annual growth rate (CAGR) of 44.8%.
b. Rapid Growth in the Chinese Market
According to IDC, the Chinese enterprise AI agent market is expected to reach approximately RMB 19 billion in 2025, with a projected CAGR of more than 110% between 2025 and 2028. This projected growth rate is substantially higher than the global average, reflecting the increasing role of AI agents in enterprise digital transformation in China.
c. Accelerating Commercialization
In the first half of 2025, 371 AI agent-related projects were awarded, of which 305 publicly disclosed projects had a combined contract value of RMB 1.016 billion. This represents a 128% increase from the RMB 445 million recorded during the same period in 2024. The sharp increase in commercial activity indicates that AI agents are moving beyond the proof-of-concept stage toward broader procurement and enterprise deployment.
| Metric | Data |
|---|---|
| Total projects awarded in H1 | 371 |
| Publicly disclosed projects | 305 |
| Total value of publicly disclosed projects | RMB 1.016 billion |
| YoY growth vs. H1 2024 | 128% (H1 2024: RMB 445 million) |
3.2.1 The Impact of Open Source on AI Agent Development
a. Early-Mover Advantage in Open-Source Foundation Models
The international open-source community gained an early lead in open-source foundation models for AI agents. Meta's Llama family has evolved continuously through the Llama 3.x and 4.x generations, becoming one of the key foundations for AI agent development outside China. Although Anthropic's models are primarily closed-source, the company introduced the Model Context Protocol (MCP) as a fully open protocol, making it a major milestone in agent tool interoperability in 2025.
The open-source ecosystem has also been expanding around standardized Skills—modular, reusable capabilities that allow agents to flexibly invoke different functions. Together with MCP, these developments help address ecosystem fragmentation and high integration costs when agents interact with external environments, contributing to a more open, interoperable, and collaborative agent ecosystem.
b. Foundation Models Accelerate the Evolution of Open-Source AI Agent Frameworks
In 2025, the availability of open-source foundation models accelerated the rapid evolution and large-scale adoption of AI agents. Empirical findings from the Measuring Agents in Production (MAP) study indicate that agents deployed in real-world production environments exhibit a strong emphasis on engineering practicality and operational reliability:
Deterministic workflows with tightly controlled autonomy: Enterprises generally have very low tolerance for model hallucinations. 68% of production agents execute no more than 10 autonomous steps before triggering human verification. Many production systems have moved away from conventional ReAct-style loops toward graph-based architectures that explicitly constrain agent behavior and improve execution stability.
Multi-model routing as a mainstream architecture: 59% of production-grade agents adopt multi-model routing. Open-source models are commonly assigned to different tasks according to their capabilities and cost profiles. Lightweight models such as Qwen 2.5 8B can handle relatively simple tasks, including intent recognition and basic tool calls, while more capable reasoning models such as DeepSeek R1 are used for demanding workloads such as code refactoring and complex planning. This approach helps balance cost and performance.
A new approach to performance trade-offs on the application side: 66% of complex reasoning applications can tolerate minute-level latency in exchange for more reliable outputs. When models are upgraded, 70% of business use cases rely on static and dynamic prompt optimization, while fewer than one-third perform fine-tuning of model weights. Meanwhile, 74% of agent applications use a human-in-the-loop approach for evaluating system performance.
As the open-source ecosystem continues to expand, competition among AI agent frameworks is intensifying. Clear differences are emerging in areas such as ease of use, governance and control, and enterprise-grade capabilities. At the same time, validating the real-world effectiveness of these frameworks in production has become a common challenge across the industry.
c. International Open-Source Framework Ecosystem
International open-source frameworks are leading the development of agent orchestration and multi-agent collaboration:
LangChain / LangGraph: LangChain has a 55.6% developer adoption rate, based on statistics from 542 projects compiled by Upwork in 2025, making it one of the most widely used AI agent frameworks. LangGraph focuses on stateful workflows and uses graph-based structures to manage task state.
AutoGen (Microsoft): With approximately 21.2K GitHub stars, AutoGen is among the leading frameworks for multi-agent conversational systems. Its key features include support for multi-turn interactions and mechanisms for incorporating human feedback.
CrewAI: A role-based multi-agent collaboration framework built around the concept of treating agents as a “team of microservices,” offering a differentiated approach to multi-agent orchestration.
LlamaIndex: Focuses on the data retrieval layer, providing agents with structured access to enterprise data.
n8n: Positioned as a self-hosted combination of Zapier and AI agent orchestration, n8n provides more than 400 integration nodes, a visual workflow editor, and dedicated AI agent nodes, while supporting self-hosted deployment. The project has more than 100K GitHub stars.
d. Domestic Open-Source Framework Ecosystem
Chinese open-source frameworks have developed distinctive strengths in low-code AI agent development, RAG, and enterprise deployment:
Dify: With approximately 136K GitHub stars, Dify ranks among the leading AI agent frameworks globally. It provides a visual RAG workflow designer and dynamic routing capabilities, while allowing traditional NLP tasks to be decomposed into reusable atomic modules. Developers can construct complex business workflows through a visual drag-and-drop interface, making Dify a prominent platform for production-grade AI application development.
Coze Studio (ByteDance): Positioned as an “AI Agent IDE,” Coze Studio had approximately 19.4K GitHub stars as of January 2026 and is licensed under Apache 2.0. It integrates major model services and supports the visual construction of complex agents that use multiple tools and knowledge bases, making it suitable for enterprises building private, self-hosted agent platforms. Its companion project, Coze Loop (approximately 5.2K stars), provides prompt version management and automated evaluation, while Eino (approximately 11.5K stars) addresses the need for a Go-based framework for LLM application development. Eino adopts a three-layer architecture consisting of components, orchestration, and agents, allowing Go developers to build AI applications using familiar development patterns.
MaxKB: An enterprise-grade agent platform centered on RAG. As of November 2025, it had more than 19.4K GitHub stars, over 750,000 cumulative downloads, and more than 1,000 enterprise users, with deployments across industries including education, healthcare, and manufacturing. Its key strength is its support for a progressive path from basic RAG-based question answering to workflow automation and ultimately agent-based applications.
RAGFlow: A RAG engine focused on deep document understanding, integrating agent capabilities to provide knowledge-driven AI responses. As of April 2026, RAGFlow had reached approximately 77.2K GitHub stars. It is particularly suited to enterprise RAG applications that require high-precision document understanding and support for documents with complex formats.
Overall, the Chinese open-source framework ecosystem exhibits several distinct characteristics. Dify has established a strong position in full-featured AI application development. ByteDance's ecosystem, consisting of Coze Studio, Coze Loop, and Eino, spans the development environment, quality-management layer, and application framework, forming a relatively comprehensive open-source agent toolchain. MaxKB focuses on deep RAG optimization for enterprise knowledge-management and question-answering scenarios, while RAGFlow specializes in deep document understanding.
Chinese enterprises also tend to place greater emphasis on compatibility with local industry ecosystems and private deployment. Compared with many international frameworks, these projects are more closely aligned with the institutional requirements, industry workflows, and domestic technology ecosystems commonly encountered in Chinese government and enterprise environments.
e. Comparison of Open-Source AI Agent Frameworks
In 2025, the open-source AI agent framework ecosystem became highly diverse, with projects adopting increasingly differentiated technical approaches. Based on statistics covering 542 projects, LangChain had a 55.6% adoption rate, while CrewAI (9.5%) and AutoGen (5.6%) focused primarily on multi-agent collaboration, and LlamaIndex (7.1%) specialized in data retrieval.
Among Chinese frameworks, Dify (120K stars) led the low-code agent platform category, while MaxKB (19.4K stars) focused on enterprise RAG and workflow orchestration. RAGFlow specialized in deep document understanding, n8n (100K+) focused on workflow automation combined with AI, and Coze built a cloud-based agent ecosystem around its open-source components. LangGraph (15.1K) and AutoGen (21.2K) have developed distinct strengths in stateful workflows and conversational agents, respectively.
Framework selection should ultimately be based on the specific application scenario. For example, LangChain is commonly used for customer-service and question-answering applications; MaxKB is well suited to internal enterprise knowledge Q&A; AutoGen can be used for intelligent development assistants; and CrewAI can be considered for collaborative content-generation workflows.
| Project | Core Focus | Key Differentiators | License | Typical Users |
|---|---|---|---|---|
| LangChain | LLM application development framework | Modular components, chain-based execution, seamless integration with LangGraph | MIT | LLM application developers, enterprise AI teams |
| LangGraph | Stateful agent workflow framework | Graph-based state management, checkpointing and recovery, loop support | MIT | Financial risk management, workflow automation teams |
| AutoGen | Conversational multi-agent framework | Multi-turn dialogue optimization, human feedback, customizable conversation patterns | MIT | AI application developers, research teams |
| CrewAI | Role-based multi-agent collaboration | “Agents as a microservice team,” task delegation and coordination | MIT | Content production, marketing automation teams |
| LlamaIndex | Data retrieval and RAG framework | Extensive data connectors, multiple indexing structures, deep LLM integration | MIT | Data-intensive applications, RAG development teams |
| n8n | Workflow automation + AI agents | 400+ integrations, visual orchestration, self-hosted deployment | Sustainable Use License | Enterprise DevOps, workflow automation teams |
| Dify | LLM application development / agent platform | Low-code, full-stack capabilities, 120K+ GitHub-star community | Apache 2.0 | Developers, product managers, small and medium-sized enterprises |
| MaxKB | Enterprise knowledge-base Q&A agent | One-click enterprise deployment, user-friendly console, deep RAG capabilities | GPLv3 | Enterprise IT, knowledge management teams |
| RAGFlow | Deep RAG engine + agents | Advanced document understanding, layout analysis, OCR | Apache 2.0 | Document-intensive industries, compliance teams |
| Coze (including open-source components) | Cloud-based agent platform + open ecosystem | Plugin ecosystem, bot distribution channels, multimodal capabilities | Mixed licensing | Developers, content creators, small and medium-sized enterprises |
f. A Paradigm Shift in the Global Open-Source Ecosystem: OpenClaw
In February 2026, the open-source AI agent project OpenClaw surged to the top of GitHub's global trending rankings. Launched in November 2025, the project surpassed 200,000 stars within just 84 days and eventually reached approximately 236,000 stars, making it the second-most-starred project on GitHub after React. Its growth rate was reported to be approximately 18 times faster than that of Kubernetes.
OpenClaw's key innovation lies in shifting AI from a “question-and-answer” model toward an “agentic execution” model. It can autonomously manage email, schedule calendars, execute shell commands, and orchestrate multi-step workflows, reportedly saving users an average of more than 10 hours per week. Its architecture is centered around a Gateway, while its ecosystem layer includes the ClawHub skills marketplace, which has more than 5,700 community-developed skills and more than 1.5 million cumulative downloads.
OpenClaw has also triggered a broader industry response. Chinese technology companies including Zhipu AI, Tencent, Huawei, Alibaba, ByteDance, and Xiaomi have introduced OpenClaw-inspired products, while major cloud platforms have launched deployment services to support similar agent-based applications.
3.2.2 AI Agent Technical Architecture and Core Components
By 2025, AI agent architectures had converged toward a relatively standardized design pattern. The core architecture typically consists of a large language model (LLM), task planner, memory management system, and tool-calling module, with tool use serving as one of the most critical capabilities.
According to Volcano Engine's Agent Technology Landscape, mainstream agent architectures contain five major components:
- Task Planner: Decomposes high-level goals into executable tasks.
- Skill Executor: Carries out specific operations and actions.
- Memory Manager: Stores and manages contextual information.
- Tool Caller: Connects the agent to external systems and services.
- Multi-Agent Coordinator: Enables coordination and collaboration among multiple agents.
Development Languages and Tool Ecosystem: In 2025, the open-source AI agent ecosystem was primarily centered on Python, with increasing support for multi-language development. An analysis of 542 AI agent development projects on Upwork found that more than half (52%) used Python as the primary development language. Its extensive ecosystem—including TensorFlow, PyTorch, LangChain, and Hugging Face—has made Python a common environment for model inference and agent orchestration.
Production deployments typically combine Python with other programming languages. Node.js (17%) and Go (12%) also appeared frequently, particularly in applications requiring large-scale real-time APIs and highly concurrent workloads.
Memory and Database Systems: Memory and database systems serve as foundational components of AI agents, and the open-source ecosystem offers a broad range of technology choices. Among the 133 projects that explicitly referenced memory capabilities, Pinecone (22.6%) led as a managed “cloud memory” solution. Open-source alternatives such as Weaviate (16.5%), Qdrant (4.5%), and Milvus (4.5%) are also gaining adoption, particularly among teams seeking greater control over costs and data.
At the same time, the combination of PostgreSQL and pgvector (18.8%) demonstrates how traditional database systems are adapting to AI workloads. Redis (8.3%) and MongoDB (4.5%) have taken a similar approach by adding vector search capabilities to their existing database platforms.
3.2.3 AI Agent Industry Applications and Production Deployment
In 2025, open-source AI agent technologies made the transition from experimental environments to production applications, with large-scale deployments emerging across key sectors including finance, manufacturing, healthcare, and education.
According to industry research cited in the source material, AI agents have achieved adoption rates of more than 70% in intelligent customer service and approximately 60% in data analytics, making these two of the most mature application areas.
In the financial sector, one bank used LangGraph to build an intelligent risk-control system that analyzed transaction data through multi-node collaboration, reportedly increasing its anomaly detection rate by 40%. In software development, Alibaba's Tongyi DeepResearch has been validated in several Alibaba applications, including the combination of map navigation and local services in Amap's agent-based experience, as well as Tongyi Farui's retrieval of authoritative legal precedents, statute matching, and integration of professional legal opinions.
Enterprise automation has emerged as an important area for open-source AI agent technologies. In October 2024, Microsoft announced the integration of 10 autonomous AI agents into Dynamics 365. These agents were designed to automate workflows across customer service, sales, finance, warehousing, and other business functions.
The agents support OpenAI's o1 model and are designed to perform complex cross-platform business processes with a degree of autonomous learning and execution. One reported case study found that U.S. telecommunications company Lumen could save approximately USD 50 million per year through AI agents, equivalent to the productivity of 187 full-time employees.
Such commercial deployments have contributed to growing demand for open-source AI agent solutions and further accelerated development of the open-source ecosystem.
Open-source AI agent frameworks can provide both cost advantages and efficiency gains in enterprise applications. According to enterprise-user research cited in the source material, a marketing automation system built with CrewAI used a combination of “market analyst + content generator + advertising optimizer” agents and reportedly increased an e-commerce company's conversion rate by 22%.
In another reported case, a technology company used AutoGen to build a DevOps agent, increasing code-review efficiency by 50% while reducing cloud resource costs by 35%.
These cases illustrate the potential enterprise value of AI agent solutions built on open-source frameworks and their ability to generate measurable business outcomes.
It is also important to note that open-source and closed-source solutions are not necessarily direct substitutes; instead, they can form a complementary ecosystem. Academician Kai-Fu Lee has argued that although closed-source models currently maintain a somewhat larger share of commercial applications, the balance may change significantly over the next one to two years. He has also emphasized that open-source and closed-source models should not necessarily be viewed as opposing alternatives, but rather as approaches that can coexist within a balanced commercial model.
Industry data cited in the source material indicates that in 2024, open-source and closed-source models accounted for approximately 50% each of enterprise LLM adoption globally. However, enterprises using open-source models reportedly had a 78% rate of secondary development, compared with 12% for closed-source models. This suggests that open-source solutions may offer particular advantages in highly customized enterprise scenarios.
3.3 Challenges and Future Trends for AI Agents
Although the open-source AI agent ecosystem made substantial progress in 2025–2026, it continues to face challenges across several dimensions.
Data Bottlenecks: The shortage of high-quality reasoning datasets has become an important constraint on further improvements in agent capabilities. Academician Kai-Fu Lee has noted that the proportion of high-quality reasoning datasets—such as derivation chains from academic papers and engineering debugging logs—that are openly available remains below 15%, significantly lower than the availability of general-purpose text data.
Addressing this challenge may require new data-sharing mechanisms. One potential approach is to use blockchain-based incentive mechanisms to reward data contributions and encourage broader data sharing, potentially moving open-source data toward a new stage of development.
Computing and Inference Capacity: The growing adoption of AI agents is driving a sharp increase in token consumption. According to data from China's National Data Administration, China's average daily token consumption was approximately 100 billion tokens in early 2024, but had exceeded 30 trillion tokens per day by the end of June 2025, representing more than a 300-fold increase over roughly 18 months.
IDC projects that the number of active AI agents worldwide could reach 2.216 billion by 2030, while annual token consumption could increase from 0.0005 petatokens in 2025 to 152,000 petatokens, representing an increase of more than 300 million times.
Such exponential growth would place substantial pressure on the inference-computing supply chain—including chips, servers, liquid cooling, and power infrastructure—and is likely to become an important area for further technological innovation.
Trust and Adoption: IDC's analysis suggests that Chinese AI models have crossed a “technology gap” but have yet to fully overcome a “trust gap.” According to the cited analysis, overseas enterprises' reluctance to adopt Chinese AI models is driven primarily by concerns about long-term support, security, and regulatory compliance rather than model performance.
Open-source development, low-cost deployment, and local infrastructure capabilities may provide potential pathways for addressing these concerns. Luo Fuli has also argued that agent frameworks can significantly raise the practical performance ceiling of domestic open-source models that have not yet fully matched leading closed-source models. According to her assessment, domestic open-source models are already approaching the performance of the latest Claude models in many use cases while maintaining a relatively reliable baseline of performance.
Security: The 2026 Future Industries Research Report from the China Center for Information Industry Development (CCID) highlights new security and governance challenges associated with autonomous AI agent platforms such as OpenClaw. These challenges require comprehensive governance mechanisms spanning technical architecture, privacy protection, and behavioral auditing.
Five Future Trends:
The open-source AI agent ecosystem is likely to develop along five major directions.
Trend 1: Multimodal Integration
Multimodal capabilities will significantly expand the application scope of AI agents, enabling them to evolve beyond text-only interaction toward integrated processing of text, images, audio, and video.
Alibaba's open-source vision model Qwen2.5-VL can function directly as a vision agent, performing reasoning and dynamically using tools to complete complex multi-step tasks on computers and smartphones. IDC has also identified multimodal integration as an important characteristic of enterprise-grade AI agent applications.
Trend 2: Deeper Autonomous Collaboration
With the emergence of reasoning models and the adoption of Reinforcement Fine-Tuning (RFT), an increasing number of LLM-based agents can learn and explore autonomously within specific domains.
Combining the autonomous learning and exploration capabilities traditionally associated with reinforcement-learning agents with the broader capabilities of general-purpose agents—including task execution, user interaction, and complex problem solving—could drive AI agents toward greater levels of autonomy.
Trend 3: Edge Deployment and Edge Intelligence
The continued expansion of OpenClaw across PCs, smartphones, and wearable devices is accelerating the development of on-device AI.
Xiaomi has introduced miclaw on smartphones based on the OpenClaw architecture, encapsulating smartphone system capabilities into more than 50 system tools and ecosystem services. Huawei has disclosed that Xiaoyi Claw, based on HarmonyOS, is currently in beta and can assist users with tasks such as document editing, PowerPoint creation, and automated email replies, while supporting cross-device collaboration.
The OpenClaw community has also announced plans to develop a version for smart glasses based on available wearable-device development tools.
This combination of on-device intelligence and cloud computing is expected to address enterprise requirements for real-time responsiveness, privacy protection, and network bandwidth efficiency.
Trend 4: Standardization and Protocol Convergence
The global agentic AI ecosystem is developing along two parallel paths: open standards and the MCP ecosystem. Their continued convergence may help drive agent technologies toward greater maturity.
A unified agent protocol—potentially analogous to HTTP as a communication standard for the web—could eventually emerge. Standardization would reduce integration costs between different AI agent systems and facilitate cross-platform and cross-ecosystem collaboration.
At the same time, the restructuring of the token economy is becoming an important driver of LLM commercialization and adoption, potentially shifting the agent ecosystem from technology-driven competition toward a model driven by both technology and commercial value.
Trend 5: Deeper Human–AI Collaboration
Academician Kai-Fu Lee emphasized at the COP conference that the biggest future opportunities may lie in the relationship between humans and machines, particularly human–computer interaction. Over the past several decades, companies that successfully captured key interfaces between humans and machines have often played an important role in shaping technology markets.
Natural-language interaction represents another milestone in human–computer interaction. Both chatbots and AI agents are contributing to this evolution.
While 2025 was characterized by intense competition among foundation models, the center of gravity of AI development in 2026 has increasingly shifted toward agent-based systems, with agent workloads driving rapidly increasing token consumption. Lu Yanxia, Research Director at IDC China, has identified stronger agent capabilities as an important development direction for foundation models in 2026, including applications such as deep research, intelligent office automation, and AI coding assistants.
Based on the analysis above, enterprises and developers adopting open-source AI agent technologies should consider several areas:
Prioritize open-source frameworks with active communities, comprehensive documentation, and strong maintenance capabilities.
Evaluate whether to build systems in-house or leverage existing solutions based on the specific requirements of the application.
Prioritize data privacy and security management and ensure compliance with applicable regulations.
Monitor emerging trends such as multimodal AI and edge computing and build the necessary technical capabilities in advance.
These practices can help organizations adapt to the rapid evolution of AI agent technologies and make effective use of the innovation enabled by the open-source ecosystem.
4. AI Ethics, Security, and Governance: From Soft Principles to Hard Constraints
In the broader history of technological development, the period from 2025 to early 2026 represents an important turning point for China's open-source artificial intelligence (AI) ecosystem. As foundation models evolved from simple text-based interaction toward multimodal generation and autonomous agent execution, both their inherent security exposure and their potential societal risks expanded significantly.
During this period, AI governance in China and internationally moved beyond a primarily principle-based approach toward a more formalized framework involving mandatory standards, judicial precedents, defensive technical architectures, and national data-security reviews. Within China's open-source ecosystem, this has contributed to a governance approach emphasizing early compliance, legal safeguards, data sovereignty, and security-by-design.
4.1 Macro-Level Policy and Compliance Frameworks
As open-source foundation models become more widely available, regulators in China and other jurisdictions have accelerated efforts to balance technological innovation with ethical and security requirements. Compliance has increasingly become a prerequisite for the deployment and commercialization of open-source models.
International Standards and Regulatory Divergence: The Open Source Initiative (OSI) formally established the Open Source AI Definition (OSAID) in 2025, drawing a distinction between Open Weights and genuinely open-source AI systems.
Meanwhile, the EU Artificial Intelligence Act (EU AI Act) entered into broader application, introducing transparency and other requirements for general-purpose AI (GPAI) models.
Traceability and Filing Requirements in China: Compared with the international emphasis on transparency, Chinese AI governance places greater emphasis on traceability and clearly assigning responsibility.
The source text states that China's national standard GB 45438-2025, Methods for Labeling Artificial Intelligence-Generated Synthetic Content, which took effect in September 2025, requires underlying technologies to incorporate implicit metadata labeling and digital signatures with cryptographic protection. Under this framework, open-source models without the required mechanisms may face restrictions on commercialization.
The continued implementation of China's Provisions on the Administration of Deep Synthesis of Internet-based Information Services also clarifies the allocation of regulatory responsibility. When developers use open-source foundation models to provide online services, they may effectively assume the role of a service provider and become subject to relevant algorithm-filing requirements. Open source therefore does not automatically exempt an organization from regulatory obligations.
4.2 Intellectual Property and Judicial Precedents: Defining the Legal Boundaries of Open-Source Fine-Tuning
Copyright disputes involving data scraping, model training, and AI-generated content remain a major legal concern for the open-source community. According to the source material, a second-instance judgment issued by the Shanghai Intellectual Property Court in April 2026 in the “Medusa” AI model copyright case established an important judicial precedent for the open-source ecosystem.
Clarifying Copyright Infringement Standards: The court reportedly held that AI-generated content lacking substantial intellectual contribution from a human does not qualify for protection under adaptation rights. However, if a user trains a LoRA model on protected material through an open-source platform and publicly releases it, and the resulting output is substantially similar to the original work, such conduct may reproduce the original work's creative expression and constitute infringement of the right of reproduction.
A Tiered Duty of Care for Platforms: The judgment also established differentiated obligations based on the underlying technical architecture. For open-source hosting platforms using a “foundation model + LoRA fine-tuning” architecture, the platform has limited control over users' private training data. According to the reported judgment, platforms may avoid liability if they fulfill appropriate pre-use notification requirements and respond to infringement notices by taking measures such as removing the relevant content.
This precedent illustrates how legal systems may distinguish between different levels of technical responsibility while preserving room for innovation in open-source infrastructure.
4.3 Geopolitics and Data Sovereignty: National Security Governance in the Agent Era
If the “Medusa” case concerns copyright issues surrounding AI-generated content, the evolution of AI toward autonomous agents shifts the governance debate into the areas of national security and geopolitics.
The source material states that in April 2026, China's National Development and Reform Commission (NDRC) and relevant regulatory authorities used the Measures for Security Review of Foreign Investment to halt Meta's reported multi-billion-dollar acquisition of Manus, an advanced general-purpose AI agent with Chinese origins.
According to the source, this case illustrates a broader shift toward incorporating AI assets into national-security governance.
The Strategic Significance of AI Agents: Traditional foundation-model governance primarily focuses on harmful content generation. Advanced agents such as Manus, by contrast, can be deeply integrated into enterprise networks, autonomously call external APIs, and execute complex business workflows.
This gives agents access to enterprise digital infrastructure and sensitive operational processes, potentially elevating them from general-purpose productivity tools to assets with characteristics associated with critical information infrastructure (CII).
Protecting Data Sovereignty: Agents with system-wide orchestration capabilities may gain access to sensitive industrial data and operational workflows. From a regulatory perspective, cross-border transactions involving advanced AI assets with significant business-execution capabilities may therefore require stringent national-security and data-security review.
As AI becomes increasingly globalized, data sovereignty and physical or logical isolation are likely to become increasingly important considerations for AI companies operating across jurisdictions.
4.4 New Security Infrastructure: Sandboxing and Guardrails
The asymmetric security risks associated with open-source models, combined with increasingly stringent compliance requirements, are driving a shift away from traditional perimeter-based security toward more deeply integrated defense mechanisms.
From 2025 to 2026, open-source communities and security vendors increasingly focused on building end-to-end security architectures that operate closer to the model and agent execution layers.
Model Capability Spillover and the “Zero-Day” Challenge: As frontier models become more capable at reasoning and code generation, AI systems can increasingly interact with and potentially exploit conventional software infrastructure.
The source material states that in April 2026, Anthropic disclosed an unreleased restricted model called Claude Mythos Preview. According to the source, the model demonstrated potentially destructive capability spillover by autonomously identifying and exploiting long-standing high-severity vulnerabilities in mainstream operating systems, browsers, and open-source components such as the Linux kernel without direct human guidance.
In response to this potential asymmetric threat, Anthropic reportedly initiated Project Glasswing, working with Microsoft, Google, AWS, and the Linux Foundation to use Mythos in controlled environments to identify and remediate vulnerabilities in foundational software.
The broader implication is that, in an agentic AI era, the speed of vulnerability remediation in foundational software may need to keep pace with the increasingly automated discovery of vulnerabilities.
Integrated Multimodal Safety and Real-Time Blocking: To address covert cross-modal attacks, Chinese technology companies such as Baidu have developed integrated multimodal safety guardrails. Alibaba Cloud's open-source Qwen3Guard-Stream is designed to perform streaming safety detection during model generation, allowing potentially risky outputs to be detected and blocked at the token level while reducing the latency associated with offline moderation.
Agent Governance and Sandboxing: To mitigate unauthorized actions by agents, open-source agent frameworks increasingly incorporate mechanisms such as the Principle of Least Privilege (PoLP) and Human-in-the-Loop intervention.
These mechanisms can establish auditable controls over agent behavior and ensure that cross-system API calls remain within explicitly authorized and traceable boundaries.
4.5 The Dual-Use Challenge: Evolving Risks and Hybrid Governance in Open-Source Biosafety
In AI for Science (AI4Science), open-source AI can significantly broaden access to advanced scientific capabilities, but it also introduces new forms of non-traditional security risk.
With the release of biological foundation models such as Evo 2, which can process nucleotide sequences at the million-base scale, the barriers to predicting protein structures and generating novel genetic sequences have been substantially reduced.
According to the source material, a tabletop exercise at the 2025 Munich Security Conference (MSC) highlighted the strong dual-use nature of such technologies and the potential for malicious actors to misuse unrestricted open-source models in ways that could create biological-security risks.
In response, the open-source community is exploring hybrid governance approaches that seek to balance scientific openness with safety requirements.
Physical Filtering of High-Risk Data: Leading institutions are beginning to remove high-pathogenicity viral genome data from training datasets before releasing open-source life-science models, reducing the possibility of generating potentially harmful biological sequences from the underlying training data.
Managed Access: Under a managed-access model, the underlying model infrastructure can remain accessible to the public in accordance with principles of open science, while high-risk functions—such as pathogen-sequence generation—require verified identity credentials and integration with research-ethics review mechanisms.
This approach seeks to preserve the benefits of open scientific research while introducing additional safeguards against malicious use.
5. Open Science
In recent years, artificial intelligence (AI) has contributed to a shift toward what is sometimes described as a “fourth paradigm” of scientific discovery—moving from data-driven approaches associated with the “Keplerian” phase and high-throughput experimentation associated with the “Edison” phase toward systems capable of complex reasoning and hypothesis generation, positioning AI as a potential cognitive collaborator for scientists.
A major driver of this transition is the convergence of open-source AI and open science. Open-source models can broaden access to advanced AI capabilities and allow researchers around the world to fine-tune and deploy domain-specific models.
However, the widespread adoption of generative AI also introduces significant challenges. The black-box nature of some AI systems, data contamination, and model hallucinations can amplify existing concerns surrounding reproducibility in scientific research. At the same time, the widening global compute divide and the dual-use risks associated with open-source AI, including biological security concerns, introduce additional geopolitical challenges.
One of the most significant macro-level trends in AI in 2025 was the rapid narrowing of the performance gap between closed-source proprietary models and open-weight models. This development has implications not only for the commercial structure of the AI industry but also for the economic and computational accessibility of AI-assisted research.
To address increasingly complex scientific tasks, the research community is developing a highly specialized, matrix-like ecosystem of open-source models. Different model families are increasingly optimized for specific scientific research tasks and application domains, creating complementary capabilities across the broader scientific AI ecosystem.
| Model / Family | Organization / Release Period | Core Architecture & Parameter Scale | Key Applications and Impact on Scientific Research and Open Science |
|---|---|---|---|
| DeepSeek R1 / V3 | DeepSeek (2025) | Reinforcement learning (RL), Mixture-of-Experts (MoE), latent attention mechanisms, and distilled lightweight variants | Reshaped the global landscape of open-source AI computing. Provides advanced logical reasoning and long-chain reasoning capabilities at relatively low cost, with applications in mathematical theorem proving, complex code generation, and physics modeling. |
| Llama 3 / 4 Family | Meta FAIR (2024–2025) | Dense architecture, spanning approximately 7B to 405B parameters | Provides highly general-purpose foundation models and serves as a widely used base for fine-tuning applications in bioinformatics and large-scale academic literature mining. |
| Mistral Large 2 / Family | Mistral (2024–2025) | 123B parameters, with strong memory efficiency and multilingual processing capabilities | Performs well in edge-computing environments and research settings with constrained memory resources. Supports robust operation in large-scale enterprise R&D environments and automated analysis systems. |
| Qwen 3 / Family | Alibaba Cloud (2024–2025) | Multimodal capabilities, including image editing and general-purpose LLMs | Strong performance in general-purpose querying and image/video segmentation. Lightweight variants—including models capable of running on devices such as Raspberry Pi—help broaden access to AI-powered research tools, particularly in resource-constrained and underserved regions. |
| Tulu / OLMo Family | Allen Institute for AI and the open-source community (2025) | Lightweight models optimized for advanced academic research and compact reasoning | Capable of handling highly complex scientific research instructions, supporting accurate cross-validation of academic literature and bibliometric analysis. |
5.1 Frontier Breakthroughs and Open-Source Contributions in AI for Science
Supported by increasingly mature open-source infrastructure, the application of artificial intelligence to scientific research—commonly referred to as AI for Science (AI4Science)—has moved well beyond simple data fitting. AI is now being applied to core scientific problems, including the discovery of molecular mechanisms, inverse materials design, and Earth system modeling.
5.1.1 Generative Advances in Structural Biology, Genomics, and Drug Discovery
Life sciences are among the fields in which AI has achieved some of its deepest integration and most significant advances. Following the foundational breakthroughs of AlphaFold and AlphaFold-Multimer, which can predict protein structures and protein–protein interactions with remarkable accuracy, research at the frontier in 2025 increasingly shifted toward generative biology.
In early 2025, Evo 2, a large biological foundation model developed by a team led by Stanford University professor Brian Hie in collaboration with NVIDIA and the Arc Institute, was released as open source and widely regarded as a milestone for AI-driven biology. Whereas its predecessor was trained on approximately 300 billion nucleotides, Evo 2 was trained on a dramatically larger corpus of nearly 9 trillion nucleotides, covering DNA sequences from known domains of life, including humans, plants, bacteria, and some extinct species. More importantly, Evo 2 supports an ultra-long context window of up to one million nucleotides. At the biological level, this capability allows researchers to capture, at whole-genome scale, functional interactions between genomic regions that may be physically distant but closely coordinated in biological function.
Evo 2 is not limited to sequence prediction; it can be viewed as an “autocomplete engine for the language of life.” In a proof-of-concept study, researchers used Evo 2 to design and synthesize novel sequences capable of precisely regulating the epigenome—that is, the mechanisms that control whether DNA is accessible or inaccessible and thereby regulate gene expression. The researchers even used biological encoding to “write” terms such as “EVO2” and “ARC” through the physical arrangement of cellular structures, in a pattern analogous to Morse code. By making the tool fully open source, researchers worldwide can now use virtual queries lasting only minutes or hours to simulate genetic mutations that might otherwise require thousands of years of natural evolution. This has the potential to substantially accelerate the identification of disease-associated mutations and the development of precision genetic engineering strategies for solid tumors.
Within the pharmaceutical industry, the traditional model of closed, isolated R&D systems is increasingly giving way to platformization. In 2025, for example, pharmaceutical giant Eli Lilly launched TuneLab, opening an AI model pipeline trained on billions of R&D data points to external startups and academic researchers. Making industrial-scale AI capabilities available as a platform can significantly lower the barriers to entry for computational biology startups. At the application level, AI-enabled drug repurposing—including the use of open-source AI for molecular docking and target screening—has also produced notable results. Studies have shown that deep-learning-based analyses of molecular binding affinity identified bazedoxifene, a drug originally developed for osteoporosis, as a potent STAT3 inhibitor, suggesting its potential as an anticancer candidate for diseases such as breast cancer. Similarly, AI-based predictions and subsequent validation have indicated that cimetidine, a gastrointestinal drug, may interfere with tumor immune evasion by disrupting processes mediated by E-selectin.
5.1.2 Materials Science and the Integration of AI with High-Throughput Autonomous Experimentation
The discovery of new inorganic crystal structures is critical to advances in semiconductor chips, high-energy-density solid-state batteries, and low-carbon photovoltaic technologies. Traditional materials science relies heavily on researchers’ intuition and time-consuming trial-and-error experimentation. Google DeepMind’s GNoME (Graph Networks for Materials Exploration) project uses graph neural networks to fundamentally accelerate the search for new materials.
By the end of 2024 and into 2025, GNoME had predicted as many as 2.2 million novel crystal structures, including approximately 380,000 structures with high predicted thermodynamic stability. This increased the estimated number of stable materials known to science by nearly an order of magnitude. DeepMind released the corresponding density functional theory (DFT)-validated data as open data and integrated it into the Materials Project, one of the world's largest online materials databases. The database now contains more than 520,000 promising materials with energies within 1 meV/atom of the convex hull. These open datasets provide materials scientists worldwide with a vast pool of candidates for next-generation carbon-capture photocatalysts, thermoelectric energy converters, and transparent conductors.
However, computational prediction is only the first step. Translating AI-generated designs into physically synthesized materials remains a major bottleneck. To address this challenge, a team led by MIT professor Ju Li introduced CRESt (Copilot for Real-world Experimental Scientists) in Nature in September 2025. CRESt integrates large multimodal models with robotic laboratory systems to create an active-learning framework based on knowledge-assisted Bayesian optimization (KABO) and supported by scientific literature.
Researchers can issue instructions in natural language, while the underlying ChatGPT API automatically invokes Python routines to control high-throughput synthesis equipment. During experiments, the platform can also use an integrated vision-language model, ChatGPT-4V, to operate scanning electron microscopy (SEM) systems and monitor the synthesis process. It can perform nanoscale EDX elemental mapping at approximately 2-nm resolution and identify experimental anomalies. During a 90-day evaluation, the AI-driven system autonomously conducted 3,500 electrochemical tests across 900 chemical formulations and successfully identified a palladium-based fuel-cell catalyst with record-level durability.
5.1.3 Earth System Science and Climate and Weather Prediction
Addressing global climate change requires Earth system models capable of operating at very high spatial resolution while maintaining stable performance over long time horizons. Between 2023 and 2024, data-driven AI models began to challenge traditional numerical weather prediction systems based on partial differential equations. Models such as DeepMind’s GraphCast, which can generate a 10-day global weather forecast in approximately one minute on a single TPU, and Prithvi-Weather-Climate, an open-source foundation model developed by NASA and IBM Research, demonstrated remarkable efficiency on tasks such as cyclone-track prediction. By 2025, models that had initially appeared primarily in top-tier research publications were increasingly being integrated into the infrastructure of national meteorological agencies and commercial organizations.
However, the most important development in 2025 was not simply another improvement in benchmark scores. Instead, the field began to critically reassess the limitations of purely data-driven approaches. Research teams at Stanford and other leading institutions published a series of comprehensive evaluations that subjected leading first-generation AI weather models—including GraphCast, FourCastNet, and Pangu-Weather—to rigorous stress tests under extreme climate scenarios such as the South Asian monsoon season.
These studies revealed an important limitation: although these AI models perform exceptionally well when trained and evaluated on highly smoothed reanalysis datasets, their errors can increase nonlinearly when they are exposed directly to noisy observations from real-world weather stations. This limitation is particularly important for predicting localized extreme precipitation and mesoscale kinetic-energy spectra during the South Asian monsoon. Purely probabilistic generative models may lack an explicit representation of physical conservation laws, such as thermodynamic constraints, resulting in substantial deviations from physical reality.
These limitations have contributed to a broader paradigm shift in Earth system science: rather than relying solely on black-box models, researchers are increasingly incorporating physical priors—including fluid-dynamics equations and conservation laws—into neural-network architectures through Physics-Informed Neural Networks (PINNs) and related hybrid approaches. The emergence of such hybrid models reflects a shift in AI-based weather prediction from purely statistical forecasting toward more physically grounded scientific modeling.
5.2 The Epistemological Crisis of AI-Driven Research: Reproducibility Challenges and the Reshaping of Evaluation Benchmarks
Although AI has substantially expanded the boundaries of scientific exploration, its widespread misuse and inappropriate application in academia are also contributing to a new reproducibility crisis, raising fundamental questions about the epistemological foundations of scientific research.
5.2.1 Sources of the Crisis: Hallucinations, Data Leakage, and Attention-Mechanism Conflicts
The scientific community has long struggled with the failure to reproduce findings from observational and empirical studies. The integration of AI into scientific research can amplify these existing structural weaknesses. A growing number of papers use large language models to generate or process biomedical data, sometimes producing highly formulaic analyses. In the absence of consistent methodological standards, AI systems can generate hallucinations when processing scientific literature, producing conclusions that appear logically coherent but are inconsistent with established physical principles.
Another major concern is data leakage—the inadvertent inclusion of benchmark or test data in a model’s training set. Such leakage can lead to apparently near-perfect evaluation results while producing substantially weaker performance in real-world experiments. This issue has become particularly important in fields such as medical imaging and materials-property prediction.
At a deeper technical level, AI systems may also struggle to maintain symbolic consistency when performing complex logical reasoning or mathematical derivations. For example, when deriving the Navier–Stokes equations, the same model may produce different symbolic representations across independent runs and may even omit viscosity terms at critical steps. Some research has attributed such inconsistencies to interactions and conflicts among different attention heads within Transformer architectures: attention heads responsible for spatial relationships and those responsible for variable correspondence may operate without sufficient explicit coordination, potentially causing inconsistencies in the resulting derivation.
If an AI system cannot reliably reproduce the intermediate steps of a multi-stage reasoning process, its use as a foundation for rigorous engineering analysis and mathematical derivation remains subject to significant limitations.
5.2.2 Theoretical Responses and Open Science Protocols
In response to these challenges, researchers have begun exploring theoretical frameworks for imposing greater structure and consistency on large language models. One line of research proposed in 2025 draws an analogy with gauge theory in physics and seeks to reinterpret prompt engineering as a lower-level coordination protocol. By “anchoring” key states in the model’s reasoning process, researchers aim to constrain the model’s representational degrees of freedom and encourage different attention heads to remain aligned within a shared logical representation space. Reported empirical tests suggest that such gauge-inspired anchoring techniques can improve symbolic consistency in complex equation-discovery tasks, including fluid dynamics and chemical kinetics.
At the governance and methodological level, organizations such as the Observational Health Data Sciences and Informatics (OHDSI) community are promoting the cross-institutional sharing of standardized analytical code and data architectures. This allows laboratories in different countries to process electronic health records and scientific literature using consistent analytical procedures, reducing methodological variation and helping mitigate the impact of AI-generated errors on biomedical discovery.
At the same time, the principles of open science have encouraged new approaches to literature processing, including Living Systematic Reviews and mechanisms for publishing short, peer-reviewed replication or failed-replication reports of existing AI findings, such as the Marbles initiative. Together, these practices contribute to a more dynamic research ecosystem capable of continuously identifying and correcting errors.
Policymakers are also calling for stronger reproducibility requirements for AI research, including the preregistration of hypotheses before model training, higher standards for statistical power, and greater transparency around negative results. In healthcare, the FUTURE-AI consortium, involving experts from 50 countries, has proposed international guidelines based on six principles—fairness, universality, traceability, usability, robustness, and explainability—to govern the validation and deployment of AI systems throughout their clinical lifecycle.
5.2.3 Rebuilding Evaluation Benchmarks to Prevent Data Leakage
To address the growing problem of unreliable evaluation metrics, the open-source AI community has begun substantially redesigning capability benchmarks. Rather than relying on static test sets that models may have encountered during training, evaluation platforms are increasingly moving toward contamination-resistant, dynamically generated, and human-validated benchmarks. This approach is intended to reduce the possibility that models achieve high scores simply by memorizing benchmark questions and instead assess whether they can genuinely generalize their capabilities to new problems.
| Core Benchmark | Evaluation Focus and Technical Characteristics | Evolution and Contributions in 2024–2025 |
|---|---|---|
| SWE-bench / Verified | A benchmark for evaluating AI agents’ ability to resolve real-world software engineering tasks based on GitHub Issues. | AI systems solved only 4.4% of tasks in 2023, rising to 71.7% in 2024. To reduce the risk of data contamination, a more rigorously curated and containerized Verified version was introduced in 2024–2025, featuring 500 human-validated tasks with confirmed solution paths. |
| Humanity’s Last Exam (HLE) | A challenging academic benchmark designed to evaluate performance at the frontier of human knowledge, consisting of 2,500 multimodal questions spanning mathematics, the humanities, and the natural sciences. | The questions are designed to have objectively verifiable answers, helping distinguish genuine reasoning capabilities from artificially inflated scores caused by probabilistic generation. The benchmark places particular emphasis on rigorous reasoning across diverse domains. |
| LiveCodeBench | A dynamic, contamination-resistant evaluation platform for assessing programming and algorithmic reasoning capabilities. | The benchmark continuously incorporates newly released programming problems from platforms such as LeetCode that postdate model training cutoffs, reducing the likelihood of benchmark contamination and testing whether models can genuinely generalize rather than rely on memorized solutions. |
| AIME I/II | A benchmark for advanced mathematical reasoning based on competition-level problems from the American Invitational Mathematics Examination (AIME). | The benchmark provides a rigorous test of the long-chain reasoning capabilities of open-source reasoning models such as DeepSeek R1. Answers are restricted to integers from 000 to 999, making it difficult to obtain high scores through linguistic shortcuts or answer-format manipulation. |
5.2.4 Geopolitics, the Compute Divide, and Open Access in the Global South
While open-source AI has significantly advanced the democratization of scientific capabilities, it has also become deeply entangled in broader geopolitical competition. The global distribution of computing resources remains highly unequal, contributing to a pronounced “North–South divide” in access to and adoption of generative AI.
Telemetry data from the second half of 2025 indicated that generative AI tools had reached 16.3% of the world's population. However, the benefits of this technological expansion remained highly unevenly distributed. Among the working-age population in the Global North, 24.7% were reportedly using AI to support research and productive activities, compared with only 14.1% in the Global South, with the gap in growth rates continuing to widen.
The United Nations Development Programme (UNDP) has warned that, without strong policy intervention, AI could contribute to a new “Great Divergence.” Under this structural imbalance, developing countries risk being relegated to the lower-value segments of the AI value chain—providing low-cost training data, annotation labor, and critical mineral resources while capturing only a limited share of the projected economic benefits of AI, including the USD 1 trillion opportunity projected for regions such as ASEAN.
At the national-strategy level, the United States, China, and the European Union all introduced targeted policy initiatives in 2025 in response to the accelerating open-source AI trend. These initiatives reflect competing approaches to shaping the rules and infrastructure surrounding AI-enabled scientific research.
United States: Deregulatory Push and Concerns over Talent Retention. In July 2025, the White House released the America's AI Action Plan, whose central strategy emphasized reducing regulatory constraints to maintain U.S. technological leadership. The plan included a dedicated focus on investing in AI-enabled science and called for collaboration between the federal government and the private sector to reduce bureaucratic barriers to new-materials discovery, drug development, and scientific database construction. Notably, the plan explicitly called for encouraging open-source and open-weight AI, reflecting recognition of the role of open standards in shaping global academic and commercial ecosystems. At the same time, some U.S. policy discussions have increasingly focused on the country's domestic STEM talent pipeline. Sociological research has examined concerns about declining interest in advanced mathematics and science education in parts of the U.S. education system. The development of models such as DeepSeek has also intensified debate over the limitations of relying primarily on semiconductor export controls and industrial subsidies to maintain technological leadership.
China: Combining Technological Development with Engagement with the Global South. China has increasingly positioned open-source AI as both a tool of scientific and technological diplomacy and an engine for upgrading research capabilities. In education and research, the government has promoted an “AI + X” interdisciplinary model, encouraging universities to integrate AI with foundational disciplines such as biology, physics, and mathematics. At the international-governance level, China released the Global AI Governance Action Initiative at the World Artificial Intelligence Conference in July 2025. The initiative reaffirmed the importance of AI safety and controllability while proposing a 13-point action agenda. It also emphasized providing more inclusive access to AI computing resources and “AI+” industrial applications for countries in the Global South, promoting cross-border open-source communities, reducing technological barriers, and strengthening regulatory and capacity-building cooperation between China and ASEAN. The global adoption of open-source models such as DeepSeek has also become an important channel through which China's AI capabilities have gained international visibility.
European Union: Balancing Strategic Autonomy with Regulated Openness. In response to strategic competition between the United States and China, the European Union has pursued an approach that combines centralized coordination with strong regulatory oversight. In late 2025, the European Commission introduced initiatives focused on AI for science and applied AI and, through the Horizon Europe program, committed EUR 107 million to pilot the Resource for AI Science in Europe (RAISE) initiative. RAISE is envisioned as a pan-European virtual institution that would integrate computing capacity, specialized talent, and funding across member states, helping address fragmentation in the European research landscape. Within this framework, the European Open Science Cloud (EOSC) federation is expected to play an important role in providing high-quality, interoperable, and privacy- and ethics-compliant AI-ready datasets. Through initiatives such as Data Labs, the EU aims to ensure that AI-for-science models developed in Europe remain scientifically competitive while complying with the requirements of the EU AI Act and related data-governance and trust frameworks.
5.2.5 The Dual-Use Dilemma: Evolving Biosafety Risks and Hybrid Governance for Open-Source AI
If the compute divide represents a challenge of resource distribution in the economic sphere, the application of open-source AI to life sciences presents a distinct set of global security concerns. As AI systems become increasingly capable of predicting protein structures and generating novel biological sequences, technologies that could accelerate cancer research and the development of new antibiotics also exhibit significant dual-use potential.
At the Munich Security Conference (MSC) in 2025, the Nuclear Threat Initiative (NTI) and a group of experts conducted a high-level tabletop exercise. The exercise highlighted a significant challenge: existing international biological-weapons conventions, disease-surveillance mechanisms, and data-tracking systems may have limited ability to respond to open-source biological foundation models that can be distributed rapidly through platforms such as GitHub and Hugging Face.
A consensus study by the National Academies of Sciences, Engineering, and Medicine (NASEM), conducted in response to Executive Order 14110, likewise examined concerns that AI models trained on large publicly available pathogen-genome and omics datasets could lower the barriers to designing infectious biological threats with potentially serious public-health consequences. Historically, sophisticated work involving highly pathogenic viruses required multidisciplinary expertise and specialized laboratory infrastructure. As AI-assisted biological design becomes more accessible, however, researchers and policymakers have raised concerns that malicious actors with basic synthetic-biology capabilities could potentially use unrestricted models to explore viral mutations that evade existing vaccines or antiviral drugs.
In response to the tension between scientific openness and security, policymakers and researchers are exploring forms of hybrid governance. Completely restricting biological AI models could limit legitimate research aimed at responding to emerging diseases and could concentrate scientific capabilities in a small number of institutions. Conversely, unrestricted release without appropriate safeguards could create significant biological-security risks.
As a result, leading AI and life-science research teams have begun exploring built-in safety mechanisms within biological models. During the development of the previously discussed Evo 2 model, researchers at Stanford and the Arc Institute reportedly took a highly cautious approach to data curation, excluding viral genome sequences from its training dataset of approximately 9 trillion nucleotides. Such data-filtering measures are intended to reduce the risk that an openly available model could be prompted to generate sequences associated with novel pathogens or increased pathogen virulence.
In parallel, researchers and open-source communities are working toward standardized biological-security evaluation frameworks, including the Virology Capabilities Test (VCT). Such multimodal benchmarks can be combined with red-team exercises to assess whether a biological model could provide actionable assistance for the design or modification of synthetic pathogens before its public release.
A more mature governance approach under discussion is Managed Access. Under this model, a model's architecture, research papers, and general capability evaluations would remain publicly accessible in accordance with open-science principles. However, high-risk capabilities—such as generating specific pathogen sequences or predicting virulence-associated sites—would require additional controls, including verified identity, authorization, and research-ethics review. The goal is to preserve the openness and collaborative benefits of scientific research while establishing targeted safeguards against malicious use.
6. AI Applications
The AI application ecosystem in 2025 is undergoing a structural transformation. Enterprise AI strategies are shifting from passive adoption toward proactive ecosystem development, requiring underlying models to provide not only strong reasoning capabilities but also support highly fragmented deployment environments in the physical world.
6.1 Core Trend: The Pragmatic Reality of Agentic AI
In 2025, LLM applications continued to evolve toward agentic systems. Looking beyond the hype, large-scale empirical studies such as Measuring Agents in Production (MAP) suggest that production-grade agents exhibit a distinctly pragmatic engineering profile:
Deterministic workflows take precedence over unrestricted autonomy: Production systems generally have limited tolerance for uncontrolled model behavior. Approximately 68% of agents reportedly execute no more than 10 autonomous steps before requiring human verification. Many enterprises have therefore moved away from purely iterative ReAct loops toward explicit graph-based workflows that impose tighter constraints on agent behavior.
Multi-model routing is becoming increasingly common: Approximately 59% of production agents reportedly use routing architectures. Simple tasks such as intent classification and tool calling can be assigned to low-cost open-source models such as Qwen 2.5 8B, while computationally intensive tasks such as complex code refactoring and planning can be routed to more capable reasoning models such as DeepSeek R1.
Performance priorities are shifting toward higher latency tolerance and prompt-first optimization: For complex reasoning tasks, as many as 66% of applications reportedly tolerate response times on the order of minutes in exchange for greater reliability. As foundation models continue to improve, approximately 70% of applications rely primarily on static or dynamic prompt optimization, while fewer than one-third involve weight-level fine-tuning. In evaluation and oversight, approximately 74% rely heavily on human-in-the-loop processes.
6.1.1 The Landscape of Agent Frameworks and the “Intuition Test”
Major agent frameworks differ substantially in usability, degree of control, and enterprise-oriented capabilities, resulting in distinct architectural trade-offs across different use cases.
| Framework | Architectural Paradigm & Key Strengths | Best-Fit Production Use Cases | Developer Ecosystem Feedback |
|---|---|---|---|
| LangGraph | Models workflows as directed graphs with explicit state management and checkpointing/recovery, providing fine-grained control over agent execution. | Enterprise-critical workloads requiring high fault tolerance, multi-level approval workflows, and complex execution loops. | Widely adopted in large-scale deployments, with more than 38 million monthly downloads; however, it has a relatively steep learning curve. |
| CrewAI | Uses role-based multi-agent collaboration to decompose tasks and delegate responsibilities among agents. | Marketing campaign planning, multi-perspective code review, and content-generation pipelines. | Accessible to business users and easy to reason about, although its state-management mechanisms can become somewhat cumbersome in more complex workflows. |
| Agno | Natively integrates production-grade databases with knowledge retrieval, memory management, and tool calling. | Developing specialized multimodal assistants that require long-term memory and enterprise knowledge bases. | Clear documentation and well suited to development teams seeking an integrated solution with strong observability. |
| Google ADK | Modular and lightweight architecture with deep integration into the Google Cloud and Vertex AI ecosystem. | Large-scale concurrent task orchestration and RAG integration in Google Cloud-native environments. | Enables rapid deployment, but may introduce some degree of cloud-vendor lock-in. |
| Smolagents | A minimalist, code-first approach that enables models to “think in code” rather than relying primarily on JSON-based tool-call parsing. | Rapid local validation of tasks using lightweight open-source models from Hugging Face. | Popular among developers who prioritize minimal overhead, lightweight deployments, and transparency into the underlying execution process. |
Because underlying models can experience unexpected performance drift even after minor version updates, previously stable workflows may suddenly stop working. This exposes a key limitation of traditional benchmarks such as MMLU: they do not adequately capture reliability in dynamic, long-running tasks. This has contributed to the emergence of real-time monitoring platforms such as IsItNerfed, which track changes in model performance over time.
6.1.2 Compute Economics: The TCO Battle Across Hardware and Infrastructure
As AI moves toward widespread commercial deployment in 2025, software–hardware compatibility and total cost of ownership (TCO) are becoming decisive factors in determining which technical approaches are economically viable. The average price of high-quality open-source model APIs has reportedly fallen to approximately USD 0.83 per million tokens, compared with USD 6.03 for comparable closed-source models, fundamentally changing the economics of enterprise AI applications.
Local Compute: High-Memory Macs vs. High-Bandwidth NVIDIA GPUs
The Capacity Advantage — Mac Studio: Thanks to Apple Silicon’s unified memory architecture, Mac Studio configurations with up to 512 GB of memory overcome the conventional VRAM capacity constraint, making it possible to load heavily quantized models with hundreds of billions of parameters—such as DeepSeek R1 671B—at a relatively attractive cost. However, its memory bandwidth of approximately 819 GB/s limits generation speed to around 17–18 tokens/s.
The Bandwidth Advantage — NVIDIA RTX 5090: Although a single RTX 5090 provides only 32 GB of VRAM, its memory bandwidth of up to 1,792 GB/s enables substantially higher throughput when running models that fit within its memory capacity, such as quantized 70B-class Llama models. Generation speeds can exceed 200 tokens/s in suitable configurations, making it attractive for latency-sensitive and high-throughput workloads.
Hardware Selection Guidelines: For high-concurrency, low-latency RAG workloads and model fine-tuning, NVIDIA’s CUDA ecosystem is generally the more suitable choice. For large, complex batch workloads where latency is less critical, or deployments where minimizing upfront CapEx is a priority, high-memory Mac systems can be a cost-effective alternative.
Baseline Local VRAM Requirements for Leading Open-Source Reasoning Models (Estimated at FP16):
| Model | Base Architecture | Estimated Minimum VRAM | Approximate Enterprise Hardware Cost |
|---|---|---|---|
| DeepSeek R1 14B | Qwen-based distilled architecture | ~28 GB | ~USD 800 (high-end consumer GPU) |
| DeepSeek R1 32B | Qwen-based distilled architecture | ~64 GB | ~USD 1,600 (dual high-end consumer GPUs) |
| DeepSeek R1 70B | Llama-based distilled architecture | ~140 GB | ~USD 3,200 (multi-GPU small rack server) |
| Llama 4 Scout 17B/109B | Native MoE architecture (16 experts) | ~218 GB | ~USD 4,800 (enterprise high-density AI workstation) |
Note: Advanced quantization techniques such as W8A8 and INT4 can reduce VRAM requirements by more than 50% while retaining most of the model’s original performance.
6.1.3 Cloud Deployment and Compliance Strategy
Cloud-Native Optimization: The TPU ecosystem leverages JetStream and MaxText to enable tensor sharding across multiple hosts, while the GPU ecosystem relies heavily on inference engines such as vLLM, which uses PagedAttention to improve memory efficiency and serving performance.
Mixed Routing Pipeline: Given the substantial differences in API costs and throughput across models, enterprises can use Llama 4 Scout for large-scale document retrieval and summarization, while routing complex logical decision-making to DeepSeek R1 at the final stage of the pipeline. This model-routing strategy can help optimize the overall cost–performance trade-off.
Compliance and Geopolitical Considerations: Model evaluation is increasingly expanding beyond technical benchmarks to include legal approval processes, data sovereignty, and broader geopolitical risk assessments. Enterprises can deploy private computing infrastructure and fine-tune open-source models such as DeepSeek, Llama, and Qwen to build isolated, privately controlled AI infrastructure, reducing dependence on external APIs.
6.1.4 Evolution of the Model Ecosystem and Underlying Protocols
In 2025, the cost of model inference reportedly declined by as much as 280×, while the performance gap between leading open-source and closed-source models narrowed to less than 1.7%.
Milestone Open-Source Models
| Model | Total / Active Parameters | Context Window | Open-Source License | Key Industrial Advantages and Positioning |
|---|---|---|---|---|
| Qwen3 (235B-A22B) | 235B / 22B | Not publicly disclosed | Apache 2.0 | Strong multilingual capabilities; well suited for general-purpose knowledge management in multinational enterprises |
| Mixtral 8×22B | 141B / 44B | 64K | Apache 2.0 | Strong cost–performance efficiency, making it suitable for general-purpose computing workloads |
| DeepSeek-V3 (R1) | 671B / 37B | 128K | DeepSeek License | Strong logical reasoning, mathematical, and coding capabilities; uses reinforcement learning directly to elicit chain-of-thought reasoning without conventional SFT |
| Llama 4 Scout | 17B (16 experts) | Long context | Community License | Multimodal processing across edge and cloud environments; well suited for analyzing industrial engineering drawings |
| Grok-1 | 314B / 78.5B | 8K | Apache 2.0 | Large parameter capacity, suitable for high-density processing of private and domain-specific knowledge bases |
MCP: The “TCP/IP Moment” for AI
A key breakthrough in agentic AI is the decoupling of an agent’s internal capabilities (Skills) from its external interfaces and tools.
Model Context Protocol (MCP): Developed by Anthropic, MCP provides a standardized protocol for interaction between AI models and external systems, including databases, GitHub, and Manufacturing Execution Systems (MES).
Token-Efficient Tool Use: AI agents are increasingly adopting a code-execution-based tool-use paradigm. Instead of loading complete tool schemas into the context window, models can dynamically retrieve the schemas of specific tools when needed. This approach can substantially reduce token consumption and help prevent context-window overload.
Manufacturing: Unified Namespace (UNS) and Integrated Industrial Data
Data Integration: Manufacturing organizations are increasingly adopting a Unified Namespace (UNS) architecture based on the MQTT Sparkplug B protocol to integrate data across IT and OT environments, reducing reliance on fragmented and heavyweight integration architectures.
Predictive Maintenance (PdM): Lightweight vision-language models (VLMs), such as SmolVLM, can run at the edge and jointly analyze vibration and temperature sensor data with historical time-series data processed by specialized models. This enables high-frequency data collection and automated work-order generation based on detected anomalies.
| Key AI Use Case | Key KPI Improvements | Typical ROI | Payback Period |
|---|---|---|---|
| 24/7 Predictive Maintenance | 20–30% reduction in unplanned downtime; improved Overall Equipment Effectiveness (OEE) | 300%–500% | 6–12 months |
| Automated Visual Quality Inspection | Near-100% defect detection; significantly reduced inspection time | 200%–300% | 9–15 months |
| Intelligent Supply Chain Management | 10–20% reduction in logistics costs; improved inventory turnover | 150%–250% | 12–18 months |
| Automated Compliance Reporting | Saves employees 40–60 minutes of data-processing time per day | ~367% | ~26 months |
Healthcare: Pillar-0, a 3D Medical Imaging Foundation Model
Breaking the 2D Imaging Constraint: The open-source Pillar-0 model introduces a multi-window architecture and the Atlas hierarchical vision backbone, overcoming the conventional practice of reducing medical imaging to 2D slices. By processing 3D medical images natively, including CT and MRI scans, the model can preserve richer spatial and contrast information across the volume.
Data Privacy and On-Premises Deployment: Pillar-0 can be deployed on edge servers within a hospital's private network and integrated with PACS systems through the DICOMweb protocol. Federated learning enables privacy-preserving collaborative training across hospitals and geographic regions without requiring institutions to share raw patient data.
Energy and Digital Grids: An Intelligent Control Layer for Renewable Energy Variability
Accurate Load Forecasting: Low-cost, non-intrusive smart meters built on open-source microcontroller platforms can be combined with open-source ensemble algorithms such as CatBoost to achieve highly accurate short-term electricity demand forecasting.
Virtual Power Plants (VPPs): Through MCP-based integration with cloud-based AI agents, VPPs can coordinate tens of thousands of smart devices and trigger automated demand-response actions in near real time, helping smooth peak–valley fluctuations in electricity demand. Technologies such as zero-knowledge proofs can further enhance privacy protection and transaction-settlement security.
Precision Agriculture: Reimagining Aerial Autonomy and Edge Vision
Bridging the Sim-to-Real Gap: Open-source aerial autonomy stacks based on ROS 2, such as the aerial-autonomy-stack, support efficient hardware-in-the-loop (HIL) simulation across the full software stack, substantially accelerating algorithm development and iteration for agricultural drone fleets.
High-Performance Edge Vision with YOLO26: By eliminating computationally expensive post-processing steps such as Non-Maximum Suppression (NMS), an end-to-end architecture can significantly accelerate inference. In offline environments, edge vision models can estimate crop health and identify invasive species with high precision, reducing pesticide use and labor costs.
Deployment Bottlenecks in the Physical World and Strategic Outlook
As enterprises move toward end-to-end automation, they continue to face significant challenges in physical infrastructure and system integration:
Overcoming Power and Thermal Constraints: Power availability and thermal management are increasingly becoming critical constraints alongside compute capacity. Data centers and edge computing facilities are evolving into important components of regional energy infrastructure. Future industrial competition will increasingly depend on the ability to tightly coordinate computing capacity, power supply, and cooling through integrated energy and computing systems.
The Challenge of Edge AI Orchestration: Deploying AI applications across heterogeneous edge nodes with intermittent network connectivity—such as offshore platforms and logistics vehicles—is highly challenging. Open-source orchestration platforms designed for distributed environments could help address this problem through capabilities such as differential-weight updates, network-independent graceful degradation, and self-healing mechanisms. Building robust orchestration infrastructure for these environments will be essential for deploying AI reliably at scale.
Concluding Remarks
Looking back at 2025, China's open-source AI ecosystem underwent a significant transformation, evolving from primarily adopting and adapting existing technologies toward making increasingly visible contributions to the global AI ecosystem. The year saw DeepSeek demonstrate how advanced reasoning capabilities could be achieved at comparatively low cost, the performance gap between leading open-source and closed-source models narrow substantially, and a broader shift in development practices from prompt engineering toward context engineering.
At the same time, these advances have highlighted important limitations. The “Lost-in-the-Middle” problem in ultra-long contexts, divergent architectural approaches to multimodal integration, the difficulties exposed by real-world evaluation of embodied AI, and the reproducibility challenges facing open science all point to the same conclusion: the path toward artificial general intelligence remains long, and the open-source community must continue to balance efficiency and safety, innovation and compliance, and openness and data sovereignty.
Looking ahead from 2026, several questions will continue to shape the evolution of open-source AI: Can native multimodal architectures achieve the level of precision required for industrial applications? Can embodied AI overcome the Sim-to-Real gap? Can AI governance frameworks become sufficiently institutionalized while leaving adequate room for innovation? And how will the relationship between open-source and closed-source AI evolve?
The answers may not lie in any single technological path. Instead, they will emerge through the collective experimentation and collaboration of open-source developers around the world. China's open-source community, with its emphasis on practical engineering and rapid innovation, is making increasingly visible contributions to the global development of AI. The global adoption of projects such as DeepSeek demonstrates how technological advances can translate into widely accessible and reusable capabilities. The future of open-source AI will ultimately be shaped by the developers, researchers, and organizations who build upon and contribute to this ecosystem.
2025 COSR