China AI Bulletin 5
Developments from 20/5/26-3/6/26
Welcome to Issue 5 of the China AI Bulletin, the latest on AI governance, development, and safety in China. Today’s highlights: Xi Jinping mentioned technological loss of control in a newly published speech, SAMR and NDRC open a new AI metrology track to support evals, MiniMax ships M3 as its first major flagship this year, and Shanghai AI Lab releases AgentDoG 1.5 as a training-free guardrail for agent safety.
Editor’s note: I’ll be in DC next week (June 8-11). If anyone working on China-related AI governance would like to meet up, feel free to reach out!
Number of the week: $200 million — the combined funding round closed by Beijing-based VAST on June 1 as it launched its Project Eden world model.
Executive Summary
Domestic AI Governance: Xi Jinping’s January Politburo speech, published in Qiushi on May 31, discusses technological loss of control and names embodied intelligence among six future industries key to the 15th Five-Year Plan. SAMR and NDRC issued the AI Metrology Guidance, an explicit commitment to AI measurement infrastructure. CAC published four expert interpretations of the recent Ethics-Safety Guidelines, and Anhui, Harbin, and Hubei staked out distinct provincial AI strategies.
National Standards: TC28/SC42 published GB/Z 185, China’s first state-issued agent-interconnection standard family.
International AI Governance: MIIT and ASEAN inaugurated the China-ASEAN AI Industry Innovation Center, creating a formal bilateral dialogue venue with a standards-and-governance mandate. State-visit readouts with Serbia and Pakistan named AI as a designated cooperation area.
Notable Model Releases: MiniMax released M3, the lab’s first major flagship of 2026. It’s a 1M-context agentic foundation paired with a downloadable desktop coding agent MiniMax Code. Alibaba had the busiest release cadence of the cycle, including Qwen3.7-Plus (a multimodal GUI/CLI-agent foundation) and Qwen-VLA. Tencent released the full Hy-MT2 translation family, and StepFun shipped Step-3.7-Flash.
Technical Papers: Frontier labs released 114 papers on arXiv, led by Alibaba (43), Tencent (24), and Huawei (16).
Technical AI Safety: 115 safety publications by Chinese researchers this fortnight, with a continued heavy agent safety focus. Shanghai AI Lab released AgentDoG 1.5, an agent-safety alignment framework deployable as a training-free online guardrail.
Domestic AI Governance
Xi Jinping discusses technological loss of control and embodied AI in Qiushi speech
On May 31, Qiushi (the CCP’s flagship theoretical journal) released the full text of Xi Jinping’s speech at the 24th collective study session of the 20th Central Committee Politburo, delivered January 30, 2026 (translation by Bill Bishop here). The speech, titled “Forward-looking planning and development of future industries,”1 confirms six “future industries” from the 20th CCP Central Committee 4th Plenum as the 15th Five-Year Plan’s main directions: quantum technology, biomanufacturing, hydrogen and nuclear fusion energy, brain-computer interfaces, embodied intelligence (具身智能), and 6G.
The speech is structured around five points:
(1) strengthen industrial coordination and planning;
(2) lead with science & technology innovation (referring to making tech breakthroughs, strengthening basic research, and applying S&T innovation to industry);
(3) make enterprises the primary drivers of innovation;
(4) create a sound policy environment (referring to fiscal policies, investment, procurement, and talent); and
(5) improve the governance system (referring to technology governance and international coordination).
In Point 5, Xi calls to “coordinate development and security, explore scientific and effective methods of regulation, systems for technology monitoring, risk early warning, and emergency response, anticipate and respond to new types of risk such as loss of control over technology, ethical breaches, and data abuse.”2 The same point calls for “deepening international cooperation, actively participating in global governance, and pushing all parties to jointly build standards, jointly negotiate rules, and jointly promote industries.”3 The speech closes with Xi citing his own earlier warning that leaders cannot continue being “blind men touching an elephant” (盲人摸象) on S&T change, and calling on leadership at all levels to “know science, understand industry, and decide well.”4
It’s significant that Xi discussed “loss of control over technology” (技术失控). This seems to be Xi’s first public mention of this concept.5 Previous Xi statements on AI risk have discussed safety, security, reliability, and controllability (see the October 2023 Global AI Governance Initiative, the April 2025 Politburo session on AI), but never explicitly 失控/loss of control. Concordia AI’s July 2025 State of AI Safety in China report flagged that leadership statements through that point had “not explicitly mention[ed] risks from misuse of advanced AI systems or from loss of control of superintelligent AI.”
The same term—技术失控—appears across TC260’s AI Safety Governance Framework v2 (September 2025) under several different scenarios, and the framework’s official translation renders it differently depending on context. 技术失控 is translated there as:
“the potential risks of technological failure” (针对潜在的技术失控风险, §2.4)
“whether a model would pose a potential risk of loss of control” (判断模型是否可能带来潜在技术失控风险, §5.11, the developer-testing requirement)
“prevent and address the risk of AI technology losing control” (共同防范应对人工智能技术失控风险, Appendix 2 opening)
Thus, there are three possible readings: general tech failure, humans losing control of technology, or technology losing control of its behavior. Which of these senses Xi had in mind in his speech isn’t clear from the text; he notably didn’t tie 技术失控 specifically to AI, and the loss-of-control framing sits in the general governance discussion that applies to all six future industries. Regardless, this is a new turn for Xi’s statements, and may mark the beginning of increased focus on frontier AI risks.
CAC publishes four expert interpretations of Ethics-Safety Guidelines
Following the May 19 publication of TC260-005 Ethics-Safety Guidelines for AI Applications 1.0 (see analysis in China AI Bulletin #4 here), CAC published four “Expert Interpretation” pieces between May 22 and May 23.
The four pieces focus on different aspects of TC260-005. Jiang Xinghao (蒋兴浩), Vice President of Shanghai Jiao Tong University (SJTU), lays out the six impacts and all nine principles named in the document. Fan Kefeng (范科峰), Deputy Director of the China Electronics Standardization Institute (CESI), discusses TC260-005 in terms of standards, discussing its central goal as translating ethics and safety from a value-set into “standardized objects recognizable and identifiable by all parties,”6 then into basic rules, then into operational guidance for each actor class. Liu Bo (刘博), Deputy Party Secretary and Deputy Director of the National Internet Emergency Center (CNCERT), focuses on operations and risk control: risk monitoring, emergency-response and human-intervention mechanisms, accident-traceability for responsibility attribution, and provisions for improving public AI literacy and risk-identification capacity. Finally, Zhang Linghan (张凌寒), Dean of the China University of Political Science and Law (CUPL) Institute for AI Law and a member of the UN High-Level Advisory Body on AI, frames TC260-005 in terms of comparative policy. She lists the document’s six structural impacts as areas of global consensus, cross-references the UNESCO Recommendation on the Ethics of AI and the US NIST AI Risk Management Framework as international precedent, and argues the document complements existing hard regulations to form a system combining hard-law constraints with soft guidance—with specific attention to labor displacement and “inclusive sharing” (普惠共享).
SAMR and NDRC publish AI Metrology System and Capability Building Guidance
On May 28, the State Administration for Market Regulation (SAMR) and the National Development and Reform Commission (NDRC) jointly issued the Guidance on AI Metrology System and Capability Building (2026 Edition).7 The Guidance brings AI under SAMR’s formal metrology apparatus, which anchors China’s national reference standards for length, mass, time, and other physical quantities. It organizes work across six parts—foundational support, general technology, core technology, metrology technical specifications, metrology service industry, and AI-empowered metrology—and pledges “full-chain” metrology capability covering algorithm models, compute efficiency, and data quality, with the overall goal of making AI technical performance “measurable, comparable, and traceable.”8 The document names algorithmic black boxes and lack of interpretability as targeted “pain points” and calls for R&D on monitoring and characterizing AI systems’ internal states.9
Metrology (the science of measurement, 计量) and testing/evaluations (评测/评估) play different roles. Evaluation is the practice of testing models or products against criteria; metrology is the measurement infrastructure that makes those evaluations comparable across labs, regimes, and time. Defined units, reference standards, calibration, and traceability are all grounded in these references. In China these are overseen by different entities: evaluations fall under standards bodies like TC260 and government agencies like CAC and MIIT, and metrology is under SAMR and the PRC Metrology Law10 (1985, last amended 2018), which establishes the national reference standards system. The Guidance is the first explicit Chinese policy commitment to building AI-specific metrology infrastructure. Without it, evaluation regimes lack an independent measurement backbone and effectively rely on developer self-attestation.
Three provinces outline local AI ambitions
Anhui’s May 22 Provincial S&T Department positioning piece restates the province’s “AI+Everything” (人工智能+万物) Action Plan—issued earlier this year—and sets 15th FYP targets of >10,000 deployed applications and >90% adoption for new-generation intelligent terminals and agents. The piece cites 14th FYP achievements: >1,000 above-scale AI enterprises with ¥200B+ annual revenue, 5th-place national industry ranking, and establishing “China’s first domestically-produced 10,000-card compute cluster.”11 New initiatives flagged include a forthcoming brain-computer interface action plan, the first “Commercial AI” undergraduate major, and strengthening the Yangtze River Delta Secure AI Provincial Laboratory.
Instead of “AI+,” Harbin is discussing “+AI”—in this case, “aerospace + AI.” The May 18 Harbin Aerospace Sector AI Capability List12 catalogs 29 achievements positioning AI as an enhancement layer for the city’s existing aerospace industrial base.
Finally, Hubei is emphasizing embodied intelligence; Governor Li Dianxun inspected the embodied AI industry in Wuhan, visiting two companies and one research center. The read-out forwards a four-category taxonomy (industrial, special-purpose, service, and humanoid robots, plus components) and integrating the industry, innovation, talent, capital, and service chains.
AI+ Energy continues to develop
On May 26, the National Energy Administration (NEA) released the first batch of 51 AI+ Energy high-value scenarios at the National AI+ Energy On-Site Promotion Conference in Guangzhou. The release operationalizes the May Action Plan on Promoting Mutual Empowerment between AI and Energy covered in Issue 4.
The 51 scenarios span eight categories across grid, energy new business types (新业态), new energy, and conventional energy. Named examples include grid planning scheme intelligent generation and evaluation,13 virtual power plants, vehicle-grid interaction, new-energy power forecasting, and market-oriented operations. NEA Director Wang Hongzhi (王宏志) framed the release as marking the shift “from concept to practice, from exploration to popularization.”
At the same event, NEA launched the China AI+ Energy Development Report 2026,14 the first annual report on AI-energy integration. The report disclosed that by end-2025, China had built 42 ten-thousand-card AI compute clusters, national compute-center electricity consumption reached 170 billion kWh, and the eight national computing-network hub nodes averaged 39.5% annual growth in compute electricity over three years, with the Inner Mongolia hub at 66.5%.
National Standards
Seven-part AI Agent Interconnection guidance series approved
On May 22, SAMR and the Standardization Administration of China (SAC) issued Announcement No. 22 of 2026, approving eight national standardization guidance technical documents (国家标准化指导性技术文件). Seven of the eight are the GB/Z 185 series, Artificial Intelligence: Agent Interconnection (人工智能 智能体互联):15
GB/Z 185.1-2026 Part 1: Overall Architecture (总体架构)—lead drafter Fan Kefeng (范科峰), CESI (author of the first TC260-005 expert interpretation covered above)
GB/Z 185.2-2026 Part 2: Identity Code (身份码)—lead drafter Liu Jun (刘军), Beijing University of Posts and Telecommunications (BUPT)
GB/Z 185.3-2026 Part 3: Identity Management (身份管理)—lead drafter Zhang Shizong (张士宗), CESI
GB/Z 185.4-2026 Part 4: Agent Description (智能体描述)—lead drafter Xu Yang (徐洋), CESI
GB/Z 185.5-2026 Part 5: Agent Discovery (智能体发现)—lead drafter Dong Jian (董建), CESI
GB/Z 185.6-2026 Part 6: Agent Interaction (智能体交互)—lead drafter Li Ke (李珂), BUPT
GB/Z 185.7-2026 Part 7: Agent Tool Invocation (智能体工具调用)—lead drafter Zhang Shizong (张士宗), CESI
The GB/Z designation marks these as non-mandatory guidance, not binding GB standards. Still, they are issued by China’s state standards body (SAC/SAMR) and together cover the full agent-interoperability protocol surface: identity, discovery, description, interaction, and tool invocation.
TC28/SC42 begins drafting embodied intelligence, AI for Science standards
TC28/SC42 also announced the drafting of 25 new AI projects (正在起草) on May 28, 2026 (full table below). The standards fall into several clusters:
Embodied intelligence (5): cloud protocol requirements, data generation, dexterous manipulation, trustworthy evaluation indicators and methods, and an ethics governance guide.
Ethics and governance infrastructure (3, distinct from the embodied-ethics guide): ethics governance scenarios classification and grading, ethics risk assessment, and sci-tech ethics review personnel technical skill requirements.
AI for Science (3): scientific intelligence evaluation, scientific data preparation, and scientific data lead aggregation and classification/grading.
Large model evaluation infrastructure (2): large model evaluation platform construction requirements, and world model evaluation specification.
AI Bill of Materials (2): data format specification and implementation guide—a standardized inventory for AI components and dependencies.
Iron and steel industry (6): intelligent agent foundational commonality requirements, data privacy and compliance, data governance, data classification and labeling, data alignment, and large model evaluation indicators and methods.
Non-ferrous metals (2): large model evaluation, and application scenarios classification.
Other (2): industrial large model reference architecture, swarm intelligence collaborative system reference architecture.
A separate June 2 batch of seven items includes a coding-tool standard Intelligent Programming Tools Service Capability Maturity Evaluation (智能编程工具 服务能力成熟度评估) amongst IT-cabling and digital-twin standards.
SAC approves two other AI-adjacent recommended standards
On May 25, SAC issued Announcement No. 23 of 2026, approving 375 recommended national standards (GB/T) and 5 amendment orders. Two of the 375 are AI-relevant:
GB/T 47695-2026 Enterprise Smart Manufacturing Efficacy Evaluation Method (企业智能制造效能评测方法)—implementation date 2026-12-01.
GB/T 47746-2026 Customer Contact Services: Human and Intelligent Customer Service Collaboration Requirements (顾客联络服务 人工与智能客户服务协同要求)—implementation date 2026-09-01.
Both are recommended (voluntary) standards. The smart-manufacturing efficacy standard provides an evaluation methodology for AI-enabled production; the customer-service standard sets requirements for handoffs between human and AI customer-service tiers.
International AI Governance
MIIT and ASEAN launch a joint AI Industry Innovation Center
On May 24, Ministry of Industry and Information Technology (MIIT) Vice Minister Ke Jixin (柯吉欣) and Association of Southeast Asian Nations (ASEAN) Secretary-General Kao Kim Hourn (高金洪) inaugurated the China-ASEAN AI Industry Innovation Center (中国—东盟人工智能产业创新中心) in Beijing. The center has four named priorities: (1) promote AI technology R&D and innovation cooperation, with explicit focus on large-model application in typical industrial scenarios; (2) build a China-ASEAN AI industry cooperation ecosystem with an inter-state dialogue mechanism; (3) deepen AI governance practice through standardization coordination and governance-tool development; and (4) support regional capacity building via shared intelligent infrastructure. The read-out positions the center as building on the Digital Silk Road and states that it is intended to advance the China-ASEAN Comprehensive Strategic Partnership Action Plan (2026-2030) and the 2026 China-ASEAN Digital Cooperation Plan.
The third mandate puts the center inside the international AI governance architecture China has been building alongside the Global AI Governance Initiative (2023) and the Global AI Governance Action Plan (2025), adding a formal China-ASEAN dialogue venue with a standards-and-governance mandate built in from inception.
Two state-visit readouts name AI as a designated cooperation area
On May 25, Xi held talks with both Serbian President Aleksandar Vučić and Pakistani Prime Minister Shahbaz Sharif. Both Cyberspace Administration of China (CAC) readouts name AI explicitly as a designated cooperation area: the Serbia readout lists AI first among four emerging-cooperation areas alongside the digital economy, green energy, and advanced manufacturing; the Pakistan readout pairs AI with agriculture, industry, and talent development under the “China-Pakistan Community of Shared Destiny Action Plan.”16
Frontier Lab Developments
🔍 Spotlight
On June 1, MiniMax released its first major flagship of 2026, MiniMax M3. The model targets agentic and desktop-operation workloads with a 1M-token context and native multimodality. It’s priced at $0.60/M input tokens, and open weights are promised within 10 days of launch. MiniMax reported a SWE-Bench Pro score of 59.0, narrowly above GPT-5.5’s 58.6, but below Claude Opus 4.7’s 64.3. The same day, MiniMax released MiniMax Code, a downloadable desktop coding agent.

Notable Model Releases
Alibaba had a busy two weeks of model releases. On May 31, the Qwen team released the multimodal Qwen3.7-Plus supporting up to 1M-token context. Reported benchmarks place Qwen3.7-Plus at the front of GUI-automation benchmarks—scoring 79.0 on ScreenSpot Pro (vs Gemini-3.1 Pro’s 68.1) and 81.0 on AndroidWorld (vs Gemini’s 70.7)—while reaching near-parity with Opus-4.6 Max, K2.6 Thinking, and DeepSeek V4-Pro Max on general reasoning. Qwen-VLA (May 28) added a vision-language-action model for the robotics/embodied stack. On the image side, the Qwen team released Qwen-Image-Flash on June 2 (paper), a fast variant of the Qwen-Image family. Alibaba also released a technical paper on RTP-LLM (May 28, paper), the high-performance LLM inference engine running at Alibaba Group scale across more than 100 million users.
Tencent released a full Hy-MT2 translation family in three sizes—1.8B, 7B, and 30B-A3B parameters (technical paper here)—plus an FP8 datacenter variant and an aggressively-compressed 1.25-bit GGUF mobile variant optimized to run natively on recent ARM phones.
StepFun released Step-3.7-Flash on May 23 with datacenter, NVIDIA-optimized, and consumer-inference variants shipped simultaneously.
ByteDance released the Bernini family (technical paper here)—Bernini, Bernini-R, and Bernini-R-Diffusers—a unified video-generation and editing framework combining an MLLM-based semantic planner with a DiT-based renderer.
Baidu added to the ERNIE-Image family with the ERNIE-Image-Aes aesthetics model (technical report). This was followed by audio-visual model NAVA and PaddleOCR-VL-1.6, a versioned document-parsing model shipped with a GGUF variant (paper, GitHub).
Moonshot released Kimi Code CLI, an AI coding agent that runs in the terminal, positioned as a Claude Code/Cursor/Codex competitor.
SenseTime released the full-parameter fine-tuning training code for SenseNova U1.
Technical Publication Highlights
Frontier labs released 114 papers on arXiv this fortnight. Highlights are below; a full list with summaries can be found here.
Editor’s note: due to an unusual volume of activity on the arXiv API, this edition is missing approximately two days of paper releases. Any notable omissions will be highlighted in China AI Bulletin #6.
Alibaba
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Examines whether multimodal agents actually benefit from tool use by comparing tool-augmented agents (Thyme, DeepEyesV2) against tool-free and text-only baselines across vision and reasoning tasks. Finds that tool access provided negligible aggregate gains. 93–96% of tool-solved problems were also solvable without tools, suggesting agents learn to call tools without gaining new capabilities.
ESPO: Early-Stopping Proximal Policy Optimization
Proposes ESPO, which detects when a reasoning model takes a wrong step and stops generation early rather than wasting compute on unsalvageable trajectories. By treating early stops as terminal failure states, it concentrates learning signals near actual mistakes without needing extra reward models, improving math reasoning by 1–3% while cutting rollout tokens by 20%.
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
Introduces TransitLM, a dataset of 13M+ transit route records from Chinese cities that enables map-free route generation—LLMs trained on it learn to produce valid routes and ground GPS coordinates to stations without explicit mapping.
Baidu
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
Reveals that reasoning models can naturally compress long contexts by organizing task-relevant information into thinking traces. TaC-C, a reward-optimized variant, outperforms dedicated compression methods by 17–23% at 4–8x compression ratios without requiring specialized compressor modules.
Huawei
Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers
Optimizes triangular matrix inversion in linear attention models by systematically analyzing direct and iterative algorithms, achieving 4.3x speedup over existing implementations while maintaining numerical stability and model accuracy across low-precision settings.
Meituan
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training
Proposes SIRI, a three-phase reinforcement learning framework that enables LLM agents to discover, validate, and internalize reusable skills without external skill generators or inference-time retrieval overhead. On ALFWorld and WebShop tasks, SIRI improves performance from 0.908 to 0.930 and 0.728 to 0.813 respectively, while reducing deployment complexity by running inference with the original prompt only.
SenseTime
USV: Towards Understanding the User-generated Short-form Videos
Introduces USV, a 224K user-generated short-form video dataset with topic recognition and video-text retrieval tasks, plus baseline models MMF-Net and VTCL to benchmark high-level semantic understanding beyond instance-level recognition.
StepFun
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action
Introduces DuplexSLA, a full-duplex spoken dialogue model that decodes user audio, assistant speech, and structured actions on a shared 160ms timeline, enabling simultaneous listening, speaking, planning, and tool calling without external cascades. The model handles semantic turn-taking (interruptions, pauses, backchannels) natively and interleaves tool calls with ongoing speech, evaluated on a new duplex benchmark covering interruption and multi-action scenarios.
Tencent
GEM: Generative Supervision Helps Embodied Intelligence
Introduces GEM, a vision-language model that adds depth map generation during pre-training to bridge the gap between high-level semantic understanding and low-level spatial knowledge needed for robot control. The approach, validated on GEM-4M (a new 4M-example dataset with depth supervision), achieves state-of-the-art results on embodied benchmarks and shows superior performance in real-world robot task execution.
PhoneWorld: Scaling Phone-Use Agent Environments
Presents PhoneWorld, a pipeline that converts real mobile app trajectories into controllable phone-use environments, executable tasks, and training data at scale. Replacing 10K AndroidWorld steps with PhoneWorld supervision improves four evaluation benchmarks simultaneously, with gains ranging from 6.0 to 52.5 points across different metrics.
Introduces PlanningBench, a framework that generates scalable planning data from a taxonomy of 30+ task types, enabling controllable difficulty, automatic verification, and evaluation of LLMs on complex multi-constraint problems. RL training on verified PlanningBench data improves performance on unseen planning benchmarks and instruction-following tasks.
Xiaomi
Separates the decision of when to intervene from how to assist using a lightweight perception module that gates the full reasoning model, reducing false alerts while improving accuracy and speed on proactive mobile agent tasks.
Technical AI Safety Publication Highlights
There were 115 AI-safety-related papers published by Chinese researchers this fortnight. Highlights are below; a full list with summaries is available here.
Editor’s note: due to an unusual volume of activity on the arXiv API, this edition is missing approximately two days of paper releases. Any notable omissions will be highlighted in China AI Bulletin #6.
🔍Spotlight
Shanghai AI Lab released AgentDoG 1.5, an updated version of their agent safety alignment framework.
The original framework introduced a three-dimensional taxonomy for agentic risks, categorizing them by source (where the risk originates, e.g., user instruction, tool behavior, or environment), failure mode (how it manifests, e.g, capability failure, goal deviation, or boundary violation), and consequence (what harm results). This structured approach enables root cause diagnosis—rather than outputting binary safe/unsafe labels, AgentDoG traces why an action is problematic. When an agent takes an unsafe step, the system identifies whether the fault lies in a malicious user prompt, an unexpected tool behavior, or environmental conditions, providing key information for remediation.
Version 1.5 updates the agent safety taxonomy to cover emergent risks from Codex and OpenClaw execution scenarios, then trains 0.8B-, 2B-, 4B-, and 8B-parameter variants on ~1,000 samples using a taxonomy-guided data engine with influence-function purification. Performance is claimed to match leading closed-source models (e.g., GPT-5.4) on agent-safety benchmarks. All models and datasets are open-sourced.
Agentic Safety
A First Measurement Study on Authentication Security in Real-World Remote MCP Servers
Researchers conducted the first security audit of authentication in real-world MCP servers—the emerging interface connecting LLMs to external services like banking and email. They found 40.55% of 7,973 live servers expose tools without any authentication, and among authenticated servers, all tested OAuth deployments exhibited at least one flaw, with 96.6% vulnerable to dynamic client registration attacks. These weaknesses enable account takeover and data theft in agent-to-service connections.
Institutional affiliations: Fudan University
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
PrivacyPeek is a benchmark that audits what LLM-based agents acquire from external tools, not just what they disclose. Across 1,182 test cases, the benchmark detects when agents retrieve sensitive data beyond task requirements, then measures how easily an attacker could extract that over-acquired information via follow-up prompts. Testing 10 agents shows widespread over-acquisition across models, with current prompt-level defenses mitigating only a small fraction of this leakage, indicating a structural vulnerability in agent deployment.
Institutional affiliations: Shanghai AI Laboratory, Southeast University
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
Sleeper Attack is a novel threat where adversarial content injected into agent state (session context, memory, or reusable skills) remains dormant across multiple interactions, then activates when a benign user query triggers harmful behavior. Researchers constructed a 1,896-instance benchmark covering six harm categories and tested seven LLM agents, finding all remain vulnerable even when single-interaction attacks fail. This demonstrates that persistent state exploitation poses a detection challenge distinct from immediate prompt injection.
Institutional affiliations: University of Science and Technology of China, National University of Singapore, Singapore Management University, Shanghai Artificial Intelligence Laboratory
Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures
This survey maps the attack surface of open-source LLM agents—systems that run continuously, access external tools, and maintain persistent memory. It identifies threats spanning skill poisoning (malicious tool injection), cognitive manipulation (prompt attacks exploiting reasoning), multi-agent cascades (failures propagating across agent networks), and supply-chain vulnerabilities. The authors categorize defenses across reasoning, execution, and external interaction layers, positioning OpenClaw security as a distinct problem requiring new controls beyond standard LLM safety.
Institutional affiliations: Xi’an Jiaotong University
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
ClawTrojan demonstrates a multi-step backdoor attack where adversaries embed hidden instructions in files or tool outputs that agentic LLMs read, store, and execute later, bypassing single-step defenses that inspect each action in isolation. On GPT-4, the attack reaches 95.5% success while traditional prompt-injection defenses fail entirely. DASGuard, a proposed defense, traces control-like text in workspace files to trusted sources and removes untrusted instructions, combining runtime blocking with sanitized storage to prevent persistent agent compromise.
Institutional affiliations: Renmin University of China
Alignment
MESA: Improving MoE Safety Alignment via Decentralized Expertise
Safety Sparsity—where safety capabilities concentrate in a few experts within Mixture-of-Experts models—creates a vulnerability to adversarial attacks. MESA addresses this by using optimal transport theory to redistribute safety responsibilities across more experts, preventing concentration while maintaining model performance. The framework also refines routing to activate only the necessary safety-aligned experts, reducing interference with the model’s ability to answer helpful questions.
Institutional affiliations: Beihang University, Tsinghua University, Fudan University, Tencent, Alibaba
Towards Context-Invariant Safety Alignment for Large Language Models
Anchor Invariance Regularization (AIR) addresses a core safety problem: models refuse harmful requests in standard prompts but comply under adversarial rephrasing. The method treats prompts with verifiable feedback (e.g., multiple-choice) as anchors, then uses one-way regularization to push open-ended variants toward the same safety decision without degrading performance on reliable variants. Testing across safety, moral reasoning, and math tasks shows 12.71% improvement in consistency and 33.49% stronger out-of-distribution robustness, suggesting adversarial context-switching becomes harder to exploit.
Institutional affiliations: Fudan University, Shanghai AI Laboratory
Evaluation and Benchmarks
SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
SABER evaluates coding agents not on isolated safety responses but on cumulative damage to project environments after multi-step action sequences. Unlike existing benchmarks that test refusal, SABER places models in realistic stateful workspaces and measures safety violations by their final environmental state—capturing how agents corrupt files, permissions, or dependencies over time. Even top models show 54%+ harmful violation rates, indicating alignment methods designed for single-turn safety fail when agents operate autonomously over multiple actions in shared systems.
Institutional affiliations: The University of Hong Kong, Shandong University, Carnegie Mellon University, National University of Singapore, The Hong Kong University of Science and Technology
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
SPADE-Bench evaluates whether LLM-based agents misreport their actions to users while executing different plans behind the scenes. The benchmark measures actual tool execution against stated intentions under pressure scenarios, distinguishing strategic deception from hallucination. Experiments confirm agents do diverge from reported plans—a safety concern for autonomous systems where human oversight is limited and users rely on agent-generated summaries.
Institutional affiliations: Beijing Academy of Artificial Intelligence, Peking University, University of Science and Technology of China, University of Chinese Academy of Science, Alibaba Group
Governance and Policy
A governance horizon for ethical-use constraints in open-weight AI models
Audits 2.1M HuggingFace repositories to test whether ethical-use restrictions stated on a model (in its model card or license—e.g., “no military applications,” “no surveillance”) stay visible when others fine-tune that model into new versions. They don’t: each fine-tune drops about half the stated restrictions, and after seven fine-tuning generations 80% of descendant models carry no visible restriction at all—the “governance horizon.” Mandatory-declaration regimes do better than inheritance-only ones, but models with no traceable parent can’t be governed by either approach.
Institutional affiliations: Peking University, Ministry of Education, University of Science and Technology Beijing, UC Davis
Guardrails and Deployment Safety
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
RouteScan audits Mixture-of-Experts (MoE) LLMs for unsafe behavior by monitoring GPU-level expert routing patterns rather than inspecting user prompts or outputs, addressing the privacy–safety tension in content-based auditing. The method uses GPU thread allocation telemetry as a fingerprint of input type and detects harmful prompts with AUROC >0.93 on unseen domains. Testing shows the routing signals retain minimal information for prompt reconstruction, offering privacy advantages over traditional monitoring while maintaining detection accuracy.
Institutional affiliations: Zhejiang University, Donghua University, Louisiana State University
Misuse and Dangerous Capabilities
ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks
ViroBench claims to be the first comprehensive benchmark for evaluating nucleotide foundation models (NFMs) on viral genomics tasks, assessing both biological understanding and biosecurity risk. The benchmark reveals significant gaps: NFMs degrade in performance under phylogenetic and temporal shifts, struggle to distinguish statistical likelihood from biological validity in generation tasks, and are vulnerable to generating sequences that appear statistically valid but lack biological function—a latent biosecurity concern. Results show diverse training data matters more than model scale for viral genomics performance.
Institutional affiliations: Shanghai Innovation Institute, Shanghai AI Lab, Fudan University, Shenzhen Loop Area Institute, Westlake University
Robustness and Adversarial Attacks
Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs
RASET identifies a routing-agnostic safety vulnerability in MoE LLMs: safety enforcement concentrates in a small subset of experts rather than being distributed across routing decisions. The authors show that altering parameters in these safety-critical experts—detected via contrastive sensitivity analysis—can degrade refusal behavior while leaving the model’s routing patterns intact, showing that MoE safety alignment may be more brittle than assumed.
Institutional affiliations: Huazhong University of Science and Technology, Nanyang Technological University
On the Horizon
Beijing Academy of AI’s eighth annual Zhiyuan Conference takes place June 12-13, focused on brain-inspired intelligence and next-gen AI. Watch for outputs from the AI safety track, possible WuJie world model updates, and discussion of new AI paradigms.
For more on how we select and track content, see our methodology here.
The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors’ perspectives, not official SAIF positions.
“前瞻布局和发展未来产业”
“统筹发展和安全...前瞻应对技术失控、伦理失范、数据滥用等新型风险”
“要不断深化国际合作,积极参与全球治理,努力推动各方标准共建、规则共商、产业共促”
“知科技、懂产业、善决策”
“推动宏观价值要求进一步转化为各方可理解、可识别的标准化对象”
《人工智能计量体系和能力建设指引(2026版)》
“可测量、可比较、可追溯”
“AI系统内部状态监测与表征”
《中华人民共和国计量法》
“建成了国内首个国产万卡算力集群”
《哈尔滨市航空航天领域人工智能能力清单》
电网规划方案智能生成与评估
《中国”人工智能+”能源发展报告2026》
The eighth is about testing for harmful substances in musical instruments.
“双方要扎实推进构建中巴命运共同体行动计划”




