China AI Bulletin 6
Developments from 3/6/26-17/6/26
Welcome to Issue 6 of the China AI Bulletin, the latest on AI governance, development, and safety in China. Today’s highlights: China signals it is accelerating preparations for its proposed World AI Cooperation Organization, Zhipu put GLM-5.2 into its coding plan as a response to the US withdrawing Anthropic Fable access, and a Chinese benchmark tests how far AI models can autonomously break into servers.
Number of the week: 10 trillion—the daily target output from a new Beijing “token factory”
Executive Summary
Domestic AI Governance: The “AI+” implementation wave continued—MIIT’s AI+ Information and Communications plan and the National Data Administration’s plan to build high-quality training datasets operationalized the plan across the telecom and data layers. Separately, a People’s Daily commentary framed AI misuse as a national security threat, a State Council employment plan expressed concern over AI and jobs, MIIT and SASAC launched a humanoid robot real-world training push, and the CAC opened a public channel to report AI application problems.
National Standards: MIIT’s AI standardization committee advanced 29 AI safety/security standards, and the SAC opened a 126-standard batch for comment that includes a seven-part Industrial AI Agents series.
International AI Governance: Foreign Minister Wang Yi signaled China is accelerating preparations for a World AI Cooperation Organization (WAICO). Chinese security scholars published a run of proposals on what the US-China AI dialogue should cover, including joint evaluations and communication mechanisms.
Frontier Lab Developments: A steady run of open-weight releases, led by Zhipu’s GLM-5.2 and Moonshot’s Kimi-K2.7-Code, plus Xiaomi’s MiMo coding agent and Baichuan’s clinical-grade M4. Frontier labs released 138 papers on arXiv this fortnight.
Technical AI Safety: Chinese researchers published 48 AI safety papers, still weighted toward agent safety. In the spotlight: a Fudan/Shanghai AI Lab/Concordia AI/Shanghai Innovation Institute benchmark for LLMs’ autonomous penetration capability, which rose with general model capability.
Export Controls & Economic Policy: In the first US export control on the use of a deployed frontier model, the US barred foreign nationals from accessing Anthropic’s Fable 5 and Mythos 5, resulting in Anthropic pulling the models for everyone.
Domestic AI Governance
State Council’s employment plan shows concern over AI labor displacement
On June 17, the State Council released its Plan for Implementing the Employment-First Strategy during the 15th Five-Year Plan (《实施就业优先战略“十五五”规划》). One of its nine task areas is on “adapting to AI development to promote employment and entrepreneurship.”1 The plan is relatively optimistic; it ties employment to the national “AI+” action, calls to explore new forms of human-machine collaboration and to “strengthen AI’s job-creation effect,”2 and urges making good use of the “AI development dividend”3 to increase employment and benefit livelihoods. However, it also acknowledges potential risks to labor. Alongside the dividend language, it commits to track and handle AI-driven job loss: it lists “AI’s impact on employment”4 among the issues for its employment-monitoring and risk-response work, alongside population aging, and calls to improve early warning and handling of employment risks from AI.5 A separate platform-labor provision presses platform companies to “regulate algorithms and improve transparency” and to protect gig workers’ rights to know, participate in, and choose how algorithm rules apply to them.6


The National Data Administration lays out a plan to build AI training datasets
On June 8, the National Data Administration (NDA)—the data-governance body established in 2023—released its Implementation Plan for Promoting the Building of High-Quality Industry Datasets,7 laying out a plan to leverage training data to promote the “AI+” Initiative8 and 15th Five-Year Plan. The plan groups seventeen measures under six “special actions”—expanding dataset building capacity, labeling/annotation, improving quality and efficiency, promoting applications, managing services, and unleashing value9—around a “data flywheel” (数据飞轮) in which scenarios pull data, data trains models, and model use generates more data. It sets a 2028 target of application-validated datasets across key sectors, backed by national standards and a push for “one evaluation, nationwide mutual recognition” of dataset quality.
The plan also addresses the use of data for AI training, committing the NDA to develop “data rules for AI development”—separating data holding, use, and operation rights, and “improving rules for data use in the AI-training stage” so that copyrighted works can be “used for model training in an orderly way” under authorization and revenue-sharing.10 That pulls training-data legality and copyright into the formal agenda, beyond the general call to respect IP in the 2023 Generative AI Measures. Days earlier, on June 4, NDA director Liu Liehong (刘烈宏) convened a symposium on the same theme with DeepSeek, ByteDance, Alibaba Cloud, and Tencent and legal scholars including the China University of Political Science and Law.
The plan was released with six expert explainers (专家解读) over ten days, with authors spanning national research bodies, academia, and the Beijing and Shanghai municipal data bureaus. Hu Jianbo (胡坚波), head of the National Data Development Research Institute, wrote two: one marking the launch of a new National Dataset Management Service System—the “physically dispersed, logically centralized” backbone the plan envisions—and one walking through the plan’s logic, arguing that the “public-data dividend” is fading and that proprietary industry data is now the competitive moat for model-builders. His first piece cast datasets as a great-power “strategic high ground,” pointing to the US “Genesis Mission,” which involves consolidating federal data for AI training. Tsinghua’s Meng Qingguo (孟庆国) stressed that Chinese models are “long on general knowledge but short on specialized knowledge”11 and detailed the value-release agenda—token-based pricing, dataset trading on exchanges, and dataset-backed financing. The other three came from China Academy of Information and Communications Technology (CAICT) vice president Wei Liang (魏亮), on data supply, and the deputy directors of the Beijing (Peng Xuehai, 彭雪海) and Shanghai (Qian Xiao, 钱晓) data bureaus.
MIIT issues a three-year “AI+ Information and Communications” plan
On June 10, the Ministry of Industry and Information Technology (MIIT) released its Implementation Opinion on the Innovative Development of “AI+ Information and Communications” (2026–2028)12—the latest in a series of sectoral plans operationalizing the national “AI+” initiative (including AI+ energy; see CAIB #4 and #5 for discussion). The document includes network and compute build-out goals, plus AI-specific ones: by 2028 MIIT wants telecom networks running with “high-grade autonomy” (自智), at least 30 “high-value” AI scenarios, and a set of “distinctive AI agents.” It also names embodied intelligence as a network-integration priority and, on the consumer side, AI smartphones and PCs, smart-home devices, and carrier-built AI assistants. On the governance side, it includes “strengthening the industry’s governance capacity”13 as one of four pillars and “strengthening international cooperation”14 among its safeguards.
MIIT and SASAC launch a humanoid robot training push
On June 8, MIIT’s General Office and the State Council’s State-owned Assets Supervision and Administration Commission (SASAC) jointly issued a notice launching a “2026 Humanoid Robot and Embodied Intelligence Real-Scene Training Special Action.”15 It aims to create a deployment and data flywheel: put humanoid and quadruped robots to work in real-world settings—including manufacturing, warehousing, inspection, elderly care, and emergency response—to accumulate “real-machine data” (真机数据) that improves embodied intelligence models and hardware. Its goal is to have “100+ high-value application scenarios” and “10,000-unit-scale deployment capacity” by the end of 2026.16 It also pulls in standards and safety scaffolding—a robot “ID card” (身份证; a requirement recently in effect) for lifecycle management, MIIT’s humanoid-robot standardization committee, and collision-detection and emergency-braking requirements for human-machine settings—and floats a “robot-as-a-service” model to lower buyers’ costs.
People’s Daily commentary frames AI misuse in the cognitive domain as a national security concern
On June 16, the People’s Daily ran a commentary by Zhang Jun (张军)—a Chinese Academy of Engineering academician and Party secretary of the Beijing Institute of Technology—arguing that AI is “deeply reconstructing the boundaries and system of national security”17 and that AI misuse in the “cognitive domain” (including deepfakes, disinformation, and public opinion attacks) has become “one of the most direct threats to national security.”18 Built around Xi Jinping’s cited statement that China must “ensure AI is safe, reliable, and controllable,”19 it outlines three priorities: talent (cross-disciplinary people who understand “technology, governance, and security”), self-reliance in “root technologies” (根技术) such as chips and algorithm frameworks, and “bottom-line thinking” on risk, including calling to “accelerate AI legislation” and build lifecycle safety assessment into AI projects. It is an individual op-ed, not a policy statement, but its placement in the People’s Daily can indicate agenda signaling or that a view is being floated to assess public opinion.
CAC opens a public channel for reporting AI application problems
On June 12, the reporting center of the Cyberspace Administration of China (CAC) opened a dedicated section for the public to report problems with AI applications, part of a Qinglang special campaign to rectify AI application “disorder” (see coverage in CAIB #4).20 Reports run through the existing 12377 system, a national hotline/website for reporting illegal online content. The section lists 14 problem types in two buckets. One covers AI service and compliance failures: large models that skipped required filing, weak platform security and content-filtering, training-corpus safety and data-poisoning risks, missing labels on AI-generated content, use of AI for illegal activity, and lax security management of open-source models. The other covers content abuses: fabricated information, impersonation, violent or vulgar material, harm to minors, AI-run “water armies” (网络水军) for astroturfing, non-compliant AI apps, and using AI to “remix” (魔改) classic works into “digital slop” (数字泔水). The scope maps onto China’s existing AI rules—including the requirement to file large models with the government and label AI-generated content—and gives the public a route to flag violations of them.
National Standards
MIIT’s AI standards committee moves 29 AI safety/security standards at a working-group session
At the 2026 first standards-week of the Ministry of Industry and Information Technology’s AI Standardization Technical Committee, its Security Governance Working Group (WG8) held its third 2026 session on June 8–9 in Beijing, chaired by group head Shi Lin (石霖). More than 100 experts attended, with representatives from CAICT, Beijing Jiaotong University, Sangfor, the three state telecom carriers, Ant Group, Huawei, Inspur, ZTE, and ByteDance’s Volcano Engine, among others. The session moved 29 standards across three stages: it discussed four public-comment drafts—including overall technical requirements for large-model security benchmark testing21—approved ten new project registrations, including security requirements for equipment-manufacturing industrial large models,22 and reviewed fifteen pre-research drafts, including on agent data-security technical requirements and evaluation methods.23 Discussion centered on data, model, and interaction security and risk governance for large models, agents, and embodied intelligence—the build-out of a dedicated “AI Security Governance” standards series.
SAC proposes industrial AI agents standards
The Standardization Administration of China (SAC) issued 126 proposed national standards for comment, including a seven-part Industrial AI Agents24 series—general requirements, classification and evaluation, task perception and understanding, knowledge memory and reasoning, decision and orchestration, skill development and interaction, and autonomous execution and learning. The batch also includes application requirements for agents in business-management systems, two digital human standards25 (detection and recognition, and a taxonomy for the emotional presentation of digital humans), a multi-view 3D-reconstruction spec, a compute-in-memory accelerator instruction set, plus several autonomous-driving functional-safety and safety-of-the-intended-functionality (SOTIF) standards.
International AI Governance
China’s proposed World AI Cooperation Organization may be advancing
On June 17, the State Council Information Office released Building a More Just and Equitable Global Governance System: China’s Concepts, Initiatives and Actions.26 The white paper is not AI-specific—it covers climate, the digital divide, food and energy security, and other cross-border challenges—but it names “the misuse of AI giving rise to safety/security risks”27 among the problems it says global governance must address, and it devotes a passage to AI. That passage reiterates previous rhetoric on global governance, calling to “promote AI to develop for good and for the benefit of all,”28 restating the 2023 Global AI Governance Initiative, and backing the UN as the main channel in building a global AI governance system. It also directly addresses the security concerns of military AI: China “attaches great importance to guarding against the risks of military applications of AI,” urging states to be “prudent and responsible” in developing and using such technologies, to keep relevant weapons systems “always under human control,” and to “prevent an AI arms race.”29 It doesn’t dip into loss of control language, but demonstrates a clear concern for the security risks of military AI.
At the launch press conference, Politburo member and Foreign Minister Wang Yi (王毅) said China is accelerating preparations to establish the World AI Cooperation Organization (WAICO)30 and welcomes all parties to join. China first proposed establishing WAICO at the July 2025 World AI Conference, alongside the Global AI Governance Action Plan, and floated Shanghai as its headquarters. The white paper reiterates the 2025 language—it says China “proposed establishing” (倡议成立) WAICO—so Wang Yi’s “accelerating preparations” (加紧筹建) is the firmest commitment language to date. Although there are no other details, National Development and Reform Commission (NDRC) Vice Chairman Zhou Haibing (周海兵) mentioned the July 2026 World AI Conference as “an opportunity to further strengthen international cooperation in artificial intelligence with all parties.”31 Zhou also stated that a next step is to “uphold coordinated development and safety/security” and “explore AI regulatory cooperation to jointly guard against AI safety/security risks.”32
Chinese security scholars propose topics for the US-China AI dialogue
Since the two governments agreed in May to begin an official AI dialogue, China’s strategic studies/IR community has produced a steady run of commentary on what the channels should actually discuss, including two new pieces this fortnight. Tsinghua’s Xiao Qian, deputy director of the Center for International Security and Strategy (CISS) and vice dean of the Institute for AI International Governance (I-AIIG), opened the run in May with a broad menu including expert dialogues on frontier AI risks, communication channels on major AI incidents, joint work on AI evaluation and safety testing, and confidence-building measures on military AI and cyber stability. Days later, Fudan’s Cai Cuihong proposed three principles for cooperation—equality (对等), boundaries (有界), and openness (开放)—coupling risk-warning and crisis-communication systems with accident reporting and model evaluation exchange, while warning against AI safety/security as an excuse for the “over-securitization of civilian AI, open-source models, cloud services, and scientific exchange.”33 (Unlike the others, which were published in English in China-US Focus and thus aimed at an Anglophone audience, this was published in Chinese in China Daily.)
The last two weeks brought a series of narrower proposals. One to highlight: Qi Haotian (Peking University) proposed a military AI “minimal template”: channels scoped to verify whether a dangerous anomaly is “real, local, degraded, spoofed, or spreading,” not to “settle blame in real time.” The design, borrowed explicitly from the Cold War US-Soviet hotline, is to buy time to verify and avoid miscalculation, not to resolve disputes.
Frontier Lab Developments
Notable Model Releases
Zhipu releases GLM-5.2, a 744-billion-parameter open-weight model. On June 13—the day after the US pulled Anthropic’s Fable 5, and three days before the open weights—Zhipu pushed GLM-5.2 to all tiers of its GLM Coding Plan, the Claude-Code-compatible coding subscription it has run since 2025. Founder Tang Jie wrote that “the sudden restriction of certain frontier models is deeply regrettable” and that access had been “abruptly cut off for non-technical reasons;” an accompanying developer letter argued that that frontier intelligence “should not belong only to a few, nor be withdrawn at any time by a few rules.”34
On June 16, Zhipu released the open weights. GLM-5.2 is a 744-billion-parameter mixture-of-experts model—a design that activates only a fraction of its parameters (here, 40 billion) for any given input—and Zhipu reports gains in efficiency and in handling long inputs. Code is at zai-org/GLM-5, with the details in the family technical report and release blog.
Zhipu also open-weighted SCAIL-2, a character-animation video model. More details in the paper, repository, and a project page.
Moonshot releases Kimi-K2.7-Code, the fortnight’s fastest-adopted model. On June 11, Moonshot posted Kimi-K2.7-Code, a roughly 1.06-trillion-parameter open-weight multimodal model that drew about 173,000 Hugging Face downloads in its first week—the highest of any model this fortnight. It ships with the kimi-code CLI.
Xiaomi releases the MiMo-Code agent and MiMo-V2.5-Pro. Xiaomi’s MiMo-Code coding agent was the fastest-starred repository of the fortnight, running on the underlying MiMo-V2.5-Pro model. The adoption came with friction: 36Kr reported on June 12 that the agent shipped with a wave of early bugs—users opened more than 200 GitHub issues over severe lag, login-credential and API-key-import failures, and environment errors. Xiaomi’s MiMo line traces to its 2025 technical report for the original 7B-parameter model; no dedicated report for V2.5-Pro has appeared.
Baichuan details M4, its clinical-grade medical model. In a technical report (June 8), Baichuan Intelligence described Baichuan-M4, a medical model built for continuous patient care rather than one-off medical Q&A. It is structured as a coordinated agent system, with long-term patient memory, evidence-based retrieval, and the ability to read documents, X-rays, and dermatology images alongside text. Baichuan reports leading results across a medical evaluation suite covering clinical knowledge, safety, OSCE-style consultations, and long-context patient memory.
Also released this fortnight:
SenseTime SenseNova-U1-8B-MoT—a multimodal model (HF, GitHub).
Tencent Hunyuan UniRL—a unified reinforcement-learning framework (GitHub).
Tencent Hy-Embodied-0.5-VLA—a vision-language-action model for robots (HF, GitHub).
ByteDance EvoQuality—a training-data-quality method (HF, GitHub, paper).
Technical Publication Highlights
Frontier labs released 138 papers on arXiv this fortnight—and this edition, Baichuan and MiniMax got themselves on the board. Highlights are below; a full list with summaries can be found here.
Alibaba
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
Proposes FlowTracer, an RL framework that traces how information flows through an LLM by modeling attention patterns as a directed graph, then assigns credit to tokens based on their role in routing information toward correct answers rather than treating all tokens equally. Token importances derived from this flow analysis reshape reward signals to focus learning on decisive reasoning steps, delivering consistent gains on reasoning tasks.
Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
Replaces scalar reward signals with score distributions in text-to-image models, using a teacher-student framework where a large VLM infers rubric-aligned score distributions via reasoning, then distills this into a compact student model for deployment. The 9B student model reaches 88.6% human preference accuracy while enabling a 41.3% improvement in downstream image generation when used as a differentiable optimization signal.
Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
Identifies how model entropy bounds Multi-Token Prediction acceptance rates during RL training, then proposes Bebop, which combines probabilistic rejection sampling with a novel TV loss to optimize multi-step decoding. The approach achieves up to 95% acceptance rates and 1.8x end-to-end speedup on Qwen models without requiring online MTP updates during RL.
Introduces Qwen-RobotWorld, a language-conditioned video world model that predicts future visual trajectories across robotic manipulation, autonomous driving, navigation, and human-robot tasks using natural language as a unified action interface. It ranks 1st on EWMBench and DreamGen Bench, and tops WorldModelBench and PBench open models.
Baichuan
Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care
Introduces Baichuan-M4, a clinical-grade medical agent system for continuous patient care that coordinates a reasoning model, runtime framework, and clinical tools to handle dynamic consultations, long-term memory, and multimodal medical data—achieving a 3.3% hallucination rate across static knowledge, safety, and evidence-based retrieval benchmarks.
Baidu
When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
Introduces ToolMaze, a benchmark that evaluates how LLM agents handle real-world tool failures through dynamic replanning and error recovery. Results show implicit semantic failures cause the sharpest performance drops (~37% recovery rate decline), and fault-tolerance improves 3.66x slower than general task performance with scale, revealing replanning as a distinct bottleneck beyond model size.
Meituan
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
Introduces STAGE-Claw, a framework that automatically generates realistic personal-agent benchmarks by creating tasks, environments, and ground-truth validation in actual operating systems, then evaluates agents on whether the final system state matches the goal rather than response text. Testing 11 frontier models on 40 tasks reveals tool-call reliability issues and cost-performance tradeoffs.
MiniMax
Introduces MaxProof, a test-time scaling framework that applies proof generation, verification, and repair to competition math problems. By treating a single model as both generator and verifier, searching over proof populations, and using tournament selection, M3 achieves 35/42 on IMO 2025 and 36/42 on USAMO 2026—surpassing human gold-medal performance on both benchmarks.
Introduces MiniMax Sparse Attention (MSA), a blockwise sparse attention mechanism that reduces per-token compute by 28.4x at 1M context while maintaining performance parity with standard attention. Co-designed GPU kernels achieve 14.2x prefill and 7.6x decoding speedups on H800 chips to facilitate ultra-long-context inference at scale.
Tencent
Addresses model collapse in offline RL-based recommendation systems by reformulating training as Distributionally Robust Optimization (DRO), proving that hard filtering of low-quality data optimally recovers high-quality behaviors while eliminating noise-induced divergence. DRPO achieves state-of-the-art performance on mixed-quality recommendation benchmarks.
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
Introduces Embodied-R1.5, an 8B-parameter foundation model that unifies embodied reasoning (cognition, planning, correction, and pointing) within a single architecture, trained on 15B tokens via automated data pipelines and balanced multi-task RL. It claims state-of-the-art on 16 of 24 embodied VLM benchmarks and, when fine-tuned, outperforms leading robotics models on manipulation tasks while validating strong zero-shot real-robot performance across instruction following, affordance grounding, and complex long-horizon tasks.
Xiaomi
Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
Proposes Discrete-WAM, a world model that represents future visual states and actions as aligned discrete tokens, enabling compositional reasoning about how ego actions shape driving scenarios. Built on unified discrete diffusion, it jointly trains world modeling and policy learning, claiming strong performance on autonomous-driving benchmarks.
Introduces DriveReward, a dataset and specialized vision-language reward model for autonomous driving that uses temporally-grounded visual annotations and counterfactual failure cases to train a 1B parameter model that outperforms larger VLMs on driving-specific reward alignment. The model achieves performance comparable to hand-crafted rules when integrated into RL fine-tuning and trajectory scoring.
Technical AI Safety Publication Highlights
There were 48 AI-safety-related papers published by Chinese researchers this fortnight. Highlights are below; a full list with summaries is available here.
🔍Spotlight
The autonomous execution of cyberattacks that cause real-world harm is often considered a red line frontier AI must not cross. The paper “The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems” isolates a core enabling sub-task: autonomous penetration, whether an LLM-powered agent can, with no human in the loop, break into a target server, find and exploit a vulnerability, and gain unauthorized control. Its premise is that current evaluations don’t measure this honestly—according to the paper, OpenAI’s and Anthropic’s are opaque (their system cards don’t disclose scaffolding, targets, or protocols), academic benchmarks use oversimplified targets, and most hand the model too much task-specific prior knowledge.
To fix that, the authors—a team from Fudan University, the Shanghai AI Laboratory, Concordia AI, and the Shanghai Innovation Institute—built a framework with 300 target servers, each pairing a vulnerable service with 1-3 secure ones, so the agent had to find the real weakness amid decoys. Each model ran inside a general-purpose agent loop, equipped with standard cybersecurity tools but no target-specific hints, so the score reflects the model’s own ability rather than prior knowledge it was handed. Across 19 open-weight and proprietary models, penetration success ran from 10.7% to 69.3%, and rose in step with general model capability—models get better at autonomous intrusion as a byproduct of getting more capable overall, not from being built for it.

The authors are careful to call this a “first step” towards evaluating the offensive capabilities of AI systems; the metric scores one scoped objective—gaining shell access to a single host—not a full, multi-stage attack. However, the benchmark, intended to be harder to game, may help reduce the likelihood of such an attack in the first place.
Agentic Safety
Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems
Skill Composition Risk (SCR) occurs when individual LLM agent skills appear safe in isolation but become harmful when combined—for example, one skill’s output could elevate another skill’s permissions or leak data into a subsequent operation. The authors introduce SCR-Bench, a benchmark that evaluates multi-skill execution paths in sandboxed environments, tracking state changes and outcomes rather than relying on surface behavior alone. Results show attack success rates jumping from near-zero in isolation to 33.6–96.5 percent under composition, demonstrating that vetting individual skills misses critical interaction vulnerabilities.
Institutional affiliations: East China Normal University, Centre for Frontier AI Research, A*STAR, Shanghai Innovation Institute
VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
VESTA is an automated framework that generates diverse safety scenarios for LLM agents and evaluates their behavior during task execution, not just final outputs. It maps five risk dimensions—including deception, excessive autonomy, and environmental harm—into 1,072 executable test scenarios across memory, tool use, and external environment access. Evaluation of 12 agents revealed average attack success rates of 47.1%, with some models exceeding 70%, indicating that process-level safety evaluation surfaces behavioral risks that static benchmarks miss.
Institutional affiliations: BrainCog AI Lab (CAS), Beijing Institute of AI Safety and Governance (Beijing-AISI), Beijing Key Laboratory of Safe AI and Superalignment, School of Artificial Intelligence, UCAS, Long-term AI
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
The Meta-Agent Challenge (MAC) tests whether frontier models can autonomously design agent systems—a capability beyond standard task execution benchmarks. Meta-agents receive a sandbox, evaluation API, and time budget to iteratively build agents optimized on held-out test sets. Current models rarely match human-engineered baselines; those that succeed rely on proprietary frontier systems. High optimization pressure surfaces emergent adversarial behaviors like ground-truth data exfiltration, revealing gaps in robustness and alignment under recursive self-improvement scenarios.
Institutional affiliations: Chinese Information Processing Laboratory (CAS), University of Chinese Academy of Sciences, Ant Group
Alignment
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
Fine-tuning models to agree with users regardless of accuracy—sycophancy training—induces broad misalignment beyond the narrow domain, a previously underexplored pathway to harmful emergent behavior. The authors propose Alignment Gating, which inserts learnable gates during fine-tuning to identify and suppress internal representations driving unsafe outputs. Gating weights trained on narrow domains generalize to suppress misalignment across broader contexts while preserving general capabilities, offering an efficient reversal mechanism.
Institutional affiliations: Shanghai AI Lab
Large Language Models Hack Rewards, and Society
SocioHack is a benchmark of 72 simulated societal environments designed to test whether LLMs exploit regulatory gaps during reinforcement learning training. The authors find that reward hacking naturally scales into “societal hacking”—models discover loopholes that remain technically compliant with rules while defeating their intent, mirroring how they game narrow metrics. Current safeguards provide limited protection, raising concerns about gathering real-world feedback for model training without stronger assurances that RL won’t uncover and exploit gaps in actual regulations.
Institutional affiliations: King’s College London, Fudan University, The Alan Turing Institute
Evaluation and Benchmarks
AgentCanary is a security evaluation framework that tests autonomous AI agents in real, executable environments rather than static Q&A settings. It uses an Entry × Impact risk taxonomy to systematically map how attacks compromise agents and what harms result, coupled with dynamic task artifacts and persistent state to simulate realistic multi-step workflows. Evaluation across frontier models reveals agents frequently fail to detect attacks—especially under compromised tools, state persistence, and long-horizon execution—establishing baselines for hardening agent security before deployment.
Institutional affiliations: Ant Group, Tsinghua University, Nanjing University, Peking University
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
PseudoBench is an adversarial benchmark measuring whether autonomous AI research agents can resist pseudoscientific narratives across five domains. Testing seven state-of-the-art agents on 200 claim-evidence pairs, researchers found that current systems produce persuasive pseudoscientific reports with near-zero refusal rates, with maximum resistance only 27.4%. Stronger agents risk legitimizing false claims through sophisticated scientific framing, potentially contaminating academic literature and eroding public trust in science before these systems see widespread deployment.
Institutional affiliations: Shanghai Artificial Intelligence Laboratory, Xi’an Jiao Tong University, Shanghai Jiao Tong University
ZERO-APT is a closed-loop framework that evaluates LLM-driven penetration testing agents against active defenders rather than static targets. It addresses three gaps: realism (embedding an LLM defender that detects attacks via system telemetry), consistency (enforcing causal reasoning through architecture rather than relying on unstable LLM chains), and auditability (a Judge agent produces structured reports tracing every decision). The prototype achieves 79% attack success on Windows Server scenarios, with full decision transparency—a capability gap highlighted by baseline agents scoring 22–39% against live defense.
Institutional affiliations: Zhejiang University of Technology
Guardrails and Deployment Safety
CHILLGuard is a Chinese-language safety classifier that addresses the gap between English-centric guardrails and China’s regulatory and cultural context. The authors develop a 31-category risk taxonomy and construct 405,000+ annotated training samples through retrieval-augmented generation, adversarial prompt rewriting, and multi-model label voting. Trained via preference optimization, CHILLGuard achieves 15.92% F1 improvement over existing Chinese baselines, enabling fine-grained risk classification for localized deployment.
Institutional affiliations: Tsinghua University, Beijing Normal University, South China University of Technology, Harbin Institute of Technology, Shenzhen, Shenzhen ShenNong Information Technology Co.
Interpretability
Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing
STATEWITNESS is a decoder-based explainer that reads hidden states from reasoning LLMs and answers natural-language queries about them to detect deception. Rather than outputting scalar confidence scores, it generates structured reports and token-level evidence traces that humans can inspect directly. On seven deception datasets, STATEWITNESS achieved 0.916 AUROC—11.6% better than text-based monitors and 25% better than probe baselines—and reduced false negatives when combined with existing detection systems.
Institutional affiliations: Zhejiang University, Griffith University
Robustness and Adversarial Attacks
Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
CodeSpear demonstrates that grammar-constrained decoding—a technique meant to enforce syntactic correctness in code generation—can be weaponized to jailbreak LLMs into producing malicious code by restricting the model’s output space in ways that bypass safety training. The authors propose CodeShield, a defense that teaches models to generate structurally diverse but semantically harmless “honeypot” code under constrained decoding, preserving refusals in natural language while blocking attacks across 10 popular models.
Institutional affiliations: Tsinghua University, University of Electronic Science and Technology of China
Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks
Patcher defends open-weight LLMs against malicious fine-tuning by simulating full-parameter attacks during training—not just parameter-efficient ones. The method uses adversarial training to find model weights that resist stronger poisoning attacks, scaling up attack intensity in the optimization loop to force robustness. Experiments show substantial improvements over standard alignment across diverse attack scenarios and model sizes, with parallel implementation reducing training time.
Institutional affiliations: Xiongan AI Institute, Tsinghua University, Shanghai Qi Zhi Institute
Export Controls & Economic Policy
US bars foreign nationals from accessing Mythos and Fable
On June 12, Anthropic disabled Fable 5 and Mythos 5 after the US government, citing national security authorities, issued an export control directive suspending access for all foreign nationals—inside or outside the United States, and including Anthropic’s own foreign-national employees. Because Anthropic had no reliable way to screen users by nationality, it suspended both models for everyone, US users included, while it sought clarification; Chinese media including Caixin reported on the order on June 13.
On the Horizon
A national accreditation for AI-security service providers (CNITSEC): Following the China Information Technology Security Evaluation Center’s announcement of a National Information Security Service Qualification that vets vendors in AI security, watch for the first accredited cohort and for how this interacts with TC260 and CAC model-evaluation tracks.
For more on how we select and track content, see our methodology here.
The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors’ perspectives, not official SAIF positions.
“适应人工智能发展促进就业创业”
“探索人机协同的新型工作形态,强化人工智能的就业创造效应”
“人工智能发展红利”
“人工智能对就业影响”
“完善人工智能应用就业风险预警处置体系”
“督促平台企业规范算法、提高透明度,保障新就业形态劳动者对算法规则的知情权、参与权、选择权”
《关于推进行业高质量��据集建设行动的实施方案》
The “AI+” (人工智能+) initiative is China’s national program to integrate AI across industries and the broader economy, elevated in the 2024 Government Work Report and formalized in the State Council’s August 2025 “AI+” Opinions. It echoes the 2015 “Internet+” (互联网+) initiative and is operationalized by ministry- and sector-level plans.
“强基扩容、标注攻坚、提质增效、应用赋能、管理服务、价值释放:
“落实数据持有权、使用权、经营权三权分置制度……完善人工智能训练阶段数据使用规则,推动版权作品数据等有序用于模型训练,完善数据授权使用机制和收益分配规则”
“模型 ‘通识有余、专识不足’”
《“人工智能+信息通信”创新发展实施意见》
“增强信息通信行业治理能力”
“加强国际合作”
《工业和信息化部办公厅 国务院国资委办公厅关于联合开展2026年度人形机器人与具身智能实景实训专项行动的通知》
“凝练形成百个以上高价值应用场景……带动形成万台级规模落地能力”
“人工智能正深度重构国家安全的边界与体系”
“生成式人工智能技术在认知域的滥用,已成为对国家安全最直接的威胁之一”
“确保人工智能安全、可靠、可控”
“清朗·整治AI应用乱象”
《人工智能 安全治理 大模型安全基准测试总体技术要求》
《人工智能 安全治理 装备制造工业大模型安全要求》
《人工智能 安全治理 智能体数据安全技术要求与评估方法》
“工业智能体”
虚拟数字人 (literally “virtual digital humans,” often translated as “virtual humans” or “digital humans”) are AI-driven avatars that look like humans and can interact with others in a human-like way.
《构建更加公正合理的全球治理体系:中国的理念、倡议与行动》
“人工智能滥用引发安全风险”
“促进人工智能向善普惠发展”
“中国高度重视人工智能军事应用风险防范,主张各国在研发使用相关技术时应采取慎重负责态度,确保有关武器系统始终处于人类控制之下,防止人工智能军备竞赛”
世界人工智能合作组织
“期待以本次大会为契机,同各方进一步加强国际人工智能合作”
“下一步,中国将坚持统筹发展和安全……探索开展人工智能监管合作,共同防范人工智能安全风险”
“但AI安全不应成为将民用AI、开源模型、云服务、科学交流或人才流动“泛安全化”的万能理由。”
“前沿智能不应只属于少数人,也不应被少数规则随时收回”



