Chinese Academy of Sciences Launches AoCang Scientific Corpus
The Chinese Academy of Sciences launched the AoCang S&T Corpus on July 17 at the 2026 World Artificial Intelligence Conference in Shanghai. Led by its National Science Library and built with nearly 100 institutions, the resource draws on more than 320 petabytes of scientific data, 150 million papers, 120 million patents and 5 million books to supply AI-ready material across scientific disciplines.
The Chinese Academy of Sciences launched the AoCang S&T Corpus on July 17 at the 2026 World Artificial Intelligence Conference in Shanghai. The National Science Library of the Chinese Academy of Sciences led the project with nearly 100 participating institutions and experts.
The official academy announcement describes AoCang as a national scientific-data infrastructure designed to turn heterogeneous research material into forms that AI systems can process. Its source repositories exceed 320 petabytes of scientific data and include 150 million research papers, 120 million invention patents and 5 million scientific books. The material spans text, images, charts, formulas, spectra, waveforms and other scientific formats.
How the corpus is organized
AoCang uses a three-stage refinement model covering foundational, feature-rich and knowledge-oriented corpora. Its current architecture combines one cross-disciplinary national collection, seven specialist collections and application-specific datasets. The academy says the coverage includes mathematics and physics, chemistry and chemical engineering, astronomy and space science, Earth science, life science and other fields.
The academy also says AoCang data has already been supplied to its Panshi scientific foundation model and is beginning to support general-purpose and domain models from organizations including ByteDance, Alibaba, iFlytek and Ant Group. Those statements describe deployment activity reported by the project operator; the announcement does not provide an independent benchmark, licensing terms or a public-access schedule.
What it means for AI research
The practical significance is the attempt to make scientific information more consistent and machine-readable before model training or retrieval. Scientific material is difficult to use at scale because it is distributed across institutions and encoded in formats that general web corpora do not handle well. A curated infrastructure can improve provenance and domain coverage, but its value will depend on access, documentation, quality controls and measured downstream performance.
AoCang is therefore best read as a substantial infrastructure launch rather than proof of a model-performance breakthrough. The scale figures establish the project's ambition; reproducible evaluations will be needed to show how much the curated data improves scientific reasoning in specific systems.
Key Points
- 1The Chinese Academy of Sciences launched AoCang on July 17 with nearly 100 participating institutions and experts.
- 2The corpus draws on repositories exceeding 320 petabytes of scientific data plus 150 million papers, 120 million patents and 5 million books.
- 3The operator reports early model integrations, but has not published an independent benchmark, licensing terms or a public-access schedule.
Scoring Rationale
A national-scale scientific corpus with broad institutional participation could improve AI-for-science data infrastructure, but access terms and independently measured model gains remain undisclosed.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

