Artificial intelligence is driving materials research toward a new data−driven paradigm of AI for Materials (AI4M), with data being a key foundation for the success. Both the Genesis Mission program of U.S. and the AI for Science Strategy of U.K. identify the collaborative operation of high−quality research datasets and supercomputing as national strategic infrastructure. China's Overall Construction Plan for the New Materials Big Data Center also clearly aims to integrate national data resources to accelerate materials research using AI. Currently, the lack of high−quality, AI compatible "ready−to−use" data is considered as the biggest bottleneck in AI4M. It should be recognized that datasets are only a component of a larger whole of materials information since any dataset merely reflects the scientific understanding of its creators. Limitations and bias are inevitable. The datasets that are effective today may not necessarily be so in the future. This paper systematically analyzes new requirements that AI imposes on the content and form of materials data, and proposes a future−oriented materials data infrastructure. It should be structured like a "tree" with a comprehensive information data pool as the root and high−quality datasets/corpora as the leaves, which allows for sustainable evolution. By standardizing root resources in advance, raw data can be reused and repurposed to meet the evolving needs of specialized models and general large language models. A modular and assemblable data model is designed to address the standardization challenges caused by the diversity and variability of material data. Building a national−level materials big data infrastructure is a long−term endeavor, crucial to the future development of China's new materials industry. Forward−looking planning is required—not only to meet the current needs of creating quality datasets, but more importantly, to aim on potential future demand on data sources, leaving enough room for development in the years to come. Therefore, material data infrastructures with public welfare and service attributes should have an 'integrated pool and warehouse' structure, catering to both datasets and comprehensive materials information pools, to balance the present with the future.
| 科 Family | 属数 Number of genus | 种数 Number of species | 占总种数比例 Percentage of total species (%) | 属 Genus | 种数 Number of species | 占总种数比例 Percentage of total species (%) |
|---|---|---|---|---|---|---|
| 鹅膏菌科Amanitaceae | 2 | 11 | 5.26 | 鹅膏菌属 Amanita | 10 | 4.78 |
| 小菇科 Mycenaceae | 2 | 12 | 5.74 | 丝盖伞属 Inocybe | 5 | 2.39 |
| 多孔菌科 Polyporaceae | 8 | 14 | 6.70 | 蜡蘑属 Laccaria | 5 | 2.39 |
| 红菇科 Russulaceae | 3 | 23 | 11.00 | 小皮伞属 Marasmius | 6 | 2.87 |
| 小菇属 Mycena | 11 | 5.26 | ||||
| 光柄菇属 Pluteus | 5 | 2.39 | ||||
| 红菇属 Russula | 17 | 8.13 | ||||
| 栓菌属 Trametes | 5 | 2.39 |