Speaker
Description
The increasing adoption of artificial intelligence (AI) and machine learning (ML) in Earth and environmental sciences places new demands on scientific data infrastructures that extend beyond traditional FAIR (Findable, Accessible, Interoperable, Reusable) principles. While FAIR provides a strong foundation for data sharing and reuse, it does not fully address requirements such as provenance transparency, bias documentation, uncertainty quantification, target annotation, and machine-actionable metadata that are critical for AI-driven workflows. In this contribution, we present a practical framework for assessing and improving AI readiness within distributed Earth and Environment data infrastructures, developed in the context of the Helmholtz E&E DataHub. The framework was derived through a review of existing AI-readiness concepts, analysis of community needs, and discussions among data providers, infrastructure architects, and AI practitioners. It defines a set of criteria organized into five dimensions: data preparation, data quality, documentation, access, and supervised-learning targets. Rather than prescribing specific technologies or maturity scores, the framework emphasizes transparency and fitness for purpose, enabling users to evaluate whether datasets are suitable for particular AI applications. The applicability of the criteria is illustrated through use cases ranging from analysis of environmental observations to data-driven forecasting. We argue that AI readiness is inherently use-case dependent and cannot be reduced to a single metric. Instead, transparent documentation of data characteristics provides a common basis for communication between data producers, infrastructure operators, and AI developers, supporting the development of interoperable and AI-capable scientific data spaces.