Speaker
Description
Artificial Intelligence (AI) is central to scientific research, but model performance is fundamentally constrained by data readiness. The FAIR principles govern findability, accessibility, interoperability and reusability of research data, but not dataset properties that matter for AI like task-specific completeness, feature relevance, label quality or class imbalance. Existing frameworks offer data readiness levels, consolidate dimensions and metrics, identify domain-specific pipelines and link readiness to fairness; tools evaluate published datasets against selected criteria and fairness; metadata standards offer standardized documentation and machine readability. But a shared definition of data readiness for AI, comprehensive, machine-actionable criteria and guidance across the entire data lifecycle as well as domains is still missing. The Data2AI Compass closes this gap. It delivers a definition of data readiness for AI and machine-readable criteria catalog developed by an expert working group across domains. The Catalog spans four dimensions (data quality, task fit, governance & documentation, intended use), pairs a domain-agnostic core with domain-specific extensions and is complemented by modular training materials, enabling transfer across domains and actionability in everyday research practice. The Data2AI Companion operationalizes the catalog as automated, versioned tests. Datasets are read through modality adapters (tabular, image, text). Users specify their intended use (e.g. domain, task, model class), per-metric results are visualized, and a conversational interface explains the findings and proposes remediation. Results export as a readiness report, machine-readable Croissant metadata file and Hugging Face dataset card. A modular metric class and YAML-based metric configuration enable community contribution.