8–9 Oct 2026
GEOMAR - Standort Ostufer / GEOMAR - East Shore
Europe/Berlin timezone

Data2AI Companion: A tool assessing data readiness for AI applications

8 Oct 2026, 17:15
1h 30m
5-1.214 - ATLANTIK / ATLANTIC - Linke Seite – Großer, unterteilbarer Konferenzraum (GEOMAR - Standort Ostufer / GEOMAR - East Shore)

5-1.214 - ATLANTIK / ATLANTIC - Linke Seite – Großer, unterteilbarer Konferenzraum

GEOMAR - Standort Ostufer / GEOMAR - East Shore

20
Show room on map

Speaker

Luca Greiner (Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany)

Description

Artificial Intelligence (AI) is central to scientific research, but model performance is fundamentally constrained by data readiness. The FAIR principles govern findability, accessibility, interoperability and reusability of research data, but not dataset properties that matter for AI like task-specific completeness, feature relevance, label quality or class imbalance. Existing frameworks offer data readiness levels, consolidate dimensions and metrics, identify domain-specific pipelines and link readiness to fairness; tools evaluate published datasets against selected criteria and fairness; metadata standards offer standardized documentation and machine readability. But a shared definition of data readiness for AI, comprehensive, machine-actionable criteria and guidance across the entire data lifecycle as well as domains is still missing. The Data2AI Compass closes this gap. It delivers a definition of data readiness for AI and machine-readable criteria catalog developed by an expert working group across domains. The Catalog spans four dimensions (data quality, task fit, governance & documentation, intended use), pairs a domain-agnostic core with domain-specific extensions and is complemented by modular training materials, enabling transfer across domains and actionability in everyday research practice. The Data2AI Companion operationalizes the catalog as automated, versioned tests. Datasets are read through modality adapters (tabular, image, text). Users specify their intended use (e.g. domain, task, model class), per-metric results are visualized, and a conversational interface explains the findings and proposes remediation. Results export as a readiness report, machine-readable Croissant metadata file and Hugging Face dataset card. A modular metric class and YAML-based metric configuration enable community contribution.

Authors

Luca Greiner (Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany) Mr Nico Fuhrberg (1Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany) Sonja Hänzelmann (Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany)

Presentation materials

There are no materials yet.