Data Science Symposium 2026

Europe/Berlin
5-1.214 - ATLANTIK / ATLANTIC - Linke Seite – Großer, unterteilbarer Konferenzraum (GEOMAR - Standort Ostufer / GEOMAR - East Shore)

5-1.214 - ATLANTIK / ATLANTIC - Linke Seite – Großer, unterteilbarer Konferenzraum

GEOMAR - Standort Ostufer / GEOMAR - East Shore

20
Show room on map
Gauvain Wiemer (GEOMAR), Timm Schoening
Description

Welcome to the 2026 Data Science Symposium, organized on October 8th & 9th at GEOMAR in Kiel. This 11th Data Science Symposium is the next annual event of a series of symposia organized by Helmholtz Earth & Environment centres. We invite you to join us for presentations, workshops and lively networking discussions on the overarching topics of AI, data science and data infrastructures in the earth sciences.

Wen plan to host talks, poster presentation and reserve ample time for in-depth discussions during breaks and targeted workshop slots. Bring your topic to be discussed, advanced or solved with peers from Earth & Data Sciences. We reserve ample time before and after the main event (e.g. to continue the strategic "AI in POF V" discussions") and offer focussed mini-workshop time slots during the main event running from noon-to-noon (e.g. to discuss "What makes data AI-ready?", "What's the hardest problem you can't solve right now?", "Where does AI actually help in your workflow", and more).

At the DSS2026, we welcome all Earth Data Science topics and will organize the sessions around this years foci AI and Earth Systems and AI-ready data as Infrastructure. We plan for an on-site meeting only.

Registration and abstract submission will open in late May. For registration or abstract submission you need to login in with your institute account (Click on "Single Sign-on / Shibboleth").

  • Thursday 8 October
    • Open Workshops - Suggest your workshop topic!
    • 12:00
      Lunch
    • Welcome
    • Session 1.1 - AI and Earth Systems (Talks)
      • 1
        BenthicAI: From Underwater Imagery to AI-Ready Spatial Observations of Elusive Benthic Fauna

        Underwater imaging generates large volumes of observational data, but converting these images into reliable, structured scientific observations remains challenging. This is particularly true for benthic organisms that retreat into the sediment when disturbed. The razor clam Ensis provides a representative case, as individuals can rapidly burrow, making direct surveys difficult and potentially biased.

        We present BenthicAI, a data-science workflow combining computer vision, photogrammetry, and spatial data processing to detect and spatially reference elusive benthic fauna. ROV imagery is used to develop and evaluate a YOLO-based object detection pipeline that identifies Ensis from persistent surface features such as paired siphon openings, enabling non-invasive detection without triggering a behavioral response.

        To preserve the spatial context of these detections, image sequences are processed photogrammetrically to generate georeferenced orthomosaics and local 3D reconstructions of the seabed. AI-based detections can thereby be linked to a common spatial reference system, enabling spatial analysis and relative size estimation.

        The workflow addresses key challenges of underwater computer vision, including variable illumination, turbidity, sediment appearance, and image quality. By combining automated detection with spatial reconstruction, BenthicAI demonstrates how complex underwater imagery can be transformed into structured, spatially consistent, and potentially AI-ready observation data. The approach provides a scalable basis for integrating deep learning into marine observation workflows and future autonomous underwater vehicle deployments.

        Speaker: Judith Fischer
      • 2
        Targeted optical mapping for the verification of potential unexploded ordnance locations in multibeam echosounder data

        The Baltic Sea contains an estimated 300,000 tons of unexploded ordnance (UXO), presenting ongoing environmental and safety challenges.
        While multibeam echosounder (MBES) surveys are the standard for detecting these hazards due to their high positional accuracy, the acoustic signatures of munitions often resemble geological features or debris, necessitating manual visual verification.
        We present an autonomous three-stage framework designed to automate this process.
        The system integrates on-device MBES processing with a GPU-accelerated geomorphometric classification dictionary and a Vision Transformer (ViT) priority classifier, both trained on expert-annotated data.
        This pipeline picks and prioritizes potential targets and automatically generates survey routes optimized for current environmental conditions for a camera-equipped autonomous underwater vehicle (AUV), which then autonomously performs a photogrammetric survey of each target.
        Field testing in the Baltic Sea demonstrates that this system is capable of identifying and revisiting relevant potential targets, as verified by human experts. These results suggest that linking automated acoustic detection with targeted optical mapping can improve the efficiency of large-scale maritime clearance operations and reduce the need for human intervention in the survey process.

        Speakers: Jochen Mohrmann (GEOMAR Deep Sea Monitoring), Tim Benedikt von See (GEOMAR Deep Sea Monitoring), Valentin Buck (GEOMAR Deep Sea Monitoring)
      • 3
        Potential of Epi-Fluorescence Microscopy for the Analysis of Impurities' Behaviour in Ice

        The location of ice impurities in its microscopic structure, particularly within its grain boundaries, significantly affects the preservation of climatic signals and the deformation behaviour of ice. Ice-impurity interactions also influence the microstructure of ice crystals. Understanding this interplay is highly relevant to interpretations of ice-core records and our understanding of climate change and past Earth conditions. Studying impurities in solid ice presents a significant analytical challenge, and a holistic understanding of ice-impurity interactions requires data acquisition from multiple approaches. Mapping chemical impurities in ice with laser ablation inductively-coupled plasma mass spectrometry (LA-ICP-MS) is one such approach, but it is costly and time-consuming. Here, we combine light, polarised and fluorescence microscopy to provide fast insights into impurity relocation processes in ice, expanding the 2D spatial perspective provided by LA-ICP-MS. Fluorescence microscopy can complement studies of ice formation and metamorphism, particularly by enabling real-time investigation of impurity localisation. To our knowledge, it has not yet been applied specifically to the study of ice impurities. Thus, we built a custom-made epifluorescence setup to follow the behaviour of fluorescent particles simulating ice impurities. We performed image analysis to assess particle behaviour in terms of movement patterns, velocities and allocations relative to grain boundaries in development. We successfully tracked fluorescent particles embedded in artificial ice under laboratory conditions at -15C and -30C, and over time. By enabling fluorescence-based particle tracking during ice ageing, we expand the available approaches for investigating ice-core impurities in terms of their properties and introduce a new methodology with high potential for examining the impacts of impurity allocation on the evolving microscopic structure of ice during ageing.

        Speaker: Dr Susana Marcela Simancas Giraldo (AWI)
      • 4
        Multiple Coordinated Views and Data Visualization for Exploring and Understanding Sedimentological Data

        Modern sensors and analytical methods produce data at higher spatial and temporal resolution while adding new variables, data types, and processing steps. This is particularly evident in geosciences, where information from different locations, instruments, sampling campaigns, and analytical workflows needs to be considered together. The central challenge is therefore no longer a lack of data, but our ability to structure this complexity so that humans can efficiently explore multidimensional relationships.
        Users must be able to find relevant information, select suitable subsets, understand the context of measurements, and recognize relationships among variables and data products. Complex datasets may remain underused when these tasks require substantial time or technical expertise.
        A strategy for addressing this challenge is the use of Multiple Coordinated Views (MCV). An MCV interface combines several linked visual representations of the same dataset. Selecting or filtering data in one view updates the corresponding information in the others, allowing users to examine data from different perspectives while retaining their spatial and scientific context.
        We implemented such an interface on top of a structured environmental database and demonstrate its application using a case study from the Wadden Sea. The interface links spatial information, sedimentological measurements, scientific analyses, and underlying processing workflows. Users can move from spatial patterns to individual samples, compare variables across representations, filter views simultaneously, and trace analytical results back to their source data and methods. We show how these functions reduce the need to manually combine separate files and software tools and make relationships among locations, observations, variables, and analytical results more accessible. The MCV approach thereby supports both scientific interpretation and the practical reuse of complex environmental data.

        Speaker: Ivo Brunnenkant (IfG Kiel)
      • 5
        Agentic Implementation of a Lab-Scale Optical "Remote-Sensing" Workflow

        Laboratory-scale rotating-tank experiments are used for demonstration purposes and for investigating scaling laws in geophysical fluid dynamics. Quantitative studies using affordable outreach-focused rotating tank setups are challenging and usually depend on manual measurements and narrative-based protocols prepared by the exerimenters. This talk outlines a processing pipeline for videos of a rotating tank which allows for optical quantitative investigation of dye concentration, velocity, and trajectories of submerged particles in a 'Weather in a Tank' rotating table (MIT). The pipeline supports detection and correction of rotation and perspective, and provides object tracking optical-flow diagnostics. It emits transformed and corrected video output, as well gridded numerical datasets of dye concentration, of velocity and vorticity estimates, and tabular numerical datasets of particle trajectories. The talk also outlines the mainly agentic-AI-based development process of the end-to-end processing pipeline. It highlights challenges and surprises related to general scientific knowledge and geometric reasoning of the coding agent.

        Speaker: Willi Rath (GEOMAR)
      • 6
        Get a Grid? – Solving Geospatial Data Challenges while Creating a Circum-Polar Permafrost Disturbance Inventory from Earth Observation Data

        We present our work on the latest version of AWIs Database of AI-detected Retrogressive Thaw Slumps (DARTS) using a pan-arctic coverage from Sentinel-2 Quarterly Mosaics (~20 Mio km² annually) and auxiliary data sources. Retrogressive Thaw Slumps (RTS) are abrupt disturbances in ice-rich permafrost that dramatically change topography, hydrology, biogeochemistry, and local ecosystems by permafrost thaw, erosion, and massive sediment mobilization. These features increase in abundance and size in many Arctic regions and a spatio-temporal understanding of their dynamics is crucial for assessing impacts of climate change and disturbances on permafrost ecosystems.

        We cover lessons learned during the creation of the DARTS dataset – building on large amounts of Earth Observation data – and the relevant processing pipelines, as well as the implementation in a Deep Learning inference workflow. In this presentation, we provide our insights into data acquisition, usage of geospatial Python libraries as well as the efficient and performant processing and storage on different systems/infrastructure. Data processing required transfer and storage of multiple terabytes of data and used AWIs High Performance Compute cluster Albedo as well as compute resources provided through HAICORE (JUWELS Booster at the Jülich Supercomputing Center). Efficient provisioning data for Artificial Intelligence processing included utilization of modern file formats for vector (GeoParquet) and multidimensional raster (ZARR/Icechunk) data.

        We also highlight how Geospatial Data Science in polar regions faces numerous challenges specific to the data formatting and georeferencing. This is particularly true for the acquisition and processing of data with circum-polar coverage in mind, concerning choice of appropriate georeference systems due to coordinate wrapping and raster distortion.

        Speaker: Jonas Küpper (Alfred-Wegener-Institut, Bremerhaven)
    • 15:00
      Coffee Break
    • Session 1.1 - AI and Earth Systems (Talks)
      • 7
        Data Science at AWI: Applied Artificial Intelligence for Polar & Marine Research

        The Alfred Wegener Institute has accumulated a large and heterogeneous body of marine and polar observation data including multi-decadal time series from long-term observatories, passive acoustic recordings, environmental DNA sequencing, and seafloor imagery.
        We use AI to turn this data into knowledge. For example, we use detection models to extract different types of whale call events from multi-year passive acoustic recordings, simplifying the analysis of seasonal marine mammal presence. We developed MetaParse to segment seafloor scenes containing many co-occurring, visually entangled species and derive descriptive metadata from image content alone; it is built so that the approach transfers to segmentation problems beyond seafloor imagery. We also built BioCUDA for NMR imaging of marine organisms, which yields stacks of blurred, low-contrast greyscale slices. BioCUDA assembles these into a three-dimensional volume and segments organs and muscle tissue, making morphological change under changing climate conditions measurable.
        Each application runs into the same limit: a model is only as good as the data behind it, and FAIR compliance alone does not make data usable for machine learning and AI. We are therefore developing the Data2AI companion, which combines a deterministic quality-control engine, computing a broad set of readiness metrics with an agentic layer that explains the resulting assessment and recommendations on how to act on it.
        We also run a set of agents that support submission into the systems that manage our data: a sample management system, a sequence and metadata management system, and a chatbot for documentation.
        We present results from MetaParse, Whale Call Detection and BioCUDA as well as the Data2AI companion and other agents.

        Speaker: Sonja Hänzelmann
      • 8
        ML ocean and sea-ice modeling at GEOMAR: 1. RánCast, a global 3D ocean emulator trained on kilometer-scale simulations

        Machine learning is reshaping how we simulate the Earth system: data-driven weather models now rival operational forecasts at a fraction of their cost. Whether the same holds for the ocean is far less clear, as it is a slowly evolving, strongly forced component whose simulations have to stay stable far beyond weather forecast horizons. At GEOMAR we follow two complementary routes to machine-learning based ocean and sea-ice modeling; this first of two talks presents the route that replaces the numerical model entirely with a data-driven emulator.

        A key insight from recent advances in numerical ocean modeling is that small-scale processes are not merely a detail: their representation is necessary even for a correct simulation of large-scale features such as the Gulf Stream path. This motivates training on kilometer-scale simulation data rather than coarser reanalysis: the training signal must contain the small-scale variability that shapes large-scale structure, even if the emulator itself operates at reduced resolution. Two tensions then arise: capturing the stochastic imprint of these processes to produce realistic mesoscale variability, while maintaining rollout stability over months to years.

        We present RánCast, a global 3D ocean emulator trained on 30 years of ocean data from the coupled FESOM–IFS nextGEMS simulation, regridded from kilometer-scale to 1.5° resolution. The architecture adapts ArchesWeather to the ocean domain, predicting sea surface height, potential temperature, salinity and horizontal velocities as daily snapshots. We evaluate two additions aimed at improving multi-month to multi-year rollout stability. First, atmospheric surface forcing (2-m temperature, surface pressure, precipitation, 10-m winds) from the IFS component of the same simulation is re-injected as encoder input at every autoregressive step, providing a physically grounded external constraint that anchors the ocean trajectory and preserves mesoscale variability where forcing dominates. Second, a learnable residual scaling coefficient per variable and depth level replaces the fixed skip connection, letting the model adapt the persistence–increment balance to the dynamics at each depth.

        Speaker: Nils Hutter
      • 9
        ML ocean and sea-ice modeling at GEOMAR: 2. Learning sea-ice thermodynamics with a hybrid numerical–ML model

        Machine learning is reshaping how we simulate the Earth system, and data-driven emulators of the ocean and sea ice are becoming stable and skillful. Yet purely data-driven models lack physical interpretability with respect to individual dynamical and thermodynamic processes, and their generalizability to a changing climate remains uncertain. At GEOMAR we therefore also follow a second, complementary route: keep the physics and learn only what we do not know. This talk presents that route for sea ice.

        Sea ice plays a central role in the climate system by regulating exchanges of heat and momentum between ocean and atmosphere. Representing its evolution numerically is challenging, as growing model complexity increases computational cost while key processes remain unresolved and their parametrizations poorly constrained. We investigate a hybrid framework that bridges numerical and data-driven modeling: a machine-learning-enabled numerical sea-ice thermodynamic model. The end-to-end differentiable implementation of zero-layer column thermodynamics in Python allows sensitivities with respect to model parameters to be computed directly. This enables gradient-based parameter optimization and, more importantly, the description of individual parametrizations by process-emulating neural-network components, jointly trained and evaluated with the numerical model against snow and ice thickness observations from ice mass-balance buoys.

        These low-complexity components remain physically interpretable owing to their explicit input–output relationships and local pointwise operation, in contrast to high-dimensional ML models. The hybrid model improves the representation of snow and melt processes while retaining the stability and interpretability of the numerical backbone, a compromise between physical constraint and data-driven flexibility. We discuss what governs whether such hybrid models remain stable and physically plausible—in particular training strategy and imbalanced observational data—and close with an outlook on sea-ice physics learned directly from observations.

      • 10
        Towards a Holistic Topological Analysis Across Foodwebs

        Foodwebs, as interaction networks of species, are fundamental to understanding ecosystem function. To date, their complex structure and dynamics have been the subject of intensive research. However, most studies have focused on specific properties such as degree distributions, connectance, or species traits in order to identify universal patterns across foodwebs. Holistic investigations that go beyond these specific properties are still lacking.

        Here we present a novel analytical workflow based on normalized Laplacian spectra and dimensionality reduction to analyze foodweb topology in a more holistic manner. Laplacian spectra, a well‑established descriptor in network science, have rarely been applied in ecological network research. Our proposed workflow provides an interpretative framework and shows initial indications that it enables (i) the description of generic topological patterns, (ii) the inference of plausible development trajectories, and (iii) the assessment of sampling artifacts across empirical foodwebs.

        Our results demonstrate that normalized Laplacian spectra are a powerful tool for the holistic analysis of foodweb topology. Moreover, our proposed workflow could offer a promising blueprint to study other complex, data‑driven networks in ecology and beyond.

        Speaker: Joel Habedank (GFZ Helmholtz-Zentrum für Geoforschung)
      • 11
        Making well-grounded science-based decision with DASF-based AI Agents

        Scientific decision-making in complex systems requires interactive systems, such as digital twins, that can combine domain knowledge, data-intensive workflows, and transparent computational methods. We present an approach for building AI agents that support well-grounded, data-based decisions by embedding conversational capabilities directly into scientific web applications. The approach combines the Data Analytics Software Framework (DASF), a secure message-broker-based remote procedure call system, with large-language-models, to expose scientific tools, workflows, and data products as AI agents.

        Rather than providing a detached chatbot beside a domain application, DASF-based AI agents operate within the application context. They can trigger analyses, parameterize workflows, orchestrate computations close to data and high-performance computing resources, and return results as domain-native outputs such as maps, plots, tables, and dashboards. This enables conversational interaction while preserving scientific traceability, institutional security requirements, and established user interfaces.

        As an exemplary implementation, we demonstrate the concept with iseapower.hereon.de, a decision-support tool for offshore wind farm maintenance. In this setting, the AI agent assists users in exploring maintenance-relevant scenarios, initiating analytical workflows, and interpreting results in the operational context of the platform. DASF enables secure and asynchronous execution without exposing infrastructure through inbound internet-facing ports, while the conversational agent provides lightweight tool access.

        The resulting architecture turns conversation into an additional interaction modality for scientific decision-support systems. The approach generalizes to knowledge-transfer scenarios in which robust, transparent, and context-aware AI support is needed for well-grounded science-based decisions.

        Speaker: Philipp Sommer (Helmholtz-Zentrum Hereon)
    • Group photo
    • Session 1.3 - AI and Earth Systems (Posters)
      • 12
        Automated identification and classification of fossil pollen from sediments

        Fossil pollen analysis is a key tool for reconstructing past vegetation and climate, but conventional microscopic identification and counting are time-consuming, require substantial taxonomic expertise, and can be affected by observer bias. We develop an automated workflow combining multispectral imaging flow cytometry (MIFC) with convolutional neural networks (CNNs) for high-throughput and reproducible identification of fossil pollen from sediment samples. MIFC rapidly acquires multispectral images of individual particles, providing the basis for CNN-based image recognition and taxonomic classification.
        Our workflow follows a hierarchical classification strategy. A first CNN distinguishes pollen grains from non-target sediment particles, including charcoal, debris, and Lycopodium marker spores. A subsequent CNN assigns detected pollen grains to taxonomic classes. As a palaeoecological test case, we apply this approach to fossil pollen samples from Lake Emanda, eastern Siberia, which provide a Late Quaternary record with conventional pollen counts for validation of automated classification.
        Currently, we compare CNN-classified pollen abundances with conventional microscopic pollen counts at family level to assess quantitative agreement between approaches. In parallel, image-feature analyses investigate factors limiting classification performance, including particle orientation, focal plane, and image quality. The resulting benchmark image datasets will be curated and stored using OMERO-based image data management, supporting standardized metadata, accessibility, and future model development.
        Ultimately, we aim to develop a CNN-based system that supports or partially replaces microscopic pollen identification and counting, enabling faster, scalable, and more reproducible reconstructions of past vegetation and climate.

        Speaker: Kathleen Stoof-Leichsenring
      • 13
        Data2AI Companion: A tool assessing data readiness for AI applications

        Artificial Intelligence (AI) is central to scientific research, but model performance is fundamentally constrained by data readiness. The FAIR principles govern findability, accessibility, interoperability and reusability of research data, but not dataset properties that matter for AI like task-specific completeness, feature relevance, label quality or class imbalance. Existing frameworks offer data readiness levels, consolidate dimensions and metrics, identify domain-specific pipelines and link readiness to fairness; tools evaluate published datasets against selected criteria and fairness; metadata standards offer standardized documentation and machine readability. But a shared definition of data readiness for AI, comprehensive, machine-actionable criteria and guidance across the entire data lifecycle as well as domains is still missing. The Data2AI Compass closes this gap. It delivers a definition of data readiness for AI and machine-readable criteria catalog developed by an expert working group across domains. The Catalog spans four dimensions (data quality, task fit, governance & documentation, intended use), pairs a domain-agnostic core with domain-specific extensions and is complemented by modular training materials, enabling transfer across domains and actionability in everyday research practice. The Data2AI Companion operationalizes the catalog as automated, versioned tests. Datasets are read through modality adapters (tabular, image, text). Users specify their intended use (e.g. domain, task, model class), per-metric results are visualized, and a conversational interface explains the findings and proposes remediation. Results export as a readiness report, machine-readable Croissant metadata file and Hugging Face dataset card. A modular metric class and YAML-based metric configuration enable community contribution.

        Speaker: Luca Greiner (Alfred-Wegener-Institute for Polar and Marine Research, Data Science, Bremerhaven, Bremen, Germany)
      • 14
        Empowering Earth & Environment Research: The Helmholtz Information and Data Science Platforms

        Discover the full spectrum of data science support and expertise at the Helmholtz booth, where five Helmholtz-wide platforms showcase their combined strength to accelerate research—particularly in the Earth and Environment field. These platforms operate across the entire Helmholtz Association to advance AI, imaging, metadata, research software and training. Our booth serves as a one-stop hub for information and data science exchange and support.

        What We Offer:
        - Helmholtz AI offers tailored consulting, computing resources, and funding opportunities to democratise and accelerate the application of artificial intelligence in science, supporting a wide range of use cases from foundation models to domain-specific solutions.
        - HIDA offers a rich portfolio of training, networking, and funding opportunities, fostering the next generation of data scientists and supporting upskilling across all career stages.
        - HIFIS provides and brokers federated cloud services hosted across Helmholtz centers and delivers training and support on using those services as well as on Research Software Engineering.
        - Helmholtz Imaging provides access to cutting-edge imaging infrastructure, funding opportunities, and expert consulting, enabling researchers to unlock new insights from complex earth and environmental data and beyond.
        - HMC provides internationally aligned data management concepts, tools, services and training to support the Helmholtz Association, ensuring your data is findable, accessible, interoperable, and reusable (FAIR) for collaborative research and advanced science.

        Visit our booth to explore how we can support your research, connect with experts, and discover new opportunities in funding, training, infrastructure, and community engagement.

        Speakers: Katharina Kriegel (Helmholtz Imaging, DESY), Lena Heel
      • 15
        Forecasting and Observing the Open-to-Coastal Ocean for Copernicus Users (FOCCUS): AI-Enhanced Coastal Monitoring and Forecasting for the Digital Ocean

        Advancements in monitoring and predicting our coastlines is vital for coastal protection in the face of climate change. Novel developments in Digital Ocean technology are driving these advancements, including Artificial Intelligence (AI) enhanced numerical models and coastal observations leveraging data fusion and high-resolution products. The EU Horizon FOCCUS project (Forecasting and Observing the Open-to-Coastal Ocean for Copernicus Users; foccus-project.eu) builds on this technology and the Copernicus Marine Environment Monitoring Service to advance its coastal dimension by improving existing capabilities and developing innovative coastal products. Collaboration between the project’s 19 partners from 11 countries with their own Member State Coastal Systems (MSCS) and users gives the project a pan-European perspective.

        Endorsed by the UN Ocean Decade’s CoastPredict program, FOCCUS enhances the coastal capabilities of Copernicus Marine through three key pillars: i) developing novel high-resolution coastal data products by integrating multi-platform observations (remote sensing and in-situ), AI algorithms, and advanced data fusion techniques for improved coastal monitoring; ii) developing advanced hydrology and coastal models including a pan-European hydrological ensemble for improved river discharge predictions, and improving member state coastal systems by testing new methodologies in MSCS production chains while taking advantage of stochastic simulation, ensemble approaches, and AI technology; and iii) demonstrating innovative products and improved co-produced services that address both environmental and societal challenges, enhancing the performance and societal relevance of coastal ocean forecasting systems. FOCCUS emphasizes codevelopment and collaboration with end-users, policy-makers, and local communities, supporting informed and advanced decision-making processes for sustainable coastal management and climate adaptation strategies in a changing world.

        Speaker: Kelli Johnson (Hereon)
      • 16
        From local observations to regional habitat maps: a multimodal data approach for monitoring Baltic Sea brown-algal reefs

        Brown-algal reefs are important marine habitats and potential contributors to Blue Carbon storage, yet their abundance and spatial distribution are difficult to assess over large areas. Optical observations provide detailed information on algal occurrence and morphology, but their spatial coverage is limited. To overcome this scaling challenge, we combine optical data from autonomous underwater vehicles (AUVs) and towed camera systems with hydroacoustic data.
        Within a six-year monitoring project, two research cruises per year are conducted at three sites within German Baltic Sea Marine Protected Areas. The surveys generate thousands of images per cruise, complemented by video, hydroacoustic and biological sampling data. Together, these datasets form a large, heterogeneous and temporally repeated dataset capturing seasonal and interannual variability.
        As the project is in its initial phase, at the current project stage the focus is on data acquisition, exploration and the development of an efficient data workflow. A key challenge is the management and analysis of the diverse and high-volume datasets. Automated image-detection methods are being investigated to support the future analysis of optical data. In parallel, we are exploring how optical observations can be related to hydroacoustic information to enable future extrapolation from localized, high-resolution observations to larger spatial scales. This work will provide the basis for a habitat-modelling framework as the project progresses.
        We present the monitoring concept, data-acquisition strategy and data-integration workflow, together with example datasets and preliminary observations. The poster highlights the challenges and opportunities of integrating heterogeneous marine sensor data and exploring approaches for their automated analysis and future use in scalable habitat monitoring.

        Speaker: Anne Hennke
      • 17
        Identifying numerical instabilities: Leveraging neural networks for numerical noise detection

        Ocean models have to balance efficiency and accuracy. The general assumption is that the higher the resolution, the more realistic the simulated flow. However, the actual grid spacing is not the physically realistic and reliable resolution of a model configuration. Numerical artifacts due to choices of solvers, numerical schemes or parameterizations can affect the accuracy and reliability of the simulated flow. At intermediate model resolutions of a few to tens of kilometers, viscosity parameterizations play a critical role in maintaining model stability and in dampening numerical modes that may otherwise result in numerical artifacts. In this study, we present a machine learning method to identify numerical noise in configurations of the FESOM2 unstructured grid ocean model. We train a vector-quantized variational autoencoder (vqAE) on surface vorticity data from an idealized channel configuration, interpolated onto a regular grid using nearest neighbour interpolation. The reference configuration employs sufficiently high viscosity to produce a smooth, largely noise-free vorticity field optimized for the setup. After training, we then test the autoencoder on simulations with (temporally) reduced viscosity. Thus, we identify regions of intensified noise that are revealed after reconstruction. To target noise at specific scales, particularly after interpolation to a finer regular grid, we conduct sensitivity tests with the image Euclidean distance metric (IMED) and the feature similarity index (FSIM). These methods aim to enhance spatial robustness and apply a targeted edge detection algorithm, respectively, improving the detection of anomalies. The vqAE robustly identifies and flags intensified noise, therefore enabling automated detection of potential numerical instabilities that may degrade the simulated flow without causing obvious model blow-ups. It can ultimately act as a warning flag to modelers for model tuning and before analyzing or releasing model data.

        Speaker: Stephan Juricke
      • 18
        Intraseasonal prediction of monthly storminess with ACE2 atmospheric emulator and Random Forests

        This research explores the predictability of seasonal storminess in the North Sea and East China Sea using machine learning and the ACE2 weather model emulator, focusing on stratospheric and upper-tropospheric influences on winter storms. Using ERA5 reanalysis data (1940–2024), a storminess index based on storm frequency was developed to examine links with large-scale atmospheric fields.
        For the North Sea, storminess was predicted using ACE2 and a Random Forest model. ACE2 is a machine-learning weather emulator developed by the Allen Institute for AI and trained on ERA5. ACE2 simulations show that lower stratospheric temperatures and stronger winds on 1 December increase emulated mean January surface wind speeds across much of the North Sea. The Random Forest model achieved its highest skill when predicting January storminess from December fields, with correlations of 0.55–0.60.
        Predictability from both models follows the seasonal cycle of polar-vortex intensity. Skill peaks in winter, when stratosphere-troposphere coupling is strongest, with a lead time of 4-6 weeks. North Sea storminess is neither significantly correlated with storminess elsewhere nor persistent from month to month, suggesting regionally specific drivers operating on subseasonal timescales.
        Overall, the results show that stratospheric conditions play an important role in North Sea winter storminess and that machine learning can improve subseasonal prediction in this region.
        The East China Sea component is ongoing. Preliminary results indicate that a November sea-surface temperature and sea-level pressure pattern resembling La Niña is linked to December storm numbers. Strong relationships are also found between Southern Hemisphere high-latitude stratospheric temperature, zonal wind, and December storm activity. Together with evidence that La Niña favors a positive SAM, these results suggest a pathway linking La Niña, November SAM conditions, and enhanced December storminess in the East China Sea.

        Speaker: Proshonni Aziz (Helmholtz-Zentrum hereon)
      • 19
        iSEA: Intelligent Seafloor & Animal Image Annotator

        Artificial intelligence (AI), particularly computer vision algorithms like YOLO, has become standard for automated organism detection. However, AI complexity remains a barrier for non-specialist researchers, limiting adoption in marine biodiversity studies. This project presents iSEA (Intelligent Seafloor & Animal Image Annotator), an interactive Python-based GUI that integrates YOLO for automated detection and classification of marine fauna in underwater images. The intuitive interface allows researchers to train custom models without programming expertise, reducing analysis time while improving dataset quality through human-in-the-loop refinement. Real-time processing enables use aboard research vessels as a decision support tool during expeditions. A seafloor classification module using InceptionV3 enables dual-mode operation, allowing simultaneous fauna detection and habitat characterization within the same video transect. The platform was tested on ROV footage from deep-sea coral habitats in Brazil's Campos Basin and the models trained were evaluated using standard metrics (F1, mAP, precision, recall). Designed to be user friendly, iSEA bridges the gap between technical AI knowledge and practical marine science applications, accelerating image analysis and ultimately enhancing conservation strategies for vulnerable ecosystems.

        Speaker: Raphaela Neves Lopes
      • 20
        Learning Where It Matters: Instance-Wise Feature Attribution for Antarctic Sea Ice Drivers

        Antarctic Sea Ice (ASI) shows strong regional and interannual variability shaped by a complex interplay of atmospheric and oceanic processes. As the climate system redistributes energy poleward to offset radiative imbalances, remote conditions can influence ASI through large-scale circulation patterns, or teleconnections. Established teleconnection indices link these patterns to ASI variability via winds, heat fluxes, and moisture transport, but their spatial aggregation limits their ability to explain regional sea ice extremes — a limitation compounded by evidence that the dominant drivers differ from one extreme event to the next, underscoring the need for methods that can adapt to individual events rather than assuming a single fixed pattern governs ASI variability as a whole.
        We propose an instance-wise feature attribution framework to address this gap. Rather than relying on predefined indices, we convert global climate fields into superpixels to preserve regional coherence, then train a selector network jointly with a predictor network so that each prediction is paired with a learned mask identifying the regions most responsible for it. This approach is designed to surface non-linear, spatially distributed, and event-specific drivers of ASI variability that static indices cannot capture.

        Speaker: Nina Öhlckers (Alfred-Wegener-Institut, Universität Bremen)
      • 21
        Memento 2.0: Revision of the global synthesis product for N2O and CH4 data

        The ocean plays a key role in mitigating climate change, with N2O and CH4 acting as potent greenhouse gases within the ocean-atmosphere system. Robust, well-curated data are essential to reduce uncertainties and refine predictions of N2O and CH4 fluxes (Bange 2022), leading GOOS (Global Ocean Observing System) to designate N2O as an Essential Ocean Variable (EOV). Building on this, MEMENTO (MarinE MethanE and NiTrous Oxide database) compiles global spatio-temporal datasets of N2O and CH4 from the open ocean and coastal waters to advance our understanding of ocean-atmosphere gas exchange dynamics.
        However, Lange et al. (2023) evaluated ocean biogeochemistry data synthesis products using a data readiness level concept and identified gaps in MEMENTO regarding data management and FAIR (Findable, Accessible, Interoperable, Reusable) compliance. To close these gaps, we use FROST (Fraunhofer IOSB), a data management platform implementing the OGC Sensor Things API (STA), and align our metadata as closely as possible with STAMPLATE (SensorThings API Metadata ProfiLes for eArTh and Environment), a Helmholtz initiative extending the STA core model with metadata on data quality, provenance, and sensor description to enable harmonized, interoperable time-series data (Lorenz, 2024). Specifically, we plan to improve MEMENTO through the enhanced transparency and verifiability of data harmonization and quality control, completion of metadata, increased automation, visualization masks for dataset availa-bility, and the integration of additional datasets.
        As a result, we expect a substantially higher readiness level for MEMENTO 2.0 in terms of data management and information products, along with an improvement in FAIR compliance compared to the previous assessment, which, in turn, potentially enables refined N2O and CH4 flux estimates.

        References
        Bange HW (2022): https://doi.org/10.3389/fmars.2023.1078908.
        Lorenz, C (2024): https://doi.org/10.5281/zenodo.13151659.

        Speaker: Daniela Niemeyer-Teuteberg
      • 22
        O2A WORKSPACE - Scientific Collaboration Made Simple

        Your Managed Platform for Data-Driven Research

        Research today rarely stays within one institution, one dataset, or one infrastructure. As it spans providers and partners, the tools that support it tend to fragment along the same lines: every institution provides its own environments, connects them to its own data by hand, and enforces its own access rules, with little of that effort carrying over from one collaboration to the next.

        We are developing O2A WORKSPACE, a new platform within the O2A data flow framework, to turn that fragmentation into one shared ecosystem. Still in its concept phase, it is designed around two sides at once.
        On the one hand researchers get a single login into project-based workspaces populated with ready-to-use services, JupyterHub and RStudio among them, pre-connected to institutional and curated data sources such as O2A INGEST and PANGAEA, so that working in them means working with findable, accessible, and reusable data from the outset, with resource use shown transparently as credits rather than buried in an invoice.
        On the other hand infrastructure providers keep full autonomy over their own systems: they define a service once as a standardized template, and O2A WORKSPACE deploys and orchestrates it across partner institutions, absorbing the operational burden of updates and provisioning. Providers keep control over the quality of what runs on their infrastructure and gain usage metrics for capacity planning.

        Positioned between data acquisition and data publication, O2A WORKSPACE is meant to be where processing and analysis actually happen within the O2A data flow. Collaboratively and FAIR by design, without either side having to manage infrastructure to make that possible. This poster presents the concept as designed so far, illustrated through example project, and the path toward a first working prototype.

        Speaker: Nico Harms (AWI)
      • 23
        SEA2LAND Navigator-GPT

        Coastal and marine planning is difficult due to complex regulations, large and fragmented datasets, and the involvement of multiple authorities across local, regional, and national levels. The SEA2LAND Navigator tool was developed to support this process by organizing governance responsibilities and providing access to a wide range of relevant resources. Given the breadth and depth of information it brings together, users may benefit from additional support when exploring the tool and applying it to specific planning questions. To address this, we introduce SEA2LAND Navigator-GPT, an AI chatbot integrated into the SEA2LAND Navigator website. The chatbot allows users to ask natural-language questions and receive clear, step-by-step guidance supported by user data and official documents. By simplifying access to information and guiding users through complex planning tasks, Navigator-GPT improves usability, supports transparent decision-making, and helps bridge the gap between data, policy, and practical coastal planning.

        Speaker: Harsh Grover (Hereon)
      • 24
        Shipboard ADCP Data acquired by the DAM Project “Marine Data – Research Vessels”: Machine Learning for Automated Seafloor Detection based on Acoustic Backscatter

        [Shortened! Full abstract as pdf]
        Shipboard ADCPs provide continuous upper-ocean velocity profiles and are indispensable for oceanographic field studies. In shelf regions and coastal waters, a particularly time-consuming part of post-processing is editing the velocity time series to remove contamination near the seafloor caused by sidelobe interference from the bottom echo. Existing automated approaches are often insufficiently robust under real-world conditions — particularly given acoustic interference from other simultaneously operated instruments — so reliable quality control still largely depends on manual inspection, limiting timely data provision as raw data volumes grow.
        Within the DAM (German Marine Research Alliance) project "Marine Data – Research Vessels", we aim to develop a two-stage machine learning solution that automates bottom-signal editing using the same signals an experienced analyst relies on: beam-averaged echo intensity and along-track velocity. A convolutional neural network first locates the seafloor by jointly evaluating both signals across consecutive pings; a gradient-boosted classifier then flags velocity cells contaminated by bottom interference.
        A key contribution of this project is the compilation of a large, high-quality reference dataset: more than 50 manually processed shipboard ADCP cruises, comprising around 4,500 raw data files with expert-derived bottom and contamination masks. Models are trained and evaluated at cruise level to avoid data leakage, and each prediction is accompanied by a confidence score so that ambiguous cases can be flagged for expert review rather than processed silently.
        The resulting module will be integrated into the OSADCP toolbox — the standardized workflow developed within this DAM project for FAIR provision of ocean current data from the German research fleet — with the longer-term goal of enabling autonomous, near-real-time processing on board, as a service to the wider oceanographic community.

        Speaker: Robert Kopte (Institut für Geowissenschaften, CAU Kiel)
      • 25
        Spatial Downscaling of Pollen-Based Vegetation Reconstructions into Paleo Land Cover Maps

        Past land cover dynamics provide important context for understanding long-term ecosystem responses to climatic change, but paleoecological records are typically sparse, irregular, and site-based. Here, we present a spatial downscaling framework that transforms pollen-derived vegetation reconstructions into gridded paleo land cover maps. The pipeline first translates pollen-based vegetation composition into land cover class frequencies, then distributes these classes within pollen source areas using environmental similarity, and finally applies a U-Net-based interpolation approach to generate spatially continuous reconstructions. We apply the framework to Alaska and western Canada, producing 300 m land cover reconstructions across 15 land cover classes and 119 time slices. The source-area product preserves the direct link to pollen records and contains approximately 2.56 billion reconstructed pixels, while the U-Net-based product extends these reconstructions to continuous spatial coverage across the full study region. Reconstructed land cover trends show broad stability during the late Holocene, with stronger changes around 10 ka BP, including declining evergreen needle-leaved forest cover toward older time slices and corresponding changes in sparse vegetation and shrubland. Relative uncertainty estimates indicate higher uncertainty in forested regions and increasing uncertainty with age, reflecting both class ambiguity and reduced paleo-record support. An example analysis of elevational forest dynamics demonstrates the potential of the dataset for spatially explicit paleoecological applications. The framework provides a bridge between site-based pollen records and landscape-scale analyses, while emphasizing that reconstructed maps represent model-based estimates rather than direct observations of past environments.

        Speaker: Laura Schild (Alfred-Wegener-Institut, Potsdam)
      • 26
        Unsupervised Embeddings of In-Situ Particle Data for Mesoscale Eddy 3D Characterization

        Mesoscale eddies modulate vertical particulate organic carbon (POC) flux, yet linking eddy dynamics to in-situ biological observations at scale remains a major data-engineering bottleneck. Satellite altimetry (e.g., SWOT, AVISO Eddy Atlas) and in-situ particle imaging (e.g., Underwater Vision Profiler, UVP) differ substantially in format, resolution, and archival structure.

        We present the UVP extension of the Eddy Hunter System, a data-fusion framework under development at GEOMAR. Built on a PostGIS-indexed pipeline, the framework spatiotemporally co-locates UVP profiles with eddy trajectories drawn from a multi-mission altimetry eddy atlas. Profiles are retrievable by profile ID, single eddy observation, or complete eddy-track lifetime.

        Each UVP cast provides depth-resolved particle-size distributions across 30 Equivalent Spherical Diameter (ESD) classes (40 µm to >26 mm). We standardize pressure onto a unified depth grid and stack profiles into fixed-shape tensors, yielding machine-learning-ready arrays.

        Building on this pipeline, we propose encoding each profile into a low-dimensional embedding via a compact neural network and finally identify particle-spectrum clusters.
        We will evaluate cluster membership against eddy polarity, distance to eddy center, and eddy age using permutation and mixed-effects models, testing whether mesoscale eddies imprint a detectable signature on particle size distributions, and how that signature varies from eddy core to periphery.

        Speaker: Federico Scarscelli (GEOMAR)
      • 27
        Vibing to Visibility: Using the Bot to Shed Light into the Data Abyss

        Research institutes often host databases containing petabytes (PB) of valuable scientific data. This data can originate from different times, locations, devices, and research groups, and may be in different states of curation. Getting a visual overview of such data can therefore be challenging and time-consuming. This poster shows how AI can support the planning, building, and debugging of a complete workflow for visualising research data. Using an AI assistant, ideas were discussed, suitable tools were explored, and the setup was implemented through vibe coding. The workflow combines Docker, SQLite, and Grafana to turn data from an institutional database into an interactive dashboard. In addition to the technical workflow, the poster reflects on the challenges and learnings encountered along the way. The time and effort required are also compared to provide a practical perspective on the use of AI assistance for research data workflows. The aim is to share hands-on experiences and encourage others to explore how AI assistants can support them in working with complex research data.

        Speaker: Barbara Glemser
      • 28
        Whale vocalization detection on noisy data

        Analysis of passive acoustic monitoring (PAM) data is a cost-effective method to study whale presence. The resulting recordings have a strong imbalance between short call events and long periods of background noise, making automated whale call detection challenging. Also, the annotation of the whale calls is often imprecise, which complicates model training. State-of-the-art methods convert audio clips to spectrograms and use computer vision models, such as You Only Look Once (YOLO), for the detection and classification of whale calls.
        We base our research on the BioDCASE[1] challenge dataset and reference method. The dataset contains around 188,000 spectrograms generated from 1,880 hours of acoustic recordings, containing blue and fin whale calls. The method uses YOLO11 and has a reported F1-score of 0.44 (Recall = 0.32, Precision = 0.62), which is not sufficient for operational use. Therefore, we evaluate different ways to improve the performance of detection models: edge label filtering, which removes the labels from the spectrogram edges; bounding box inflation, which enlarges the annotations along the time and frequency axes; additional reannotation of the data; and testing different YOLO architectures. For this, we reannotated 1,000 spectrograms randomly sampled from the dataset and fine-tuned the model by continuing to train the last five layers of the baseline model. We compare this to fine-tuning on the original labels and YOLO11s with other YOLO models. Instead of using the original BioDCASE data split, which assigns entire site-year subsets to either the training or the test set, we split each subset separately into both to avoid a batch effect in the experiments.
        Although fine-tuning and edge label filtering have shown promising preliminary results, the more accurate experiments are still ongoing, and the results will be presented at the Symposium.

        [1] BioDCASE Challenge 2026 Task 2: https://biodcase.github.io/challenge2026/task2

        Speaker: Marie Vogler (Zentrum für Bioinformatik, Universität Hamburg)
    • 19:00
      Dinner
  • Friday 9 October
    • Session 2.1 - AI-ready data as Infrastructure (Talks)
      • 29
        ODIS2ODV: Bridging Ocean Data Discovery and Scientific Analysis through FAIR and AI-Ready Metadata

        The Ocean Data Information System is an international and diverse federation of data systems coordinated through the International Oceanographic Data and Information Exchange of UNESCO’s Intergovernmental Oceanographic Commission, sharing data and information about their holdings and capabilities over the Web using linked open data norms. “Nodes” in the ODIS Federation share metadata using Schema.org semantics and JSON-LD serialisation, to ensure cross-domain and -sector interoperability. While ODIS improves the findability and interoperability of digital assets in general, the visibility of tools for analysis and visualization of ocean data is still underexplored. This contribution presents ODIS2ODV, an open-source project that bridges the linked data ecosystem of ODIS with the application-centric data ecosystem of Ocean Data View . ODIS2ODV introduces a lightweight JSON-LD mapping specification describing how tabular oceanographic datasets can be converted into ODV Generic Spreadsheet collections. The mapping preserves semantic information such as measured variables, units, quality flags, and controlled vocabulary identifiers, making datasets both FAIR across the ODIS Federation, and the far wider ecosystems using schema.org/JSON-LD, including those powering agentic AIs. The presentation introduces the concept using a small tutorial dataset and then demonstrates a case study based on a selected primary-production dataset from the Hawaii Ocean Time-series program. A dedicated converter validates the JSON-LD description and automatically transforms datasets into ODV collections. In the opposite direction, the converter can generate ODIS-compliant JSON-LD from existing ODV datasets, providing a pathway for publishing historical scientific datasets as FAIR, machine-readable resources. ODIS2ODV demonstrates how semantic metadata can bridge data discovery, scientific analysis, and AI-driven applications discovering assets over standard Web architectural patterns.

        Speaker: Sebastian Mieruch
      • 30
        Elbe River Basin Water Body Knowledge Graph: Integrating Modern and Historical Data

        Climate and land-use change are reshaping the relationships between hydrological processes, water quality, and biodiversity in river ecosystems. Effective adaptation requires an integrated understanding of these interactions, yet water management often relies on fragmented, localized, and sector-specific datasets, potentially leading to suboptimal decisions. In particular, hydrological and biological connectivity must be considered across spatial, temporal, and administrative boundaries.
        Data describing water bodies span geospatial properties, hydrological characteristics, chemical and biological status, management actions, and remote-sensing observations. Elemental and genetic analyses of sediment cores, together with archival records, can further extend these data into the past and reveal long-term environmental change.
        This project aims to integrate these diverse sources for water bodies throughout the Elbe River basin. It comprises three stages: (1) consolidating contemporary datasets and extracting historical information from administrative documents into validated, structured formats; (2) constructing a knowledge graph that connects water bodies, observations, environmental processes, and management activities; and (3) developing an accessible interaction layer for retrieving, exploring, and interpreting the resulting data and their context.
        By supplementing existing water databases with evidence from sediment cores and archives, the project will establish a connected, long-term information resource. This resource is intended to support more comprehensive management decisions, facilitate cross-domain analyses, and improve predictive assessments of how river and lake ecosystems respond to environmental change.

        Speaker: Laura Schild (Alfred-Wegener-Institut, Potsdam)
      • 31
        HMG NetCDF 2.0 – Building a Community-Driven Metadata Guideline for FAIR and AI-Ready NetCDF Data

        NetCDF is one of the most widely used scientific data formats in E&E sciences, yet metadata quality, completeness, and interoperability remain highly heterogeneous across institutions, disciplines, and data infrastructures. While standards such as CF Conventions, ACDD, and DataCite provide important foundations, their implementation often varies significantly, limiting data reuse, interoperability, and integration into automated and AI-driven workflows.

        The HMG NetCDF initiative addresses this challenge by developing community-driven metadata guidelines for the Helmholtz E&E domain. Following the first guideline release, HMG NetCDF 2.0 focuses on transforming the guidelines into a living standard supported by an active community, implementation tools, and sustainable governance.

        This talk presents the vision and planned activities of HMG NetCDF 2.0. We introduce its three pillars: guideline evolution, technical implementation, and community building. The initiative extends metadata recommendations with emerging concepts such as provenance, data quality indicators, and AI-readiness criteria. In parallel, a modular validation ecosystem is being developed to support automated metadata quality assessment and integration into research workflows, repositories, and continuous integration environments. Strong emphasis is placed on user-centred design, community engagement, and reusable metadata profiles that lower adoption barriers for researchers, data managers, and infrastructures.

        Beyond enhancing metadata quality and interoperability for NetCDF, HMG NetCDF 2.0 aims to establish a foundation for machine-actionable and AI-ready data ecosystems. The initiative will assess the transferability of metadata profiles, validation approaches, and governance mechanisms to other scientific data formats, such as Zarr and HDF, and to frameworks including RO-Crate and the Helmholtz Knowledge Graph, contributing to metadata harmonisation across infrastructures and disciplines.

        Speakers: Björn L. Saß (Helmholtz-Zentrum Hereon), Dr Romy Fösig (Karlsruhe Institute of Technology)
      • 32
        Marine Data - Research Vessels: AI-readiness underway

        What makes research data AI-ready? AI-readiness is often considered a property of individual datasets. For heterogeneous observational data, however, it emerges from the infrastructure and processes connecting data acquisition, metadata, quality control, processing, publication and reuse.
        The Marine Data – Research Vessels project (formerly “Underway” Research Data), coordinated by the German Marine Research Alliance (DAM), addresses this challenge for underway data obtained from permanently installed scientific sensors aboard German research vessels. Building on previous phases, it establishes standardised processes and infrastructures across the data chain – from acquisition at sea and transfer to shore, through quality control and metadata management, to FAIR publication, international data services and reuse. This demonstrates how AI-readiness can become an infrastructure capability rather than something added retrospectively to datasets. Published bathymetry data already support AI-based identification of geomorphological structures on the seafloor.
        Readiness is underway, not finished. The current phase advances the infrastructure through standardisation, digitalisation and automation. Automated sensor metadata flows, integration of OSIS and O2A REGISTRY, persistent identifiers and defined metadata sources will make measurement context machine-actionable and scalable. Existing underway datasets will be evaluated for AI-readiness and suitability for machine learning with data science experts. Based on this assessment, AI-supported methods for quality control, error detection and classification will be developed and integrated into processing workflows to improve efficiency and scalability.
        These developments are pursued with project partners, including AWI, GEOMAR and Hereon, combining expertise in marine observations, data infrastructures, data management and data science.

        Speaker: Marcus Krüger (Deutsche Allianz Meeresforschung)
      • 33
        FAIR enough – a Data Journey from Observations to Analysis and Archives

        The Earth System sciences are disciplines that generate tremendous and ever-growing amount of data in many heterogeneous forms. Thus, tools and frameworks to manage these data are becoming increasingly important in order to promote, foster, and safeguard the “FAIR Guiding Principles for scientific data management and stewardship” [1]. The O2A data flow framework [2] presents a holistic approach to handle (meta-)data, from observation through analyses and archives. It offers solutions for metadata management, and for the collection, provision, monitoring, collaborative analysis/processing, exploration, and publication of research data. To ensure integrity, provenance, and interoperability, O2A makes use of standardised interfaces (e.g. STA, WMS, CSW), well-known persistant identifiers (ORCID, ROR, DOI, …), and facilitates the use of controlled vocabularies (like NERC, QUDT, AMETSOC). In addition to the graphical user interfaces, REST APIs enable machine interaction, ranging from private scripts to mobile apps. Hence, O2A allows flexible, time-saving, and sustainable working in a FAIR manner. In this talk, we would like to demonstrate the practical benefits it brings in the everyday work of scientists and data managers. We will explore how the O2A bagpipe is played and highlight some examples of its applications in practice.

        [1] https://doi.org/10.1038/sdata.2016.18
        [2] https://doi.org/10.1007/s12145-026-02204-9

        Speakers: Nico Harms (AWI), Noemi Ruegg (AWI)
      • 34
        Beyond FAIR: A Framework for AI-Ready Data in Earth and Environmental Science

        The increasing adoption of artificial intelligence (AI) and machine learning (ML) in Earth and environmental sciences places new demands on scientific data infrastructures that extend beyond traditional FAIR (Findable, Accessible, Interoperable, Reusable) principles. While FAIR provides a strong foundation for data sharing and reuse, it does not fully address requirements such as provenance transparency, bias documentation, uncertainty quantification, target annotation, and machine-actionable metadata that are critical for AI-driven workflows. In this contribution, we present a practical framework for assessing and improving AI readiness within distributed Earth and Environment data infrastructures, developed in the context of the Helmholtz E&E DataHub. The framework was derived through a review of existing AI-readiness concepts, analysis of community needs, and discussions among data providers, infrastructure architects, and AI practitioners. It defines a set of criteria organized into five dimensions: data preparation, data quality, documentation, access, and supervised-learning targets. Rather than prescribing specific technologies or maturity scores, the framework emphasizes transparency and fitness for purpose, enabling users to evaluate whether datasets are suitable for particular AI applications. The applicability of the criteria is illustrated through use cases ranging from analysis of environmental observations to data-driven forecasting. We argue that AI readiness is inherently use-case dependent and cannot be reduced to a single metric. Instead, transparent documentation of data characteristics provides a common basis for communication between data producers, infrastructure operators, and AI developers, supporting the development of interoperable and AI-capable scientific data spaces.

        Speaker: Tobias Weigel (Helmholtz-Zentrum Hereon)
    • 10:30
      Coffee Break
    • Session 2.2 - AI-ready data as Infrastructure (Posters)
      • 35
        A Sustainable Spatial Data Infrastructure for the Automated Integration of Distributed Research Data

        The growing demand for discoverable and accessible research data and metadata, driven by FAIR principles and user requirements for data portals, repositories, and search engines, has led to an increasing popularity of interactive, particularly map-based, data exploration. However, aligning technical realities with custom user visions in the design of sustainable services for interactive map viewers remains challenging.

        To overcome these challenges, the O2A Spatial framework has been developed, enabling rapid and low-effort deployment and configuration of classic Spatial Data Infrastructure (SDI) components such as storage, databases, geo web servers, and catalogues. Additionally, it facilitates the creation and curation of data products compiled from diverse sources, including the PANGAEA repository, the Observations to Analysis and Archives (O2A) pipeline, Sensor Observation Services, and data provided by scientists directly. Hence, simple metadata harmonisation is offered.

        Publicly available Standard Operating Procedures and data exchange specifications guide scientists and institutions in providing their data products as standard-compliant OGC web services, further contributing to their FAIR status.

        This modular, scalable, and highly automated SDI has been developed and operated at the Alfred Wegener Institute for over a decade, continuously improving and providing map services for GIS clients and portals, including the Marine Data and Earth Data Portals. Long-term maintainability is ensured through the use of common open-source technologies, established geodata standards, containerization, and extensive automation. The modularity of O2A Spatial and SDI components ensures flexibility and future expandability. Embedded within O2A, development and operation are financially and staff-wise secured in the long term.

        Speaker: Peter Konopatzky (AWI)
      • 36
        AMPBA & sedaECHO: Modular Snakemake Workflows for Reproducible Amplicon and Ancient DNA Metagenomic Analysis

        High-throughput sequencing is now central to biodiversity assessment and paleoecological reconstruction, both relying on multi-step pipelines whose choices shape ecological interpretation. Existing solutions are either flexible but poorly documented scripts, or standardized but rigid, infrastructure-heavy platforms, leaving little room for adaptable tools. We address this gap with two complementary, Snakemake-based workflows for different sequencing strategies.

        AMPBA (Accessible Metabarcoding Platform for Biodiversity Analysis; https://helmholtz.software/software/ampba-workflow) addresses amplicon metabarcoding for biodiversity assessment and monitoring. Because markers differ in taxa, evolutionary rate, and resolution, AMPBA decouples feature generation from any fixed biodiversity unit: denoising, clustering, and cooccurrence collapse each build a new unit (ASVs, swarm/OTU clusters, cASVs/cOTUs) from whichever precedes it, while NUMT filtering, decontamination, and replicate merging refine rather than redefine it; taxonomic assignment switches independently between marker-specific classifier/database pairs.

        sedaECHO (https://gitlab.awi.de/data-science-team/sedaecho) addresses shotgun metagenomic analysis of sedimentary ancient DNA (sedaDNA), enabling high-resolution reconstruction of past ecosystems by capturing taxa that rarely appear in the fossil record. It ensures reproducibility through explicit rule definitions, pinned software versions, and traceable provenance, integrating established tools for taxonomic and functional classification and damage pattern authentication, and allowing multiple methods to run side by side under identical conditions.

        Both pipelines share the same design philosophy: locally installable, workflow-managed, and modular, reducing setup complexity and enabling systematic comparison of methods. Together, they support standardized, reproducible analysis of environmental and ancient DNA across ecological and paleoecological research.

        Speaker: Stefan Neuhaus
      • 37
        Benefits when using the O2A cruise planning and preparation recommendations for data accessibility

        Preparing data management ahead of a cruise through O2A REGISTRY’s mission management pays off well beyond the campaign itself. Standardized item sets and persistent identifiers assigned during cruise planning carry through into post-cruise data handling, making file-based data — such as Multibeam, Oceanographic, Weather, Seismic, and Timeseries data — together with their metadata, considerably easier to find, access, and reuse. Here we present examples from recent campaigns illustrating these gains in data availability and comfort.

        Speaker: Maximilian Betz
      • 38
        Cruise Planning From a Research Data Management Perspective: Mission Management with O2A Registry

        Every campaign aboard a research vessel is facing unique challenges and uncertainties, nonetheless the chances for standardization and harmonization in data flows are manyfold to come FAIR (meta)data a bit closer. Hence, we consequently make use of the mission management system of O2A REGISTRY to i) create sets of shipborne standard items, ii) (semi-)automatize the cruise preparation from a data management perspective, iii) benefit from PIDs of items for subsequent systems and processes, iv) create a self-service for principle investigators and their equipment. Here we present our experiences from the introduction for a selection of vessels, what controlled vocabularies have to do with it, and why everyday work with DSHIP can be challenging.

        Speakers: Noemi Ruegg (AWI), Norbert Anselm
      • 39
        Futureproofing Sensor and Spatial Data Infrastructure for Growing Data Demands

        The Helmholtz Coastal Data Center (HCDC) operates and maintains the GeoHub, a data portal enabling Hereon scientists to share information internally, with external scientists and the public. It hosts data products, exposed through OGC-compliant endpoints and web applications such as maps, dashboards, and data-centric websites.

        In the context of increasing pressures on coastal and climate systems, the GeoHub contributes to scientific collaboration and informed decision-making by facilitating easy access to crucial environmental data.

        The smooth operation of this data portal and its underlying infrastructure requires more than monitoring uptime. End-to-end observability includes assessing data and metadata quality, controlling ingestion pipelines, and ensuring access to responsive endpoints and reliable visualisations. These aspects are increasingly important as data are consumed not only by scientists and applications, but also by automated workflows and AI-based systems.

        Ongoing developments in creating Digital Twin products and AI applications for coastal and climate science increase demands on infrastructure, while simultaneously requiring AI-ready data that are consistent, discoverable, and machine-readable.

        In this poster, we present our approach to futureproofing the GeoHub, ensuring the availability and reliability of our data infrastructure and discuss improvements to meet continuously increasing demands.

        Speaker: Markus Benninghoff (Helmholtz-Zentrum Hereon)
      • 40
        instrugram

        The project instrugram aims to simplify and standardize the registration of proper instrument metadata and the inclusion of PIDs. This will be done with a templating system based on the already existing PIDINST scheme, to prepare complete metadata sets for various instrument types. A software component to register metadata from those templates at exchangeable target systems will be developed.

        Speaker: Sylvia Reißmann
      • 41
        Introducing a Platform for Permafrost Essential Climate Variable Data

        The Global Terrestrial Network for Permafrost (GTN-P) is the primary international programme dedicated to sustained, long-term monitoring of permafrost. Recently, the network was accredited by the Global Climate Observing System (GCOS) as an Affiliated Network, reflecting the importance of permafrost as one of the 55 Essential Climate Variables (ECVs) defined by GCOS for characterizing Earth’s climate. Bringing together researchers and institutions from more than thirty countries, GTN-P collects in-situ data on the Permafrost ECV quantities Permafrost Temperature (PT) and Active Layer Thickness (ALT), with Rock Glacier Velocity (RGV) to be included soon. These long-term observations are fundamental to understanding how permafrost responds to a warming climate and to assessing the future greenhouse gas contributions expected from permafrost thaw.

        The new GTN-P Data Platform (https://data.gtn-p.org) was launched earlier this year following extensive community consultation, marking a major step forward in data management and accessibility. It enables users to explore data availability, access standardized metadata, and visualize temperature-depth time series, while a public API provides programmatic access to data and metadata and facilitates interoperability with other systems and tools. The underlying infrastructure is hosted at the AWI Computing and Data Center and designed for long-term reliability, scalability, and maintainability. Persistent identifiers (PIDs) provide stable references for PT and ALT sites, datasets, and related resources. Starting next year, annual data compilations will be published through the PANGAEA World Data System as DOI-citable data products, with all contributors credited as authors to ensure recognition and ownership remain with the original data providers. Together, these features establish a reliable and accessible infrastructure for the long-term management, sharing, and use of global permafrost data.

        Speaker: Tillmann Lübker
      • 42
        MetaParse

        Imaging systems now acquire images faster than their contents can be annotated by hand, making cataloguing the bottleneck and leading to increasing costs with collection size. Existing methods only address a part of the problem: unsupervised segmentation and embedding group visually similar objects but don’t attach labels to the group; supervised detectors return interpretable classes while demanding exhaustive per object annotation; vision-language models describe entire scenes without reference to object-level detections. We therefore couple all three into a single pipeline and evaluate on seafloor imagery.
        Objects are first segmented and embedded with pretrained foundation models and then clustered without labels. A domain expert reviews a small, fixed number of representative crops from each cluster, and every accepted cluster is promoted to a trained detector. Because that number is fixed, annotation effort per category no longer scales with how many instances the collection happens to contain. Finally, a vision-language model captions each image, and the captions are cross-checked against the detections.
        Clustering runs to completion on collections of several hundred images. Detectors bootstrapped from cluster-level review require an order-of-magnitude less annotation effort than a fully supervised baseline trained on exhaustive per-object annotation. A systematic accuracy comparison is left to future work. The captioning stage has so far been assessed qualitatively; quantitative evaluation is in progress. The pipeline was tested on a single imaging domain.

        Speaker: Stefan Pinkernell (Data Science, Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research)
      • 43
        O2A SAMPLES: An Interoperable Framework for FAIR, Traceable, and Sustainable Sample Management

        Field research generates diverse and valuable samples under demanding conditions, yet their long-term usability is often constrained by fragmented data management and heterogeneous workflows. The O2A Sample Management System (O2A SAMPLES) addresses these challenges by providing a sustainable, interoperable platform for traceable, FAIR-compliant, and AI-ready sample metadata across disciplines. By integrating established research infrastructures, including the O2A REGISTRY, PANGAEA, DSHIP, and the Earth Data Portal, the system enables standardized sample registration, storage management, loans, and transparent availability for collaboration. Automated metadata enrichment from connected infrastructures reduces manual effort while improving metadata completeness and consistency. In addition, end-to-end traceability and linking to Nagoya Protocol documentation support legal compliance. Standardized workflows, standard operating procedures (SOPs), and QR code-based tracking enhance reproducibility and accessibility throughout the sample lifecycle. As a sustainable long-term service, O2A SAMPLES avoids singular, short lived solutions by establishing a unified digital framework that improves sample discoverability, interoperability, and cross-institutional collaboration from field acquisition to digital repository for environmental research.

        Speaker: Maren Rebke (AWI)
      • 44
        O2A SEQUENCES: A FAIR Platform for End-to-End Metabarcoding Data Management

        Metabarcoding, the sequencing of marker genes from environmental samples to identify the or-ganisms present, is central to biodiversity and environmental research, and its data volumes are growing quickly. Turning these data into a lasting, shareable resource depends on keeping se-quences, their metadata, and analyses connected and archive-ready, which remains difficult to do consistently across groups and tools.

        O2A SEQUENCES (https://sequences.o2a-data.de) provides this missing layer. Embedded in AWI's O2A (Observations to Analysis and Archives) framework, it manages raw sequences from ingestion to archive submission. Researchers upload their sequences and complete a standardized set of metadata fields aligned with the HARMONise schema (https://harmonise.awi.de), which builds on the MIxS standard. The platform runs standard, pre-configured workflows reproducibly, and it lowers the barrier to archiving by assembling raw se-quences together with their metadata into submission-ready files that follow European Nucleotide Archive (ENA) and German Federation for Biological Data (GFBio) conventions.

        Offered as software-as-a-service, O2A SEQUENCES makes metabarcoding data FAIR, AI-ready, and reusable across centers, providing a foundation for sustainable research-data man-agement that extends to other environmental data types.

        Speaker: Alex Savchik (Alfred Wegener Institute, Helmholtz Centre for Polar and Marine Research (AWI))
      • 45
        OSIS – advancing ocean science metadata

        Over the past decade, German marine research centers have increasingly collaborated to improve interoperability between their locally developed information systems and to establish shared infrastructure. Today, the German Marine Research Alliance (DAM) coordinates this collaboration and the joint development of a federated marine information and data infrastructure.
        One component of this infrastructure is OSIS, the Ocean Science Information System developed at GEOMAR. OSIS serves as a central integration point for expedition-related metadata, connecting information on planned, ongoing and completed expeditions, scientific activities, instruments and equipment, responsible personnel, and resulting data publications. It supports researchers, technical staff, cruise planners, and data managers in documenting and tracking the information flow from expedition planning and individual observations through to data publication and backwards. The interlinking of metadata and data throughout their entire research cycle is openly accessible through a web interface and a public API, allowing OSIS to serve both human users and other information systems.
        Rather than replacing specialized systems, OSIS connects them within a federated information infrastructure. Participating institutions share responsibility for keeping information up to date through curation and automated synchronization with authoritative source systems. In particular, information on instruments, their deployment, maintenance, and status is exchanged with REGISTRY, the sensor management system operated by AWI within its observation-to-archive (O2A) framework. The interaction with DSHIP aboard research vessels integrates operational vessel data, underway observation and science activities into data provenance records. We will present recent enhancements in OSIS that further advance metadata management, interoperability and data provenance documentation across the ocean science research lifecycle.

        Speaker: Carsten Schirnick (GEOMAR Helmholtz Centre for Ocean Research)
      • 46
        Persistent Identifiers in Practice: GEOMAR’s PID Landscape

        Persistent Identifiers (PIDs) are key to FAIR and reproducible research. At GEOMAR Helmholtz Centre for Ocean Research Kiel, we use a range of PIDs, including ORCID for persons, ROR for organizations, IGSN for samples, and DOIs for publications, software and datasets, as well as Handles for the later. Implementation varies across systems, with many manual processes still manual and lacking missing standardized workflows. As part of the Helmholtz DataHUB Earth and Environment we plan to implement a common PID workflow to enhance interoperability.
        Here we present an overview of GEOMAR’s PID landscape, covering systems, responsibilities, and tools, and highlights challenges such as limited automation, platform interoperability, and gaps in policy and documentation. We also outline ongoing improvements, including API-based registration and coordinated workflows as well as support and training.
        By sharing our PID structures, we aim to engage the community, gather feedback, and identify opportunities to enhance interoperability and efficiency in institutional PID management.

        Speaker: Gregor Börner (GEOMAR)
      • 47
        Polarsternchen - Demonstrating O2A Near Real-Time Data Flows

        A LEGO research vessel demonstrating how near real-time data flows through O2A. Polarsternchen is a hands-on model closely following the path that data from research vessels take through the O2A dataflow framework, from metadata and setup description through automated data ingest to data monitoring.
        The demonstration will include the Mobile Event Log App for O2A Registry, showcasing its QR code feature, enabling scientists to record configuration changes and events in the field.

        Speaker: Noemi Ruegg (AWI)
      • 48
        Project-oriented Knowledge Graphs for Metadata Alignment Across Repositories

        Knowledge graphs enable the structuring and linking of scientific metadata from heterogeneous sources, supporting advanced search and analysis. Their effectiveness, however, depends on the level of metadata consistency and standardization. Graphs linking high-quality topical information can also be very useful for training tailored LLMs and AI models for specific purposes. With this activity, we aim to identify and extract topical datasets from repositories, improve and harmonize discipline-specific information in the metadata, turn them into graphs, and make them available as training data for AI applications. For this purpose, we explore an approach to building knowledge graphs from project metadata in scientific repositories. We use projects as a linking layer to integrate and align datasets, as well as their associated methods, instruments, keywords, and geographic descriptions. Metadata is extracted and normalized from structured fields and textual descriptions and then subjected to semantic categorization. The poster presents the methods applied, challenges encountered, and results, including statistics on harvested metadata, category coverage, and observed differences in metadata representation across repositories. We also discuss the potential of project-oriented knowledge graphs for improving the quality and interoperability of scientific metadata.

        Speaker: Stanislav Malinovschii (GEOMAR (HMC))
      • 49
        SimShare - The business card for your Earth-System Simulation

        Numerical model simulations are an essential part of earth system research. They help to reconstruct and to forecast the state of the earth system with varying complexity and resolution in space and time. In order to accomplish this task in a sustainable and verifiable manner it is imperative to adhere to the FAIR principles : Findable, Accessible, Interoperable, Reuse. While the output data of such simulations is often quite good documented, this is rarely the case for their very foundation, which is the code, the build process and the runtime environment. Modern policies of today's' journals often require the code behind model simulations to be published together with the analysis at least in the form of code repositories or fixed archive files (e.g. tar-balls or similar). But this only ensures in parts the reproducibility (e.g. no documentation on input data), maybe the interoperability (sheer access to (parts of) the code, but different formats: git, svn, tar balls, zip archives,...) and maybe in parts the accessibility (repo access permission; might be restricted due to license). It falls short of the findability and full accessibility (code and input files). The SimShare metadata fills this gap, by retaining that information that cannot be retracted easily or at all from the code automatically. The file format is YAML and is therefore machine readable and can be used to build searchable databases and re-build simulations

        Speaker: Dr Markus Scheinert (GEOMAR)
    • Closing
    • 12:45
      Lunch
    • Open Workshops - Suggest your workshop topic!