Our domain experts from Area A (Synthesis), Area C (Theory and Computations), and E (Use Cases) will present recent developments in NOMAD that enable FAIR research data management in computational and experimental materials research. Their talks will highlight practical approaches to implementing the FAIR data principles across condensed-matter physics.
Research Data Management for High-throughput DFT Calculations Using NOMAD Oasis
Speaker: Vikrant Chaudhary
Date & Time: Tuesday, March 10; 10:30–10:45
Location: BEY/0138
Session: FM 6: Focus Session: Materials Discovery I – Material informatics
Abstract:
NOMAD Oasis is an open-source, locally installable version of the central NOMAD software [nomad-lab.eu] [1]. Importantly, it allows users to add custom extensions such as data schemas and workflows. In this work, we showcase two representative cases implemented in our local NOMAD Oasis: (I) high-throughput spin Hall conductivity calculations for 4486 2D materials [2], and (II) medium-throughput bulk-photovoltaic effect in 549 experimentally available two-dimensional hybrid perovskites. In both cases, data generation relies on complex computational workflows involving multiple software packages, including VASP, Wannier90, and WannierBerri. We implement a customized parser for Wannier90, and a locally developed parser for WannierBerri, with the resulting data organized based on NOMAD’s built-in workflow functionalities. These examples show how NOMAD Oasis can simplify large-scale computational data management with custom parsers, automated workflows, unrestricted storage and uploads, and enhanced privacy. It also ensures compliance with the FAIR principles, making data Findable, Accessible, Interoperable, and Reusable.
[1] M. Scheidgen et al., JOSS 8, 5388 (2023).
[2] Li, F. et al., arXiv:2509.13204 (2025).
Finding Interoperable Datasets in Diverse Databases via Provenance and Similarity Analysis
Speaker: Martin Kuban
Date & Time: Tuesday, March 10; 10:45–11:00
Location: BEY/0138
Session: FM 6: Focus Session: Materials Discovery I – Material informatics
Abstract:
Collecting data from different sources has the potential to significantly increase the amount of available data for data-driven discovery. However, different producers of data use distinct methods and setups, e.g., approximations and parameters used in computational data, to achieve the best data quality for the properties that are studied in a specific project. Bringing these data together requires to understand the impact of the method and setup on the accuracy and precision of the produced data.
In order to do so, two key requirements must be fulfilled: First, the provenance and metadata of each data point need to be recorded. This can be achieved by leveraging the NOMAD infrastructure[1, 2], an ecosystem of parsers, schemas, and workflow tools, to extract rich metadata and provenance information. Second, using this information, similarity metrics can be used to identify data that achieve similar precision besides distinct computational setups[3]. We showcase our approach on different examples using data from NOMAD.
[1] Draxl and Scheffler, MRS Bulletin 9 (2018), 676-682.
[2] Scheidgen et al., Journal of Open Source Software 8 (2023), 5388.
[3] Kuban, et al., MRS Bulletin 47 (2022), 991–999.
Building a FAIR Community around Parsing
Speaker: Nathan Daelman
Date & Time: Wednesday, March 11; 10:45–11:00
Location: SCH/A216
Session: MM 22: Data-driven Materials Science: Big Data and Workflows III
Abstract:
NOMAD [nomad-lab.eu][1, 2] is an open-source data infrastructure for materials science data. One of its most praised features is how NOMAD allows for direct ingestion of various software output formats. This gives data producers access with minimal effort to the whole toolkit infrastructure system regardless of their choice of simulation code. As the NOMAD community extends into related scientific disciplines, parsing procedures should grow alongside and empower casual users to contribute too. To this end, I will be presenting two new parsing frameworks: (i) Mapping Annotation which connects code-specific formats to the NOMAD interoperable schema, while gracefully handling syntatic concerns; (ii) an agentic LLM interface for hooking up third-party parsers via the Model Context Protocol (MCP). Finally, I will highlight how both approaches fit into NOMAD Plugins and NOMAD Actions.
[1] Scheidgen, M. et al., JOSS 8, 5388 (2023).
[2] Scheffler, M. et al., Nature 604, 635-642 (2022).
Support for Self-driving Labs within the NOMAD Ecosystem
Speaker: Sarthak Kapoor
Date & Time: Thursday, March 12; 17:15–17:30
Location: BEY/0127
Session: AKPIK 6: AI Methods for Physics and Materials Science
Abstract:
Self-driving laboratories (SDLs) rely on robust digitization, structuring, and analysis of experimental data. We present NOMAD [nomad-lab.eu] [1] as a comprehensive research data management and workflow ecosystem that addresses the challenges inherent to emerging SDLs. The NOMAD ecosystem supports direct interfacing with lab instruments and addresses the transformation of instrument outputs into machine-actionable formats, a key requirement in SDLs, through a flexible schema system that allows researchers to represent raw data as standardized entries based on community-developed or laboratory-specific definitions. NOMAD Actions provide a robust framework for defining, executing, and monitoring sophisticated analysis and decision-making SDL workflows, such as ML pipelines and Bayesian optimization strategies. Moreover, NOMAD's workflow storage framework facilitates detailed provenance tracking, along with tools for navigating workflow graphs. Together, these capabilities position NOMAD as a foundational toolkit for realizing scalable, reliable, and FAIR SDLs.
[1] Scheidgen, M. et al., JOSS 8, 5388 (2023).
FAIR and Flexible Workflow Support within the NOMAD Infrastructure
Speaker: Joseph F. Rudzinski
Date & Time: Thursday, March 12; 17:45–18:00
Location: BEY/0127
Session: AKPIK 6: AI Methods for Physics and Materials Science
Abstract:
NOMAD [nomad-lab.eu] [1, 2] is an open-source, community-driven research data infrastructure designed for modern physics. It provides FAIR-compliant storage, management, and analysis for diverse computational and experimental materials science data, and its modular, plugin-based architecture enables low-barrier extensions for adjacent and interdisciplinary domains. Here we present NOMAD*s workflow capabilities as a foundation for scalable and AI-ready data pipelines. A general workflow schema supports both standardized and custom workflows that record detailed provenance and link heterogeneous data streams. Standardized workflows enable powerful search, visualization, and automation features, while custom workflows support agile, project-specific digitalization. Workflow entries can be created via Python-based plugins, a YAML workflow specification, or the NOMAD ELN interface, ensuring accessibility for researchers with varying technical backgrounds. Combined with a toolkit for high-throughput interfacing, NOMAD provides a robust and sustainable digital infrastructure across physics subdisciplines.
[1] Scheidgen, M. et al., JOSS 8, 5388 (2023).
[2] Scheffler, M. et al., Nature 604, 635-642 (2022).