Exploration

ResearchFeatured

Discovery

DiscoverSourcesQuality

Analysis

Working setReviews
Flow StudioTeamConcept
Settings

Partners

  • AI AlliancePrime
  • BrightQueryBuilds Meridian
  • OpenMinedFunded partner
  • MLCommonsFunded partner
  • Hugging FaceDeployment platform
See the full consortium and what each partner wires

Meridian is the discovery layer for research data, built by BrightQuery within the AI Alliance.

hybrid · semantic + lexical · 20 datasets ranked · 1.09s

Structurecomposite1tabular1
Depthcataloged18measured2
Licensenon commercial20
Accessopen20
Formatzip11netcdf5csv4pdf1
Sourcezenodo20
clear
1-20 of 20sortrelevancemeasured firstqualitysize
tabular

Dataset Priming against Salmonella enterica affects differentially haemocyte sub-populations in Armadillidium vulgare

0.00

Pailler, Louis

207 rows × 1 cols · 22 KB · csv

1 text

These datasets were collected using a total of 585 Armadillidium vulgare females. The objective of this study was to measure survival after immune priming with Salmonella enterica , using different doses and inactivation methods, and to characterise the underlying cellular mechanisms using flow cytometry. Survival analysis according to S. enterica method of inactivation and dosage Survival after injection of a lethal dose of live S. enterica was measured after immune priming using bacteria inactivated by heat (HK) or paraformaldehyde (PFA) and two different doses (10^3 and 10^6). This allowed to collect the Dataset_survival.csv, analysed using the Script_survival_analysis.R. Dataset_survival.csv: Ind: individual identification Treatment: priming treatment that females received. C : control females, no priming injection. PBS : females primed with sterile PBS. PFA3 / PFA6 : females primed with 10^3 or 10^6 PFA-inactivated S. enterica . HK3 / HK6 : females primed with 10^3 or 10^6 heat-inactivated S. enterica . Repl: experimental replicate Status: 1 = dead, 0 = live Survival : Time at death. 168 indicates living females at the end of the experiment Time: Time elapsed between the two injections. T24 : 24hours. T7 : 7 days Haemocyte sub-populations analysis according to priming treatment Following the survival experiment, and because the A. vulgare primed with PFA6 inactivated S. enterica exhibited the highest survival rates against LD50 infection, this inactivated method and dosage were used to examine the haemocyte sub-populations of A. vulgare mounting immune priming. Haemolymph was sampled from all females (C, PBS, PFA6) either 2 days (2D-PP) or 6 days (6D-PP) after the initial priming injection, or 2 days after the LD50 injection (2D-LD50). Distinct sets of females were used for each time point. This allowed to generate the Dataset_cytometry.csv, analysed using the Script_cytometry_pca_analysis.R. Principal Component Analysis (PCA) for each observation time allowed to extract projection values of the four PCs for each individual (Dataset_pca_ind_coord.csv). Dataset_cytometry.csv: Ind: individual identification Treatment: priming treatment that females received. C: control females, no priming injection. PBS: females primed with sterile PBS. PFA6: 10^6 PFA-inactivated S. enterica . Repl: experimental replicate Box: experimental box Exp: experimental day Time: Sampling time. 2D-PP: 2-days after priming. 6D-PP: 6-days after the priming injection. 2D-LD50: 2 days after the LD50 infection. P1_percent: Percentage of the first population. P1_FSCA: Cell diameter (size) of the first population. P1_SSCA: Internal complexity (internal granularity) of the first population. Viab_P1: Viability of the first population. P2_percent: Percentage of the second population. P2_FSCA: Cell diameter (size) of the second population. P2_SSCA: Internal complexity (internal granularity) of the second population. Viab_P2: Viability of the second population. Dataset_pca_ind_coord .csv: Ind: individual identification Treatment: priming treatment that females received. C: control females, no priming injection. PBS: females primed with sterile PBS. PFA6: 10^6 PFA-inactivated S. enterica . Repl: experimental replicate Box: experimental box Exp: experimental day Time: Sampling time. 2D-PP: 2-days after priming. 6D-PP: 6-days after the priming injection. 2D-LD50: 2 days after the lethal dose infection. Dim.1: projection values on the first principal component (PC1) for each individual. Dim.2: projection values on the second principal component (PC2) for each individual. Dim.3: projection values on the third principal component (PC3) for each individual. Dim.4: projection values on the fourth principal component (PC4) for each individual.

open·CC-BY-NC-4.0·Zenodo·0% null·completeSource
composite

Data of lithium loss in the copper foil

0.00

Li, Tong · Bresser, Dominic

1 files · 100 MB · zip

open·CC-BY-NC-ND-4.0·Zenodo·completeSource
declared

Study habits of students at Universidade Atlântica

0.00

Mendes, Ana · Agonia Pereira, Luís · Vairinhos, Valter

1 files · 83 KB · zipdeclared

Dataset and commented Python program that reproduce, end to end, all the quantitative and lexical analyses of a study on study habits, self-regulation and well-being in hybrid higher education (n = 69). The associated article is currently under review; its full reference will be added upon acceptance. The package includes the anonymised questionnaire data, the reproduction script (fixed random seed), the exact library versions (requirements.txt) and full documentation of the composite indices.

open·CC-BY-NC-4.0·Zenodo·completeSource
declared

Planet4Health Project: mHM Model Runs in South African domain at 0.015625deg resolution - Soil Water Content Layers 3 & 4

0.00

Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo

70 files · 34 GB · netcdfdeclared

Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil water content layers 3 and 4, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables swc_l03: Soil water content layer 3 (150-300 mm depth) [mm] swc_l04: Soil water content layer 4 (300-500 mm depth) [mm] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Planet4Health Project: mHM Model Runs in South African domain at 0.015625deg resolution - Soil Moisture Layers 2 & 3

0.00

Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo

70 files · 34 GB · netcdfdeclared

Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil moisture layers 2 and 3, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables sm_l02: Volumetric soil moisture layer 2 (50-150 mm depth) [mm mm-1, fraction between 0 and 1] sm_l03: Volumetric soil moisture layer 3 (150-300 mm depth) [mm mm-1, fraction between 0 and 1] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Planet4Health Project: mHM Model Runs in South African domain at 0.015625deg resolution - Soil Water Content Layers 5 & 6

0.00

Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo

70 files · 33 GB · netcdfdeclared

Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil water content layers 5 and 6, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables swc_l05: Soil water content layer 5 (500-1000 mm depth) [mm] swc_l06: Soil water content layer 6 (1000-2000 mm depth) [mm] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Planet4Health Project: mHM Model Runs in South African domain at 0.015625deg resolution - Soil Water Content Layers 1 & 2

0.00

Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo

70 files · 34 GB · netcdfdeclared

Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625°. The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components for soil water content layers 1 and 2, consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ). The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch (https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Spin-up: 30-year spin-up using 1990-2019 ERA5 climatology Model version: v1.0 Setup Scope: Model run for domain 1020011530, post-processed and clipped. Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline (https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables swc_l01: Soil water content layer 1 (0-50 mm depth) [mm] swc_l02: Soil water content layer 2 (50-150 mm depth) [mm] 📫 Contact Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de Institution Helmholtz Centre for Environmental Research - UFZ, Department of Computational Hydrosystems

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

llm-conceptual-models-evaluation

0.00

Vasic, Iva · Reitemeyer, Benedikt · Fill, Hans-Georg

1 files · 849 KB · zipdeclared

These are code and supplementary material for the BIR 2026 conference paper on formal evaluation of LLM-generated conceptual models. Title: Evaluating LLM-Generated Conceptual Models: A Theoretical Approach for Formalizing the Calculation of Metrics on the Meta Level The work was financially supported by the Smart Living Lab , a joint project funded by the University of Fribourg, EPFL, and HEIA-FR.

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Planet4Health mHM Simulation Output for South Africa at 0.015625deg resolution

0.00

Modiri, Ehsan · Shrestha, Pallav Kumar · Samaniego Eguiguren, Luis Eduardo

72 files · 36 GB · netcdfdeclared

Historical Hydrological Simulations over the South African Domain (1990-2024) The mHM's simulations of the Planet4Health project This dataset contains historical hydrological simulations for the South African domain (domain 1020011530) conducted with the Mesoscale Hydrological Model (mHM) at a spatial resolution of 0.015625° . The simulation period spans 1990-2024 and was part of the Planet4Health (P4H) project, utilising the ERA5 meteorological forcing. This archive is prepared for DOI assignment and ensures long-term reproducibility. It includes relevant clipped NetCDF components (streamflow 'q', soil moisture 'sm_l01', domain mask, and uparea assets) , consistent with the infrastructure provided within the Helmholtz Centre for Environmental Research (UFZ) . The simulations were executed using a specific version of the mHM model with the SCC method for gauges, paired with the mRMv1.0 routing configuration. 🛰️ Simulation Details Model: Mesoscale Hydrological Model (mHM) Codebase: scc_for_gauges branch ( https://git.ufz.de/shresthp/mhm/-/tree/scc_for_gauges?ref_type=heads ) Spatial resolution: 0.015625° Temporal resolution: Daily Simulation period: 1990-2024 Simulation type: Historical simulation Model version: mHMv5.11.3 (Release mRMv1.0) Setup Scope: Model run for domain 1020011530, post-processed and clipped. Simulation Version: v1.0 Configuration & Modules The configuration utilises standard structural components with the SCC methodology. Modules included: Snow processes: Degree-day method Soil moisture: Feddes equation for evapotranspiration reduction Infiltration: Multi-layer Brooks-Corey-like approach Direct runoff: Linear reservoir exceedance method Potential evapotranspiration: Hargreaves-Samani method Interflow: Storage reservoir with nonlinear outflow Groundwater: Linear reservoir Routing: Adaptive time-step routing with mRMv1.0 mechanisms 📥 Input Datasets Meteorological Forcing: ERA5 (Hersbach et al., 2020) at a native input meteorological resolution of 0.25°, dynamically downscaled/mapped to model requirements. Processing Infrastructure: Tracked, processed, and validated under the Planet4Health deployment pipeline ( https://git.ufz.de/planet4health/mhm_production/-/tree/main/postproc?ref_type=heads ). Data Interfaces: Climate Data Interface version 2.2.4 (CDI) | Climate Data Operators version 2.2.2 (CDO) | NetCDF Operators version 5.1.7 (NCO). 📤 Output Variables q: Routed streamflow (discharge) [m3 s-1] sm_l01: Volumetric soil moisture of soil layer 1 (top 50 mm) [mm mm-1, fraction between 0 and 1] mask: Domain clipping structure [binary flag / dimensionless] uparea: Upstream catchment area matrix [m2] 📫 Contact For questions or collaboration inquiries, please contact: Ehsan Modiri - ehsan.modiri@ufz.de Pallav Kumar Shrestha - pallav-kumar.shrestha@ufz.de 📚 References Boeing, F. et al., 2022. Hydrol. Earth Syst. Sci. , 26, 5137-5161. Hargreaves, G.H. & Samani, Z.A., 1985. Applied Engineering in Agriculture , 1(2), pp.96-99. Hartmann, J. & Moosdorf, N., 2012. Geochem. Geophys. Geosyst. , 13(12). Hengl, T. et al., 2017. PLoS One , 12(2), e0169748. Hersbach, H. et al., 2020. QJRMS , 146(730), pp.1999-2049. Kumar, R. et al., 2013. Water Resources Research , 49(1), pp.360-379. Lehner, B. et al., 2011. Front. Ecol. Environ. , 9(9), pp.494-502. Rakovec, O. et al., 2016. J. Hydrometeorology , 17(1), pp.287-307. Rakovec, O. et al., 2022. Earth's Future , 10(3), e2021EF002394. Samaniego, L. et al., 2010. Water Resources Research , 46(5). Samaniego, L. et al., 2023. mhm-ufz/mHM: v5.13.1, Zenodo. DOI: 10.5281/zenodo.8279545 Thober, S. et al., 2019. Geosci. Model Dev. , 12(6), pp.2501-2521.

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Time Series Analysis for Environmental Data: An R-Based Introduction

0.00

Bunn, Andrew G.

1 files · 27 MB · zipdeclared

An applied, R-based introduction to time series analysis for environmental scientists and ecologists. It covers autocorrelation and stationarity, ARMA models, cross-correlation, regression with autocorrelated errors, trend detection, forecasting and reconstruction, and frequency-domain methods including wavelets, with an emphasis on building intuition and getting things done in R rather than mathematical derivation. The book originated as graduate course notes at Western Washington University.

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

brainWhiz: interactive multi-atlas exploding-brain viewer and figure tooling

0.00

Newman-Norlund, Roger

1 files · 51 MB · zipdeclared

brainWhiz is a static, single-page Three.js viewer for multi-atlas neuroimaging figures. It renders brain parcellations in 3D, colors regions by per-region CSV values or voxelwise .nii statistical maps (auto-matching a CSV to its atlas), draws DTI / resting-state connectivity, shows native-resolution slices and mosaics, and composes publication-ready figure panels — entirely in the browser. Bundled atlases, NeuroQuery task maps, and connectivity are third-party data with their own terms, for non-commercial research; cite the original sources.

open·CC-BY-NC-4.0·Zenodo·completeSource
declared

AMLNet - Synthetic Anti-Money Laundering Benchmark Dataset, Version 2.0

0.00

HUDA, SABIN

1 files · 729 MB · csvdeclared

AMLNet - Synthetic Anti-Money Laundering Benchmark Dataset, Version 2.0 Relation: Is new version of identifier: 10.5281/zenodo.16736515 AMLNet is a synthetic anti-money laundering benchmark dataset created for machine learning evaluation. This Version 2.0 release is associated with the paper: "AMLNet: A Knowledge-Guided Synthetic Benchmark for Machine Learning Evaluation in Anti-Money Laundering" The dataset contains a fixed synthetic benchmark instance generated using the AMLNet framework. It includes approximately 1.09 million transactions over a 195-day simulation period, from 13 October 2025 to 27 April 2026. The benchmark contains 1,411 suspicious transactions, corresponding to a suspicious rate of approximately 0.13%. The dataset is fully synthetic. The accounts, transactions, timestamps, locations, balances, metadata, labels, and customer activity patterns do not correspond to real customers, real institutions, or real banking activity. AMLNet was designed to support machine learning experiments under rare-event anti-money laundering conditions. The dataset includes ordinary transactions and suspicious transaction sequences representing structuring, layering, and integration patterns. Suspicious activity is embedded within ordinary account activity to make the detection task more realistic and non-trivial. CONTENTS This release includes: - the fixed AMLNet Version 2.0 synthetic transaction dataset; - transaction labels; - laundering typology labels; - transaction metadata; - dataset documentation and column descriptions. Evaluation scripts will be added to this Zenodo record after publication of the associated paper. The AMLNet generator source code is not included in this release due to security concerns. DATA FORMAT The main dataset is provided as a CSV file with the following columns: - step: Sequential simulation step or transaction index. - type: Transaction type, such as BPAY, CASH_OUT, DEBIT, EFTPOS, NPP, OSKO, PAYMENT, or TRANSFER. - amount: Transaction amount in Australian dollars. - category: Transaction category, such as housing, food, transport, recreation, healthcare, education, utilities, shell company, property investment, cryptocurrency, or other. - nameOrig: Originating account or customer identifier. - nameDest: Destination account, customer, or merchant identifier. - oldbalanceOrg: Originating account balance before the transaction. - newbalanceOrig: Originating account balance after the transaction. - isFraud: Binary label used for suspicious/fraudulent transaction detection. - isMoneyLaundering: Binary AML label, where 1 indicates suspicious money laundering activity and 0 indicates ordinary activity. - laundering_typology: Laundering typology label. Values include normal, structuring, layering, and integration. - metadata: JSON-style metadata containing timestamp, location, device information, payment method, risk indicators, and typology-specific details where applicable. - fraud_probability: Model-generated or risk-score field used in selected experiments. This field may be empty for some records. - hour: Hour of transaction. - day_of_week: Day of week. - day_of_month: Day of month. - month: Month number. LABELS The dataset includes two binary label fields: - isFraud: Binary suspicious/fraudulent transaction label. - isMoneyLaundering: Binary anti-money laundering label. The laundering_typology column provides the suspicious activity type for labeled suspicious transactions. Values include: - normal - structuring - layering - integration DATASET STATISTICS - Total transactions: approximately 1.09 million - Suspicious transactions: 1,411 - Suspicious transaction rate: approximately 0.13% - Simulation period: 195 days - Simulation dates: 13 October 2025 to 27 April 2026 - Generated customer accounts: 10,000 - Observed graph nodes: 11,000, including customer and merchant nodes - Transaction types: BPAY, CASH_OUT, DEBIT, EFTPOS, NPP, OSKO, PAYMENT, and TRANSFER - Payment identifiers: BSB_Account, CardNumber, and PayID - Laundering typologies: structuring, layering, and integration - Geographic setting: Australian synthetic banking context SPLIT AND EVALUATION PROTOCOL The associated manuscript describes the split protocol and evaluation protocol used in the reported experiments. For transaction-level experiments, transactions are ordered chronologically to reduce temporal leakage. For node-level graph experiments, the dataset is converted into an account/entity graph, and node features are computed from aggregated transaction statistics. The manuscript describes the chronological train, validation, and test protocol used for model evaluation. Evaluation scripts will be added to this Zenodo record after publication of the associated paper. SYNTHETIC DATA NOTICE This dataset is fully synthetic. It does not contain real customer records, real account information, real banking transactions, real IP addresses, or real financial institution data. All customer identifiers, merchant identifiers, balances, timestamps, locations, transaction patterns, labels, and metadata values were generated for research purposes. The dataset should not be interpreted as a sample of actual banking activity. GENERATOR SOURCE CODE The AMLNet generator source code is not included in this release due to security concerns. Requests for access to the generator source code may be considered for legitimate research purposes, subject to identity verification, institutional affiliation, and appropriate use conditions. The associated manuscript provides the generator pseudocode, main configuration details, demographic assumptions, regulatory constraints, laundering typologies, timing rules, routing rules, split protocol, and evaluation protocol. This public release provides the fixed generated dataset instance used in the associated paper. Evaluation scripts will be added to this Zenodo record after publication. USAGE This dataset can be used for: - anti-money laundering machine learning research; - rare-event classification experiments; - transaction-level suspicious activity detection; - account-level graph classification; - feature ablation studies; - explainability experiments; - synthetic-to-real transfer learning research; - educational and academic research on financial crime detection. LICENSE This dataset is released under the Creative Commons Attribution-NonCommercial 4.0 International License. You may share and adapt the dataset for non-commercial purposes, provided that appropriate credit is given. Commercial use is not permitted without prior permission. For commercial licensing enquiries, please contact: s.huda@griffith.edu.au, hudasabin@gmail.com CITATION If you use this dataset, please cite the Zenodo dataset record and the associated paper: Huda, S., Foo, E., Newton, M. A. H., Jadidi, Z., Huda, S., Morgan, G., and Sattar, A. AMLNet: A Knowledge-Guided Synthetic Benchmark for Machine Learning Evaluation in Anti-Money Laundering. Under review. Dataset: Huda, S. et al. AMLNet Synthetic Anti-Money Laundering Benchmark Dataset, Version 2.0. Zenodo. 10.5281/zenodo.21237971 CONTACT For questions about the dataset, please contact: Sabin Huda School of Information and Communication Technology Faculty Member, Academy of Excellence in Financial Crime Investigation & Compliance Griffith University Email: s.huda@griffith.edu.au, hudasabin@gmail.com

open·CC-BY-NC-4.0·Zenodo·completeSource
declared

Dataset --- "LYNX: a deep generative model for linking spatial dynamics and cell interactions in multimodal spatial data"

0.00

Jin, Yinuo · Myers, Joshua · Rajbhandari, Presha · et al.

1 files · 20 GB · zipdeclared

Dataset overview Pre-aligned multi-modal liver dataset for sample NIH_F5 with 2D (single tissue section) and a 3D (serial sections section_01...08 ) variants. Each variant pairs two spatially co-registered modalities - Xenium (spatial transcriptomics / RNA) and DESI (mass-spectrometry imaging / metabolomics), with cross-modal aligned spatial coordinates stored in the SpatialData .zarr stores ( obsm/xenium_map , obsm/desi_map ). data/LYNX_liver_dataset/ ├── NIH_F5_2D/ # single 2D section │ ├── xenium/ # Xenium (RNA) │ │ └── cell_feature_matrix.h5 # cell × gene counts (.h5ad) │ └── DESI/ # DESI (metabolomics) │ ├── NIH_F5.h5 # procesed pixel × ion intensity matrix │ └── NIH_F5.ome.tif # raw ion-image file │ └── NIH_F5_3D/ # serial 3D stack ├── xenium/ │ └── section_{}/ # one Xenium bundle per section │ └── sdata.zarr └── DESI/ └── section_{}_sdata.zarr # aligned DESI SpatialData store per section Formats: .zarr - SpatialData/OME-Zarr stores (images + AnnData tables with aligned spatial / xenium_map / desi_map coordinates); .h5 - cell×gene (Xenium) or pixel×ion (DESI) matrices; .ome.tif - morphology / DESI ion images.

open·CC-BY-NC-ND-4.0·Zenodo·completeSource
declared

GDEE: A Structure-Based Platform for Gene Discovery and Enzyme Engineering

0.00

Souza, Caio · Correia, João

54 files · 16 MB · csvdeclared

This repository contains the second version of the metamodel-based rescoring framework for protein-ligand binding affinity prediction. The first version introduced a linear metamodel combined with random train/test splits, but did not explicitly account for structural similarity or information leakage between training, validation, and test partitions. In this updated version, we included: Python scripts for data preprocessing, model training, and evaluation Pairwise ΔΔG evaluation code and associated data Visualization scripts for generating publication figures This dataset and codebase are intended to support reproducibility and further development of metamodel-based rescoring strategies in structural bioinformatics and computational drug design.

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

Dados da Tese: Formação Docente em Inteligência Artificial para o Ensino Superior: dos Fundamentos Computacionais ao Uso Crítico e Pedagógico da IA Generativa.

0.00

Costa, Carlos Fransley Scatambulo

3 files · 614 KB · csv, pdfdeclared

Dados anonimizados referentes ao estudo na tese: Formação Docente em Inteligência Artificial para o Ensino Superior: dos Fundamentos Computacionais ao Uso Crítico e Pedagógico da IA Generativa .

open·CC-BY-NC-ND-4.0·Zenodo·completeSource
declared

Dataset for "A Knowledge-Based Multi-Agent Framework for Security Control Recommendation"

0.00

Fernández-Martínez, Carolina · Siddiqui, Muhammad Shuaib · Daza, Vanesa

1 files · 5.9 MB · zipdeclared

Dataset for "A Knowledge-Based Multi-Agent Framework for Security Control Recommendation" Authors : Carolina Fernández-Martínez (i2CAT / UPF), Shuaib Siddiqui (i2CAT), Vanesa Daza (UPF) This contains the dataset and its generating code, as used in Section 3 of the article "A Knowledge-Based Multi-Agent Framework for Security Control Recommendation", published in Elsevier's Knowledge-Based Systems in 2026. It provides a security control selection based on NIST SP 800-53 rev5 extended catalogue, Multi-Agent Influence Diagram, Game Theory and No-Regret-based utilities. Besides the dataset itself, the scripts used to correlate and curate this data from InfoSec and academic sources are provided, along with such sources and the links to their original sources. The main repository for the dataset and its code used in Section 3 as well as the code used in Section 4 can be found in GitHub . Section overview This work is structured as follows: . ├── dataset_analysis.py ├── dataset_contribution.py ├── dataset_curation.py ├── dataset_helpers.py ├── ground_truth │ ├── manual │ │ ├── csftools_stridelm.csv │ │ └── gemini3pro_secdims_impl.csv │ ├── papers │ │ ├── doi_10_1007_impl_control.csv │ │ ├── doi_10_1016_jisa_2025_104056 │ │ │ ├── doi_10_1016_jisa_2025_104056_raw.csv │ │ │ └── generation_scripts │ │ │ ├── controls_summary.xlsx │ │ │ ├── controls.xlsx │ │ │ ├── domain.py │ │ │ ├── groups.xlsx │ │ │ ├── main.py │ │ │ ├── patterns.gml │ │ │ ├── README.md │ │ │ ├── requirements.txt │ │ │ ├── software.xlsx │ │ │ ├── technique.xlsx │ │ │ ├── ttp-control.xlsx │ │ │ ├── ttps.xlsx │ │ │ └── utils.py │ │ ├── doi_10_1093_cybsec_tyaf020_mapping.csv │ │ └── doi_10_1093_cybsec_tyaf020_scores.csv │ └── standards │ ├── Cybersecurity_Framework_v2-0_Concept_Crosswalk_800-53_5_2_0_draft.csv │ └── NIST_SP-800-53_rev5_catalog.json ├── output │ └── dataset_curated.csv └── README.md The ground truth contains both manual mappings, academic papers and InfoSec standardised data: ground_truth : hosts sources used for the data curation. manual : data extracted manually, whether directly checking sources or iteratively requested to an LLM. csftools_stridelm.csv : manually extracted data from CSF tools indicating the contribution of each control subfamily to mitigate a given STRIDE-LM threat. Each value follows a comma-separated format (e.g. "0,2,3,4,8,12") or use -1 if there is no contribution. gemini3pro_secdims_impl.csv : LLM-parsed data from CSF tools , requesting Gemini 3 Pro to extract data on the security control subfamilies: (1) whether these can be SW-implementable, (2) their coverage to the different Security Dimensions and (3) a text-based justification regarding such coverage. papers : doi_10_1007_impl_control.csv : dataset post-processed from that provided by paper with DOI:10.1007/s10664-025-10649-7 . Basically, this CSV assigns numeric codes to the column "Related to implementation-level feature? (yes/no)" from the tab "SP800 53 rev. 3 (technical cont" of the "2) Systematic Review - Security Standards.xlsx" file in that dataset. This can tabke the following values: 0 (if not SW-implementable), 1 (if SW-implementable), -1 (if undefined in the original dataset) or -2 (if the security control subfamily is not even present in the original dataset). doi_10_1016_jisa_2025_104056 : dataset provided by paper with DOI:10.1016/j.jisa.2025.104056 . generation_scripts : minor modifications to the original scripts to generate their dataset. See README.md inside. doi_10_1093_cybsec_tyaf020_mapping.csv : dataset post-processed from that provided by paper with DOI:10.1093/cybsec/tyaf020 . This CSV contains the table from "Appendix A" of the "Appendix A - D.docx". doi_10_1093_cybsec_tyaf020_scores.csv : dataset post-processed from that provided by paper with DOI:10.1093/cybsec/tyaf020 . This CSV contains the table from "Appendix C" of the "Appendix A - D.docx". standards : Cybersecurity_Framework_v2-0_Concept_Crosswalk_800-53_5_2_0_draft.csv : NIST resource that maps CSF 2.0 subcategories to security control subfamilies from SP 800-53 rev5. NIST_SP-800-53_rev5_catalog.json : NIST SP 800-53 rev5 catalogue of security control subfamilies as obtained from the full catalogue in JSON format. The output folder contains the generated, curated dataset by default. Upon running the scripts below, more files will follow. Generating the dataset and ancillary files 1. Curated dataset The curated dataset is generated under "output/dataset_curated.csv" after running the following script. This file is required for the other scripts. python3 dataset_curation.py 2. Summaries, statistics and figures The analysis on the dataset extracts statistics (in .csv and .tex files) and generates figures (in .pdf and .png) from the dataset: output dataset_curated_ciatunp.{csv,tex} : table with number of control families contribute to each Security Dimension. dataset_curated_stridelm.{csv,tex} : table with number of control families contribute to each STRIDE-LM threat. dataset_curated_score_summary.{csv,tex} : statistics for minimum, average, mode, maximum, standard deviation and inter-quartile range per control family. figures : dataset_curated_ciatunp_contribution_implementable_cats_ids.{pdf,png} : distribution of the contribution of SW-implementable control families (axis Z) and subfamilies (axis Y) towards security dimensions (axis X). dataset_curated_ciatunp_contribution_total_cats_ids.{pdf,png} : distribution of the contribution of all kinds of control families (axis Z) and sub families (axis Y) towards security dimensions (axis X). dataset_curated_score_contribution_total_cats_ids.{pdf,png} : distribution of the score (axis X, in deciles) for all kinds of control families (axis Z) and subfamilies (axis Y). dataset_curated_stridelm_contribution_implementable_cats_ids.{pdf,png} : distribution of the contribution of SW-implementable control families (axis Z) and subfamilies (axis Y) towards mitigating types of STRIDE-LM threats (axis X). dataset_curated_stridelm_contribution_total_cats_ids.{pdf,png} : distribution of the contribution of all kinds of control families (axis Z) and subfamilies (axis Y) towards mitigating types of STRIDE-LM threats (axis X). python3 dataset_analysis.py Besides this, the following script quantifies the contribution of the academic datasets and sources used during the process, generating these files: output dataset_curated_score_summary_contribution_ds_imp_{all,top}.tex : table comparing the score of each of the top SW-implementable control subfamilies across the curated dataset ("Total" column) and the datasets used from academic papers (other columns). dataset_curated_score_summary_contribution_ds_tot_{all,top}.tex : table comparing the score of each of the top control subfamilies of all kinds across the curated dataset ("Total" column) and the datasets used from academic papers (other columns). dataset_curated_score_stats_imp_{all,top}.tex : table with statistics on the amount and average score for both all and top SW-implementable control subfamilies. dataset_curated_score_stats_mt0_imp_{all,top}.tex : table with statistics on the amount and average score for both all and the top SW-implementable control subfamilies whose score is more than 0. dataset_curated_score_stats_tot_{all,top}.tex : table with statistics on the amount and average score for both all and top control subfamilies of all kinds. dataset_curated_score_stats_mt0_tot_{all,top}.tex : table with statistics on the amount and average score for both all and the top control subfamilies of all kinds whose score is more than 0. dataset_curated_summary_{all,top}.xlsx : sheet files with multiple tabs to determine grouping and statistical data, such as the score of the top security control subfamilies and the score of their counterparts in the used datasets. Tabs "contribution_ds_tot" and "contribution_ds_imp" are the most relevant, performing these calculation for all kinds and SW-implementable security control subfamilies, respectively. In all cases, the first file considers all control subfamilies, whereas the second considers the top 20 ones. Note that this script has a specific pre-requirement that must be installed to generate the excel file. sudo apt install python3-openpyxl python3 dataset_contribution.py Licence This work is dual-licenced according to the type of resource: Datasets: CC BY-NC 4.0 Code: GNU AGPL 3.0 Funding This work was supported by the grants COALESCE-6G PID2024-163028OB-I00, funded by MICIU/AEI/10.13039/501100011033/FEDER, EU; and AEI-PID2021-128521OB-I00, funded by the Spanish Recovery, Transformation and Resilience Plan through the European Union (Next Generation).

open·CC-BY-NC-4.0·Zenodo·completeSource
declared

Effective enzyme engineering through modeling kinetic parameter changes upon mutations with geometric deep learning

0.00

Yuan, Qianmu · Zhu, Mingming

1 files · 456 MB · zipdeclared

This ZIP file contains: 1. The DeltaCata-DB dataset preprocessing scripts and the processed dataset. 2. The source code and the trained model checkpoints of DeltaCata.

open·CC-BY-NC-ND-4.0·Zenodo·completeSource
declared

Pragmastat: Pragmatic Statistical Toolkit

0.00

Akinshin, Andrey

1 files · 2.1 MB · zipdeclared

This manual presents a toolkit of statistical procedures that provide reliable results across diverse real-world distributions, with ready-to-use implementations and detailed explanations. The toolkit consists of renamed, recombined, and refined versions of existing methods. Written for software developers, mathematicians, and LLMs.

open·CC-BY-NC-SA-4.0·Zenodo·completeSource
declared

A meta-analysis resolves the huntingtin interactome into coactivator losses and a robust proteostatic and synaptic gain network

0.00

Seefelder, Manuel Thomas

1 files · 147 MB · zipdeclared

This record contains the derived data and analysis code accompanying the manuscript: Seefelder, M. A meta-analysis resolves the huntingtin interactome into coactivator losses and a robust proteostatic and synaptic gain network (2026). The study integrates four previously published huntingtin (HTT) affinity-proteomics datasets and contrasts wild-type and polyglutamine-expanded HTT within a single Bayesian differential-interactomics model (BayesInteractomics), assigning every protein a calibrated, condition-dependent interaction call. Of 4,338 proteins evaluated, 275 are condition-dependent, describing a bidirectional remodelling of the HTT interactome - a loss of transcription-activation coactivators (Mediator, the ASCOM H3K4-methyltransferase, CREBBP, CDK9) and a gain of proteostatic and synaptic contacts (the 26S proteasome, HSP70 chaperones, synaptic and actin-cytoskeletal networks), around an intact chaperonin-HAP40 core. Contents Source differential-interactome table (differential_results.xlsx, all.csv, HTT_unchanged.csv) - per-protein posterior interaction probabilities, differential calls and false-discovery rates for wild-type and mutant HTT. Call-specific protein lists (gene-symbol lists for each differential class). Derived over-representation / enrichment results. Per-figure source data for all main and supplementary figures. A snapshot of the analysis and figure-generation code used to produce all results and figures. A snapshot of the BayesInteractomics.jl framework version used for interaction scoring. Raw data provenance This is a re-analysis of published datasets. The primary (raw) affinity-proteomics data are not re-hosted here and remain available from the original publications and their associated repositories: Greco et al. (2022), Justice et al. (2025), Sap et al. (2021) and Gutiérrez-García et al. (2023). Reproducibility Running the deposited code against the deposited derived data regenerates the figures and tables of the manuscript. The BayesInteractomics method is developed openly at https://github.com/ma-seefelder/BayesInteractomics.jl and described in full in a companion methods paper (Seefelder, submitted).

open·CC-BY-NC-ND-4.0·Zenodo·completeSource
declared

Code and data for: Forecasting ecological trajectories from ecological dynamic regimes to improve resilience analysis

0.00

Sánchez-Pinillos, Martina · Fortin, Marie-Josée · Messier, Christian · et al.

1 files · 173 MB · zipdeclared

Code and data for: Sánchez-Pinillos, M., Fortin, M.-J., Messier, C., Kneeshaw, D. 2026. Forecasting ecological trajectories from ecological dynamic regimes to improve resilience analysis. Methods in Ecology and Evolution.

open·CC-BY-NC-4.0·Zenodo·completeSource

Select a result to see its full details here: the measured structure, quality, and the loader, without leaving your search.