Yang, Qingliu
6 files · 8.0 MB · bzip2
hybrid · semantic + lexical · 20 datasets ranked · 3.23s
Yang, Qingliu
6 files · 8.0 MB · bzip2
Dataset Description This dataset contains 3D lightning location results, DALMA and FALMA waveform for two Energetic Compact Stroke (ECS) events. location results are included: HF3D_1732785151.dat - 3D lightning locations for the ECS leader A flash. HF3D_1734785454.dat - 3D lightning locations for the ECS leader B flash. The timestamp 1734785454 and 1732785151 corresponds to the occurrence time of the lightning flash in Japan Standard Time. File format and parameters The first row contains the lightning occurrence time. Column descriptions: Time (ms) - time relative to the lightning source. X, Y, Z (m) - 3D spatial coordinates relative to ground level. The origin (0, 0, 0) corresponds to latitude 36.76°N and longitude 136.76°E. FALMA and DALMA waveform ECSLeaderA_DALMA_waveform.bz2 is DALMA waveform of Leader A. ECSLeaderA_FALMA_waveform.bz2 is FALMA waveform of Leader A. ECSLeaderB_DALMA_waveform.bz2 is DALMA waveform of Leader B. ECSLeaderB_FALMA_waveform.bz2 is FALMA waveform of Leader B. This dataset allows analysis of the spatial and temporal development of these two ECS flashes.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Robert Koch-Institut
8 files · 223 MB · pdf, tsv, xzdeclared
Im Datensatz "SARS-CoV-2-Sequenzdaten aus Deutschland" umfasst vollständige Virusgenomsequenzen sowie zugehörige Metadaten aus bundesweit erhobenen Proben. Die Proben werden in Zusammenarbeit vom IMSSC2-Labornetzwerk, dem Nationalen Referenzzentrum für Coronaviren an der Charité sowie dem RKI sequenziert und bioinformatisch analysiert. Der Datensatz ermöglicht eine fundierte molekular-epidemiologische Analyse der SARS-CoV-2-Ausbreitung in Deutschland und stellt eine zentrale Ressource für Forschung und öffentliche Gesundheitsüberwachung dar.
Safaei, Sina · Rezaee, Parham · Ghysels, An
4 files · 26 GB · xzdeclared
The files provided here contain examples of 1D, 2D, and coarse-grained systems that demonstrate Hamiltonian exchange within the Hamiltonian Replica Exchange Transition Interface Sampling (HRETIS) framework. The coarse-grained (MARTINI 3) path sampling simulations show the permeation of glutamine (P2-P5) and methionine (P2-C6) through a DPPC membrane at 323 K using GROMACS. These simulations can be resumed from the existing "restart.toml" files located in the following directories: CG/path_sampling_set_2/RETIS_main_1/infretis/ CG/path_sampling_set_2/HRETIS_75_1/infretis/ CG/path_sampling_set_2/RETIS_helper_1/infretis/ The command to continue a run using the infretis package (https://github.com/infretis/infretis) is: infretisrun -i restart.toml The simulations can also be started from scratch by copying the initial paths from: CG/path_sampling_set_2/initial_paths/ into the "infretis/load" directory of each simulation. After that, the run can be started with: infretisrun -i infretis.toml Additional information about running the 1D and 2D simulations, the directory structure, and the simulation input files is provided in the "readme.txt" file.
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Barreto, Bruno · Eisencraft, Marcio
3 files · 2.6 GB · parquetdeclared
This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.
Saldanha, Raphael
8 files · 601 MB · parquet, zipdeclared
Description This deposit contains annual, municipality-level datasets derived from the Brazilian Primary Health Care Information System (SISAB). The files combine two complementary data sources: Public SISAB Saúde report downloads from the Atendimento/Visita production report. CID-10 and CIAP-2 attendance data obtained from SISAB through requests under the Brazilian Access to Information Law (Lei de Acesso à Informação, LAI). The datasets are organized as tidy annual files in CSV (Zipped) and Parquet format. They are intended to support reproducible analysis of primary care production, procedures, evaluated problems/conditions, and CID/CIAP-coded attendances across Brazilian municipalities. The public SISAB report datasets are stratified by competence month, state, municipality, DataSUS age group, SISAB sex category, and the selected report category. For each competence month and report type, the extraction combines 36 stratified SISAB downloads: 18 age groups by 2 sex values. Monthly files are merged into yearly files, completing missing combinations of observed competence, municipality, age group, sex, and category with valor = 0 . The LAI dataset contains yearly CID-10 and CIAP-2 attendance counts by competence month, municipality, code type, and code. When multiple valid LAI files cover the same competence, the processing pipeline selects the file with the largest number of data rows, using file size and request folder order as tie-breakers. Provenance columns identify the selected LAI request and source file. Variables SISAB Saúde Produção Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : Age group. sexo : SISAB sex category, Masculino or Feminino . tipo_producao : production type from the SISAB report. valor : count reported by SISAB. SISAB Saúde Procedimento Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . procedimento : procedure from the SISAB report. valor : count reported by SISAB SISAB Saúde Condição Avaliada Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . condicao_avaliada : evaluated problem or condition from the SISAB report. valor : count reported by SISAB. SISAB LAI CID/CIAP Columns: ano_competencia : competence year. competencia : competence month in YYYYMM format. competencia_date : first day of the competence month. co_municipio_ibge : municipality IBGE code. tp_codigo : code type, CID or CIAP . codigo : CID-10 or CIAP-2 code. qt_atendimentos : number of attendances. source_request : selected LAI request folder. source_file : selected source CSV file. Methods The public SISAB report files were generated with the sisab_scrapper processing pipeline. For each month, report type, age group, and sex value, the pipeline downloads the all-Brazil municipality report from SISAB, validates the returned CSV, preserves raw cache files for resumable runs, and writes a sorted monthly tidy dataset. The yearly merge validates required columns, expected age groups, expected sex values, category values, and month gaps unless explicitly allowed. The CID/CIAP files were generated with the sisab_lai processing pipeline. The pipeline imports CSV files received through LAI requests, detects the real CSV header after any SQL*Plus preamble, validates candidate files, resolves overlapping requests by competence, standardizes old and new schemas into one tidy table, and exports annual CSV and Parquet files together with audit reports. Sources - SISAB public reports, Ministry of Health, Brazil: https://sisab.saude.gov.br/ - SISAB LAI files obtained through Brazilian Access to Information Law requests. - Processing code for public SISAB report data: https://github.com/rfsaldanha/sisab_scrapper - Processing code for LAI CID/CIAP data: https://github.com/rfsaldanha/sisab_lai Notes - Counts are aggregated administrative records and should be interpreted in light of SISAB reporting practices, data quality, and changes in municipal reporting coverage. - Municipality boundaries, names, and coding practices may vary over time. - Public SISAB report datasets are completed with zero values only for combinations defined by observed municipalities, observed competencies, all expected age groups, both expected sex values, and observed report categories within the yearly merge. - LAI CID/CIAP data preserves selected source-file provenance through source_request and source_file . - This deposit corresponds to an individual year. Deposits for other years are published separately.
Anonymous, Anonymous
8 files · 25 GB · xzdeclared
This dataset is available at doi: 10.5281/zenodo.21125584 and builds on top of the slightly smaller dataset previously published at doi: 10.5281/zenodo.19842646 . In this archive, we provide the complete data and code needed to reproduce the experiments discussed in the paper Benchmarking Permutation-based Encodings and Metaheuristics for the Traveling Tournament Problem . The Traveling Tournament Problem (TTP) asks us to organize a tournament where N teams compete with each other in a fair and efficient way. We focus on the double round robin (2RR) flavor of this problem, where each team plays twice against each other team, once at home and once away, i.e., at the stadium of the other team. Here, efficient means to reduce the total travel length summed up over all teams. The fairness aspect is embodied in constraints that prevent any team from having more than three consecutive home and away games, from two teams playing each other twice in a row, and that enforce a compact 2RR schedule. We approach the TTP using two different permutation-based encodings under diverse metaheuristics, ranging from randomized local search (RLS), RLS with Frequency Fitness Assignment (FFA), Simulated Annealing (SA), two hybrid algorithms combining local search with FFA, and NSGA-II. The experiments take a very long time to conduct, so there are several different steps that can be taken separately to reproduce them to different degrees. All the data elements are packaged in files of the format tar . xz . Under Linux, you can unpack them using tar -xf archive.tar.xz where archive is to be replaced with the archive name. Under Windows, you would probably use a tool like WinRAR . In this updated dataset, we only provide additional runs for some setups and a slightly updated evaluator program. Most of the runs of the algorithms as well as the algorithm implementations have not changed and are already given in doi: 10.5281/zenodo.19842646 . We use the moptipy package to run the experiments. For the experiment, we use some algorithms already provided in that package. We use the moptipy -API for them and we also use it to evaluate the raw data generated by the experiments and to create the plots. We use the moptipyapps package, which provides the benchmark instances, data structures, and components needed for the TTP. Therefore, we provide the related versions of these packages in the file software.tar.xz , although they are available both on GitHub and PyPI. Just to be sure. In our paper, we present two experiments. The main "Experiment 1" applies the algorithms to the TTP. The side-dish "Experiment 2" investigates the runtimes and results of the two different encodings. Experiment 1 The raw data of the experiments can be found in folder main/results presented compressed as two files, main_results_1.tar.xz and main_results_2.tar.xz . They contain the data for the first and second encoding, respectively. Only main_results_2.tar.xz is part of this new dataset, the rest of the raw data is already provided in doi: 10.5281/zenodo.19842646 . This data is quite big. We have over 11'000 files that together make up over 80.1 GB of data in total. You can unpack the files and they will create the folders main/results/gameEncoding_1 and main/results/gameEncoding_2 , respectively. The folder structure below these root folders follows the pattern algorithm/instance/algorithm_instance_randseed.txt . Since we use moptipy to execute our experiments, the folder structure and the exact structure of the log files, which contain all improving steps that the algorithms too, the algorithm setup, the system configuration, as well as the final best discovered solution, follows the specification given in: T. Weise and Z. WU: Replicable Self-Documenting Experiments with Arbitrary Search Spaces and Algorithms. Genetic and Evolutionary Computation Conference Companion (GECCO'2023) Companion, July 15-19, 2023, Lisbon, Portugal, pages 1891-1899. New York, NY, USA: ACM. doi: 10.1145/3583133.3596306 . Expanding upon that, we generate all random seeds in a deterministic manner per instance and all algorithms use the same seeds. This makes cheery picking impossible and our experiments can be replicated also using different implementations of algorithms even in different programming languages. The source code of the algorithm implementations is part of moptipy for the algorithms except for the temperature schedule used with SA, which is provided in folder main/sources packaged as main_sources.tar.xz . There are two experiment execution scripts: run_encoding_1.py for the first encoding and run_encoding_2.py for the second encoding. They can also be found in that folder. With them, you can exactly replicate the contents of the results folder (but your log files would contain other time stamps, system configurations, etc.). Be aware that doing this would take a very very long time. The experiment script will execute the runs in a random order. So you could go and start it, wait for some runs to complete, and compare them with the corresponding log files of the same name and path in our main/results folder. Their content, at least the improving steps and final solutions, should be the same. Then you could start the experiment again, get some other runs, and do the same. This way, you can verify that our data is genuine. Of course, you need to have the right version of the open source packages moptipy and moptipyapps for that, which we will describe later on. If you have all the data in the main/results folder, you can execute the scripts in the main/evaluator folder (using again the right version of moptipy and moptipyapps ), which are contained in file main_evaluator.tar.xz . Then this will produce the figures and tables that we are having in the paper. This output is contained in folder main/evaluation provided in main_evaluation.tar.xz , but you can also re-create it by yourself. Algorithms The following algorithms are implemented and used in the main experiment: RLS : The randomized local search flips swaps two elements of a permutation. The algorithm is implemented as class RLS . The search operator is given in class Op1Swap2 . The results are given in folder main/results/*/rls_swap2 . FRLS : The randomized local search algorithm RLS , but using Frequency Fitness Assignment. Using this algorithm is like using RLS above, but we replace class RLS with class FEA1plus1 . The results are in folder main/results/*/fea1p1_swap2 . EAFEA A : A hybrid algorithm executing the RLS and the FRLS in an alternating fashion: First one objective function evaluation (FE) of the RLS , then one of the FRLS , then again one of the RLS , and so on. If the FRLS discovers a solution with frequency-value 1, it will copy it to the RLS , thereby overwriting this algorithm's current solution. This is implemented in class EAFEAA . The results are in folder main/results/*/eafeaa_swap2 . EAFEA B : A hybrid algorithm executing the RLS and the FRLS in an alternating fashion: First one objective function evaluation (FE) of the RLS , then one of the FRLS , then again one of the RLS , and so on. If the FRLS discovers a solution which is better or equally good as the current solution of the RLS -strand, it will copy it to the RLS , thereby overwriting this algorithm's current solution. This is implemented in class EAFEAB . The results are in folder main/results/*/eafeab_swap2 . Simulated Annealing (several setups): Several optimized setups of the classical Simulated Annealing, using adaptive temperature schedules. This is implemented in class SimulatedAnnealing . The adaptive temperature schedules are implemented in a class provided directly in the experiment execution scripts previously mentioned. The results are in the folders of the format main/results/*/sa**_swap2 , where ** stands for the different adaptive temperature schedule definitions. Benchmark Problems We apply the algorithms to the compact 2RR instances of RobinX , which can be found at https://robinxval.ugent.be/RobinX/travelRepo.php and are also implemented in moptipyapps We generally have at least 4 runs for each algorithm on each benchmark problem. For some problems and algorithm combinations, we have more runs for historical reasons. However, to ensure absolute fairness between the algorithms, we only use 4 runs in this case, too, in our evaluation. More precisely, for each problem, we use the data generated by the runs using the same 4 random seeds for each algorithm. The runs take a long time, since we provide them with a budget of 10 9 objective function evaluations (FEs). 4 runs are not much, but over all the instances and based on the numerical results, the outcomes are still quite clear. However, for the sake of completeness, for the interested reader, and maybe for our own future work, we include all runs for all algorithms that we have. This means that, if for some algorithm/instance combination we actually have more runs, we include those in the archives main_results_1.tar.xz and main_results_2.tar.xz (depending on to which encoding they belong). Even though they are not used in the evaluation. Using them in the evaluation would also not change the conclusions in any way, as far as we can see here. Still, fairness is fairness and we try to do a good job to be fair here. External Libraries To run and to evaluate our experiments, you need to install the open source library moptipy . For evaluating, you may use a later, more current version. The required library and its dependencies can be installed from PyPI , via pip install moptipy moptipyapps . To be on the safe side, we included the libraries in the archive software.tar.xz , which unpacks into software . We include the current state of the libraries on GitHub, as well as several historical versions that were used in our experiments. Since our experiments span multiple years, the software has developed further - the code design remains such that even the versions today would produce exactly the same results as those used in the experiments. Reproducing the Experiment We include the following parts of our experiment: The folder main/sources (packaged as main_sources.tar.xz ) contains the Python sources needed to run the experiment. The folder main/results (packaged as main_results_1.tar.xz and main_results_2.tar.xz ) contains the results folder, i.e., all the log files that will be produced if the experiment is executed. If you would run the experiment again, you would get exactly these files, but maybe with different time stamps and different system configuration data. But the objective values/objective function evaluation indices would be exactly the same. Running the experiment takes a long, long time. The folder main/evaluator (packaged as main_evaluator.tar.xz ) contains the evaluator code. These programs are executed after the experiment completes. They load the results from the results folder and extract the evaluation-relevant data. They produce the figures and tables in our paper. The folder main/evaluation is the output folder of the evaluator. It contains the figures and tables in our paper. Reproducing the Raw Data / Log Files To reproduce the raw data, you need to perform the following steps. Unpack main_sources.tar.xz to some folder, let's call it A . Open a terminal and cd into folder A/main/sources . Make sure that you have moptipy and moptipyapps of a compatible version installed. If not, you can either do pip install moptipy moptipyapps . Ideally you should do this in a virtual environment. Discussing how to set up virtual environments and install packages in them goes beyond the scope here, so we refer to https://docs.python.org/3/library/venv.html . Run python3 run_encoding_1.py . Stop it when you have sufficiently many runs with the first encoding. Run python3 run_encoding_2.py . Stop it when you have sufficiently many runs with the second encoding. This will automatically create a folder A/main/results , which will eventually have the same contents as provided by the archives main_results_1.tar.xz and main_results_2.tar.xz . Notice, though, that this takes a very long time even for two runs per setting. We run 1'000'000'000 FEs per run, for several algorithms, over many problems... Reproducing the Evaluation Steps If you either already have the raw data reproduced or have unpacked our main_results_1.tar.xz and main_results_2.tar.xz files, then you can reproduce the evaluation. Unpack evaluator.tar.xz to some folder, let's call it A . Open a terminal and cd into folder A/main/evaluator . Make sure that you have moptipy and moptipyapps of a compatible version installed. If not, you can either do pip install moptipy moptipyapps . Ideally you should do this in a virtual environment. Discussing how to set up virtual environments and install packages in them goes beyond the scope here, so we refer to https://docs.python.org/3/library/venv.html . Run the evaluator.py scripts in the folder A/main/evaluator , in order to create figures and tables. If you perform these steps, then the output collected in your folder A/main/evaluation should be exactly the same as the data we provide in archive evaluation.tar.xz . It should contain the figures and tables provided in our paper. Experiment 2 The second experiment in our work is basically structured exactly like the first one. Here, the goal simply is to figure out how long the two different encodings need and how many different solutions they produce. This experiment can be executed and evaluation in about a hour or two. We include the complete experiment in archive encoding.tar.xz which unpacks the folder encoding . Therein are the files needed to execute the experiment (in folder sources ), the results/output (in folder results ), the scripts to evaluate the results (in folder sources again), and the evaluation result, i.e., data and figures used in our paper (in folder evaluation ). To reproduce this experiment, you can follow the steps discussed above for the first experiment.
Anonymous Authors
15 files · 2.2 GB · parquet, rar, torchdeclared
Economic losses caused by extreme climate are not confined to the locations where events occur, but can propagate across regions through physical and economic linkages. Yet existing climate-impact assessment methods remain poorly suited to tracing how shocks spread across space and reshape the geography of economic loss. Here we develop a mechanistically informed multi-scale spatiotemporal autoregressive graph neural network model to quantify spatially cascading climate impacts. The model couples scale-specific, physically structured spatiotemporal autoregressive processes through an adaptive gating mechanism, allowing heterogeneous cross-scale interactions to be learned from data. Model estimation is achieved through tailored graph convolutional neural networks that are mathematically equivalent to spatiotemporal autoregressive models, enabling scalability while preserving transparent parameter interpretation. Monte Carlo simulation experiments show that the model accurately recovers true parameters and distinguishes between scale-dependent processes. Applying the framework to extreme precipitation, we find that large-scale upwind-to-downwind cascades driven by atmospheric circulations dominate aggregated economic losses. A one-standard-deviation increase in log extreme precipitation is associated with a 0.19 percentage-point decline in economic growth rate at the large scale, with 62.3% of the loss arising from spatial cascades. These findings highlight the need for transboundary risk governance that incorporates spatial cascading into climate-extremes monitoring and early-warning. Description of the uploaded file Monte Carlo simulation code data_generator_factors.py: Data generation script for multi-scale Monte Carlo simulation experiments. sarnn_model.py:Implementation of the proposed MS-STARGNNs model architecture definition. train.py:Training pipeline script for the MS-STARGNNs model. Data and spatial weights matrices for empirical analysis global_panel_1deg_std.parquet:Standardized Large-scale (1°) datase; global_panel_2km_std.parquet:Standardized Small-scale (2 km) dataset. W_global_2km_knn8.pt:Small-scale spatial weights matrix based on 8-nearest neighbors (KNN8). W_Large-scale:A large-scale spatial weights matrix derived from moisture transport pathways (2005-2021) mapping_1deg_to_2km.parquet: Correspondence file mapping large-scale (1°) grid cells to small-scale (2km) grid cells. ipcc_region_mapping_coarse_8regions.parquet: Mapping file linking Large-scale grid cells to the 8 IPCC AR6 reference regions. ipcc_region_mapping_fine_8regions.parquet: Mapping file linking Small-scale grid cells to the 8 IPCC AR6 reference regions. Code model.py: Core architecture definitions for empirical analysis. Provides the base classes and computational layers engineered to handle real-world geospatial complexities. All subsequent training scripts import modules from this file. STARGNNs.py: Implementation of the single-scale baseline. Serves as a reference point for evaluating the efficacy of cross-scale feature fusion. MS-STARGNNs_fixed.py: Configuration script for the MS-STARGNNs model utilizing fixed autoregressive coefficients. MS-STARGNNs.py: Configuration script for the MS-STARGNNs model utilizing annually varying autoregressive coefficients. MS-STARGNNs_8 regions.py: Executes the MS-STARGNNs model with decoupled regional parameters, loading unique autoregressive weights and βvectors for each IPCC region. Implements null value handling for regions lacking observational data (e.g., Antarctica).
Kim, Kyurhi
10 files · 9.6 GB · bzip2, zipdeclared
This repository contains cis-xQTL mapping results, colocalization analysis results, and transcriptome-wide association study (xTWAS) weights and association test statistics for six transcriptomic modalities generated from bulk RNA-seq data of dorsolateral prefrontal cortex (DLPFC) tissue from the ROS/MAP cohorts (n = 1,035). The RNA trait tables (BED format) used for cis-xQTL mapping and xTWAS model training were generated using the Pantry pipeline but are not included in this repository. Colocalization and xTWAS analyses were performed using the publicly available Alzheimer's disease (AD) dementia GWAS summary statistics from Bellenguez et al. ( Nature Genetics , 2022).
Allendorf, Daniel · Bläsius, Thomas · Leonhardt, Alexander · et al.
2 files · 12 GB · bzip2, zipdeclared
About this Repository This repository contains the software, datasets, and experimental data to reproduce the experiments in the above mentioned article. Please refer to the README file for more details and instructions. Article Abstract We consider a maximum entropy edge weight model that allows for negative weights. Given a graph Gand possible weights W typically consisting of positive and negative values, the model selects edge weights w ∈ W^m uniformly at random from all weights that do not introduce a negative cycle. We propose an MCMC process and show that it converges to the required distribution. We then engineer an implementation of the process using a dynamic version of Johnson's algorithm in connection with a bidirectional Dijkstra search as well as an innovative resampling method. We empirically study the performance characteristics of these novel sampling algorithms as well as the output produced by the model. Dataset Most of the input data (graph data) is generated dynamically via random graph models. In addition to the result data from the experiments, unew.data.tar.bz2 also contains trimmed US road networks used for the ROAD dataset in the paper. Code The code is developed at https://codeberg.org/lukasgeis/unew --- you may want to check there for updates.
OpenFF, YDS
18 files · 334 MB · bzip2, csv, pngdeclared
Generated by yammbs-dataset-submission: https://github.com/openforcefield/yammbs-dataset-submission
David, Cédric
19 files · 370 KB · netcdf, parquetdeclared
RAPID comes along with a set of test files based on a synthetic experiment called the Sandbox. More information is available in SANDBOX.md .
Zobel, Dominik
1 files · 3.2 GB · xzdeclared
The files in this archive accompany the paper "GPU Porting of the ICON Ocean Model: Performance and Possibilities for Submesoscale Climate Simulations" from Dominik Zobel, Nils Brüggemann, Helmuth Haak, Leonidas Linardakis and Peter Korn (submitted to JAMES 2026). The directory structure is as follows - 01_source_code : ICON code version used to conduct the experiments - 02_configuration : Details about configuring and building ICON - 03_scripts_and_results : Scripts, log files and results from relevant runs
Diehl, Martin · Eisenlohr, Philip · Roongta, Sharan · et al.
2 files · 56 MB · xz, zipdeclared
DAMASK® is a unified multi-physics crystal plasticity simulation package. The solution of continuum mechanical boundary value problems requires a constitutive response that connects deformation and stress at each material point. This problem is solved in DAMASK on the basis of crystal plasticity using a variety of constitutive models and homogenization approaches. However, treating mechanics in isolation is no longer sufficient to study emergent advanced high-strength materials. In these materials, deformation happens interrelated with displacive phase transformation, significant heating, and potential damage evolution. Therefore, DAMASK is capable of handling multi-physics problems. Following a modular approach, additional field equations are solved in a fully coupled way using a staggered approach.
Tekpinar, Mustafa
1 files · 85 MB · bzip2declared