Altenhoff, Adrian
hybrid · semantic + lexical · 39 datasets ranked · 2.89s
Altenhoff, Adrian
6 files · 100 MB · hdf5
OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Petitpierre, Remi · di Lenardo, Isabella · Hudson, Polly · et al.
4 files · 1.4 GB · geopackage, zipdeclared
The dataset provides a large-scale vector layer of historical building footprints for the Greater London area at the end of the 19th century. It comprises 1,299,029 individual building footprints, extracted automatically from the Ordnance Survey five-feet-to-the-mile (1:1,056) map. The original maps were surveyed between 1891 and 1895 and published between 1893 and 1896, covering approximately 450 km² of urban and suburban London. The dataset was produced using a deep learning-based semantic segmentation pipeline, achieving approximately 97% precision and 95% recall in building detection. Source The source maps were digitised by the National Library of Scotland at 400 ppi and manually georeferenced in collaboration with the David Rumsey Map Collection. The extraction relies on 753 georeferenced map sheets, provided that six sheets are missing from the original archive. Extraction methodology The dataset was generated through an automated workflow: 262 image patches were annotated manually with three classes: (i) regular buildings; (ii) compound buildings; (iii) building boundaries A Mask2Former [1] semantic segmentation model is used. For training and inference, we followed the approach described by [2]. At inference time, boundary predictions are supplemented by the predictions of a specialist model, trained on cadastral plans [3]. Polygon geometries are extracted, based on predicted building contours, and vectorised. Compound buildings components are merged. Content Two versions of the dataset are provided: london_buildings_1891-96_raw.gpkg Raw output from the automated extraction pipeline (after vectorisation and merging) london_buildings_1891-96_corr_v1.gpkg Minimally corrected version including manual adjustments for major structures (e.g. railway stations, monuments, large buildings) Each geometry has a field buil_class , taking values in ['regular','compound'] , relative to the cartographic representation of building classes and the associated processing approach. In corr_v1, manually added geometries take the value NULL . In addition, we relase the labeled images used for training the segmentation model: annotations.zip ZIP folders containing the training and validation data The archive contains subfolders images , containing the tif image samples and labels , corresponding to the semantic class indexes, in png format. labelmap.txt details the class index codes. Data format Format: Geopackage Coordinate reference system: WGS 84 / Pseudo-Mercator (EPSG:3857) Temporal coverage Survey period: 1891-1895 Publication period: 1893-1896 Data creation: 2024-2026 Descriptive statistics Number of building footprints: 1,300,831 (raw), 1,299,040 (corr_v1) Coverage area: 450 km² Detection performance (raw): Precision: 97% Recall: 95% Use and reuse potential This dataset supports research in: urban history historical GIS urban morphology economic and social history It is particularly suited for studying long-term urban change and fine-grained spatial patterns in industrial-era cities. Related publication This dataset is described in a data paper submitted to the Journal of Open Humanities Data Corresponding author Remi Petitpierre Email: remi.petitpierre@epfl.ch Funding The research was supported by the College of Humanities at EPFL and the European Union Horizon Europe Programme (Grant No. 101233051). License This dataset is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Citation If you use this dataset, please cite: @misc{london_footprints_petitpierre_2026, author = {Petitpierre, R{\'{e}}mi and di Lenardo, Isabella and Hudson, Polly and Herold Hendrik and McDonough, Katherine and Hecht, Robert and Vaienti, Beatrice and Fleet, Christopher}, title = {{A Layer of Late Victorian London: The Building Footprints from the 1:1,056 Ordnance Survey Map (1891-1896)}}, year = {2026}, publisher = {EPFL}, url = {https://doi.org/10.5281/zenodo.19497434}} Limitations Lower accuracy for very small buildings Occasional segmentation errors in dense or complex areas Minor inconsistencies near map sheet boundaries Liability The authors assume no liability for the use of this dataset. References Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., & Girdhar, R. (2022). Masked-attention Mask Transformer for Universal Image Segmentation . arXiv. https://doi.org/10.48550/arXiv.2112.01527 Petitpierre R. (2026) Generalizable Multiscale Segmentation of Heterogeneous Map Collections . arXiv. https://doi.org/10.48550/arXiv.2603.05037 Petitpierre, R., di Lenardo, I., & Rappo, L. (2024). Revealing the Structure of Land Ownership through the Automatic Vectorisation of Swiss Cadastral Plans . Digital History Switzerland. https://doi.org/10.13140/RG.2.2.26632.33281
Zhang, Yunwei
6 files · 40 GB · hdf5, zipdeclared
Files related to INVNET, a deep learning model for surface wave dispersion spectrum inversion in geophysics.
Geisler, Jan · Rakhimberdiev, Eldar · Boom, Michiel P. · et al.
29 files · 124 MB · csv, shapefiledeclared
1. Many migratory birds now reach their Arctic breeding grounds earlier in order to keep pace with advancing springs and shifting nutrient peaks, either by departing earlier from non-breeding grounds or by travelling faster. For dark-bellied brent geese, there is limited potential to travel faster, as their migration to the Siberian breeding grounds is already among the fastest of Arctic geese and swans. Earlier departure would require reaching departure body mass earlier, either through a faster accumulation of energy stores during spring staging or via adjustments earlier in the annual cycle. 2. We examined long-term shifts in spring staging phenology and changes in winter and spring body mass trajectories of brent geese at the population level, with particular emphasis on the effects of winter temperature on body mass and spring body mass on departure timing. 3. We used more than five decades of body mass measurements from individuals caught in the United Kingdom and France, and in the Dutch Wadden Sea to reconstruct changes in spring and winter mass trajectories, respectively. These data were combined with over five decades of migration counts in the Netherlands and more than two decades of counts in Denmark to quantify changes in spring staging phenology. 4. We found that brent geese have not shifted their spring arrival in the Wadden Sea but have advanced departure timing. Furthermore, brent geese were heavier during and after milder winters, and have changed mass trajectories over recent decades. They no longer lose mass during winter and the second spring staging phase, and fuelling rates in the first spring staging phase have declined. Annual variation in body mass was not related to annual departure timing. 5. These results suggest that milder winters have relaxed energetic constraints and improved body condition in brent geese throughout the non-breeding season. Our findings highlight the importance of considering the full annual cycle when assessing how animals with limited capacity to adjust migration timing or speed respond to global change.
Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.
37 files · 1.6 GB · hdf5, tiff, torchdeclared
This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.
Marques da Silva, Gabriel · de Freitas Junior, Adirson Maciel · Feitosa, Flávia da Fonseca
2 files · 500 MB · geopackage, xlsxdeclared
Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) Este conjunto de dados contém o recorte territorial do banco de dados preditor consolidado para o município de São Paulo (no formato GeoPackage: 04_final_consolidado_31_completo_sp.gpkg ) e seu respectivo dicionário e resumo de variáveis (no formato Excel: Dicionário Modelo SP.xlsx ). O banco de dados é estruturado em uma grade celular regular de alta resolução com células de 50 metros por 50 metros (50m x 50m) . Cada célula de 2.500 m² representa uma unidade territorial urbana que agrega variáveis morfológicas urbanas, informações socioeconômicas e escores resultantes de um modelo de aprendizado de máquina (Machine Learning) preditor de assentamentos informais/precários. 1. Visão Geral do Recorte de São Paulo (SP) Total de feições (células de análise) : 388.536 Dimensão das células : 50m x 50m (2.500 m² por célula) Área urbana total analisada : Aproximadamente 971,34 km² (composta pela união das células da grade dentro do município) Total de variáveis analisadas : 111 variáveis originais no GeoPackage, sendo 109 colunas ativas com dados (não-nulos) para São Paulo. Qualidade do preenchimento : Excelente cobertura. Das 109 variáveis ativas, 108 apresentam 100% de cobertura espacial para o município de São Paulo. A única variável com preenchimento parcial é a ranking_candidato (com 93,42% de cobertura). 2. Estrutura do Dicionário (Dicionário Modelo SP.xlsx) O arquivo de dicionário e resumo é composto por 5 colunas que descrevem as variáveis do GeoPackage: Variable (Nome da Variável): O nome técnico da coluna na tabela de atributos do GeoPackage (ex: ID , prob_fcu , ibge_mediapopc ). Type (Tipo do Dado): Indica a natureza do dado físico (ex: object , float64 , int16 / int32 , geometry ). Non_Null_Count (Registros Válidos): Quantidade de feições em São Paulo que possuem dados válidos nesta coluna, desconsiderando campos nulos, vazios ou informados como ausentes (como "None" ou "null"). Coverage_Percentage (Taxa de Cobertura): O percentual de feições de São Paulo que possuem a variável preenchida. Sample_Value (Valor de Amostra): Um exemplo real do dado contido na coluna retirado da primeira linha válida encontrada, servindo para referência rápida sobre o formato da informação. 3. Categorias de Variáveis no Banco de Dados As variáveis presentes no banco de dados preditor de São Paulo dividem-se nos seguintes blocos temáticos: A. Variáveis de Localização e Resolução Territorial : Identificam as células geograficamente e associam os recortes às divisões oficiais (ex: ID , id_rg2017_cd_mun , id_rg2017_mun_nome , id_100m , id_200m , id_grid_1km ). B. Métricas de Classificação do Modelo Preditor : Valores resultantes da aplicação do modelo de inteligência artificial sobre o território (ex: prob_fcu , rank_class , ranking_total , ranking_candidato ). C. Variáveis de Importância (Feature Importance) : O modelo registra quais fatores foram os mais determinantes para a classificação de cada polígono específico (ex: top1_feat a top16_feat , top1_score a top16_score , top1_val a top16_val ). D. Variáveis Socioeconômicas e de Infraestrutura (Censo IBGE) : Indicadores estatísticos herdados dos setores censitários do IBGE (ex: ibge_mediapopc , ibge_mediadomc , m_ibge_renddppo , p_ibge_esginade , p_ibge_lixoinade , p_ibge_naocalca , p_ibge_naoilupub ). E. Variáveis Morfológicas Urbanas (GBA) : Métricas físicas de ocupação do solo e edificações (ex: gba_num_edif , gba_edif_por_ha , idade_ocupacao_menor ). 4. Agradecimentos e Financiamento Este estudo foi desenvolvido no âmbito do Centro de Estudos da Favela (CEFAVELA) (projeto FAPESP nº 2022/12259-8). A.M.F.J. agradece à Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP) pela bolsa de pós-doutorado vinculada a este projeto (processo 2025/03115-0). G.M.S. agradece à Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP) pela bolsa de Treinamento Técnico (TT/FAS) vinculada a este projeto (processo 2025/24201-2). 5. Como Citar (How to Cite) ABNT : SILVA, Gabriel Marques da; FREITAS JUNIOR, Adirson Maciel de; FEITOSA, Flávia da Fonseca. Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) . Zenodo, 2026. DOI: 10.5281/zenodo.21268382. APA : Silva, G. M., Freitas Junior, A. M., & Feitosa, F. F. (2026). Dicionário e Resumo do Banco de Dados Preditor – Recorte São Paulo (SP) . Zenodo. https://doi.org/10.5281/zenodo.21268382 6. Links Relacionados (Related Links) Centro de Estudos da Favela (CEFavela): https://cefavela.ufabc.edu.br/ Projeto Revelando Favelas (Arcabouço Metodológico): https://cefavela.ufabc.edu.br/revelando-favelas-arcabouco-metodologico-para-identificacao-e-caracterizacao-de-favelas/
Vines, Jose I.
1 files · 3.0 GB · hdf5declared
Pre-computed, resolution-broadened (R=1500) stellar atmosphere spectra cache for the astroARIADNE SED fitting package. Contains 7 model grids (Phoenix v2, BT-Settl, BT-NextGen, BT-Cond, Castelli & Kurucz 2004, Kurucz 1993, Coelho 2014) resampled to a common logarithmic wavelength grid (0.125-4.629 µm). This cache eliminates the need to download the full ~770 GB model libraries for SED plotting.
Kavanagh, Jack · Anthony, Patrick
6 files · 11 MB · geojson, geopackage, shapefiledeclared
A historical map showing a segment of the boundaries of New Spain in c. 1800. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
Lopez Dubon, Sergio · Sgarabotto, Alessandro · Lanzoni, Stefano
11 files · 38 MB · hdf5, npydeclared
This record contains the trained autoencoder, extracted encoder, and processed world/real-river latent-space reference cloud associated with the manuscript *A data-driven approach to discern the curvature spectral complexity of compound meander bends*. The full autoencoder is provided to support reconstruction-based validation and reproducibility of the learned representation. The extracted encoder is provided for inference and future software tools. It maps preprocessed 64 × 64 single-channel curvature-spectrum images to the two-dimensional latent space used to analyse meander shape complexity and skewness. The file `world_latent_cloud.npy` contains the two-dimensional latent coordinates of the world/real-river meander dataset used as the reference background cloud in the manuscript latent-space figures. This file is a processed latent-coordinate dataset only; it does not contain raw satellite imagery, raw centreline geometries, or training images. The release includes model weights, architecture files, model summaries, export metadata, the world/real-river latent cloud, example inference scripts, a validation script, environment files, and a minimal example input. The models should only be applied to curvature-spectrum images generated consistently with the preprocessing workflow described in the associated manuscript. Main files included in this release are: - trained_autoencoder.h5: full trained autoencoder. - encoder_only.h5: extracted encoder in HDF5/Keras format. - encoder_only.keras: extracted encoder in native Keras format. - model_architecture.json: full autoencoder architecture. - encoder_architecture.json: encoder architecture. - model_summary.tx and encoder_summary.txt: layer summaries. - world_latent_cloud.npy: world/real-river reference latent-space cloud. - world_latent_cloud_metadata.json: metadata for the world/real-river latent-space cloud. - model_card.md: intended use, inputs, outputs, limitations, and citation guidance.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
Kavanagh, Jack · Anthony, Patrick
6 files · 9.3 MB · geojson, geopackage, shapefiledeclared
A historical map of the boundaries of Prussia in c. 1795. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Francescone, Marco
14 files · 380 KB · shapefile, zipdeclared
This dataset contains original geomorphic mapping of surface fault traces along the Dixie Valley Fault (DVF) range front and piedmont zone, central Nevada, USA. Traces were mapped directly from a 1-m bare-earth lidar digital elevation model (DEM), using hillshade and slope-raster visualizations. This dataset accompanies the manuscript: Francescone, M., et al. (in review), LiDAR-Based Fault-Scarp Analysis and Rupture Hazard Assessment: Earthquake Scenarios of the Dixie Valley Fault System (Nevada, USA). See Section 3.1 ("Fault Trace Mapping") of the manuscript for full methodological details
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Barreto, Bruno · Eisencraft, Marcio
3 files · 2.6 GB · parquetdeclared
This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.
Dey, Hemal · Shao, Wanyun
14 files · 5.5 MB · csv, jpeg, shapefiledeclared
Despite the proliferation of social vulnerability assessment methodologies, selecting the most appropriate model remains a critical challenge due to inter-model variability. To explore the inter-model variability, this study systematically investigated inter-algorithmic and inter-classification variability to assess how methodological design influences outcomes.
Langermann, Florence · Blaschta, Stephanie · Dietze, Klara · et al.
3 files · 731 KB · geopackagedeclared
Beschreibung (Deutsch) Diese Datenpublikation enthält archäologische Geodaten aus den Forschungsprojekten "The Cultic Centre of the Sun-God of Heliopolis (Egypt)" und „Eclipse and Mutation: The end of the sun temple of Heliopolis" zu den ergrabenen und aufgefundenen Befunden im altägyptischen Heiligtum in Heliopolis im heutigen Kairo. Die Befunde wurden innerhalb einer Reihe von Grabungskampagnen zwischen 2012 bis 2024 ergraben und dokumentiert. Diese Datensammlung stellt eine Zusammenstellug dieser Befunde dar und soll als Grundlage für eine Gesamtkartierung dienen. An der Zusammenstellung der Daten waren maßgeblich die drei verantwortlichen Ausgräberi:innen Stephanie Blaschta , Klara Dietze und Florence Langermann beteiligt. Für die Erstellung der Datenstrukturen und Schemata und dem Zusammenführen der Daten war Michael Schleier verantwortlich. Das Projekt wurde von der Deutschen Forschungsgemeinschaft (DFG) sowie weiteren akademischen und privaten Förderinstitutionen unterstützt. Weitere Informationen: Projektwebseite DAI Datendokumentation des i3mainz, Hochschule Mainz Description (English) This data publication contains archaeological geodata from the research projects "The Cultic Center of the Sun-God of Heliopolis (Egypt)" and "Eclipse and Mutation: The end of the sun temple of Heliopolis" on the excavated and discovered features in the ancient Egyptian sanctuary in Heliopolis in present-day Cairo. The features were excavated and documented during a series of excavation campaigns between 2012 and 2024. This data collection represents a compilation of these features and is intended to serve as the basis for an overall mapping. The three excavators responsible, Stephanie Blaschta , Klara Dietze and Florence Langermann , played a key role in compiling the data. Michael Schleier was responsible for creating the data structures and schemas and merging the data. The dataset is published under a CC BY-SA 4.0 license . The project is funded by the German Research Foundation (DFG) and a wide network of institutional and private sponsors. Further information: Project website at DAI Datendokumentation of the i3mainz, Mainz University of Applied Sciences