Lee, Kanghwi
hybrid · semantic + lexical · 26 datasets ranked · 4.19s
Trained variational autoencoder (VAE) checkpoints and per-vocalization latent representations for three developing zebra finches (183K-274K vocalizations each, 40-101 days post-hatch), accompanying the Interspeech 2026 paper "Trajectory Variance: An Unsupervised Measure of Developmental Vocal Plasticity in Birdsong." With the code at https://github.com/hwiora/trajectory_variance , these files reproduce the paper's evaluation tables. The raw audio recordings are not included; they are available from the authors on request and will be released in a future study. See ZENODO_README.md for details.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Fry, Lauren · Seglenieks, Frank · Shrestha, Narayan Kumar · et al.
3 files · 22 MB · sevenzipdeclared
In response to flooding on Lake Ontario and the St. Lawrence River in 2017 and 2019, the U.S. and Canadian governments directed the Great Lakes - St. Lawrence River Adaptive Management (GLAM) Committee to expedite review of Lake Ontario outflow regulation Plan 2014 ahead of its usual 15-year review cycle. The dataset includes (1) a baseline record based on data from 1961-2020, (2) a set of stochastic scenarios that incorporates extremes that may not be in the historical record but are plausible under current hydroclimate conditions, and (3) a set of climate change scenarios to support evaluation of the plan under plausible decadal-scale hydroclimate changes. All data required to simulate Lake Ontario water levels, outflows, and downstream water levels are provided for each scenario. The data have been compressed into the following files: Baseline.7z - Baseline hydroclimate record from 1961-2020 Climate.7z - 8 future climate hydroclimate records Stochastic.7z - 500 stochastic hydroclimate records Within each compressed file are subdirectories that contain the files for each water supply sequence. The following table lists the files that are provided for each of the water supply sequences. For each file there is a description of the variable contained in the file, the location of the data in the file, and the unit of the data in the file if applicable. Filename Variable Location Units dpmi_flw_cms_qm48_na.csv Flow Des Prairies and Mille Iles River cms rich_flw_cms_qm48_na.csv Flow Richelieu River at Rapides Fryer cms stfr_flw_cms_qm48_na.csv Flow St. Francois River at Chut Hemming cms stmc_flw_cms_qm48_na.csv Flow St. Maurice River at La Gabelle cms ont_nbs_cms_qm48_na.csv Net Basin Supply Lake Ontario cms ont_nbs_cms_qm48_na_spinup.csv Net Basin Supply Lake Ontario cms eri_flw_cms_qm48_na.csv Outflow Lake Erie cms eri_flw_cms_qm48_na_spinup.csv Outflow (one year spinup) Lake Erie cms slon_flw_cms_qm48_na.csv SLON flow St. Lawrence River (Lake St. Louis) cms slon_flw_cms_qm48_na_spinup.csv SLON flow (one year spinup) St. Lawrence River (Lake St. Louis) cms stl_icw_cms_qm48_na.csv Ice/weed retardation St. Lawrence River cms stl_tde_m_qm48_na.csv Tidal Signal St. Lawrence River m ont_mlv_m_qm48_na_spinup.csv Water Level (one year spinup) Lake Ontario m stl_ics_xx_qm48_na.csv Ice status indicator St. Lawrence River N/A bati_rgh_xx_qm48_na.csv Ice/weed roughness factor Batiscan N/A card_rgh_xx_qm48_na.csv Ice/weed roughness factor Cardinal N/A corn_rgh_xx_qm48_na.csv Ice/weed roughness factor Cornwall N/A intw_rgh_xx_qm48_na.csv Ice/weed roughness factor International Tail Water N/A irhw_rgh_xx_qm48_na.csv Ice/weed roughness factor Iroquois Head Water N/A irtw_rgh_xx_qm48_na.csv Ice/weed roughness factor Iroquois Tail Water N/A jty1_rgh_xx_qm48_na.csv Ice/weed roughness factor Montreal Jetty No. 1 N/A lspr_rgh_xx_qm48_na.csv Ice/weed roughness factor Lake St. Pierre N/A lstd_rgh_xx_qm48_na.csv Ice/weed roughness factor Long Sault Dam N/A morr_rgh_xx_qm48_na.csv Ice/weed roughness factor Morrisburg N/A ogde_rgh_xx_qm48_na.csv Ice/weed roughness factor Odgensburg N/A ptcl_rgh_xx_qm48_na.csv Ice/weed roughness factor Pointe - Claire N/A sahw_rgh_xx_qm48_na.csv Ice/weed roughness factor Saunders Head Water N/A sorl_rgh_xx_qm48_na.csv Ice/weed roughness factor Sorel N/A summ_rgh_xx_qm48_na.csv Ice/weed roughness factor Summerstown N/A triv_rgh_xx_qm48_na.csv Ice/weed roughness factor Trois Rivières N/A vare_rgh_xx_qm48_na.csv Ice/weed roughness factor Varennes N/A na_fst_xx_qm48_na.csv Forecast indicator N/A N/A
Gomułka, Kamil · Woźniak, Piotr · Krzeszowski, Tomasz
1 files · 6.4 GB · sevenzipdeclared
======================= License ======================= This dataset is made available under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. ======================= Summary ======================= DATASET FOR EVALUATING THE EFFECTIVENESS OF AI-GENERATED DATA FOR VIDEO-BASED ACTION RECOGNITION This repository contains the dataset accompanying the paper "Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition". The dataset was curated to investigate the impact of incorporating AI-generated synthetic video data into the training process of deep learning models (such as I3D and TimeSformer) for human action recognition. It is particularly focused on evaluating cross-domain generalization and addressing the synthetic-to-real distribution shift using domain adaptation techniques. It is obligatory to cite the following paper in every work that uses the dataset: K. Gomulka, P. Wozniak, T. Krzeszowski: Evaluating the Effectiveness of AI-Generated Data for Video-Based Action Recognition, Electronics, MDPI, 2026. ======================= Data description ======================= The dataset focuses on 15 target action categories from the HMDB51 benchmark: Draw Sword, Fall, Flick Flack, Handstand, Jump, Kick, Pick, Punch, Run, Shoot Bow, Shoot Gun, Sit, Throw, Walk, and Wave. The repository consists of three main data subsets: * Grok Synthetic Dataset (Grok): Contains 1,500 synthetic video sequences (100 videos per class) generated using the Grok Imagine 1.0 model. Resolution is 560x560 pixels with a duration of 6 seconds per clip. * Meta Synthetic Dataset (Meta): Contains 1,480 synthetic video sequences (80 in the 'throw' class, 100 videos per other classes) generated using the Meta AI Vibes model. Resolution is 624x624 pixels with a duration of 5 seconds per clip. * HMDB51 Subset (hmdb51_org): A specific subset of the real-world HMDB51 dataset containing 2,315 video sequences distributed across the 15 target action classes, providing a direct real-world baseline. The repository also includes sample .csv files defining the train, validation, and test splits. In the example configuration, 60% of the HMDB51 dataset is allocated to training, 20% to validation, and 20% to testing, with all synthetic Grok videos appended exclusively to the training set. ======================= Dataset structure ======================= * Grok/ - directory containing 1,500 synthetic videos generated by Grok Imagine 1.0. * Meta/ - directory containing 1,480 synthetic videos generated by Meta AI Vibes. * hmdb51_org/ - directory containing 2,315 real-world videos from the HMDB51 dataset. * ghtrain_hval/ - directory containing example .csv files defining the train, validation, and test splits for the baseline HMDB + Grok setup. * train.csv - training set paths and labels. * val.csv - validation set paths and labels. * test.csv - testing set paths and labels. * README.md - detailed documentation including dataset overview, configuration details, and instructions for training the TimeSformer model using the provided data splits. ======================= Generation Methodology ======================= The synthetic videos were generated using a structured, combinatorial prompt engineering framework resulting in 100 unique prompts per class. Each prompt combined predefined descriptions of the action, environmental context (e.g., vintage film style, urban outdoor), and camera framing to ensure high scene diversity. While the generated sequences preserve essential motion semantics, they also may contain inherent generative artifacts.
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
Morais, Thomaz
1 files · 12 GB · sevenzipdeclared
This repository contains a Brazilian Portuguese speech corpus developed as part of the project "Incorporating Regionalisms and Accents into Speech Synthesis Models". The corpus was created to support research on speech synthesis, regional linguistic variation, and accent adaptation in text-to-speech systems, with particular attention to speech varieties from Paraíba and Northeastern Brazil. The corpus was built by organizing, processing, and integrating speech data from multiple sources, including CoLingPB, the Sotaque Brasileiro dataset, and additional collaborative speech recordings collected for this project. The data was curated and standardized to support experiments involving fine-tuning of pre-existing speech synthesis checkpoints, evaluation of regional accent preservation, and comparison between base and adapted text-to-speech models. This dataset is intended for academic and research purposes, especially for studies related to Brazilian Portuguese speech synthesis, regional accents, computational linguistics, speech processing, and deep learning approaches for text-to-speech. Users of this corpus should cite the original data sources and related works used in its construction and contextualization. Suggested citations in BibTeX format: @mastersthesis{morais2026regionalismos, author = {Morais, Thomaz Diniz Pinto de}, title = {{Incorporando Regionalismos e Sotaques em Modelos de Síntese de Fala}}, year = {2026}, school = {Universidade Federal de Campina Grande}, address = {Campina Grande, PB, Brazil}, type = {Master's thesis}, note = {Programa de Pós-Graduação em Ciência da Computação} } @misc{stein2015colingpb, author = {Stein, Cirineu Cecote and others}, title = {{Corpus Linguístico da Paraíba (CoLingPB)}}, year = {2015}, institution = {Universidade Federal da Paraíba}, address = {João Pessoa, PB, Brazil}, url = {https://www.cchla.ufpb.br/colingpb/}, note = {Corpus Linguístico da Paraíba} } @dataset{milan2021sotaquebrasileiro, author = {Milan, Gabriel Gazola}, title = {{Sotaque Brasileiro}}, year = {2021}, publisher = {Zenodo}, version = {2021-09-06}, doi = {10.5281/zenodo.5466945}, url = {https://doi.org/10.5281/zenodo.5466945} } @mastersthesis{batista2019sotaques, author = {Batista, Nathalia Alves Rocha}, title = {{Estudo sobre identificação automática de sotaques regionais brasileiros baseada em modelagens estatísticas e técnicas de aprendizado de máquina}}, year = {2019}, school = {Universidade Estadual de Campinas}, address = {Campinas, SP, Brazil}, type = {Master's thesis}, url = {https://hdl.handle.net/20.500.12733/1636122} }
Majumdar, Sayantan · Smith, Ryan G. · ReVelle, Peter · et al.
3 files · 82 GB · sevenzip, zipdeclared
AZ-Hydro: Historical and Projected Arizona Annual Water Use, 1896-2099 A 2 km-resolution gridded dataset of Arizona groundwater and surface-water withdrawals, irrigation consumptive use, and pumping-induced surface-water capture, spanning 204 years (1896-2099) with quadrature-combined uncertainty bands. Companion data archive for Majumdar et al. (in prep., Nature Scientific Data ) and Majumdar et al. (in prep., AGU Earth's Future ). Graphical Abstract: https://github.com/montimaj/az-hydro/blob/main/docs/images/Graphical_Abstract_Fig1.png Patch release over v1.0.0. The dataset itself is unchanged - every gridded prediction, consumptive-use, surface-water-capture, uncertainty, per-well, and CAP shortage-scenario product in `az-hydro-headline.7z` and `az-hydro-data.7z` is identical to v1.0.0. This version updates only figures, map labeling, and documentation. What changed - Map & figure fixes - corrected map orientations and axis / colorbar / legend labels across the era-mean, σ-attribution, surface-water-capture, and CAP shortage-scenario figures. Withdrawal and consumptive-use HUC12 intercomparison: difference-map orientation/footprint fixes, one-column scatter layout, and decade axis-tick de-cluttering; added public-supply volume-difference maps. Regenerated era-mean and σ-attribution figures and the graphical abstract. - Documentation - README corrections across the source and archive READMEs: basin-count wording (52 basin polygons spanning Arizona's 51 ADWR groundwater basins), disk-space requirements, citation/reference fixes, funding acknowledgment, and data-paper title. Added the pipeline data-harmonization figure. - Source code (`az-hydro-1.0.1.zip`) - matches the tagged GitHub release; now also includes the AZ-Hydro Explorer Earth Engine App (`gee/azhydro-visualizer.js`) and the GEE asset-upload scripts already deployed with v1.0.0. What's in this deposit File Size Contents Audience az-hydro-headline.7z ~8.7 GB Focused subset: published per-pixel/per-year predictions (6-band augmented rasters with σ + CV + SNR + 95 % CI), per-well GeoParquet, SW capture, aggregated time series, CAP shortage scenario outputs, validation against USGS/ADWR/Reitz, era-mean and trend spatial figures. Most users want this. Reviewers, downstream researchers using the published product az-hydro-data.7z ~80 GB Full reproducibility archive: all of the above plus raw inputs (GEE tiles, ADWR meter records, well registry, GW basin / AMA-INA / CAP / SRP / streamflow / USBR vectors, statewide WTD), Step 2 cross-validation outputs, intermediate predictor stacks, per-component σ rasters (σ_MACA / σ_Model / σ_Irr / σ_LULC / σ_GW / σ_USBR / σ_CU). Anyone reproducing the full pipeline from scratch az-hydro-1.0.1.zip ~77 MB Source code release (Python pipeline, GEE export scripts, documentation). Same content as the GitHub repository at the tagged release. Anyone running the pipeline Inside each .7z, see Data/HEADLINE_README.md (headline archive) or Data/README.md (full archive) for a complete per-directory inventory and external-source citations. Important: how to unpack the .7z files The .7z format is not openable by macOS Archive Utility. Use one of: macOS - Keka or The Unarchiver Windows - 7-Zip , WinRAR , or Bandizip Linux - p7zip (e.g. apt install p7zip-full), then 7z x az-hydro-data.7z LZMA2 + solid-block compression yields ~35 % ratio (80 GB compressed from ~224 GB raw; 8.8 GB compressed from ~40 GB raw). Methods overview AZ-Hydro is a four-step physics-constrained ML pipeline: XGBRF prediction of total annual water-use depth per pixel, trained on per-well ADWR meter records (1984-2024) with 16 predictor bands (climate, ET, Peff, irrigation fraction, well density, canal density, water-rights density, etc.). Density-ratio partition decomposing the total prediction into Irrigation/Non-Irrigation × GW/SW × CU using era-mapped factors anchored to USGS Circulars 1950-2015 and ADWR Annual Reports 2016-2024. Six-component quadrature uncertainty quantification : σ_MACA (5 GCMs) + σ_Model (10 XGBRF seeds, t-corrected) + σ_LULC (4 USGS FORE-SCE scenarios) + σ_USBR (5 CMIP3 Upper-Colorado streamflow members, t-corrected) + σ_GW (5 recent ADWR Well Registry snapshots, t-corrected) + σ_CU (analytic propagation through Irrigation Efficiency). Per-pixel SW Capture Fraction and Volume with σ_GW propagation to quantify pumping-induced streamflow depletion. Key validations 2016 ADWR Total : model 6.72 MAF vs ADWR ~7.0 MAF (within -0.28 MAF) 2017 ADWR : 6.81 vs 7.0 MAF; GW share 44.9 % vs 41 % (within 4 pp) 2015 USGS GW pumping : 2.96 vs USGS 3.09 MAF (within -0.13 MAF) 2019-2020 ADWR irrigation share : 73.8 % vs 74 % (essentially exact) WestWater (2026) CAP shortage scenarios : AZ-Hydro Basic Coordination cumulative ΔGW = 7.24 MAF vs WestWater Fig 4 anchor 8.0 MAF (within -9 %); Extreme Shortage = 13.08 MAF vs Fig 4 anchor 8.7 MAF (gap reflects AZ-Hydro's no-regulatory-ceiling framing). CAP delivery shortage scenario sweep Eight scenarios (Baseline_900kAF, DCP Tier 0/1/2a/2b/3, WestWater Basic Coordination, Extreme Shortage) re-partitioned 2026-2099. Cumulative additional GW pumping over 2027-2060: DCP Tier 3 = 10.7 MAF, Basic Coordination = 7.24 MAF, Extreme Shortage = 13.08 MAF. Spatial maps (basin choropleth, per-pixel cumulative ΔGW, σ_cum context, basin/pixel signal-to-noise) included for both 2027-2060 (WestWater anchor) and 2027-2099 windows. External datasets required to reproduce the pipeline (not redistributed here) The pipeline reads four USGS ScienceBase data products that you must download separately: USGS NHM withdrawals ( Haynes et al. 2023 ) - irrigation withdrawals + efficiency by HUC12, 2000-2020 USGS NHM CU / IE reanalysis ( Martin et al. 2023 ; Martin et al. 2025 ) - irrigation consumptive use + Peff by HUC12 USGS public-supply reanalysis ( Luukkonen et al. 2023 ; Alzraiee et al. 2024 ) - public-supply withdrawals by HUC12 USGS Reitz historical ET / Peff ( Reitz et al. 2023 ; Reitz, Sanford & Saxe 2023 ) - 800 m gridded irrigation ET, 1980-2018 Bundled here directly: ADWR Well Registry, ADWR Meter Data, GW basin / AMA-INA / CAP / SRP / streamflow / USBR vectors, HarDWR water-rights shapefile (Lisk et al. 2024) , GRAIN canal network (Suresh et al. 2026) , and the Ma et al. 2026 statewide WTD TIFs . See Data/README.md inside each .7z for exact filename / sub-directory naming required. Citation Data archive (this deposit): Majumdar, S., Smith, R.G., ReVelle, P., Hasan, M.F., & Wogenstahl, C. (2026). AZ-Hydro - Historical and Projected Arizona Annual Water Use: Software, Input Data, Models, Raster and Well Package Predictions, and Validation at 2 km Resolution (1896-2099). Zenodo. https://doi.org/10.5281/zenodo.19057936 Companion papers: Majumdar, S., Smith, R.G., ReVelle, P., Hasan, M.F., & Wogenstahl, C. (2026). Freshwater withdrawals, irrigation consumptive use, and surface water capture for Arizona, 1896-2099 . In prep. for Nature Scientific Data . Majumdar, S., Smith, R.G., ReVelle, P., Hasan, M.F., & Wogenstahl, C. (2026). Where Arizona's Water Goes: Declining Agricultural Dominance and Rising Urban Demand Drive a Two-Century Shift in Withdrawal Patterns (1896-2099). In prep. for AGU Earth's Future . License Data archives (az-hydro-data.7z, az-hydro-headline.7z): CC-BY-4.0 Source code (az-hydro-1.0.0.zip): BSD 3-Clause "Revised" (see LICENSE inside the zip) External datasets bundled here (HarDWR, GRAIN, WTD, ADWR products) retain their original upstream licenses Acknowledgments This work was supported by NASA (Grant numbers 80NSSC21K0979 and 80NSSC23K1453) and U.S. Army Corps of Engineers (Grant number W912HZ25C0016). We thank the open-source software and data communities, the OpenET consortium, and the Arizona Department of Water Resources for making their resources and datasets publicly available, and Google Earth Engine for compute and storage support. S.M. and P.R. acknowledge Dr. Justin L. Huntington, Christopher Pearson, Charles G. Morton, Blake A. Minor, Dr. Samapriya Roy at the Desert Research Institute, and Dr. David Ketchum at the University of Montana for their contributions to related projects that informed this work. We also thank Rahel Pommerenke at Colorado State University for presenting preliminary results from this work at the 2025 ESA Living Planet Symposium. The views expressed herein are those of the authors and do not necessarily reflect those of the funding agencies. Related links Live web app - AZ-Hydro Explorer : https://azhydro.projects.earthengine.app/view/azhydro-explorer (interactive GEE App: year slider 1896-2099, side-by-side category compare, click-driven pixel/basin/sub-basin/well time series with 95 % CI, CAP scenario × window dropdowns) GitHub repository : https://github.com/montimaj/az-hydro Issue tracker / bug reports : https://github.com/montimaj/az-hydro/issues Nature Scientific Data preprint (when available): (link) AGU Earth's Future preprint (when available): (link)
Woodgate, William
1 files · 24 MB · sevenzipdeclared
Light- and CO 2 response curves using the modified LI-6800 were collected from 8 leaves at the long-term Tumbarumba research site in April 2022 (Leuning et al., 2005). Combined gas exchange, PAM fluorescence, reflectance and transmittance (400-650 nm), and spectral fluorescence both forward- and back-scattered (660-850 nm) observations were collected for each light and CO 2 level. See Github repo for processing tools: https://github.com/wwoodgate/li6800_ms
Meekes, Lisa · Tabaro, Francesco · Bexkens, Michiel · et al.
41 files · 8.2 GB · csv, fasta, pdfdeclared
This record contains the Python software for PEPTiGEN, a tool for generating tryptic peptides from prokaryotic gene sequences and their variants, and the associated antimicrobial resistance (AMR) peptide database. The database is provided as an SQL file and a CSV file containing all genes and predicted peptides. The README file contains explanation of the PEPTiGEN tool. The SQL database schema files contains both the database schema of the SQL database used in the PEPTiGEN analysis as the database schema of the AMR peptide datbase.
Barreto, Bruno · Eisencraft, Marcio
3 files · 2.6 GB · parquetdeclared
This dataset provides a fixed benchmark dataset for stellar atmospheric parameter estimation from Sloan Digital Sky Survey Data Release 12 (SDSS DR12) optical stellar spectra. The dataset is organized into three predefined Parquet splits: 30,000 spectra for training, 5,000 spectra for validation, and 15,000 spectra for testing. Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, fixed-length processed spectral features, source identifiers, basic metadata, and catalog stellar-parameter labels with their associated uncertainties. The supervised regression targets are the adopted catalog stellar atmospheric parameters: effective temperature (Teff, in K), metallicity ([Fe/H], in dex), and surface gravity (log g, in dex). The dataset also includes relevant observational and catalog information such as SDSS plate, MJD, fiber identifier, sky coordinates, signal-to-noise ratio, adopted radial velocity, raw flux, logarithmic wavelength grid, inverse variance, pixel mask, and processed flux features. This release is intended to support machine-learning research on stellar spectroscopy, including regression models for atmospheric parameter estimation, benchmark comparisons, uncertainty-aware evaluation, and experiments using either processed fixed-length spectra or native observed-frame spectral arrays.
Saldanha, Raphael
8 files · 601 MB · parquet, zipdeclared
Description This deposit contains annual, municipality-level datasets derived from the Brazilian Primary Health Care Information System (SISAB). The files combine two complementary data sources: Public SISAB Saúde report downloads from the Atendimento/Visita production report. CID-10 and CIAP-2 attendance data obtained from SISAB through requests under the Brazilian Access to Information Law (Lei de Acesso à Informação, LAI). The datasets are organized as tidy annual files in CSV (Zipped) and Parquet format. They are intended to support reproducible analysis of primary care production, procedures, evaluated problems/conditions, and CID/CIAP-coded attendances across Brazilian municipalities. The public SISAB report datasets are stratified by competence month, state, municipality, DataSUS age group, SISAB sex category, and the selected report category. For each competence month and report type, the extraction combines 36 stratified SISAB downloads: 18 age groups by 2 sex values. Monthly files are merged into yearly files, completing missing combinations of observed competence, municipality, age group, sex, and category with valor = 0 . The LAI dataset contains yearly CID-10 and CIAP-2 attendance counts by competence month, municipality, code type, and code. When multiple valid LAI files cover the same competence, the processing pipeline selects the file with the largest number of data rows, using file size and request folder order as tie-breakers. Provenance columns identify the selected LAI request and source file. Variables SISAB Saúde Produção Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : Age group. sexo : SISAB sex category, Masculino or Feminino . tipo_producao : production type from the SISAB report. valor : count reported by SISAB. SISAB Saúde Procedimento Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . procedimento : procedure from the SISAB report. valor : count reported by SISAB SISAB Saúde Condição Avaliada Columns: competencia : competence month in YYYYMM format. uf : Brazilian state abbreviation. ibge : municipality IBGE code. municipio : municipality name. faixa_etaria : age group. sexo : SISAB sex category, Masculino or Feminino . condicao_avaliada : evaluated problem or condition from the SISAB report. valor : count reported by SISAB. SISAB LAI CID/CIAP Columns: ano_competencia : competence year. competencia : competence month in YYYYMM format. competencia_date : first day of the competence month. co_municipio_ibge : municipality IBGE code. tp_codigo : code type, CID or CIAP . codigo : CID-10 or CIAP-2 code. qt_atendimentos : number of attendances. source_request : selected LAI request folder. source_file : selected source CSV file. Methods The public SISAB report files were generated with the sisab_scrapper processing pipeline. For each month, report type, age group, and sex value, the pipeline downloads the all-Brazil municipality report from SISAB, validates the returned CSV, preserves raw cache files for resumable runs, and writes a sorted monthly tidy dataset. The yearly merge validates required columns, expected age groups, expected sex values, category values, and month gaps unless explicitly allowed. The CID/CIAP files were generated with the sisab_lai processing pipeline. The pipeline imports CSV files received through LAI requests, detects the real CSV header after any SQL*Plus preamble, validates candidate files, resolves overlapping requests by competence, standardizes old and new schemas into one tidy table, and exports annual CSV and Parquet files together with audit reports. Sources - SISAB public reports, Ministry of Health, Brazil: https://sisab.saude.gov.br/ - SISAB LAI files obtained through Brazilian Access to Information Law requests. - Processing code for public SISAB report data: https://github.com/rfsaldanha/sisab_scrapper - Processing code for LAI CID/CIAP data: https://github.com/rfsaldanha/sisab_lai Notes - Counts are aggregated administrative records and should be interpreted in light of SISAB reporting practices, data quality, and changes in municipal reporting coverage. - Municipality boundaries, names, and coding practices may vary over time. - Public SISAB report datasets are completed with zero values only for combinations defined by observed municipalities, observed competencies, all expected age groups, both expected sex values, and observed report categories within the yearly merge. - LAI CID/CIAP data preserves selected source-file provenance through source_request and source_file . - This deposit corresponds to an individual year. Deposits for other years are published separately.
Hagger, Thomas · Hassanzadeh, Mohammadreza · Urbonavicius, Aidas · et al.
3 files · 24 GB · sevenzip, zipdeclared
Raw experimental data and analysis scripts. The repository includes cathodoluminescence, electron microscopy, photoluminescence, and complementary characterization data organised by figure, together with supporting analysis scripts where applicable. The full electron microscopy data is given in TEM.7z, while the figures used in the manuscript and SI are given in the Raw_Data folder
Huang, Yicong · Shamsnia, Ali · Chen, Mengze · et al.
10 files · 26 GB · sevenzipdeclared
Paper : Intrinsic timing, not temporal prediction, underlies ramping dynamics in visual and parietal cortex, during passive behavior Authors: Huang Y, Shamsnia A, Chen M, Wu S, Stamm T, Medico S, Najafi F The data include two-photon calcium imaging from distinct excitatory and inhibitory cell types (excitatory, VIP, and SST neurons) in the visual (VIS) and parietal (PPC) cortices of mice passively exposed to auditory and visual stimuli. Number of imaged neurons: - Excitatory: ~26,600 - VIP: ~3,900 - SST: ~3,100 File details and code: Details on the contents of each file, how to use the data for analysis, and how to reproduce the paper's results are provided in: 2p_imaging/passive_interval_oddball_202412 at main · najafi-laboratory/2p_imaging Files: PPC data: YH01VT, YH02VT, YH03VT, YH14SC, YH16SC. VIS data: YH17VT, YH18VT, YH19VT, YH20SC, YH21SC. File names ending in: VT contain simultaneous imaging data from excitatory neurons and VIP interneurons. SC contain imaging data from SST interneurons.
Kharyuk, Pavel
4 files · 1.6 GB · sevenzipdeclared
This repository contains data collected under the following study: P.Kharyuk, S.Matveev, I.Oseledets. Exploring specialization and sensitivity of convolutional neural networks in the context of simultaneous image augmentations, arXiv:2503.03283 . Corresponding source code repository: https://github.com/kharyuk/activation_sa CNNs: ILSVRC: 10.5281/zenodo.18097911 Places365: 10.5281/zenodo.18098133 0_models.7z: copy of the visual transformer models (ViT-b-16, Swin-t, MaxViT-t) used in the research (reference: https://docs.pytorch.org/vision/main/models.html ) 1_sensitivity_values.7z: sensitivity values (Sobol indices, Shapley values) computed for the first experimental series (ViT-b-16, Swin-t). In addition, logs and npz-files containing the sampled parameters were packed. To be used within the jupyter notebooks, all .hdf5 files should be placed into the 'results' directory of the source code repository. 2a_predictions.7z: these files include the top-5 class predictions and corresponding classifying layer's outputs used for evaluating Table 1 (ViT-b-16, Swin-t, MaxViT-t). SupplementaryS3.7z : Supplementary material containing visualized sensitivity maps (Vit-b-16, Swin-t).
YANG, YIYAN
56 files · 11 GB · sevenzip, zipdeclared
The data are used in Github Readme and PHILM Step-by-Step Tutorial . The provided archives include: * `example.zip`: Example dataset used in the README. * `tutorial_data.zip`: Healthy human stool sample dataset used in the step-by-step tutorial. * `PHILM_input_data.7z`: Contains the input files required by PHILM. These files should be placed in the `data/` directory. * `PHILM_model_results.7z`: Contains the trained model and prediction metrics, including `PHILM_best_model.pth`, `PHILM_best_params.yaml`, `PHILM_predict_test.ft.metrics`, `PHILM_predict_val.ft.metrics`, and `PHILM_predict_train.ft.metrics`. * `PHILM_interactions.7z`: Contains the inferred interaction results, including `PHILM_interactions.tsv`, `PHILM_interactions.pvalues.tsv`, and `PHILM_interactions.pvalues.filtered.tsv`. * `perm_<start number-end number>_scores.tsv.7z`: * `perm_<start-end>_scores.tsv.7z`: These archives contain the permutation-derived PHILM scores used for empirical p-value calculation. Because storing all 1,000 permutation results in a single file would be very large, the results were divided into 50 compressed files. After decompressing all `.7z` files, organize the permutation results into the required PHILM directory structure by running: `python organize_perm_files.py --input "perm_*-*_scores.tsv" --outdir permutation_null`. This command will generate files with the following structure: `permutation_null/perm_*/PHILM_perm_*.raw_gradient.tsv`. Before running Step 5.3, move or place the generated `permutation_null/` directory under the `results/` directory.
Anonymous Authors
15 files · 2.2 GB · parquet, rar, torchdeclared
Economic losses caused by extreme climate are not confined to the locations where events occur, but can propagate across regions through physical and economic linkages. Yet existing climate-impact assessment methods remain poorly suited to tracing how shocks spread across space and reshape the geography of economic loss. Here we develop a mechanistically informed multi-scale spatiotemporal autoregressive graph neural network model to quantify spatially cascading climate impacts. The model couples scale-specific, physically structured spatiotemporal autoregressive processes through an adaptive gating mechanism, allowing heterogeneous cross-scale interactions to be learned from data. Model estimation is achieved through tailored graph convolutional neural networks that are mathematically equivalent to spatiotemporal autoregressive models, enabling scalability while preserving transparent parameter interpretation. Monte Carlo simulation experiments show that the model accurately recovers true parameters and distinguishes between scale-dependent processes. Applying the framework to extreme precipitation, we find that large-scale upwind-to-downwind cascades driven by atmospheric circulations dominate aggregated economic losses. A one-standard-deviation increase in log extreme precipitation is associated with a 0.19 percentage-point decline in economic growth rate at the large scale, with 62.3% of the loss arising from spatial cascades. These findings highlight the need for transboundary risk governance that incorporates spatial cascading into climate-extremes monitoring and early-warning. Description of the uploaded file Monte Carlo simulation code data_generator_factors.py: Data generation script for multi-scale Monte Carlo simulation experiments. sarnn_model.py:Implementation of the proposed MS-STARGNNs model architecture definition. train.py:Training pipeline script for the MS-STARGNNs model. Data and spatial weights matrices for empirical analysis global_panel_1deg_std.parquet:Standardized Large-scale (1°) datase; global_panel_2km_std.parquet:Standardized Small-scale (2 km) dataset. W_global_2km_knn8.pt:Small-scale spatial weights matrix based on 8-nearest neighbors (KNN8). W_Large-scale:A large-scale spatial weights matrix derived from moisture transport pathways (2005-2021) mapping_1deg_to_2km.parquet: Correspondence file mapping large-scale (1°) grid cells to small-scale (2km) grid cells. ipcc_region_mapping_coarse_8regions.parquet: Mapping file linking Large-scale grid cells to the 8 IPCC AR6 reference regions. ipcc_region_mapping_fine_8regions.parquet: Mapping file linking Small-scale grid cells to the 8 IPCC AR6 reference regions. Code model.py: Core architecture definitions for empirical analysis. Provides the base classes and computational layers engineered to handle real-world geospatial complexities. All subsequent training scripts import modules from this file. STARGNNs.py: Implementation of the single-scale baseline. Serves as a reference point for evaluating the efficacy of cross-scale feature fusion. MS-STARGNNs_fixed.py: Configuration script for the MS-STARGNNs model utilizing fixed autoregressive coefficients. MS-STARGNNs.py: Configuration script for the MS-STARGNNs model utilizing annually varying autoregressive coefficients. MS-STARGNNs_8 regions.py: Executes the MS-STARGNNs model with decoupled regional parameters, loading unique autoregressive weights and βvectors for each IPCC region. Implements null value handling for regions lacking observational data (e.g., Antarctica).
Matuszewska-Mach, Eliza · Konieczny, Igor · Matysiak, Jan · et al.
45 files · 46 GB · sevenzipdeclared
This dataset contains raw and processed mass spectrometry data generated for the study: "Honeybee (Apis mellifera) venom and its bioactive peptides tertiapin and apamin modulate the proteome of SARS‑CoV‑2‑infected cells." The aim of the study was to characterize host proteome remodeling in human ACE2‑expressing HEK293 (hACE2) cells infected with SARS‑CoV‑2 spike‑pseudotyped lentiviral particles and treated with: Whole honeybee venom (HBV) Apamin (APA) Tertiapin (TERT) Proteomic profiling was performed using nano‑liquid chromatography coupled to MALDI‑TOF/TOF tandem mass spectrometry (nanoLC‑MALDI‑TOF/TOF MS/MS). Experimental Design hACE2 cells were: Infected with SARS‑CoV‑2 spike pseudotyped lentiviral particles (MOI 1.33) Treated with: HBV (1 µg/mL and 2 µg/mL) Apamin (0.5, 1, 2 µg/mL) Tertiapin (0.25, 0.5, 1, 2 µg/mL) Compared to infected untreated control cells Each condition was analyzed in: Two biological replicates Two technical replicates Mass Spectrometry Workflow Protein extraction: TAP buffer with protease and phosphatase inhibitors Digestion: Trypsin (in-solution digestion, modified Pierce protocol) Peptide cleanup: ZipTip C18 nanoLC system: Easy‑nLC II (Bruker Daltonics) Fractionation: Proteineer‑fc II fraction collector Matrix: HCCA (α‑cyano‑4‑hydroxycinnamic acid) Target: AnchorChip 384 (Bruker Daltonics) Instrument: UltrafleXtreme MALDI‑TOF/TOF (Bruker Daltonics) Acquisition mode: Reflectron mode MS/MS analysis: TOF/TOF fragmentation Raw Data Files (RAW) The repository includes: Complete Bruker MALDI‑TOF/TOF raw data folders (.mcf format) All nanoLC fractions (384 spots per sample) Raw files were acquired using: FlexControl 3.4 FlexAnalysis 3.4 These files preserve the original vendor format required for reanalysis. Identification / Result Files (RESULT) Protein identification was performed using: Software: Mascot v2.4.1 Database: NCBInr Taxonomy: Homo sapiens The dataset includes: Mascot search result files Protein identification lists Organism Homo sapiens (human HEK293 ACE2-expressing cells) Instrument Bruker UltrafleXtreme MALDI‑TOF/TOF mass spectrometer