Oberbauer, Barbara · Hahnel, Ulf J. J. · Gluth, Sebastian
hybrid · semantic + lexical · 2029 datasets ranked · 5.86s
Oberbauer, Barbara · Hahnel, Ulf J. J. · Gluth, Sebastian
1 files · 100 MB · zip
The raw data set used for Study 1 of the publication Computational Mechanisms of Attribute Translations was publicly made available by Mertens et al. (2020) on the Open Science Framework . The data set of Study 2, all preprocessed behavioral data, parameter estimation results as well as simulated behavioral data are available here. These data allow reproducing all results figures provided with this paper.
Kinast, Alexander · Falkner, Dominik · Bögl, Michael · et al.
1 files · 100 MB · zip
Benchmark data set explained in the Eurocast 2026 paper: "Towards Less Energy Reliance - A Computational Study on Benchmark Instances for Selected Locations in Austria"
Riddiford, Lauren · Scagnoli, Valerio
1 files · 100 MB · zip
Generating unconventional spin-orbit torques with patterned phase gradients in tungsten thin films
Magnouloux, Eva · Lemercier, gilles · Fronchot, Céline · et al.
1 files · 100 MB · zip
MD Simulations Data and topologies for the study of Mucin1 Beta in a lipid menbrane
Baltic Marine Environment Protection Commission
1 files · 100 MB · zip
A unified dataset of 630 species distribution models covering the whole Baltic Sea, including the Danish Straits and Kattegat. The dataset contains distribution models and maps for all fish, macrophytes, and invertebrates for which sufficient observation data within the region were available. The predicted distribution maps are provided in 250 m resolution and include several outputs intended to support spatial planning in the region, including spatial confidence estimates and expert evaluation scores for each species. Importantly, the dataset was produced using a harmonized methodology suitable across a wide range of species, allowing users to combine and compare different species and species groups for subsequent spatial analyses.
Li, Xuankun · Tu, Yuezheng · Plotkin, David · et al.
5 files · 100 MB · zip
Raw molecular data for 90 newly-sequenced specimens of Eupterotidae (and related families of Lepidoptera) used to generate a molecular phylogeny for the study "Phylogeny and biogeographic history of monkey moths (Lepidoptera: Eupterotidae)" (currently under peer review as of May 2026). File names contain the sequence ID and taxonomic information for each specimen. Data files are in fastq format and have been compressed; there are two files per specimen (labelled "R1" and "R2"). Data files have been organized into five .zip files, based on family-group taxonomy, as follows: Eupterotidae: Eupterotinae (25 specimens, 50 files) Eupterotidae: Ganisa Group (15 specimens, 30 files) Eupterotidae: Janinae (27 specimens, 54 files) Eupterotidae: Striphnopteryginae (15 specimens, 30 files) Other Lepidoptera families (outgroup taxa): Anthelidae, Bombycidae, Lasiocampidae, Saturniidae (8 specimens, 16 files)
Lolicato, Fabio · Javanainen, Matti
1 files · 100 MB · zip
This repository contains the data, simulation files, and scripts used in the study "High-Throughput Characterization of Transmembrane Helix Partitioning in Membrane Domains." The repository includes: Scripts used to extract FASTA sequences from the Orientations of Proteins in Membranes database (OPM). Scripts and input files used to generate peptide systems and prepare the molecular dynamics simulations. Simulation parameter files, topology files, coordinate files required to reproduce the simulations. Final structure files ( md.gro ) for each simulation system. The deposited material is intended to ensure transparency, reproducibility, and reuse of the computational workflow and simulation datasets associated with this work.
Obermayer, Benedikt
1 files · 100 MB · zip
Processed data for Plumbom et al. "circVDJ-seq for T cell clonotype detection in single-cell and spatial multi-omics" - annotation: contains meta-data for different single-cell datasets - cellranger_count: contains count matrices (h5 objects) and fragment files for ATAC - cellranger_vdj: contains `cellranger vdj` output (metrics_summary.csv and filtered_contig_annotations.csv) - dandelion: contains results of re-processing of cellranger vdj output with dandelion - mixcr: contains circVDJ-seq and spatialVDJ processing results using MiXCR - objects: contains processed Seurat objects for lymph node Visium data - spaceranger: contains spaceranger output for neuroblastoma and lung/lymph node visium runs - trust: contains circVDJ-seq and MAS-ISO-seq processing results using TRUST4
Martín Frechina, Sara · Dura, Esther · Torres-Sospedra, Joaquín
2 files · 100 MB · zip
This repository contains the UVIndoorLoc-BLE-AoA&IQ , a collection of In-phase and Quadrature (I/Q) signal samples and Angle-of-Arrival (AoA) measurements based on the Bluetooth Low Energy (BLE) 5.1 specification. The project focuses on providing a BLE 5.1 AoA dataset by collecting data in three distinct real-world scenarios at the Escola Tècnica Superior d'Enginyeria (ETSE) of the University of Valencia (UV). It was designed to evaluate the impact of anchor topologies and environmental complexity on positioning accuracy. Dataset Features and Format The dataset provides the following parameters for every advertising event, organized into three stages of processing: 1. Raw Data (raw_data/) Individual files per anchor containing the original output from the u-blox firmware: timestamp : Sub-millisecond synchronization reference. ch : BLE channel index used for the transmission. rssi : Received Signal Strength Indicator in dBm. pac : Sequential packet counter. ms : Internal uptime of the NINA-B411 module. az_ublox , el_ublox : Firmware-estimated Azimuth ( $\alpha$ ) and Elevation ( $\epsilon$ ). iq_data : Sequence of 164 values (82 I/Q pairs) captured during the Constant Tone Extension (CTE). 2. Syncronized Data (processed_iq_fusion/) Merged streams from the four anchors using a temporal matching algorithm: timestamp , channel : Shared event identifiers. rssi_AnchorX , az_AnchorX , el_AnchorX , iq_AnchorX : Individual metrics for each of the 4 anchors. ax_AnchorX , ay_AnchorX : Global coordinates of each anchor for the specific topology. config : Topology identifier (Face-to-Face or 45º Inward). 3. Final Dataset (final_dataset/) Cleaned and labeled datasets ready for positioning algorithms and machine learning training: id : Unique identifier for the measurement session. az_ublox_anchorX , el_ublox_anchorX : Cleaned firmware angles. rssi_anchorX , iq_anchorX : Signal metrics for multi-anchor fusion. real_x , real_y , real_z : Precise Ground Truth (GT) coordinates measured via laser distance meter. Experimental Scenarios The dataset distribution is as follows: Scenario Training Points Evaluation Points Total Area 1.a Laboratory 28 18 $59.286\text{ m}^2$ 1.b Corridor 26 21 $135.235\text{ m}^2$ 2. Meeting Room 26 30 $187.392\text{ m}^2$ 3. Auditorium 25 20 $152.170\text{ m}^2$ Hardware Specifications Receivers : 4 u-blox ANT-B10 antenna boards (8-element patch array). Gateways : 4 Raspberry Pi 5 units synchronized via Chrony (NTP). Transmitter : u-blox C211 beacon (100 ms advertising interval).
Bhat, Shreeharsha G · Mahajan, Daanish · Jain, Chirag
7 files · 8.0 MB · gzip, zip
Panngenome graph files (GFA format) used for benchmarking panbubble and hairpin detection in the paper Billi: Provably Accurate and Scalable Bubble Detection in Pangenome Graphs . Includes minigraph pangenome graphs and pangene gene graphs. Note: The HPRC Minigraph-Cactus pangenome graphs (hprc-v1.1-mc-chm13 and hprc-v2.0-mc-chm13) and chromosome-level graphs (chrX and chrY) are not included due to their large size. The chromosome-level graphs (chrX, chrY) are available as .vg files and must be converted to GFA format using vg. Direct download links: hprc-v1.1-mc-chm13.gfa.gz: https://s3-us-west-2.amazonaws.com/human-pangenomics/pangenomes/freeze/freeze1/minigraph-cactus/hprc-v1.1-mc-chm13/hprc-v1.1-mc-chm13.gfa.gz hprc-v2.0-mc-chm13.gfa.gz: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.gfa.gz chrX.vg: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.chroms/chrX.vg chrY.vg: https://human-pangenomics.s3.amazonaws.com/pangenomes/scratch/2025_02_28_minigraph_cactus/hprc-v2.0-mc-chm13/hprc-v2.0-mc-chm13.chroms/chrY.vg Vg command to convert .vg to .gfa format: vg view input.vg > output.gfa
Personnettaz, Paolo · Cébron, David
1 files · 100 MB · zip
Data and code from "Magnetohydrodynamic drag on an oscillating sphere in a rotating spherical cavity".
Ziqing, Zhang
1 files · 100 MB · zip
Motivation: Toxicity assessment is essential for drug development and clinical safety. However, a prevalent phenomenon in computational toxicology is the trade-off between predictive performance and mechanistic interpretability, where conventional machine learning offers greater transparency but deep learning typically achieves higher accuracy. This limitation hinders the adoption of models in regulatory decision-making, where understanding the biological reason for toxicity is as vital as the prediction itself. Results: We present ToxKAN, an adverse outcome pathway (AOP) knowledge-guided Kolmogorov-Arnold framework that integrates chemical structures and gene expression profiles for mechanistic drug toxicity prediction. Extensive evaluations demonstrate that ToxKAN outperforms state-of-the-art methods in binary toxicity prediction and achieves competitive or superior, ranking among the highest reported accuracy on fine-grained multi-label pathological phenotype prediction. Critically, the model identifies biologically coherent hierarchical reasoning paths, prioritizes core toxicity genes in low-signal settings where conventional differential expression analysis fails, and learns latent representations that reflect mechanism-aware compound organization. These results indicate that mechanism-aware architecture design can simultaneously advance predictive performance and biological interpretability. Availability and implementation: The source code and processed datasets are publicly available at https://github.com/shuangquanZW/ToxKAN and Edit upload | Zenodo .
Genie3 Team
1 files · 100 MB · zip
MotifBench results for Genie3 (https://www.biorxiv.org/content/10.64898/2026.05.01.722168v1). Three independent scaffold generation runs are included, each with directional scale set to 0.1. Scaffolds generated by Genie3 team (contact: yl4599@columbia.edu). Re-evaluated by Zhuoqi Zheng (contact: h2knight@sjtu.edu.cn).
Haghshenas Haghighi, Mahmud · Piter, Andreas
1 files · 100 MB · zip
Overview This repository contains the datasets used in the SARvey and and InSAR Explorer tutorial notebooks. Tutorial notebooks The tutorial notebooks are available in the SARvey Tutorials GitHub repository . The dataset contains modified Copernicus Sentinel data 2014-2026, processed by ESA.
Genie3 Team
1 files · 100 MB · zip
MotifBench results for La-Proteina (https://openreview.net/pdf?id=RDerF20JYT). Scaffolds generated by Genie3 team (contact: yl4599@columbia.edu). Re-evaluated by Zhuoqi Zheng (contact: h2knight@sjtu.edu.cn).
Genie3 Team
1 files · 100 MB · zip
MotifBench results for FrameFlow (https://openreview.net/forum?id=fa1ne8xDGn). Scaffolds generated by Genie3 team (contact: yl4599@columbia.edu). Re-evaluated by Zhuoqi Zheng (contact: h2knight@sjtu.edu.cn).
Yang, Qingliu
6 files · 8.0 MB · bzip2
Dataset Description This dataset contains 3D lightning location results, DALMA and FALMA waveform for two Energetic Compact Stroke (ECS) events. location results are included: HF3D_1732785151.dat - 3D lightning locations for the ECS leader A flash. HF3D_1734785454.dat - 3D lightning locations for the ECS leader B flash. The timestamp 1734785454 and 1732785151 corresponds to the occurrence time of the lightning flash in Japan Standard Time. File format and parameters The first row contains the lightning occurrence time. Column descriptions: Time (ms) - time relative to the lightning source. X, Y, Z (m) - 3D spatial coordinates relative to ground level. The origin (0, 0, 0) corresponds to latitude 36.76°N and longitude 136.76°E. FALMA and DALMA waveform ECSLeaderA_DALMA_waveform.bz2 is DALMA waveform of Leader A. ECSLeaderA_FALMA_waveform.bz2 is FALMA waveform of Leader A. ECSLeaderB_DALMA_waveform.bz2 is DALMA waveform of Leader B. ECSLeaderB_FALMA_waveform.bz2 is FALMA waveform of Leader B. This dataset allows analysis of the spatial and temporal development of these two ECS flashes.
Kovaliov, Michael
1 files · 100 MB · parquet
Pribytova, Ekaterina
4 files · 100 MB · zip
Li, Tong · Bresser, Dominic
1 files · 100 MB · zip