Robert Koch-Institut · AKTIN-Notaufnahmeregister
86 rows × 8 cols · 10 KB · pdf, tsv, zip
4 numeric · 3 categorical · 1 text
hybrid · semantic + lexical · 107 datasets ranked · 2.87s
Robert Koch-Institut · AKTIN-Notaufnahmeregister
86 rows × 8 cols · 10 KB · pdf, tsv, zip
4 numeric · 3 categorical · 1 text
Der Datensatz "Daten der Notaufnahmesurveillance" wird durch das Robert Koch-Institut und das AKTIN-Notaufnahmeregister bei Krankenhausaufnahmen bereitgestellt. Der Datensatz beinhaltet aggregierte Routinedaten aus deutschen Notaufnahmen zur syndromischen Überwachung von akuten Erkrankungen. Dazu zählen grippeähnliche Erkrankungen (ILI), Coronavirus-Erkrankungen (COVID-19), akute respiratorische Erkrankungen (ARE), gastrointestinale Infektionen (GI) und schwere akute respiratorische Infektionen (SARI). Dabei wird der relative Anteil dieser Erkrankungen an der Gesamtzahl der Notaufnahmevorstellungen sowie die berechneten Erwartungswerte und Prädiktionsintervalle ausgewiesen. Die Daten sind nach Notaufnahmetypen und Altersgruppen aggregiert. Damit bietet der Datensatz eine wertvolle Ressource für die Forschung im Bereich der Notfallmedizin und der Überwachung akuter Gesundheitsereignisse in Deutschland.
Kantor, Rose · Shakya, Migun · Ruth, Nelson · et al.
2,095 rows · 907 KB · fasta, tsv
A virus genome database representing 21,015 near-complete virus genomes collected from untargeted ultra-deep RNA/DNA combined sequencing of wastewater. Sequence data was provided by the CASPER consortium and raw data may be found on NCBI SRA under bioprojects PRJNA1247874 and PRJNA1198001. Data underwent read trimming, rRNA and human read removal, de novo assembly, and selection of high-quality viral contigs. Contigs were clustered at 95% identity and 85% query coverage to dereplicate. Chimera-checking required at least two independent assemblies of the same viral genome or presence of the genome in another reference database. Annotation made use of RdRpCATCH, geNomad, checkV, BLASTN against NCBI core-nt, and RNAVirHost. The RdRp fasta files contain representative RdRp sequences identified through homology to major RdRp reference databases and clustered at 90% sequence identity over 75% sequence coverage. Included sequences contain all three conserved RdRp motifs (A, B, and C) arranged in either the canonical ABC configuration or the permuted CAB configuration.
Robert Koch-Institut · AKTIN-Notaufnahmeregister
5 files · 10 KB · pdf, tsv, zip
Der Datensatz "Daten der Notaufnahmesurveillance" wird durch das Robert Koch-Institut und das AKTIN-Notaufnahmeregister bei Krankenhausaufnahmen bereitgestellt. Der Datensatz beinhaltet aggregierte Routinedaten aus deutschen Notaufnahmen zur syndromischen Überwachung von akuten Erkrankungen. Dazu zählen grippeähnliche Erkrankungen (ILI), Coronavirus-Erkrankungen (COVID-19), akute respiratorische Erkrankungen (ARE), gastrointestinale Infektionen (GI) und schwere akute respiratorische Infektionen (SARI). Dabei wird der relative Anteil dieser Erkrankungen an der Gesamtzahl der Notaufnahmevorstellungen sowie die berechneten Erwartungswerte und Prädiktionsintervalle ausgewiesen. Die Daten sind nach Notaufnahmetypen und Altersgruppen aggregiert. Damit bietet der Datensatz eine wertvolle Ressource für die Forschung im Bereich der Notfallmedizin und der Überwachung akuter Gesundheitsereignisse in Deutschland.
Altenhoff, Adrian
6 files · 100 MB · hdf5
OMAmer - tree-driven and alignment-free protein assignment to subfamilies OMAmer is an alignment-free protein family assignment method designed to avoid overly specific subfamily predictions and to scale efficiently to phylogenomic databases containing thousands of genomes. It relies on an innovative approach that uses evolutionarily informed k-mers for alignment-free mapping to ancestral protein subfamilies. This dataset provides precomputed OMAmer databases derived from the Hierarchical Orthologous Groups in the OMA Browser . We aim to update these databases with every new OMA Browser release. Each OMAmer database is built using the latest version of the OMAmer package available at the time of the corresponding OMA Browser release. The dataset includes databases for different subsets of the species taxonomy. In most cases, we recommend using the LUCA.h5 database, which contains information from all species in the OMA database. The subset-specific databases are mainly useful when disk space is limited. The release May2026 is based on the OMA Browser release May 2026 which comprises 2983 species. We used OMAmer version 2.1.0 to build these databases.
guo, na · pan, jiachen · li, tiantian · et al.
14 files · 100 MB · csv, rar, tsv
ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR
Manghi, Paolo
2 files · 8.0 MB · gzip, tsv
Using gene-level and species-level shotgun metagenomics, we provide the first characterization of the rural, Bolivian microbiome; we identified microbial genes which strongly correlate (rho>0.45) with arsenic in urine, and that overall contribute to substantiate that the gut microbiome helps tolerate arsenic via a evict-out-of-house mechanism. Mediation analysis, followed by phylogenetic investigation of metagenomic-assembled genomes, further strengthens this observation. This study elucidates the role of the microbiome in helping to tolerate arsenic-rich environments, and paves the way for probiotic interventions that may mitigate the effects of this toxic metal. This repository contains 2,478 HQ MAGs from the MicroToxBol cohort in fasta format and a descriptive table comprising taxonomic annotation, quality-checks, and coverage estimation.
Xu, Tao
3 files · 100 MB · rar, xlsx
The deepmd_data dataset comprises energy data for 375,781 structures and force data for more than 45 million atoms, generated from six hierarchical active learning iterations. The init directory stores the non-periodic molecular configurations established at the initialization stage. The iter directories contain structures labeled by both AIMD and DPMD-FP methods for each iteration. Each folder is named according to the scheme of solvent molecule index followed by lithium salt designation. The finetuned_data folder contains SEI reaction simulation data used for fine-tuning the MLFFs. Comprehensive details regarding the solvent molecules are available in the solvent_with_smiles.xlsx file.
Yuezong Wang · Yu Niu · Jiqiang Chen
2 files · 100 MB · rar, zip
This dataset contains raw experimental images, video sequences of real test samples, simulated microscopic image data from the manuscript " A robust method for microscopic 3D shape restoration via shape-from-focus" , as well as focal volume datasets collected before vibration simulation, after vibration simulation, and post anti-vibration processing. All provided data enable the validation of conclusions and reproducibility of experimental results obtained via the computational pipeline proposed in this paper.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Daugavietis, Jānis
6 files · 3.3 GB · jpeg, rardeclared
Skala Bünte, Süki Vagonu Hallē [2026.07.08.] Koncerta beigu telefona foto/ video. SKALA BÜNTE X SÜKI FACE2FACE https://fb.me/e/7dycfN2gI Details Event by John Dow Vagonu Hall Public · Anyone on or off Facebook 8TH OF JULY VAGONU HALLE TWO BANDS TWO BACKLINES MOSHPIT IN THE MIDDLE ONCE IN A LIFETIME FACE2FACE MASSACRE PROVIDED BY SKALA BÜNTE & SÜKI 🔪🔪🔪 DOORS 19:00 7€
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Walsh, Calum · Srinivas, Meghana · Stinear, Timothy · et al.
100 files · 2.7 GB · gzip, tsvdeclared
GROND (Genome-derived Ribosomal OperoN Database) A quality-checked and publicly-available database of 16S-ITS-23S RRNA operon sequences and their constituent 16S and 23S genes. Based on GTDB release R232.
Zhang, Yunwei
6 files · 40 GB · hdf5, zipdeclared
Files related to INVNET, a deep learning model for surface wave dispersion spectrum inversion in geophysics.
Murilo Coelho · de Sousa Amâncio, Francisco Dione · Paixao, Matheus · et al.
3 files · 742 KB · rardeclared
This repository contains the replication package for the paper "More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions", accepted at the 40th Brazilian Symposium on Software Engineering (SBES 2026), São Paulo, Brazil. The package includes: (i) the complete survey instrument and sanitized participant responses; (ii) qualitative coding artifacts, including the codebook, category consolidation, and classification analysis; (iii) inter-rater reliability materials (Cohen's Kappa = 0.81); (iv) quantitative datasets and statistical analysis outputs (SPACE composite scores, Cronbach's alpha, Kruskal-Wallis, Mann-Whitney U, and Dunn's post-hoc tests); and (v) supporting literature review material. All participant data were anonymized prior to disclosure. The survey materials are in Portuguese, the language of data collection.
Robert Koch-Institut · AKTIN-Notaufnahmeregister
5 files · 24 MB · pdf, tsv, zipdeclared
Der Datensatz "Daten der Notaufnahmesurveillance" wird durch das Robert Koch-Institut und das AKTIN-Notaufnahmeregister bei Krankenhausaufnahmen bereitgestellt. Der Datensatz beinhaltet aggregierte Routinedaten aus deutschen Notaufnahmen zur syndromischen Überwachung von akuten Erkrankungen. Dazu zählen grippeähnliche Erkrankungen (ILI), Coronavirus-Erkrankungen (COVID-19), akute respiratorische Erkrankungen (ARE), gastrointestinale Infektionen (GI) und schwere akute respiratorische Infektionen (SARI). Dabei wird der relative Anteil dieser Erkrankungen an der Gesamtzahl der Notaufnahmevorstellungen sowie die berechneten Erwartungswerte und Prädiktionsintervalle ausgewiesen. Die Daten sind nach Notaufnahmetypen und Altersgruppen aggregiert. Damit bietet der Datensatz eine wertvolle Ressource für die Forschung im Bereich der Notfallmedizin und der Überwachung akuter Gesundheitsereignisse in Deutschland.
Tiskus, Edvinas · Tiškuvienė, Rūta · Bučas, Martynas · et al.
37 files · 1.6 GB · hdf5, tiff, torchdeclared
This record contains the labeled data, trained models, and analysis code supporting the article "Comparing a Vision Foundation Model (DINOv3) and a Task-Specific U-Net for Mapping Emergent Aquatic Vegetation from Fused UAV Multispectral and LiDAR Data" (Remote Sensing in Ecology and Conservation). Contents: - masks/ : georeferenced ground-truth segmentation masks (five classes: aquatic vegetation, water, sand, other objects, background), aligned to the fused UAV orthomosaics and spanning 13 sites across nine Lithuanian waterbodies surveyed between May and August 2024. - models/ : the two final trained segmentation models, a Keras/HDF5 U-Net and a PyTorch DINOv3 model. - code/ : Python scripts for training, evaluation, the label-efficiency experiment, and full-scene prediction. The fused 9-band orthomosaics (five-band multispectral, RGB, and a LiDAR canopy height model; approximately 62 GB) are archived separately because of their size and are available from the corresponding author on request. The DINOv3 SAT-493M pretrained backbone is distributed by Meta under its own license and is not redistributed here; obtain it from the official DINOv3 release.
Abbas, Syed Hassan
1 files · 86 MB · rardeclared
Abbas H, et al. "GenixRL- " The VUS were extracted from CLinVar database (downloaded [Date: August 2025]) and scored using the GenixRL framework. This dataset provides the foundation for the VUS reclassification analysis presented in the main manuscript. The dataset is provided as a single compressed CSV file: GenixRL_VUS_Scored.csv.gz Key Coulmn Descriptions: [Variant Identifier Columns, e.g., CHROM, POS, REF, ALT]: Standard genomic coordinates for each variant. SYMBOL: The official gene symbol. GenixRL_Score: The continuous pathogenicity score generated by GenixRL, ranging from 0 (most likely benign) to 1 (most likely pathogenic). GenixRL_Classification: The tiered classification based on the manuscript's thresholds: 'Likely Benign': Score < 0.53 'Likely Pathogenic': Score >= 0.53 and < 0.709 'High-Confidence Pathogenic': Score >= 0.709 [Other relevant columns]: The file also includes intermediate scores from other predictors and allele frequencies used for validation.
Cao, Shiyuan · Li, Yaning
8 files · 18 MB · csv, rardeclared
This repository provides the supporting materials for the manuscript "A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion." The archived materials include evaluation summaries, final tracking outputs, environment records, dataset mapping files, protocol reproduction materials, intermediate summary files, and revision evidence used to support the reported SportsMOT validation results. The reported formal results were obtained from the real detector, real frame-level semantic feature extraction, confidence partitioning, cascaded association, and TrackEval evaluation pipeline. The original public benchmark datasets, including SportsMOT, TeamTrack, SoccerNet Tracking, MOT17, and MOT20, are not redistributed in this repository and should be obtained from their original providers.
Nie, Yong · Huang, Bo
1 files · 1.5 KB · rardeclared
BI tree