guo, na · pan, jiachen · li, tiantian · et al.
hybrid · semantic + lexical · 80 datasets ranked · 1.02s
guo, na · pan, jiachen · li, tiantian · et al.
ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR
Xu, Tao
3 files · 100 MB · rar, xlsx
The deepmd_data dataset comprises energy data for 375,781 structures and force data for more than 45 million atoms, generated from six hierarchical active learning iterations. The init directory stores the non-periodic molecular configurations established at the initialization stage. The iter directories contain structures labeled by both AIMD and DPMD-FP methods for each iteration. Each folder is named according to the scheme of solvent molecule index followed by lithium salt designation. The finetuned_data folder contains SEI reaction simulation data used for fine-tuning the MLFFs. Comprehensive details regarding the solvent molecules are available in the solvent_with_smiles.xlsx file.
Yuezong Wang · Yu Niu · Jiqiang Chen
2 files · 100 MB · rar, zip
This dataset contains raw experimental images, video sequences of real test samples, simulated microscopic image data from the manuscript " A robust method for microscopic 3D shape restoration via shape-from-focus" , as well as focal volume datasets collected before vibration simulation, after vibration simulation, and post anti-vibration processing. All provided data enable the validation of conclusions and reproducibility of experimental results obtained via the computational pipeline proposed in this paper.
Kovaliov, Michael
1 files · 100 MB · parquet
Omukuti, Rodney
1 files · 8.0 MB · vcf
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Daugavietis, Jānis
6 files · 3.3 GB · jpeg, rardeclared
Skala Bünte, Süki Vagonu Hallē [2026.07.08.] Koncerta beigu telefona foto/ video. SKALA BÜNTE X SÜKI FACE2FACE https://fb.me/e/7dycfN2gI Details Event by John Dow Vagonu Hall Public · Anyone on or off Facebook 8TH OF JULY VAGONU HALLE TWO BANDS TWO BACKLINES MOSHPIT IN THE MIDDLE ONCE IN A LIFETIME FACE2FACE MASSACRE PROVIDED BY SKALA BÜNTE & SÜKI 🔪🔪🔪 DOORS 19:00 7€
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Geisler, Jan · Rakhimberdiev, Eldar · Boom, Michiel P. · et al.
29 files · 124 MB · csv, shapefiledeclared
1. Many migratory birds now reach their Arctic breeding grounds earlier in order to keep pace with advancing springs and shifting nutrient peaks, either by departing earlier from non-breeding grounds or by travelling faster. For dark-bellied brent geese, there is limited potential to travel faster, as their migration to the Siberian breeding grounds is already among the fastest of Arctic geese and swans. Earlier departure would require reaching departure body mass earlier, either through a faster accumulation of energy stores during spring staging or via adjustments earlier in the annual cycle. 2. We examined long-term shifts in spring staging phenology and changes in winter and spring body mass trajectories of brent geese at the population level, with particular emphasis on the effects of winter temperature on body mass and spring body mass on departure timing. 3. We used more than five decades of body mass measurements from individuals caught in the United Kingdom and France, and in the Dutch Wadden Sea to reconstruct changes in spring and winter mass trajectories, respectively. These data were combined with over five decades of migration counts in the Netherlands and more than two decades of counts in Denmark to quantify changes in spring staging phenology. 4. We found that brent geese have not shifted their spring arrival in the Wadden Sea but have advanced departure timing. Furthermore, brent geese were heavier during and after milder winters, and have changed mass trajectories over recent decades. They no longer lose mass during winter and the second spring staging phase, and fuelling rates in the first spring staging phase have declined. Annual variation in body mass was not related to annual departure timing. 5. These results suggest that milder winters have relaxed energetic constraints and improved body condition in brent geese throughout the non-breeding season. Our findings highlight the importance of considering the full annual cycle when assessing how animals with limited capacity to adjust migration timing or speed respond to global change.
Murilo Coelho · de Sousa Amâncio, Francisco Dione · Paixao, Matheus · et al.
3 files · 742 KB · rardeclared
This repository contains the replication package for the paper "More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions", accepted at the 40th Brazilian Symposium on Software Engineering (SBES 2026), São Paulo, Brazil. The package includes: (i) the complete survey instrument and sanitized participant responses; (ii) qualitative coding artifacts, including the codebook, category consolidation, and classification analysis; (iii) inter-rater reliability materials (Cohen's Kappa = 0.81); (iv) quantitative datasets and statistical analysis outputs (SPACE composite scores, Cronbach's alpha, Kruskal-Wallis, Mann-Whitney U, and Dunn's post-hoc tests); and (v) supporting literature review material. All participant data were anonymized prior to disclosure. The survey materials are in Portuguese, the language of data collection.
Abbas, Syed Hassan
1 files · 86 MB · rardeclared
Abbas H, et al. "GenixRL- " The VUS were extracted from CLinVar database (downloaded [Date: August 2025]) and scored using the GenixRL framework. This dataset provides the foundation for the VUS reclassification analysis presented in the main manuscript. The dataset is provided as a single compressed CSV file: GenixRL_VUS_Scored.csv.gz Key Coulmn Descriptions: [Variant Identifier Columns, e.g., CHROM, POS, REF, ALT]: Standard genomic coordinates for each variant. SYMBOL: The official gene symbol. GenixRL_Score: The continuous pathogenicity score generated by GenixRL, ranging from 0 (most likely benign) to 1 (most likely pathogenic). GenixRL_Classification: The tiered classification based on the manuscript's thresholds: 'Likely Benign': Score < 0.53 'Likely Pathogenic': Score >= 0.53 and < 0.709 'High-Confidence Pathogenic': Score >= 0.709 [Other relevant columns]: The file also includes intermediate scores from other predictors and allele frequencies used for validation.
Cao, Shiyuan · Li, Yaning
8 files · 18 MB · csv, rardeclared
This repository provides the supporting materials for the manuscript "A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion." The archived materials include evaluation summaries, final tracking outputs, environment records, dataset mapping files, protocol reproduction materials, intermediate summary files, and revision evidence used to support the reported SportsMOT validation results. The reported formal results were obtained from the real detector, real frame-level semantic feature extraction, confidence partitioning, cascaded association, and TrackEval evaluation pipeline. The original public benchmark datasets, including SportsMOT, TeamTrack, SoccerNet Tracking, MOT17, and MOT20, are not redistributed in this repository and should be obtained from their original providers.
Nie, Yong · Huang, Bo
1 files · 1.5 KB · rardeclared
BI tree
Nie, Yong · Huang, Bo
1 files · 9.1 KB · rardeclared
Alignments for phylogen
Kavanagh, Jack · Anthony, Patrick
6 files · 11 MB · geojson, geopackage, shapefiledeclared
A historical map showing a segment of the boundaries of New Spain in c. 1800. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Sarma, R N
2 files · 88 MB · csv, vcfdeclared
This dataset contains the genotype and phenotype data generated for a genome-wide association study (GWAS) of agronomic traits in a diverse panel of 105 Indian mungbean ( Vigna radiata L. Wilczek) accessions . The dataset comprises a filtered genome-wide SNP dataset in Variant Call Format (VCF) and the corresponding Best Linear Unbiased Predictor (BLUP) values for the measured agronomic traits.
Njie, Adama · Torkayesh, Ali E · Venghaus, Prof. Dr. Sandra
10 files · 1.6 GB · csv, gzip, parquetdeclared
Structured, speaker-attributed corpus of all German Bundestag plenary session transcripts ( Plenarprotokolle ) from the first legislative period to the present (WP01-WP21, September 1949 - April 2026). Every attributed speech is extracted from the official PDFs published by the Deutscher Bundestag under open data policy and linked to the speaker's name, parliamentary role, party affiliation, and gender. Scale: 4,611 sessions · 1,033,723 speeches · 4,205 identified MdBs · 76 years of parliamentary debate Dataset files speeches.parquet - one row per attributed speech: speaker name, role, party, gender, stammdaten_id, full German text (~1 GB) persons.parquet - one row per MdB: cross-session identity linking all name variants via stammdaten_id; canonical name, birth date, career span, total speeches. Use this - not speakers.parquet - for person-level analysis sessions.parquet - one row per plenary session: date, city, Wahlperiode, source PDF hash, extraction engine speakers.parquet - name-string index: one row per unique name string as extracted from the transcripts. Useful for understanding extraction quality; not suitable for person-level aggregation (the same politician often appears under several name variants across sessions) parties.csv - reference table of 31 German parliamentary parties, 1949-present speeches.csv.gz - CSV fallback for Stata and Excel users (same columns as speeches.parquet) datapackage.json - Frictionless Data schema with column descriptions and foreign key constraints Cross-session identity The same politician often appears under different name strings across sessions (e.g. "Schmidt", "Dr. Schmidt", "Frau Dr. Schmidt"). Cross-session person linkage is provided via stammdaten_id , matched against the official Bundestag Stammdaten biographical XML. The persons.parquet table aggregates all name variants for the same MdB into one row with correctly summed speech counts, career span, and birth date. Coverage: ~98.5% of speeches are linked to a stammdaten_id; the remaining ~1.5% are ambiguous surname-only attributions or speakers not in the Stammdaten. Coverage and sources Source PDFs are the official Stenografische Berichte downloaded from the Bundestag open-data portal (bundestag.de). Party-share normalisation in the corpus statistics uses official seat counts per Wahlperiode sourced from the Federal Returning Officer (Bundeswahlleiter, bundeswahlleiter.de). Two PDF generations are covered: scanned and OCR'd documents (WP01-WP09, Bonn era, 1949-1987) and born-digital documents (WP10-WP21, 1987-present). The engine column in sessions.parquet flags whether pdftotext (born-digital) or pdfminer (OCR fallback) was used; this is the primary data-quality indicator for NLP use. Speaker attribution Each speech is attributed using four patterns extracted from the transcript format: presiding officers (Präsident/in, Vizepräsident/in), regular members (name + party), government officials (name + Bundeskanzler/in, Bundesminister/in, etc.), and procedural roles (Berichterstatter/in, etc.). The party field is null for ~60% of speeches - this is expected, as presiding officers and ministers are not identified by party in the transcript. Gender annotation & distribution Gender is derived by matching speaker names against the official Bundestag Stammdaten biographical XML (all MdBs since 1949), with fallbacks for role title, honorific prefix, manually researched overrides, and a gender_guesser first-name heuristic. The gender_source column distinguishes stammdaten (authoritative, 83%), role_title (gendered job title in attribution, 6.4%), title_prefix (Frau/Herr honorific, 0.5%), manual (historically researched, 2.9%), and inferred (name-based heuristic, 4.7%). Gender distribution: Female 26.6% · Male 73.4% · Unknown 0.0%. Data quality All speeches pass automated validation: zero null speaker names, zero sequence gaps, zero CID artefacts, zero party-misclassified-as-Bundesland errors. Eight sessions with conflicting source PDFs were deduplicated (first lexicographic occurrence retained). 252 non-person names incorrectly accepted by the parser (table headers, legislative terms, agenda fragments) are excluded at build time via a curated exclusion list. OCR sessions (WP01-WP09) may contain Unicode replacement characters (U+FFFD); the engine field identifies these sessions. Licence CC BY 4.0. The underlying Plenarprotokolle are official government documents of the Deutscher Bundestag and are in the public domain.
abdulwahab, samaa · aduallah, mahmood z. · Sallomi, Adheed H.
34 files · 2.7 GB · csv, gzip, parquetdeclared
Intrusion-detection research on Internet Protocol version 6 (IPv6) remains bottlenecked by the scarcity of labelled, protocol-aware flow datasets. Existing machine-learning IDS benchmarks are overwhelmingly IPv4-centric, and the few IPv6 corpora that have been released target narrow attack families or rely on small academic testbeds that cannot be re-created by third parties. We present IPv6-CyberBench, a reproducible eight-phase pipeline that constructs a large, protocol-aware translated-flow corpus by harmonising CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019, applying deterministic IPv4→IPv6 address translation (6to4, NAT64, Teredo, EUI-64), synthesising 27 IPv6-specific flow features grouped in six protocol families, enforcing nineteen RFC-derived constraint categories together with temporal address dynamics, and rebalancing the long-tailed class distribution with a feature-group-conditioned per-class Wasserstein-GAN-GP augmenter and a SMOTE-KDE fallback selected per class by a formal decision rule. We scope the contribution honestly: because the seed corpora are IPv4 captures, the resulting 2,285,774-record benchmark is a translated-flow corpus suitable for training and evaluating flow-level IPv6 IDS classifiers on flooding, brute-force, scan, web-attack and infiltration traffic under IPv6 protocol-header semantics, and for studying IPv6-specific feature engineering and address dynamics in a reproducible setting. It is not a substitute for protocol-native IPv6 attack capture, and we explicitly exclude ICMPv6 Neighbour-Discovery flooding, SEND flooding, NDP exhaustion and extension-header covert-tunnelling from the threat model. The benchmark is evaluated on four axes - fidelity (Kolmogorov-Smirnov, MMD, Fréchet feature distance), utility (stratified 5×5 nested cross-validation over six classifier families including CNN-LSTM and LightGBM), privacy (Shokri-style membership-inference advantage AUC), and external fidelity against a 24 h anonymised CAIDA IPv6 trace (equinix-chicago, US backbone) and a MAWI samplepoint-F trace (WIDE backbone, Tokyo, Japan). The full pipeline, the hyper-parameter manifest, the RFC-constraint manifest, the reproduction scripts, and the 2,285,774-record benchmark are released unconditionally on Zenodo under CC BY 4.0; the dataset and pipeline are openly available at https://doi.org/10.5281/zenodo.19503446 (CC BY 4.0).
Kavanagh, Jack · Anthony, Patrick
6 files · 9.3 MB · geojson, geopackage, shapefiledeclared
A historical map of the boundaries of Prussia in c. 1795. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Costes, Jean C.
28 files · 8.0 GB · fitsdeclared
Stellar activity is the main barrier to detecting and/or confirming low-mass/long-period (and Earth-analogue) planets using radial-velocity (RV) measurements. Searching forreliable indicatorsthat bettertrace magnetic activity may be key for both distinguishing more clearly between stellar and planetary signals, and for probing the underlying physics occurring on the stellar surface. In this work we have compared observations taken for magnetically active and inactive stellar phases over multiple time-scales to study the spectral imprint due to varying stellar activity. This serves as a proof-of-concept demonstration of a technique (named Spectral Ratio Analysis, SRA) that can be used to isolate activity-driven changes directly in the stellar photospheric absorption lines where RVs are measured. Using 14 relatively quiet and well sampled G- and K-type starsthatshow stellar activity cycles, we identified hundreds of activity-sensitive spectral features. Reducing this variability information into two global metrics - amplitude and velocity shift - uncovers potential evidence of a decoupling of the photospheric and chromospheric responsesto stellar activity in earlier-type stars. Additionally, potential signatures of the variations in the magnitude of the suppression of the convective blueshift throughout the activity cycle were observed via SRA. Finally, we show that these SRA indicators better capture RV variability than classical activity proxies, such as the chromospheric log R HK index and other cross-correlation function-based parameters such as bisector span and full width at half-maximum, by up to a factor of two. The direct link between photospheric line behaviour and stellar-induced RV variability offers a promising path for improving astrophysical noise mitigation. In these files you will find for each star a .fits file in which the seasonal and nightly difference spectra and its associated SDATemplate are given per order. Their respective errors have been added as well. Finally, the wavelength for each order, as well as the seasonal and nightly bjd of each observation and its measured logRHK are availbale in these files. Additionally, a .txt file have been added for each star showing the atlas of all the SDA detected features. The wavelength of each feature, its amplitude, its width and its potentially associated lines obtained from VALD are given in these files.