hybrid · semantic + lexical · 93 datasets ranked · 2.94s
guo, na · pan, jiachen · li, tiantian · et al.
14 files · 100 MB · csv, rar, tsv
ADM_LSIR is a physics-inspired laparoscopic aerosol degradation dataset for aerosol-aware surgical image analysis and image restoration. The v1.0.0 release contains: - 21,916 clean clinical laparoscopic frames (clean/) - 9,562 real intraoperative aerosol-degraded frames (degraded/) - 36,052 simulated aerosol masks, including 19,701 smoke-like masks and 16,351 trajectory masks (mask/) - Blender simulation/cache materials (ADM_LSIR_Blender_simulation_files_v1.0.rar) - metadata_quality_report_v1.0.csv - recommended_splits_v1.0.csv - video_mapping_v1.0.csv - parts_manifest.txt - checksums_v1.0.tsv - release_manifest_v1.0.json All released clinical frames are de-identified and stored as lossless PNG files. Filenames use anonymized video identifiers, e.g., C-V##-####.png for clean frames and D-V##-####.png for degraded frames. The recommended split is defined at the source_video_id/public_video_label level to reduce leakage across frames from the same source video. The public video labels in video_mapping_v1.0.csv provide privacy-safe source-video identifiers (video1-video19). The Blender archive documents the smoke and trajectory mask simulation setup and supports reuse, but it is not a guaranteed exact per-mask reproduction package. The released pre-rendered mask library is the primary reusable dataset component. Source code for synthesis and quality screening is available at: https://github.com/SweetDeathh/ADM_LSIR
Xu, Tao
3 files · 100 MB · rar, xlsx
The deepmd_data dataset comprises energy data for 375,781 structures and force data for more than 45 million atoms, generated from six hierarchical active learning iterations. The init directory stores the non-periodic molecular configurations established at the initialization stage. The iter directories contain structures labeled by both AIMD and DPMD-FP methods for each iteration. Each folder is named according to the scheme of solvent molecule index followed by lithium salt designation. The finetuned_data folder contains SEI reaction simulation data used for fine-tuning the MLFFs. Comprehensive details regarding the solvent molecules are available in the solvent_with_smiles.xlsx file.
Yuezong Wang · Yu Niu · Jiqiang Chen
2 files · 100 MB · rar, zip
This dataset contains raw experimental images, video sequences of real test samples, simulated microscopic image data from the manuscript " A robust method for microscopic 3D shape restoration via shape-from-focus" , as well as focal volume datasets collected before vibration simulation, after vibration simulation, and post anti-vibration processing. All provided data enable the validation of conclusions and reproducibility of experimental results obtained via the computational pipeline proposed in this paper.
Kovaliov, Michael
1 files · 100 MB · parquet
Madhusudan, Gujral
6 files · 29 MB · parquetdeclared
Large language models (LLMs) are trained on massive, publicly available text datasets comprising trillions of tokens, enabling them to excel at general language tasks like next-token prediction. However, LLMs often struggle with domain-specific prompts, exhibiting reduced accuracy or generating inaccurate information (hallucinations). This is because they lack sufficient subject matter expertise. Two primary approaches exist to address this limitation for augmenting LLMs knowledge: Retrieval-Augmented Generation (RAG) and fine-tuning. This presentation focuses on fine-tuning smaller LLMs with domain-specific instruct datasets using the LoRA (Low-Rank Adaptation) technique on Gaudi hardware. We will leverage publicly available LLMs and datasets from the Hugging Face Hub for this demonstration. Though it is possible to fine tune LLMs with plain text data - sourced from documents, articles, and other materials.
Austin, Timothy · do Valle Chagas Azaneu, Marina · Roughan, Moninya
26 files · 29 MB · netcdf, pdf, pngdeclared
Data collected from a temperature mooring at Lord Howe Island maintained by UNSW Sydney and funded by Parks Australia. The mooring position is longitude = 158.97°E and latitude = -31.51°, and local depth of approximately 52 m. The data were sampled using a series of thermistors (aqualogger 520PTs) deployed on a mooring line at 4m intervals through the water column, with shallowest instrument at 13 m and deepest at 53 m. The time period spans between 14-05-2025 and 22-04-2026. IMOS standard data quality assurance and quality control processes have been followed and the data formatted following IMOS conventions. Data quality control includes automated routines and visual inspection (expert QC) and flagging of obvious errors. File are c.f. compliant NetCDF files, and file name format follows IMOS conventions and includes sampling period in the format: UNSW_Lord_Howe_Marine_Park_TZ_ yyyymmddThhmmss Z_LH050_FV01_ LH050-2511-Aqualogger-AQUAlogger-520PT16-max160m-13_END- yyyymmddThhmmssZ.
Wang, Chang · Lu, Xingcheng
2 files · 222 MB · netcdfdeclared
Machine-learning-inferred monthly anthropogenic NOx emission over the 2026 Strait-of-Hormuz disruption (global, 0.1 degree, January 2025 - May 2026). This dataset is the top-down NOx emission product underlying the companion manuscript on the 2026 Strait-of-Hormuz shipping-emission collapse. A LightGBM estimator trained on the CAMS-GLOB-ANT v6.2 inventory (2018-2024), with the observed TROPOMI NO2 column and GEOS-CF chemistry/meteorology as predictors, is applied month by month to 2025-01 through 2026-05 to infer the anthropogenic NOx emission flux. Provided as a single self-describing CF-1.8 NetCDF containing: the total anthropogenic NOx flux (kg m-2 s-1, reported as NO) and a per-pixel cross-validation uncertainty (log-space). The estimator resolves the total emission only; no sector decomposition is distributed, because over open ocean the total is essentially ship emission while on land a sector split would only re-apply the CAMS-GLOB-ANT prior shares and is not constrained by the observations. Coverage is land and ocean within +/-60 degrees latitude on a regular 0.1-degree global grid. All units and coordinates are embedded in the file.
Daugavietis, Jānis
6 files · 3.3 GB · jpeg, rardeclared
Skala Bünte, Süki Vagonu Hallē [2026.07.08.] Koncerta beigu telefona foto/ video. SKALA BÜNTE X SÜKI FACE2FACE https://fb.me/e/7dycfN2gI Details Event by John Dow Vagonu Hall Public · Anyone on or off Facebook 8TH OF JULY VAGONU HALLE TWO BANDS TWO BACKLINES MOSHPIT IN THE MIDDLE ONCE IN A LIFETIME FACE2FACE MASSACRE PROVIDED BY SKALA BÜNTE & SÜKI 🔪🔪🔪 DOORS 19:00 7€
Hackl, Jürgen
6 files · 132 MB · parquet, tiffdeclared
EuroFlood is an open, cloud-native index over the JRC/Copernicus CEMS-EFAS Satellite-Derived Flood Depth Maps for Europe (Betterle & Salamon, 2025; CC-BY-4.0) - ~3,280 satellite-derived observed flood-depth maps across Europe, 2015-2024. The bundle is a sparse Cloud-Optimized GeoTIFF encoding, per pixel, the set of flood events that inundated it, plus a combo_id -sorted GeoParquet dictionary and a small events table. Query by region and time via HTTP range reads (GDAL /vsicurl + DuckDB) to retrieve matching events, then fetch only the source depth rasters needed. Built with the open-source EuroFlood Python package ( pip install euroflood ).
Vera, Lourdes
31 files · 4.6 MB · csv, geojson, pngdeclared
Data, analysis scripts, and derived outputs reproducing the setback analysis between occupied buildings and active oil and gas wells in Karnes County, Texas, using public data (Texas Railroad Commission well locations; FEMA/ORNL USA Structures building footprints; and U.S. Census TIGER/Line block groups and 2020-2024 American Community Survey). The pipeline runs offline in Python (geopandas) and reproduces every reported distance, summary statistic, table, and figure in the associated article. This package reproduces and updates Chapter 3 of the author's doctoral dissertation: Vera, Lourdes (2022), "Environmental Data Justice in Action: Civically Valid Air Monitoring Near Oil and Gas Extraction in the Eagle Ford Shale Play," Ph.D. dissertation, Northeastern University, Boston, MA.
Miszczyszyn, Jakub · Radoń, Radosław
20 files · 2.1 GB · csv, netcdfdeclared
Monthly 1-km SPI, SPEI and SRI drought-indicator rasters for Poland (1995-2024) This dataset provides monthly gridded drought indicators for the entire territory of Poland at 1-km resolution for the period 1995-2024. Three complementary standardized indices are included: the Standardized Precipitation Index (SPI, precipitation-based), the Standardized Precipitation-Evapotranspiration Index (SPEI, based on the climatic water balance P - PET) and the Standardized Runoff Index (SRI, runoff-based). Each index is provided at five accumulation scales: 1, 3, 6, 12 and 24 months. Methods. Station indices were computed from IMGW-PIB precipitation and river-runoff observations and from AgERA5 potential evapotranspiration (used for SPEI). SPI and SRI were fitted with a gamma distribution and SPEI with a log-logistic distribution. Station values were then interpolated to a 1-km grid (EPSG:2180, PUWG 1992) by ordinary kriging. Elevation-assisted kriging (kriging with external drift) was evaluated by leave-one-out cross-validation but produced no net national improvement and was not adopted for the final product. File structure. Data are provided as CF-compliant NetCDF (CF-1.8), one file per indicator and accumulation scale (e.g. SPI_s06m_PL_1km_1995-2024_EPSG2180.nc ) . Each file has dimensions (time, y, x) with a regular monthly time axis (360 steps, 1995-01 to 2024-12) and the coordinate reference system stored as a grid_mapping variable. Months in which an index cannot be formed at the edges of the accumulation window are present as fill-valued (missing) layers, so the time axis is continuous; their completeness is documented in raster_index.csv . Contents. netcdf/ - the 15 index rasters; metadata/raster_index.csv - per-layer inventory with min/max/mean, missing-data fraction and a "present" flag; metadata/variables_dictionary.csv - variable definitions; metadata/mckee_classes.csv - the seven-class drought/wetness classification (McKee et al., 1993) with colours; CITATION.cff and checksums.md5 . Usage note. In GIS software the temporal dimension is read via the temporal/time controls (e.g. in QGIS: Layer Properties → Temporal → Dynamic Temporal Control, then the Temporal Controller); in Python the files open directly with xarray, with time parsed as dates. Coordinate reference system. EPSG:2180 (PUWG 1992 / Poland CS92). Map extents delineate the study area and do not necessarily depict accepted national boundaries.
Wyatt, James · Larsen, Karin Margretha H.
2 files · 21 MB · netcdfdeclared
A monthly dataset of volume, heat and salt transport through the Faroe-Shetland Channel. The time step is monthly, from 1993-01 to 2025-08. This data accompanies the paper 'Strengthening of the Atlantic Water Inflow through the Faroe-Shetland Channel'.
Geisler, Jan · Rakhimberdiev, Eldar · Boom, Michiel P. · et al.
29 files · 124 MB · csv, shapefiledeclared
1. Many migratory birds now reach their Arctic breeding grounds earlier in order to keep pace with advancing springs and shifting nutrient peaks, either by departing earlier from non-breeding grounds or by travelling faster. For dark-bellied brent geese, there is limited potential to travel faster, as their migration to the Siberian breeding grounds is already among the fastest of Arctic geese and swans. Earlier departure would require reaching departure body mass earlier, either through a faster accumulation of energy stores during spring staging or via adjustments earlier in the annual cycle. 2. We examined long-term shifts in spring staging phenology and changes in winter and spring body mass trajectories of brent geese at the population level, with particular emphasis on the effects of winter temperature on body mass and spring body mass on departure timing. 3. We used more than five decades of body mass measurements from individuals caught in the United Kingdom and France, and in the Dutch Wadden Sea to reconstruct changes in spring and winter mass trajectories, respectively. These data were combined with over five decades of migration counts in the Netherlands and more than two decades of counts in Denmark to quantify changes in spring staging phenology. 4. We found that brent geese have not shifted their spring arrival in the Wadden Sea but have advanced departure timing. Furthermore, brent geese were heavier during and after milder winters, and have changed mass trajectories over recent decades. They no longer lose mass during winter and the second spring staging phase, and fuelling rates in the first spring staging phase have declined. Annual variation in body mass was not related to annual departure timing. 5. These results suggest that milder winters have relaxed energetic constraints and improved body condition in brent geese throughout the non-breeding season. Our findings highlight the importance of considering the full annual cycle when assessing how animals with limited capacity to adjust migration timing or speed respond to global change.
Murilo Coelho · de Sousa Amâncio, Francisco Dione · Paixao, Matheus · et al.
3 files · 742 KB · rardeclared
This repository contains the replication package for the paper "More Productive, but at What Cost? Understanding How GenAI Shapes Developers' Work Across the SPACE Dimensions", accepted at the 40th Brazilian Symposium on Software Engineering (SBES 2026), São Paulo, Brazil. The package includes: (i) the complete survey instrument and sanitized participant responses; (ii) qualitative coding artifacts, including the codebook, category consolidation, and classification analysis; (iii) inter-rater reliability materials (Cohen's Kappa = 0.81); (iv) quantitative datasets and statistical analysis outputs (SPACE composite scores, Cronbach's alpha, Kruskal-Wallis, Mann-Whitney U, and Dunn's post-hoc tests); and (v) supporting literature review material. All participant data were anonymized prior to disclosure. The survey materials are in Portuguese, the language of data collection.
Abbas, Syed Hassan
1 files · 86 MB · rardeclared
Abbas H, et al. "GenixRL- " The VUS were extracted from CLinVar database (downloaded [Date: August 2025]) and scored using the GenixRL framework. This dataset provides the foundation for the VUS reclassification analysis presented in the main manuscript. The dataset is provided as a single compressed CSV file: GenixRL_VUS_Scored.csv.gz Key Coulmn Descriptions: [Variant Identifier Columns, e.g., CHROM, POS, REF, ALT]: Standard genomic coordinates for each variant. SYMBOL: The official gene symbol. GenixRL_Score: The continuous pathogenicity score generated by GenixRL, ranging from 0 (most likely benign) to 1 (most likely pathogenic). GenixRL_Classification: The tiered classification based on the manuscript's thresholds: 'Likely Benign': Score < 0.53 'Likely Pathogenic': Score >= 0.53 and < 0.709 'High-Confidence Pathogenic': Score >= 0.709 [Other relevant columns]: The file also includes intermediate scores from other predictors and allele frequencies used for validation.
Cao, Shiyuan · Li, Yaning
8 files · 18 MB · csv, rardeclared
This repository provides the supporting materials for the manuscript "A Confidence-Guided Sports Multi-Object Tracking Method via Jersey Semantic Fusion." The archived materials include evaluation summaries, final tracking outputs, environment records, dataset mapping files, protocol reproduction materials, intermediate summary files, and revision evidence used to support the reported SportsMOT validation results. The reported formal results were obtained from the real detector, real frame-level semantic feature extraction, confidence partitioning, cascaded association, and TrackEval evaluation pipeline. The original public benchmark datasets, including SportsMOT, TeamTrack, SoccerNet Tracking, MOT17, and MOT20, are not redistributed in this repository and should be obtained from their original providers.
Nie, Yong · Huang, Bo
1 files · 1.5 KB · rardeclared
BI tree
Nie, Yong · Huang, Bo
1 files · 9.1 KB · rardeclared
Alignments for phylogen
Kavanagh, Jack · Anthony, Patrick
6 files · 11 MB · geojson, geopackage, shapefiledeclared
A historical map showing a segment of the boundaries of New Spain in c. 1800. This map is fully open to fellow researchers and is available in multiple open source formats (SHP, GPKG, GeoJSON).
Chen, Xiao · Huang, Yibin
2 files · 237 MB · netcdfdeclared
This archive contains a global monthly marine net primary productivity (NPP) product derived from the Data-based Productivity Model (DbPM), an observationally constrained machine-learning framework developed using an expanded global compilation of in situ 14 C/ 13 C-based NPP measurements and satellite-derived environmental predictors. Original DbPM NPP product - generated directly from satellite observations and therefore containing gaps associated with missing satellite data coverage. Gap-filled DbPM NPP product - reconstructed using Empirical Orthogonal Function (EOF)-based interpolation to fill missing values caused by satellite data gaps, providing a spatially complete global monthly NPP field. Chen, X., Huang, Y., Liu, H., Cassar, N., Liu, X., Wang, W., Chai, F., Kang, J., & Huang, B. (2026). Reduced estimate of global marine primary productivity and hemispheric redistribution over the satellite era revealed by an expanded global observational database. Submitted to Communications Earth & Environment .