Mansoori, Zaeem Ahmad
hybrid · semantic + lexical · 23 datasets ranked · 2.14s
Mansoori, Zaeem Ahmad
Pretrained weights for the hybrid Transformer‑Biophysical Fusion architecture of EpiRNA. Combines a frozen DNA‑BERT backbone with a cross‑scale biophysical CNN. Trained on 11,844 human m⁶A sites (GSE63753, miCLIP) and an equal number of DRACH‑containing negative windows. Achieves AUROC 0.80, specificity 0.83, sensitivity 0.61 on the held‑out test set. Intended for offline genome‑wide prediction; the live web tool uses a faster CNN variant.
Untan, Sultan
19 KB
Raw qPCR quantification summaries (per-well Cq values and run metadata) for HOXA-10, HOXA-11, HOXA-13, and ACTB, exported directly from the BioRad CFX-96 / CFX Maestro software. These are the raw amplification run records (runs of 13-14 May 2024) underlying the relative-expression (RQ) values reported in the associated study. Each target is provided as a separate sheet (Well, Fluor, Target, Content, Sample, Cq), with a README sheet describing the contents. Standard-curve efficiency values were not generated as permanent records for these runs.
Silander, Olin
8.0 MB
Garriga-Alonso, Núria · Laromaine, Anna · Alonso-Pernas, Pol · et al.
100 MB
This dataset contains Caenorhabditis elegans ( C. elegans ) RNA sequencing data, to assess the effects of in vivo EMF exposure on C. elegans gene expression. There are 27 samples: 3 different EMF exposure conditions (Exp, Sham, Neg) and 3 C. elegans generations (G1, G2, G3), and each with 3 biological replicates. There are two FASTQ files per sample, as it is data from a paired-end sequencing. This dataset contains the following files: Sham.zip: raw FASTQ files for Sham samples Neg.zip: raw FASTQ files for Neg samples Exp.zip: raw FASTQ files for Exp samples Expression_Profile.WBcel235.gene: raw and normalized gene counts Samples_Metadata: metadata of the 27 samples
Petersen, Arnkell Jonas · Thiis, Thomas
8.0 MB
This dataset includes TMY files for the 5000 largest cities and towns in the EU, based on CERRA data from 2006-2020, and using a TMY ISO methodology. The files represents a statistically typical year meant for energy calculations of buildings, and are therefore not suited for evaluating extreme conditions. The repository contains: EPW-files - zipped collection of .epw-files The most common file format for climate data for energy simulations applicable to most programs Datasheets - a zip files of individual datasheets A collection of datasheets containing a comparison of TMYs with the data source as well as referance data The datasheets contain Norwegian text Error metrics - ZIP file Error metrics mapped to a map of Norway, for visualization purposes As well as the metrics that are mapped Selected years - a comma seperated file the selected typical meteorological months for each location Changelog - a txt file a overview of changes in the dataset since 1.0 The contents of the repository is produced by The Norwegian University og Life Science (NMBU), Department of Building- and Environmental Technology. The data is provided 'as is', without warranty of any kind, express or implied. In no event shall the authors be liable for any claim, damages or other liability. An interactive map containing the TMY3 data, but not the TMY ISO, can be found at https://www.climatedataforbuildings.eu/ If you use the datasets or tools in your work, please reference the relevant article that document the methods and data: Petersen & Thiis (2025), " Continental-Scale Assessment of Typical Meteorological Years in a Changing Climate " , Energy and Buildings, DOI: 10.1016/j.enbuild.2025.116814
Petersen, Arnkell Jonas · Thiis, Thomas
8.0 MB
This dataset includes TMY files for the 5000 largest cities and towns in the EU, based on CERRA data from 1991-2020, and using a TMY ISO methodology. The files represents a statistically typical year meant for energy calculations of buildings, and are therefore not suited for evaluating extreme conditions. The repository contains: EPW-files - zipped collection of .epw-files The most common file format for climate data for energy simulations applicable to most programs Datasheets - a zip files of individual datasheets A collection of datasheets containing a comparison of TMYs with the data source as well as referance data The datasheets contain Norwegian text Error metrics - ZIP file Error metrics mapped to a map of Norway, for visualization purposes As well as the metrics that are mapped Selected years - a comma seperated file the selected typical meteorological months for each location Changelog - a txt file a overview of changes in the dataset since 1.0 The contents of the repository is produced by The Norwegian University og Life Science (NMBU), Department of Building- and Environmental Technology. The data is provided 'as is', without warranty of any kind, express or implied. In no event shall the authors be liable for any claim, damages or other liability. An interactive map containing the TMY3 data, but not the TMY ISO, can be found at https://www.climatedataforbuildings.eu/ If you use the datasets or tools in your work, please reference the relevant article that document the methods and data: Petersen & Thiis (2026), " Continental-Scale Assessment of Typical Meteorological Years in a Changing Climate " , Energy and Buildings, DOI: 10.1016/j.enbuild.2025.116814
Petersen, Arnkell Jonas · Thiis, Thomas
8.0 MB
This dataset includes TMY files for the 5000 largest cities and towns in the EU, based on CERRA data from 1991-2020, and using a TMY3 methodology. The files represents a statistically typical year meant for energy calculations of buildings, and are therefore not suited for evaluating extreme conditions. The repository contains: EPW-files - zipped collection of .epw-files The most common file format for climate data for energy simulations applicable to most programs Datasheets - a zip files of individual datasheets A collection of datasheets containing a comparison of TMYs with the data source as well as referance data The datasheets contain Norwegian text Error metrics - ZIP file Error metrics mapped to a map of Norway, for visualization purposes As well as the metrics that are mapped Selected years - a comma seperated file the selected typical meteorological months for each location Changelog - a txt file a overview of changes in the dataset since 1.0 The contents of the repository is produced by The Norwegian University og Life Science (NMBU), Department of Building- and Environmental Technology. The data is provided 'as is', without warranty of any kind, express or implied. In no event shall the authors be liable for any claim, damages or other liability. An interactive map containing the TMY3 data, but not the TMY ISO, can be found at https://www.climatedataforbuildings.eu/ If you use the datasets or tools in your work, please reference the relevant article that document the methods and data: Petersen & Thiis (2025), " Continental-Scale Assessment of Typical Meteorological Years in a Changing Climate " , Energy and Buildings, DOI: 10.1016/j.enbuild.2025.116814
Petersen, Arnkell Jonas · Thiis, Thomas
8.0 MB
This dataset includes TMY files for the 5000 largest cities and towns in the EU, based on CERRA data from 2006-2020, and using a TMY3 methodology. The files represents a statistically typical year meant for energy calculations of buildings, and are therefore not suited for evaluating extreme conditions. The repository contains: EPW-files - zipped collection of .epw-files The most common file format for climate data for energy simulations applicable to most programs Datasheets - a zip files of individual datasheets A collection of datasheets containing a comparison of TMYs with the data source as well as referance data The datasheets contain Norwegian text Error metrics - ZIP file Error metrics mapped to a map of Norway, for visualization purposes As well as the metrics that are mapped Selected years - a comma seperated file the selected typical meteorological months for each location Changelog - a txt file a overview of changes in the dataset since 1.0 The contents of the repository is produced by The Norwegian University og Life Science (NMBU), Department of Building- and Environmental Technology. The data is provided 'as is', without warranty of any kind, express or implied. In no event shall the authors be liable for any claim, damages or other liability. An interactive map containing the TMY3 data, but not the TMY ISO, can be found at https://www.climatedataforbuildings.eu/ If you use the datasets or tools in your work, please reference the relevant article that document the methods and data: Petersen & Thiis (2025), " Continental-Scale Assessment of Typical Meteorological Years in a Changing Climate " , Energy and Buildings, DOI: 10.1016/j.enbuild.2025.116814
Pérez Gallego, Ruth · von Meijenfeldt, F. A. Bastiaan · Bale, Nicole J. · et al.
2.0 MB
Abstract Paleontological and phylogenomic observations have shed light on the evolution of cyanobacteria. Nevertheless, the emergence of heterocytes, specialized cells for nitrogen fixation, remains unclear. Heterocytes are surrounded by heterocyte glycolipids (HGs), which contribute to protection of the nitrogenase enzyme from oxygen. Here, by comprehensive HG identification and screening of HG biosynthesis genes throughout cyanobacteria, we identify HG analogs produced by specific and distantly related non-heterocytous cyanobacteria. These structurally less complex molecules probably acted as precursors of HGs, suggesting that HGs arose after a genomic reorganization and expansion of ancestral biosynthetic machinery, enabling the rise of cyanobacterial heterocytes in an increasingly oxygenated atmosphere. Subsequently, HG chemical structure evolved convergently in response to environmental pressures. Our results open a new chapter in the potential use of diagenetic products of HGs and HG analogs as fossils for reconstructing the evolution of multicellularity and division of labor in cyanobacteria. Here we supply: Supplementary Data 1. Selected cyanobacterial genomes from the PATRIC genome database (now part of the BV-BRC database). Files called 'selected_Cyanogenomes.genome_*.20220430.txt' are sourced from the PATRIC File Transfer Protocol server (ftp.patricbrc.org). 'gtdbtk.bac120.summary.tsv' is the GTDB-Tk output file, and 'qa.summary_extended.txt' the CheckM output file. Supplementary Data 2. HG biosynthetic gene clusters in selected PATRIC genomes and 14 newly sequenced genomes. The file 'islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt' contains the location of all hits to Anabaena sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: "genome | contig". The structure of a hit is as follows: "ORF number on contig | query ( e -value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]". Non-overlapping hits on the same ORF (see Online Methods) are connected with '&&&' characters. An asterisk ('*') indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with '~~~' characters. The file 'Supplementary_table.script_1.txt' contains a summary of all identified hgl islands (i.e. clusters containing at least seven unique HG biosynthesis gene hits). Supplementary Data 3. HG biosynthetic gene clusters in 255,388 prokaryotic genomes from the PATRIC genome database (now part of the BV-BRC database). The file called 'PATRIC_20230120.selection_c50_c10.txt' contains information on the selected PATRIC genomes based on data sourced from the PATRIC File Transfer Protocol server ( ftp.patricbrc.org ). The file 'all_tree_of_life_genomes.islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt' contains the location of all hits to Anabaena sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: "genome | contig". The structure of a hit is as follows: "ORF number on contig | query ( e -value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]". Non-overlapping hits on the same ORF (see Online Methods) are connected with '&&&' characters. An asterisk ('*') indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with '~~~' characters. Supplementary Data 4. Phylogeny of representative cyanobacterial genomes based on a core gene superalignment. The folder contains the files used to generate Fig. 2a. The directory 'IQ-TREE' contains the tree file and iTOL annotation files. The file 'dRep.representative_to_cluster.txt' contains the dRep clusters. Note that the manually defined subclades in the iTOL annotation file 'iTOL_annotation.manually_defined_clades.DATASET_STYLE.txt' have a different numbering from the paper: subclades 0 and 1 are the 'heterocytous sister clades', and subclades 2-10 in the annotation file are heterocytous subclades 1-9 in the paper, respectively. Supplementary Data 5. Lipid data files. The folder contains all the UHPLC-HRMS n (Orbitrap) datafiles used in this study. The directory 'CCY strains' includes 24 heterocytous cyanobacterial cultures corresponding to 23 strains grown in nitrogen-deficient media, the resulting data are shown in Supplementary Table 10. Directory 'HglT mutant' contains the datafiles used to generate Supplementary Table 15. The directory 'LEGE strains' includes the UHPLC-HRMS n (Orbitrap) and GC-MS datafiles corresponding to eight cultures of two non-heterocytous strains grown in media with and without nitrogen for 38 to 77 days, the resulting data are shown in Supplementary Tables 10, 17 and 18. Supplementary Data 6. Plasmid maps. GenBank and FASTA files of plasmids generated in this study. 'HglT deletion' directory contains the genomic region surrounding hglT in the wild-type strain and after deletion used to generate Supplementary Fig. 15. pAM5404 is shown in Supplementary Fig. 16 and p(A)RP0XX are shown in Supplementary Fig. 17. Supplementary Data 7. Phylogenies of seven hgl island genes and of a concatenated alignment of these genes. The folder contains the files used to generate Supplementary Fig. 18 (in the directory 'gene_trees_hgl_islands'), and Fig. 4 and related figures (in the directory 'gene_trees_hgl_islands_4'. The directories contain the alignments and trimmed alignments, IQ-TREE output files, and iTOL annotation files. The file 'gene_trees_hgl_islands/analysis_individual_gene_trees/explore_individual_gene_clusters.ipynb' contains the code to identify the five hgl islands that contain genes with incongruent evolutionary histories. Supplementary Data 8. Phylogeny of hglE A homologs. The folder contains the files used to generate Supplementary Fig. 24 and related figures. The file 'selected_hglE_hits.txt' contains the selected hglE A hits and the genomic cluster on which they are located. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files. All the code used in this publication including scripts used for: genome assemblies, download of genomes from public repositories, quality and contamination checks, genome analysis, construction of the phylogenetic trees, hgl island identification, etc. The shell script 'commands.sh' within each directory contains all the code used to generate the content in the directory. All the figures used in this publication including the figures in the Supplementary Information file.
Perez Gallego, Ruth · von Meijenfeldt, F. A. Bastiaan · Besseling, Marc · et al.
8.0 MB
Abstract Background Gephyrocapsa huxleyi is a coccolithophoric haptophyte widespread across most marine ecosystems that plays an important role in the carbon cycle. G. huxleyi is one of the few haptophytes known to produce long-chain alkenones (LCAs), which are lipids composed of C 35 -C 42 n -alkyl chains, with two to four trans-double bonds and a keto group at either the 2 nd or 3 rd carbon positions, which are widely used as proxies for paleotemperature reconstruction. Despite the biomarker value of LCAs, little is known about their biosynthesis, which severely limits the understanding of the regulation of their production under different conditions, and thus their predictive nature. Differences in gene expression under different conditions may help to unravel the LCA biosynthetic pathway, but require annotated reference genomes. Although a large number of G. huxleyi strains inhabiting diverse ecosystems exist, only one such genome (i.e. of strain CCMP1516) is currently available. Results To evaluate differences in LCA biosynthesis under changing conditions, we developed a tailored approach to decontaminate and assemble the genomes of two additional non-axenic G. huxleyi strains (strains CCMP1742 and CCMP2758) that produce LCAs with two to four double bonds, one of which (CCMP2758) also produces unusually short (C 35 -C 36 ) LCAs. We identified a polyketide synthase (PKS) in strain CCMP1516, which we propose as a candidate to carry out LCA formation based on the structure and composition of its modules. Additionally, we identified several CCMP1742 and CCMP2758 PKSs that were differentially expressed between temperatures and across time that closely resemble this PKS and are likely to be involved in LCA formation. Based on differences in gene expression between two different growth temperatures, we also identified a limited number of desaturases which are potentially responsible for the addition of the third and fourth double bonds in the LCAs produced by these strains. Conclusions Here, we advance the understanding of biosynthesis of LCAs, important molecules for paleotemperature reconstruction, in the marine haptophyte G. huxleyi . By assembling the genomes of two LCA-producing strains, we were able to overcome the limitations of the scarcity of reference genomes of this group, enabling the analysis of gene expression under varying growth conditions and the identification of PKSs and desaturases potentially involved in LCA biosynthesis These findings constitute a step forward in elucidating the genetic basis for LCA production and regulation in G. huxleyi , which not only advances our understanding on how LCAs are formed but also aids in the improvement of LCA-based paleotemperature reconstructions. Here we supply: Supplementary Data 1. Gas Chromatography (GC) and GC-Mass Spectrometry (GC-MS) lipid data files. This folder contains raw data for all lipid analyses. The files are organized into subfolders based on the figure they support. The dataset includes the raw GC files used for quantification and the GC-MS files used for compound identification, including the specific chromatograms that generated the representative peaks shown in Figure 2 . Where applicable, the subfolders also contains a summary table with the quantitative results extracted from the chromatograms and a sample metadata file that lists the experimental conditions (e.g., temperature, time point) and links them to the corresponding data file. Supplementary Data 2. De novo genome assemblies of G. huxleyi strains CCMP1742 and CCMP2758 generated during the evaluation of various assembly tools shown in Supplementary Figure 2. These include long read assemblers (CANU v1.8 , flye v2.8.1 , shasta v0.5.1 and wtdbg2 v2.3 with the wtpoa-cns v2.3 consenser) , the short read assembler SPAdes v3.14.1 , and hybrid assemblers that use both short and long reads (Wengan v0.2-using either DiscovarDeNovo [WenganD] or Minia3 [WenganM] as the short-read assembler- hybridSPAdes v3.14.1 , biosyntheticSPAdes v 3.14.1, and HASLR v 0.8a1. and a metagenomic assembler (OPERA-MS (v 0.8.2) in hybrid mode using MegaHIT as short read assembler. A metagenomic assembler, OPERA-MS v0.8.2 run in hybrid mode with MegaHIT, was also evaluated. Furthermore, two scaffolders, WenganD and OPERA-LG v2.0.5, were applied to two selected assemblies (hybridSPAdes and biosyntheticSPAdes). Supplementary Data 3. De novo genome and metagenome assemblies of G. huxleyi strains CCMP1742 and CCMP2758 generated at different stages of the assembly process outlined in Figure 3. The included assemblies correspond to specific steps: (1) the initial hybrid metagenomic assemblies generated using MetaSPAdes (v 3.14.1) and OPERA-MS (with SPAdes as the short-read assembler); (3) the hybrid assemblies created with SPAdes and biosyntheticSPAdes after a read-filtering step, where reads were mapped to the metagenomic assembly using Minimap2 and those mapping to contigs identified as contamination by CAT classification, coverage, and GC content were removed; (4) the scaffolded assemblies generated using OPERA-LG; and (7) the final assemblies after the identification and removal of remaining non-eukaryotic scaffolds. Supplementary Data 4. Funannotate genome annotation output files. This folder contains the "annotate_results" directories generated by Funannotate for the CCMP1742 and CCMP2758 assemblies. These assemblies were produced using MetaSPAdes as a hybrid metagenomic assembler and biosyntheticSPAdes as a hybrid genome assembler, from which reads and contigs identified as likely contamination have been removed as described in Figure 3 . Supplementary Data 5. Cell count data for cold-shock and cold-adaptation experiments in G. huxleyi strains CCMP1742 and CCMP2758. This folder contains three files: two with data from the cold-shock experiments (one for each strain, CCMP1742 and CCMP2758) and one with data from the temperature acclimation experiment. The cold-shock datasets provide a summary of cell counts and viability measurements. For each sample, cell abundance was first measured on an unstained aliquot based on chlorophyll red autofluorescence (PerCP) versus forward scatter (FSC). A separate aliquot was then stained with SYTOX Green nucleic acid stain to determine cell viability, with a formalin-killed control used to define the gating strategy for dead cells. The data includes results from both untreated samples ("FCM" tab) and live/dead tests ("FCM_Live_dead" tab). Technical replicates for these experiments are included where applicable. The temperature acclimation dataset describes the cellular abundance data for strains grown at 20 °C and 7 °C. It reports raw counts for two distinct cell clusters ("Up" and "Down"), identified by their chlorophyll red autofluorescence (PerCP) versus forward scatter (FSC), and the total cell sum of both clusters for each sample. Data are presented as individual technical replicates ("raw" tabs), the average and standard deviation of those replicates per flask ("per_flask" tabs), and the average and standard deviation for each time point across three biological replicates (except t=0, n=1). Supplementary Data 6. Source data and statistical analysis for Figures 5, 6, and 7, and Supplementary Figures 10, 15, 16, 23, 25 and 31 . This file contains the underlying dataset and the results of the statistical analyses used to generate the indicated figures. The dataset includes the underlying numerical values and categorical labels plotted in the figures, as well as the full outputs of statistical analyses-including test statistics, degrees of freedom, p-values, confidence intervals, and post-hoc comparisons where applicable. Supplementary Data 7 . RNA-seq alignment and deduplication metrics . This file contains the output metrics from the STAR aligner and UMI-tools deduplication for RNA-seq data. It includes key alignment statistics-such as the number of input reads, uniquely mapped reads (%), and reads mapped to multiple loci (%)-for strains CCMP1742 and CCMP2758 before and after deduplication, and for strain CCMP1516 (non-deduplicated). The file also includes the deduplication results from running "umi_tools dedup" command for samples of CCMP1742 and CCMP2758 strains, reporting metrics such as number of input reads, final output counts and mean and maximum number of unique UMIs per position. Supplementary Data 8. Phylogeny of desaturases . This folder contains the file used to generate Figure 9 and Supplementary Figure 33 . The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files.
Goswami, Rohit · Goswami, Ruhila
5.0 MB
Reproducibility payload for the rsx BMC Bioinformatics submission, regenerated for rsx-rs v0.2.3. The deposit contains: (1) rsx_bmc_repro_archive_20260604.tar.xz, the self-contained pixi + Snakemake workflow that clones rsx-rs at the pinned v0.2.3 tag, builds rsx, the C++ RADSex v1.2.0 reference, and the pyrsx bindings, and regenerates every figure and data table in the paper (synthetic regression suite; the four-panel literature benchmark with all 56 paired command/dataset/depth timings including the depth command; Bayesian evidence; mode and QC effects; biological unlocks, sex-system inference and candidate triage; the prior x linked-probability triage grid and heatmap; the low-depth sweep and depth-stability summary; and Python-bindings parity); and (2) the complete downloaded literature benchmark data archive (FASTQ samples, regenerated marker-table workdirs, per-dataset logs and comparisons). All benchmark timings were produced on an AMD Ryzen Threadripper PRO 3955WX workstation (16 threads) and scale with host hardware.
alanazi, yousef
438 KB
Reproducibility package for the manuscript analyzing a four-gene metabolic immune-checkpoint panel (VSIR, CD38, ENTPD1, NT5E) in lung adenocarcinoma. Contains the processed TCGA-LUAD analysis dataset (n = 497, 180 deaths), R scripts that reproduce all figures, tables, and statistics, and the generated results and figures. Raw public data (TCGA-LUAD expression and GDC phenotype from UCSC Xena/GDC; GSE68465 from NCBI GEO) are not redistributed; GSE68465 is downloaded automatically by the validation script. See README.txt for full instructions. Funded by the Deanship of Scientific Research, Northern Borders University, grant NBU-FFR-2026-2088-03.
Khan, Saad · Wang, Anthony · Desai, Rupen · et al.
8.0 MB
This resource is part of a paper titled "Mapping the spatial architecture of glioblastoma from core to edge delineates niche-specific tumor cell states and intercellular interactions" and contains spatial transcriptomic data and single-cell RNA-sequencing data for primary human glioblastoma samples. Spatial transcriptomic data was collected using the Visium and Xenium platforms. ABSTRACT : Treatment resistance in glioblastoma (GBM) is largely driven by the extensive multi-level heterogeneity that typifies this disease. Despite significant progress toward elucidating GBM's genomic and transcriptional heterogeneity, a critical knowledge gap remains in defining this heterogeneity at the spatial level. To address this, we employed spatial transcriptomics to map the architecture of the GBM ecosystem. This revealed tumor cell states that are jointly defined by gene expression and spatial localization, and multicellular niches whose composition varies along the tumor core-edge axis. Ligand-receptor interaction analysis uncovered a complex network of intercellular communication, including niche-and region-specific interactions. Finally, we found thatCD8⁺GZMK⁺T cells colocalize withLYVE1⁺CD163⁺myeloid cells in vascular regions, suggesting a potential mechanism for immuneevasion. These findings provide novel insights into the GBM tumor microenvironment, highlighting previously unrecognized patterns of spatial organization and intercellular interactions, and novel therapeutic avenues to disrupt tumor-promoting interactions and overcome immune resistance.
Robson, Soares · Alexandra, Rocha · Matheus, Camargo · et al.
8.0 MB
This repository contains the complete dataset, computational scripts, and supplementary figures/tables for the manuscript titled "Genome-wide exploration of biosynthetic gene clusters and their association to virulence in the entomopathogenic fungus Beauveria". It includes the raw sequence inputs, output tables, phylogenetic trees, and the reproducible computational pipeline (Bash and R scripts) used to generate the figures and tables presented in the study.
Nishino, Satoshi · Tominaga, Kento · Omae, Kimiho · et al.
1.5 MB
data.tar.gz ( 65.09 MB ) Contains all datasets used to generate figures in the manuscript, including annotation tables, structural similarity search results, genomic-context outputs, and intermediate processed files. 260613_unknome_R.ipynb.ipynb (6.95 MB ) R notebook used to generate all visualizations included in this study. All necessary raw datasets are provided in data.tar.gz . plot.tar.gz (1.58 MB) All plot files generated by 260613_unknome_R.ipynb.ipynb . aai_cleaned_comp50_contam10.tsv (580.95 KB) AAI calculation results for SAR11 genomes. SonicParanoid2_ortholog_groups.tsv (7.94 MB) Orthologous group assignments generated using SonicParanoid2. TMHMM2_result.tsv (14.68 MB) Predicted transmembrane regions for SAR11 proteins (TMHMM 2.0). all_SAR11_defense_finder_genes.tsv (21.27 KB) all_SAR11_defense_finder_hmmer.tsv (525.50 KB) all_SAR11_defense_finder_systems.tsv (6.90 KB) DefenseFinder outputs, including predicted defense genes, HMMER matches, and system-level classifications.
Roberts, Eric · Hoffman, Michael
8.0 MB
Minimum unique mappable read lengths: GRCh38.p14: range ≤8 bp to 65535 bp Assembly: GRCh38.p14 Range: ≤8 bp to 65535 bp Newmap version: 0.2 Files GCA_000001405.15_GRCh38_no_alt_analysis_set.unique.zip : A zip archive containing a unique.uint16 file for each sequence in the no alternative loci scaffolds analysis set for GRCh38.p14 . Methods For each position in the genome, Newmap determines the minimum length of a sequencing read starting at that position that would occur only once in the assembly. Newmap searches both the primary sequence and the reverse complement. Positions whose minimum unique mappable lengths are in [1, 8] are reported as 8. Positions whose unique lengths are greater than 65535 are reported as 0. File names All filenames begin with a sequence ID from the genome reference. The extension unique.uint16 indicates unique lengths encoded as a flat array of unsigned little-endian 16-bit integers.
Fareed, Fathmath Shaman · Singaram, Nallammai · Abdul Malek, Ahmad Zuhairi · et al.
12 KB
Accession numbers
gong, siyuan
6.3 MB
Renal ischemia-reperfusion injury (IRI) is a leading cause of acute kidney injury, associated with mitochondrial dysfunction, excessive reactive oxygen species (ROS) production, and tubular cell apoptosis. The Sigma-1 receptor (Sigma1R), an intracellular chaperone, plays a crucial role in maintaining mitochondrial homeostasis and promoting cellular survival. In this study, we found that Sigma1R expression was significantly downregulated in renal IRI. Overexpression of Sigma1R alleviated renal dysfunction, reduced apoptosis, stabilized mitochondrial membrane potential, and attenuated ROS accumulation. Mechanistically, Sigma1R modulated Rac1 activity and enhanced PINK1/Parkin-mediated mitophagy, which facilitated the removal of damaged mitochondria and the restoration of mitochondrial quality control. In vitro, similar protective effects were observed in HK-2 cells subjected to hypoxia/reoxygenation (H/R) injury. These findings demonstrate that Sigma1R exerts its renoprotective effects by regulating Rac1-mediated mitophagy and improving mitochondrial function, thereby highlighting Sigma1R as a promising therapeutic target for preventing and treating renal IRI.
Zhang, Guoqiang
47 KB
To address the challenges of ambiguous term boundaries, complex semantic hierarchies, and dispersed structural expressions in Traditional Chinese Medicine (TCM) oncology texts, this study proposes a joint optimization model based on the Bidirectional Encoder Representations from Transformers (BERT) architecture for structured modeling and semantic parsing of TCM oncology texts. This model integrates open-source medical terminology data and anonymized TCM question-answer corpora to construct a well-annotated custom corpus subset (D1: 2,000 paragraphs; D2: 6,000 paragraphs; D3: 12,000 paragraphs) for training and evaluation. The data primarily derive from the national standard TCM terminology database, abstracts of Chinese medical research papers, and Q&A content from TCM consultation forums. Begin-Inside-Outside sequence labeling and paragraph-level semantic classification annotations are applied. A BERT–Conditional Random Field–Multilayer Perceptron joint modeling architecture is designed, incorporating staged unfreezing and dynamic loss weighting mechanisms to improve model generalization. Comparative performance experiments demonstrate the superiority of the proposed model in various aspects, including inference latency (e.g., 6.798 ms per item on the D3 dataset), worst-case accuracy (91.364%), parameter efficiency, and training stability (fluctuation rate: 0.492). For the tokenization task, the model significantly outperforms baselines in metrics such as Strict-F1 (93.723 on D3), entity-level recall (93.764% on D2), and average span error (0.608 on D1). In the classification task, demonstrate that this joint modeling approach effectively supports fine-grained terminology recognition and semantic segmentation in TCM oncology texts. It provides a solid foundation for developing structured information extraction systems for TCM knowledge bases, follow-up records, and consultation notes aimed at tumor patient management and integrated clinical decision support.
Takano, Takeshi · Takano, Takeshi
9.8 KB
This dataset contains the replication data for "Neonatal social communication and single genes predict the variability of post-pubertal social behavior in a mouse model of paternal 15q11-13 duplication" paper.