Ivani, Jessica K. · Nogina, Alexandra · Zakharko, Taras
79 rows × 49 cols · 15 KB
47 categorical · 2 text
Dataset for the paper Agent-oriented modality in the languages of North America, published in the Proceedings of the 28th Workshop on American Indigenous Languages (WAIL 28), May 1-2, University of California Santa Barbara ( https://www.wailconference.org/home )
Raw qPCR quantification summaries (per-well Cq values and run metadata) for HOXA-10, HOXA-11, HOXA-13, and ACTB, exported directly from the BioRad CFX-96 / CFX Maestro software. These are the raw amplification run records (runs of 13-14 May 2024) underlying the relative-expression (RQ) values reported in the associated study. Each target is provided as a separate sheet (Well, Fluor, Target, Content, Sample, Cq), with a README sheet describing the contents. Standard-curve efficiency values were not generated as permanent records for these runs.
Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Code This repository contains the data and code required to reproduce the analyses, statistical results, and figures presented in the associated manuscript. Files are organized by function and described below. All research proposals are anonymized and labeled with non-identifying identifiers (A, B, C, ...). Reviewer identities were never provided to the authors. To prevent inadvertent disclosure, all free-text review content from both human reviewers and large language models (LLMs) has been removed; only the numerical evaluation data required to reproduce the reported analyses are included. Data files Human_raw_scores.csv Individual numerical scores assigned by human reviewers, one row per reviewer × proposal × criterion. Used to compute panel-level summary statistics and the inter-reviewer reliability metrics reported in the manuscript. Human.csv Proposal-level human reference scores (one row per proposal) used as the human benchmark against which LLM scores and rankings are compared. LLM_data_combined_clean_filtered.csv All numerical scores generated by the evaluated LLMs. Preprocessed to remove evaluations in which a model failed to return one or more required numerical scores (see Data provenance and known limitations for counts). review_criteria.txt The evaluation criteria and rating scales. See the note under Data provenance regarding the 2023 vs. 2024 criterion naming. Data dictionary Human_raw_scores.csv Column Description Applicant Anonymized proposal identifier (A, B, C, ...). Cycle Review cycle the proposal belongs to (2023 or 2024). Label Scoring criterion: Intellectual merit , Potential for impact , Collaborative Potential , or Overall ranking . Reviewer_Seq Reviewer index within a proposal (1, 2, 3, ...). Identities are unknown; this index only links a single reviewer's ratings across criteria for the same proposal, in source-file order. It is not consistent across proposals (reviewer 1 for proposal A is not reviewer 1 for proposal B). Rating Numerical score. Criterion ratings use a 1-5 scale; Overall ranking uses a 1-3 scale (3 = fund, 2 = fund with modifications, 1 = do not fund). Human.csv Column Description Name Anonymized proposal identifier (matches Applicant above). Cycle Review cycle (2023 or 2024). IM , Impact , Collab , Overall Panel-mean scores for the four criteria. Score Weighted composite panel score, computed as 0.4·IM + 0.3·Impact + 0.3·Collab , matching the weighting applied to the LLM composite scores. LLM_data_combined_clean_filtered.csv Column Description Name Anonymized proposal identifier (matches Human.csv ). Type Input given to the model: Abstract or Full_Proposal . Model Model identifier. For models with controllable reasoning depth, the tier is appended as a suffix ( _low , _medium , _high ). The portion before the first underscore is the root model. Prompt Prompting strategy ( OneShot or CoT ). Seed Requested random seed. For models that did not support seed specification at the time of execution (the reasoning models listed in the Methods), the API ignored this value and it functions only as a replicate index; output is not reproducible from it for those models. Temp Sampling temperature (0.1, 0.5, 0.9). IM , Impact , Collab Criterion ratings (1-5 scale). Overall Overall recommendation (1-3 scale). See limitation note on out-of-scale values. Score Weighted composite, 0.4·IM + 0.3·Impact + 0.3·Collab . Category Reasoning/architecture category of the model. Year Evaluation wave in which the run was performed (proposals were re-evaluated as new model generations were released); this is not the proposal's submission cycle. Use Cycle in the human files for submission cycle. Data provenance and known limitations We document the following so that users can interpret the data accurately. Two review cycles, combined. The 28 proposals come from two internal seed grant cycles: 15 from 2023 and 13 from 2024 ( Cycle column). For 2024 proposals, proposal-level means in Human.csv are the official institute panel means; for 2023 proposals they are computed from the individual ratings in Human_raw_scores.csv . Reviewers per proposal. Panels ranged from 3 to 6 reviewers. Reviewer identities were never provided; the human inter-reviewer reliability is therefore estimated with a one-way random-effects model (ICC(1,1)), which is the appropriate model when each proposal is rated by a different, unidentified set of reviewers. Six 2024 reviews not present at the individual level. For six 2024 proposals (G, R, T, W, X, Z), one reviewer's scores were submitted without written comments and are not included in Human_raw_scores.csv . For these proposals, Human.csv carries the official institute panel means, so the panel mean in Human.csv and the mean recomputed from Human_raw_scores.csv differ slightly. The reproducibility check in 06_Human_data.ipynb confirms exact agreement for all proposals with complete individual records and reports the expected small differences for these six. Out-of-scale LLM Overall ratings. A small number of responses (≈1.2% of reviews, almost entirely from gpt-3.5-turbo) rated the overall recommendation on a 1-5 scale rather than the requested 1-3 scale, in a format the parser could not distinguish. The analysis code masks values outside the valid range before computing any Overall -based result; the composite Score does not use Overall and is unaffected. Excluded LLM evaluations. Evaluations in which a model failed to return one or more required numerical scores were removed prior to analysis. The file provided here is the post-exclusion (analyzed) dataset. Code notebooks Two groups of notebooks are provided: (i) the LLM evaluation pipeline and (ii) statistical analysis and figure generation. LLM evaluation pipeline (Notebooks 1-4) Documentation of the methodology used to generate the LLM evaluations. These use synthetic examples and contain no confidential data, API credentials, or real proposal text. 01_pipeline_overview.ipynb - architecture, configuration, criteria, output format, evaluation matrix. 02_prompting_strategies.ipynb - one-shot and chain-of-thought prompting; example selection; text vs. vision input. 03_response_parsing.ipynb - regex extraction of ratings and comments; error handling; decimal ratings. 04_example_evaluation.ipynb - end-to-end workflow on synthetic data. Statistical analysis and figures (Notebooks 5-6) 05_Data_Processing.ipynb - ANOVA and effect sizes; empirical absolute score differences; ICC(2,1) and Spearman correlations vs. the human panel; Monte Carlo comparisons; scatter, slope, and bias figures. 06_Human_data.ipynb - ICC(1,1)/ICC(1,k) for human reviewers with bootstrap confidence intervals; Monte Carlo single-reviewer vs. leave-one-out panel rank agreement; consistency check of Human.csv against Human_raw_scores.csv . Reproducibility Running 05_Data_Processing.ipynb and 06_Human_data.ipynb against the included data reproduces the statistical results and figures in the manuscript. Notebooks 1-4 document the evaluation pipeline using synthetic examples. Requirements pandas numpy scipy statsmodels scikit-learn matplotlib The LLM evaluation pipeline notebooks (1-4) additionally use the packages listed in requirements.txt . Citation If you use this data or code, please cite the associated manuscript and this archive: Gorski, C., Leo, N., Gayah V. Agreement Between Large Language Models and Humans in Research Proposal Review - Data and Scripts. 2026. Zenodo. https://doi.org/10.5281/zenodo.18187034 License Data are released under CC BY 4.0; code is released under the MIT License.
We performed 10X Genomics single-cell RNAsequencing of human prepheral blood mononuclear cells from healthy control, PD and RBD patients. This dataset contains raw FASTQ files. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.
The development of social media has changed the way individuals express emotions and communicate with their social environment. One emerging phenomenon is the use of TikTok's repost feature as a means of indirect communication through the re-sharing of content perceived to represent the user's personal feelings or experiences. This study aims to determine the behavior of using TikTok's repost feature as an indirect communication medium for expressing emotions among Indonesian women aged 18-20 years. The study used a quantitative method with a descriptive design through an online survey of 200 respondents selected using a purposive sampling technique. Data were collected using a five-point Likert-based questionnaire structured based on the dimensions of the repost feature and emotional expression. The results showed that the level of utilization of TikTok's repost feature was in the high category with an average value of 76.97. In addition, the results of the hypothesis test showed a significance value of 0.000 (p < 0.05), indicating that Indonesian women aged 18-20 years significantly utilize TikTok's repost feature as an indirect communication medium for expressing emotions. These findings indicate that the repost feature not only functions as a means of content distribution, but also as a symbolic communication medium to convey emotional messages, build digital identity, gain social validation, and maintain self-image in the digital environment.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 37 KB · pdf
This dataset consists of raw sequencing snRNA-seq data and snATAC-seq data (10x Genomics Chromium Next GEM Multiome ATAC/GEX). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108 ) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors. (edited)
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 37 KB · pdf
This dataset consists of raw sequencing snRNA-seq data (10x Genomics Chromium Next GEM Single Cell 3ʹ). The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a single sequencing library. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
We performed 10X Genomics single-cell RNAsequencing of human iSPC-derived monocytes and macrophages in vitro. Cells were treated with 500 ng/mL lipopolysaccharide (LPS) and 50 ng/mL interleukin-1 beta (IL1b) for 24 hours. This dataset contains raw FASTQ files from myeloid cells, sorted into 4 groups namely monocytes (Mono) and Macrophages (Mac) non-stimulated (NS) and LPS+IL1b-stimulated cells. Sequencing was performed using NovaSeq 6000 S4 PE 100bp. Reads were processed using the 10X Genomics Cell Ranger Single Cell 2.0.0 pipeline. FASTQs generated from sequencing output were aligned to the human GRCh38 reference genome using STAR algorithm 2.7.3a.
Pérez Gallego, Ruth · von Meijenfeldt, F. A. Bastiaan · Bale, Nicole J. · et al.
2.0 MB
Abstract Paleontological and phylogenomic observations have shed light on the evolution of cyanobacteria. Nevertheless, the emergence of heterocytes, specialized cells for nitrogen fixation, remains unclear. Heterocytes are surrounded by heterocyte glycolipids (HGs), which contribute to protection of the nitrogenase enzyme from oxygen. Here, by comprehensive HG identification and screening of HG biosynthesis genes throughout cyanobacteria, we identify HG analogs produced by specific and distantly related non-heterocytous cyanobacteria. These structurally less complex molecules probably acted as precursors of HGs, suggesting that HGs arose after a genomic reorganization and expansion of ancestral biosynthetic machinery, enabling the rise of cyanobacterial heterocytes in an increasingly oxygenated atmosphere. Subsequently, HG chemical structure evolved convergently in response to environmental pressures. Our results open a new chapter in the potential use of diagenetic products of HGs and HG analogs as fossils for reconstructing the evolution of multicellularity and division of labor in cyanobacteria. Here we supply: Supplementary Data 1. Selected cyanobacterial genomes from the PATRIC genome database (now part of the BV-BRC database). Files called 'selected_Cyanogenomes.genome_*.20220430.txt' are sourced from the PATRIC File Transfer Protocol server (ftp.patricbrc.org). 'gtdbtk.bac120.summary.tsv' is the GTDB-Tk output file, and 'qa.summary_extended.txt' the CheckM output file. Supplementary Data 2. HG biosynthetic gene clusters in selected PATRIC genomes and 14 newly sequenced genomes. The file 'islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt' contains the location of all hits to Anabaena sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: "genome | contig". The structure of a hit is as follows: "ORF number on contig | query ( e -value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]". Non-overlapping hits on the same ORF (see Online Methods) are connected with '&&&' characters. An asterisk ('*') indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with '~~~' characters. The file 'Supplementary_table.script_1.txt' contains a summary of all identified hgl islands (i.e. clusters containing at least seven unique HG biosynthesis gene hits). Supplementary Data 3. HG biosynthetic gene clusters in 255,388 prokaryotic genomes from the PATRIC genome database (now part of the BV-BRC database). The file called 'PATRIC_20230120.selection_c50_c10.txt' contains information on the selected PATRIC genomes based on data sourced from the PATRIC File Transfer Protocol server ( ftp.patricbrc.org ). The file 'all_tree_of_life_genomes.islands_on_contigs.3_ORFs_in_between.expanded_island_with_nucleotide_positions.txt' contains the location of all hits to Anabaena sp. PCC 7120 HG biosynthesis genes. ORFs were predicted with Prodigal. The structure of a contig is as follows: "genome | contig". The structure of a hit is as follows: "ORF number on contig | query ( e -value; bit-score; start of alignment in query; end of alignment in query; query coverage per subject; start of alignment in subject; end of alignment in subject; subject coverage) [nucleotide position on contig start; nucleotide position on contig end; direction]". Non-overlapping hits on the same ORF (see Online Methods) are connected with '&&&' characters. An asterisk ('*') indicates that the hit is located at most three ORFs from a contig edge. Clusters of hits that are at most three open reading frames (ORFs) apart are connected with '~~~' characters. Supplementary Data 4. Phylogeny of representative cyanobacterial genomes based on a core gene superalignment. The folder contains the files used to generate Fig. 2a. The directory 'IQ-TREE' contains the tree file and iTOL annotation files. The file 'dRep.representative_to_cluster.txt' contains the dRep clusters. Note that the manually defined subclades in the iTOL annotation file 'iTOL_annotation.manually_defined_clades.DATASET_STYLE.txt' have a different numbering from the paper: subclades 0 and 1 are the 'heterocytous sister clades', and subclades 2-10 in the annotation file are heterocytous subclades 1-9 in the paper, respectively. Supplementary Data 5. Lipid data files. The folder contains all the UHPLC-HRMS n (Orbitrap) datafiles used in this study. The directory 'CCY strains' includes 24 heterocytous cyanobacterial cultures corresponding to 23 strains grown in nitrogen-deficient media, the resulting data are shown in Supplementary Table 10. Directory 'HglT mutant' contains the datafiles used to generate Supplementary Table 15. The directory 'LEGE strains' includes the UHPLC-HRMS n (Orbitrap) and GC-MS datafiles corresponding to eight cultures of two non-heterocytous strains grown in media with and without nitrogen for 38 to 77 days, the resulting data are shown in Supplementary Tables 10, 17 and 18. Supplementary Data 6. Plasmid maps. GenBank and FASTA files of plasmids generated in this study. 'HglT deletion' directory contains the genomic region surrounding hglT in the wild-type strain and after deletion used to generate Supplementary Fig. 15. pAM5404 is shown in Supplementary Fig. 16 and p(A)RP0XX are shown in Supplementary Fig. 17. Supplementary Data 7. Phylogenies of seven hgl island genes and of a concatenated alignment of these genes. The folder contains the files used to generate Supplementary Fig. 18 (in the directory 'gene_trees_hgl_islands'), and Fig. 4 and related figures (in the directory 'gene_trees_hgl_islands_4'. The directories contain the alignments and trimmed alignments, IQ-TREE output files, and iTOL annotation files. The file 'gene_trees_hgl_islands/analysis_individual_gene_trees/explore_individual_gene_clusters.ipynb' contains the code to identify the five hgl islands that contain genes with incongruent evolutionary histories. Supplementary Data 8. Phylogeny of hglE A homologs. The folder contains the files used to generate Supplementary Fig. 24 and related figures. The file 'selected_hglE_hits.txt' contains the selected hglE A hits and the genomic cluster on which they are located. The folder contains the alignment and trimmed alignment, IQ-TREE output files, and iTOL annotation files. All the code used in this publication including scripts used for: genome assemblies, download of genomes from public repositories, quality and contamination checks, genome analysis, construction of the phylogenetic trees, hgl island identification, etc. The shell script 'commands.sh' within each directory contains all the code used to generate the content in the directory. All the figures used in this publication including the figures in the Supplementary Information file.
Kim, Kwanho · Lin, Zechuan · Simmons, Sean · et al.
1 files · 37 KB · pdf
This dataset contains raw FASTQ files from the midbrain single-nucleus RNA sequencing (snRNAseq) dataset with hybrid selection for the matching PMDBS samples from the PD5D chort. The same subjects were also profiled with other omics assays including genomic DNAseq, genotyping, single-cell ATACseq, and spatial transcriptomics.
Reproducibility payload for the rsx BMC Bioinformatics submission, regenerated for rsx-rs v0.2.3. The deposit contains: (1) rsx_bmc_repro_archive_20260604.tar.xz, the self-contained pixi + Snakemake workflow that clones rsx-rs at the pinned v0.2.3 tag, builds rsx, the C++ RADSex v1.2.0 reference, and the pyrsx bindings, and regenerates every figure and data table in the paper (synthetic regression suite; the four-panel literature benchmark with all 56 paired command/dataset/depth timings including the depth command; Bayesian evidence; mode and QC effects; biological unlocks, sex-system inference and candidate triage; the prior x linked-probability triage grid and heatmap; the low-depth sweep and depth-stability summary; and Python-bindings parity); and (2) the complete downloaded literature benchmark data archive (FASTQ samples, regenerated marker-table workdirs, per-dataset logs and comparisons). All benchmark timings were produced on an AMD Ryzen Threadripper PRO 3955WX workstation (16 threads) and scale with host hardware.
Reproducibility package for the manuscript analyzing a four-gene metabolic immune-checkpoint panel (VSIR, CD38, ENTPD1, NT5E) in lung adenocarcinoma. Contains the processed TCGA-LUAD analysis dataset (n = 497, 180 deaths), R scripts that reproduce all figures, tables, and statistics, and the generated results and figures. Raw public data (TCGA-LUAD expression and GDC phenotype from UCSC Xena/GDC; GSE68465 from NCBI GEO) are not redistributed; GSE68465 is downloaded automatically by the validation script. See README.txt for full instructions. Funded by the Deanship of Scientific Research, Northern Borders University, grant NBU-FFR-2026-2088-03.
Pančíková, Alexandra · Theunis, Koen · Hulselmans, Gert · et al.
1 files · 36 KB · pdf
This dataset consists of raw sequencing snRNA-seq data using ParseBio Evercode Whole Transcriptome. The data is part of an overall set of samples derived from postmortem midbrain (n=140), cingulate cortex (n=190) and motor cortex (n=4) of healthy donors (n=114), patients with Parkinson's disease (n=75) or patients with other neurological disorder (n=1). The protocol followed to isolate nuclei from postmortem brain samples and to prepare sequencing libraries can be found below. To increase throughput and to decrease batch effects, several donors have been pooled together into a specific ParseBio barcode. To computationally demultiplex the nuclei to their corresponding donors, cellsnp-lite (version commit: aad18644adcde853c313362a856a24245c9b91f7) followed by vireo (https://github.com/single-cell-genetics/vireo/pull/108) has been used. The population VCF with the donor genotypes derived from whole genome sequencing data has been used to assign nuclei back to their donors.
data.tar.gz ( 65.09 MB ) Contains all datasets used to generate figures in the manuscript, including annotation tables, structural similarity search results, genomic-context outputs, and intermediate processed files. 260613_unknome_R.ipynb.ipynb (6.95 MB ) R notebook used to generate all visualizations included in this study. All necessary raw datasets are provided in data.tar.gz . plot.tar.gz (1.58 MB) All plot files generated by 260613_unknome_R.ipynb.ipynb . aai_cleaned_comp50_contam10.tsv (580.95 KB) AAI calculation results for SAR11 genomes. SonicParanoid2_ortholog_groups.tsv (7.94 MB) Orthologous group assignments generated using SonicParanoid2. TMHMM2_result.tsv (14.68 MB) Predicted transmembrane regions for SAR11 proteins (TMHMM 2.0). all_SAR11_defense_finder_genes.tsv (21.27 KB) all_SAR11_defense_finder_hmmer.tsv (525.50 KB) all_SAR11_defense_finder_systems.tsv (6.90 KB) DefenseFinder outputs, including predicted defense genes, HMMER matches, and system-level classifications.
This dataset contains the input sequences and the complete bioinformatic output files supporting the analysis of the Cannabis sativa Monoecy1 sex-determination locus reported in the associated manuscript. All analyses were performed on the Cannabis sativa cv. 'Pink Pepper' genome assembly GCA_029168945.1 (ASM2916894v1), chromosome X (RefSeq accession NC_083610.1), using the Dardel high-performance computing system at the PDC Center for High Performance Computing, KTH Royal Institute of Technology, Stockholm. The files document four lines of evidence used to evaluate the mechanism by which the long non-coding RNA lncREM16 (LOC133032448; transcript XR_009685085.1) silences the B3-domain transcription factor gene CsREM16 (LOC115699937; transcripts XM_030627485.2 and XM_030627490.2) within the Monoecy1 locus. Contents: Input sequences (FASTA): the 80 kb Monoecy1 genomic locus (Monoecy1_locus.fa; NC_083610.1:81,080,000-81,160,000), the two CsREM16 mRNA isoforms (CsREM16_X1.fa, CsREM16_X2.fa, CsREM16_both.fa), the lncREM16 transcript (lncREM16.fa), and the extracted 73 bp shared region from both the lncREM16 transcript and the genomic locus (shared_73bp.fa, shared_genomic.fa, both_shared.fa). Transposable element annotation (RepeatMasker v4.1.5, RMBLAST v2.14.0+, Dfam 3.7, viridiplantae): annotation tables, summary statistics, masked sequences, category files, and GFF output for both the 80 kb locus (Monoecy1_locus.fa.out, .tbl, .cat, .masked, .out.gff) and the lncREM16 transcript (lncREM16.fa.out, .tbl, .cat, .masked, .out.gff, .out.html). These files report the 33% TE content of the locus and the 37% TE content of the lncREM16 transcript, and identify the unclassified element DR2330744 at the CsREM16 exon 3 / intron 2 boundary. Sequence complementarity tests (NCBI blastn): antisense BLAST of lncREM16 against the CsREM16 mRNAs (lnc_vs_REM16_antisense.txt), against the genomic locus (lnc_vs_genomic_antisense.txt), the corresponding sense-strand orientation check (lnc_vs_genomic_sense.txt), the lncREM16 self-BLAST for inverted-repeat / hairpin detection (lnc_selfblast.txt), the alignment of the 73 bp shared region (shared_vs_lnc.txt, shared_vs_rem16.txt), and the BLAST of the 73 bp query against the NCBI nt database (blast_results.xml, blast_rid.txt). All antisense complementarity tests returned zero hits. Small RNA mapping (Bowtie v1): size-filtered small RNA reads (18-30 nt) derived from NCBI SRA run SRR25938888 (gerola_18_30.fastq.gz), and the alignment outputs in antisense and sense orientation against the CsREM16 mRNAs and the genomic locus (antisense_CsREM16.sam, sense_CsREM16.sam, antisense_monoecy1.sam, sense_monoecy1.sam) with their mapping statistics (antisense_stats.txt, sense_stats.txt, antisense_monoecy1_stats.txt, sense_monoecy1_stats.txt, overall_mapping_stats.txt, and the relaxed three-mismatch run antisense_v3.sam, antisense_v3_stats.txt). Gene records: NCBI gene-information records documenting the assembly, coordinates, and exon structure of CsREM16 and lncREM16 (rem16_gene_info.txt, lnc_gene_info.txt). Software versions and parameters are documented in the associated manuscript. The preprint / published article DOI will be added to this record under related identifiers upon availability.
Renal ischemia-reperfusion injury (IRI) is a leading cause of acute kidney injury, associated with mitochondrial dysfunction, excessive reactive oxygen species (ROS) production, and tubular cell apoptosis. The Sigma-1 receptor (Sigma1R), an intracellular chaperone, plays a crucial role in maintaining mitochondrial homeostasis and promoting cellular survival. In this study, we found that Sigma1R expression was significantly downregulated in renal IRI. Overexpression of Sigma1R alleviated renal dysfunction, reduced apoptosis, stabilized mitochondrial membrane potential, and attenuated ROS accumulation. Mechanistically, Sigma1R modulated Rac1 activity and enhanced PINK1/Parkin-mediated mitophagy, which facilitated the removal of damaged mitochondria and the restoration of mitochondrial quality control. In vitro, similar protective effects were observed in HK-2 cells subjected to hypoxia/reoxygenation (H/R) injury. These findings demonstrate that Sigma1R exerts its renoprotective effects by regulating Rac1-mediated mitophagy and improving mitochondrial function, thereby highlighting Sigma1R as a promising therapeutic target for preventing and treating renal IRI.
To address the challenges of ambiguous term boundaries, complex semantic hierarchies, and dispersed structural expressions in Traditional Chinese Medicine (TCM) oncology texts, this study proposes a joint optimization model based on the Bidirectional Encoder Representations from Transformers (BERT) architecture for structured modeling and semantic parsing of TCM oncology texts. This model integrates open-source medical terminology data and anonymized TCM question-answer corpora to construct a well-annotated custom corpus subset (D1: 2,000 paragraphs; D2: 6,000 paragraphs; D3: 12,000 paragraphs) for training and evaluation. The data primarily derive from the national standard TCM terminology database, abstracts of Chinese medical research papers, and Q&A content from TCM consultation forums. Begin-Inside-Outside sequence labeling and paragraph-level semantic classification annotations are applied. A BERT–Conditional Random Field–Multilayer Perceptron joint modeling architecture is designed, incorporating staged unfreezing and dynamic loss weighting mechanisms to improve model generalization. Comparative performance experiments demonstrate the superiority of the proposed model in various aspects, including inference latency (e.g., 6.798 ms per item on the D3 dataset), worst-case accuracy (91.364%), parameter efficiency, and training stability (fluctuation rate: 0.492). For the tokenization task, the model significantly outperforms baselines in metrics such as Strict-F1 (93.723 on D3), entity-level recall (93.764% on D2), and average span error (0.608 on D1). In the classification task, demonstrate that this joint modeling approach effectively supports fine-grained terminology recognition and semantic segmentation in TCM oncology texts. It provides a solid foundation for developing structured information extraction systems for TCM knowledge bases, follow-up records, and consultation notes aimed at tumor patient management and integrated clinical decision support.
This dataset contains the replication data for "Neonatal social communication and single genes predict the variability of post-pubertal social behavior in a mouse model of paternal 15q11-13 duplication" paper.
This dataset contains raw files from a project assessing applicability of CD90, CD105 and CD73 as signature markers for cultured murine bone marrow-derived mesenchymal stem cell as well as their utility for in situ identification of breast tumor recruited mesenchymal stem cells. The project proposes podoplanin as an effective marker for both in vitro and in vivo characterization of bone marrow derived mesenchymal stem cells. We are also exploring potential function of podoplanin expression on these cells once recruited to breast tumors.
Transcriptomes of individual neurons provide rich information about cell types and dynamic states. However, it is difficult to capture rare dynamic processes, such as adult neurogenesis, because isolation from dense adult tissue is challenging, and markers for each phase are limited. Here, Applicants developed Nuc-seq, Div-Seq, and Dronc-Seq. Div-seq combines Nuc-Seq, a scalable single nucleus RNA-Seq method, with EdU-mediated labeling of proliferating cells. Nuc-Seq can sensitively identify closely related cell types within the adult hippocampus. Div-Seq can track transcriptional dynamics of newborn neurons in an adult neurogenic region in the hippocampus. Dronc-Seq uses a microfluidic device to co-encapsulate individual nuclei in reverse emulsion aqueous droplets in an oil medium together with one uniquely barcoded mRNA-capture bead. Finally, Applicants found rare adult newborn GABAergic neurons in the spinal cord, a non-canonical neurogenic region. Taken together, Nuc-Seq, Div-Seq and Dronc-Seq allow for unbiased analysis of any complex tissue.
Methods involving analysis of pretreatment leukocyte expression profiles for prognostic assessment of seizure outcome following a treatment or medical procedure, such as stereotactic laser amygdalohippocampotomy (SLAH). In one aspect, RNA sequencing (RNA-Seq) on whole blood leukocyte samples is taken from a patient with intractable epilepsy prior to SLAH. Differential expression (DE) analysis revealed 24 significantly dysregulated genes (≥2.0-fold change, p-value <0.05, and False Discovery Rate, FDR <0.05) useful in prognostic assessment.
MASSACHUSETTS INST TECHNOLOGY et al. · granted 2024-05-28
Disclosed here is a generally applicable framework that utilizes massively-parallel single-cell RNA-seq to compare cell types/states found in vivo to those of in vitro models. Furthermore, Applicants leverage identified discrepancies to improve model fidelity. Applicants uncover fundamental gene expression differences in lineage-defining genes between in vivo systems and in vitro systems. Using this information, molecular interventions are identified for rationally improving the physiological fidelity of the in vitro system. Applicants demonstrated functional (antimicrobial activity, niche support) improvements in Paneth cell physiology using the methods.
Understanding the complex effects of genetic perturbations on cellular state and fitness in human pluripotent stem cells (hPSCs) has been challenging using traditional pooled screening techniques which typically rely on unidimensional phenotypic readouts. Here, Applicants use barcoded open reading frame (ORF) overexpression libraries with a coupled single-cell RNA sequencing (scRNA-seq) and fitness screening approach, a technique we call SEUSS (ScalablE fUnctional Screening by Sequencing), to establish a comprehensive assaying platform. Using this system, Applicants perturbed hPSCs with a library of developmentally critical transcription factors (TFs), and assayed the impact of TF overexpression on fitness and transcriptomic cell state across multiple media conditions. Applicants further leveraged the versatility of the ORF library approach to systematically assay mutant gene libraries and also whole gene families. From the transcriptomic responses, Applicants built genetic co-perturbation networks to identify key altered gene modules. Strikingly, we found that KLF4 and SNAI2 have opposing effects on the pluripotency gene module, highlighting the power of this method to characterize the effects of genetic perturbations. From the fitness responses, Applicants identified ETV2 as a driver of reprogramming towards an endothelial-like state.
SPRING VALLEY AGRISCIENCE CO LTD · granted 2023-10-24
The present invention provides a novel gene WMS conferring wheat male sterility, its anther-specific expression promoter, and uses of the same. In wheat, a well-known gene Ms2 causing dominant male sterility has been widely applied in recurrent selection in China. A RNA-seq approach was performed to reveal the anther-specific transcriptome in a pair of Ms2 isogenic lines, ‘Lumai 15’ and ‘Lumai 15+Ms2’. As a result, a WMS gene was identified showing anther-specific expression at the early stage of meiosis and only in wheat carrying the Ms2 gene. The regulation of WMS could alter plant male fertility. The promoter of WMS was found to comprise anther-specific activity. Thus, the present invention might be used to achieve anther-specific gene expression, to develop male sterility in various plant species, to establish recurrent selection in various plant species, and to assist hybrid seed production.
Techniques are provided for replacing image assays using real world data and real word evidence RNA-seq analysis for assessing biologic pathways for identifying molecular subtypes. Systems of a methods diagnose HER2 status for a patient, by identifying discordant HER2 status result between the HER2 status from immunohistochemistry (IHC) and the HER2 status from fluorescence in-situ hybridization (FISH) and diagnosing HER2 status based gene expression data.
Disclosed are methods of diagnosis of a pathogen-associated disease. These methods comprise: providing a biological sample from a human subject; determining presence, absence and/or quantity of a bacterial pathogen, a viral pathogen, or a combination thereof, by a pathogen culture, a serum antibody detection test, a pathogen antigen detection test, a pathogen DNA and/or RNA detection test, or a combination thereof; determining in the sample, expression levels of at least one endogenous gene in which aberrant expression levels are associated with infection with a pathogen, by a microarray hybridization assay, RNA-seq assay, polymerase chain reaction assay, a LAMP assay, a ligase chain reaction assay, a Southern, Northern, or Western blot assay, an ELISA or a combination thereof. The subject can be diagnosed with the disease if the subject comprises the pathogen and an aberrant level of expression of an endogenous gene.
Provided herein are compositions, kits and methods for the production of strand-specific cDNA libraries. The compositions, kits and methods utilize properties of double stranded polynucleotides, such as RNA-cDNA duplexes to capture and incorporate a novel sequencing adapter. The methods are useful transcriptome profiling by massive parallel sequence, such as full-length RNA sequencing (RNA-Seq) and 3′ tag digital gene expression (DGE).
The present invention provides a novel gene WMS conferring wheat male sterility, its anther-specific expression promoter, and uses of the same. In wheat, a well-known gene Ms2 causing dominant male sterility has been widely applied in recurrent selection in China. A RNA-seq approach was performed to reveal the anther-specific transcriptome in a pair of Ms2 isogenic lines, ‘Lumai 15’ and ‘Lumai 15+Ms2’. As a result, a WMS gene was identified showing anther-specific expression at the early stage of meiosis and only in wheat carrying the Ms2 gene. The regulation of WMS could alter plant male fertility. The promoter of WMS was found to comprise anther-specific activity. Thus, the present invention might be used to achieve anther-specific gene expression, to develop male sterility in various plant species, to establish recurrent selection in various plant species, and to assist hybrid seed production.
Techniques are provided for predicting DNA accessibility. DNase-seq data files and RNA-seq data files for a plurality of cell types are paired by assigning DNase-seq data files to RNA-seq data files that are at least within a same biotype. A neural network is configured to be trained using batches of the paired data files, where configuring the neural network comprises configuring convolutional layers to process a first input comprising DNA sequence data from a paired data file to generate a convolved output, and fully connected layers following the convolutional layers to concatenate the convolved output with a second input comprising gene expression levels derived from RNA-seq data from the paired data file and process the concatenation to generate a DNA accessibility prediction output. The trained neural network is used to predict DNA accessibility in a genomic sample input comprising RNA-seq data and whole genome sequencing for a new cell type.
Techniques are provided for predicting DNA accessibility. DNase-seq data files and RNA-seq data files for a plurality of cell types are paired by assigning DNase-seq data files to RNA-seq data files that are at least within a same biotype. A neural network is configured to be trained using batches of the paired data files, where configuring the neural network comprises configuring convolutional layers to process a first input comprising DNA sequence data from a paired data file to generate a convolved output, and fully connected layers following the convolutional layers to concatenate the convolved output with a second input comprising gene expression levels derived from RNA-seq data from the paired data file and process the concatenation to generate a DNA accessibility prediction output. The trained neural network is used to predict DNA accessibility in a genomic sample input comprising RNA-seq data and whole genome sequencing for a new cell type.
Systems and methods for in situ laser lysis for analysis of biological tissue (live, fixed, frozen or otherwise preserved) at single cell resolution in 3D. For example, a system and method for lysing individual cells in situ, including the steps of capturing a tissue sample comprising a cellular content, subjecting the tissue sample to a stream of continuous fluid flow, lysing a selected area of the tissue sample with a laser, thereby releasing at least a portion of the cellular content from the tissue sample, recovering at least one target molecule from the cellular content in the stream, and processing at least one target molecule is provided. The system collects cellular contents, performs highly multiplexed (RT-qPCR or RNA-seq), and sequentially (cell-by-cell) reconstructs a 3D spatial map of mRNA expression of the tissue with a large number of genes. A 3D spatial map of the DNA, RNA, and/or proteins can be generated for each cell in the tissue.
declared
Scholarly works are unavailable this run - the index did not respond. The other tabs are unaffected.