This dataset contains quality-controlled SNP and InDel genotype calls from whole-genome resequencing of 299 purebred Large White pigs. The dataset includes 100 individuals sequenced at 30× depth (merged genome-wide VCFs) and 199 individuals sequenced at 10× depth (chromosome-split VCFs). All variants were called against the *Sus scrofa* Sscrofa11.1 reference genome and filtered using standard GATK hard-filtering criteria: QD < 2.0, QUAL < 30.0, FS > 200.0, ReadPosRankSum < -20.0, call rate < 90%, MAF < 0.05. This data supports the findings of the study "Whole-genome resequencing identifies RASAL2 as a candidate gene for feed efficiency in Large White pigs".
Gene expression data for comparing 3 time points following injuries to the human mandible. The whole peripheral blood cells were processed for RNA sequencing (Novogene).
Hernández Miranda, Olga Andrea · Campos Contreras, Jorge Eduardo · Rosas López, Ulises · et al.
The input datasets used in this study correspond to the files associated with each analytical step implemented in the repository: https://github.com/Andrea-H-M/SystemsLevel_transcriptomics/tree/main These inputs include DEG lists from Arabidopsis thaliana , Solanum lycopersicum , Vitis vinifera , and Vanilla planifolia ; FASTA sequence files used for orthology inference; gene identifier lists; transcription factor (TF) and epigenetic regulator (EpiReg) reference tables; expression matrices and annotated count tables used for WGCNA and hub gene identification; KEGG Orthology (KO) annotation files; binary transcriptional label matrices; and module-specific gene lists used for transcriptional label integration and candidate regulator prioritization. The repository structure preserves the workflow organization described in the Methods section of the manuscript, linking each script with its corresponding input files required to reproduce orthology inference, functional enrichment, co-expression network analyses, transcriptional label integration, and candidate regulator prioritization associated with the flower-to-fruit transition and post-pollination syndrome in Vanilla planifolia .
Więch-Walów, Anna · Barton, Anna · Tracz, Michał · et al.
Intro Aim of the experiment was to identify proteins interacting with 3'UTRs of XBP1 mRNA. Potential binding partners were affinity copurified from cytosolic fraction of cellular extracts using biotinilated mRNA baits. Sample legend batch 1 241205_ZBF_AWI_3.raw; spliced XBP1 3'UTR mRNA bait; tunicamycin treated cells 241205_ZBF_AWI_4.raw; unspliced XBP1 3'UTR mRNA bait; tunicamycin treated cells 241205_ZBF_AWI_N3.raw; spliced XBP1 3'UTR mRNA bait; untreated cells 241205_ZBF_AWI_N4.raw; unspliced XBP1 3'UTR mRNA bait; untreated cells batch 2 250124_ZBF_AWI_3T_MSe.raw; spliced XBP1 3'UTR mRNA bait; tunicamycin treated cells 250124_ZBF_AWI_4T_MSe.raw; unspliced XBP1 3'UTR mRNA bait; tunicamycin treated cells 250124_ZBF_AWI_3N_MSe.raw; spliced XBP1 3'UTR mRNA bait; untreated cells 250124_ZBF_AWI_4N_MSe.raw; unspliced XBP1 3'UTR mRNA bait; untreated cells Filename legend 241206_ZBF_AWI_qip_proteins.csv; batch 1 QI output protein list 250127_ZBF_AWI_qip_proteins.csv; batch 2 QI output protein list raw.zip archives are MassLynx v4.2 raw files from each acquistion 7z archives are open-vendor format spectra (.mgf) and identifications (.mzid) from each batch RNA-bait pull-down HeLa S3 cells were cultured in 100 mm dishes (10 dishes per condition; no stress (NS) or Tm-treated). Cells were washed twice with ice cold PBS (Biowest, L0616), scraped into 1 ml PBS per dish, and collected by centrifugation (5 min, 300 × g, 4 °C). After removing PBS, the packed cell volume (PCV) was determined. Cell pellets were resuspended in five volumes of lysis buffer (10 mM HEPES, 10 mM KCl, 1 mM MgCl₂, 1 mM DTT, and 1x Protease Inhibitor cOmplete Mini (11836170001; Roche), pH 7.4) and incubated on ice for 20 min. Cells were lysed by passing the suspension 20 times through a syringe needle (G27). Lysates were clarified by ultracentrifugation (1 h, 100,000 × g, 4 °C), and protein concentration of the resulting cytoplasmic fraction was determined by A₂₈₀ absorbance. For RNA pulldown assays, 1 µg of biotinylated RNA bait was incubated with 500 µg of cytoplasmic extract for 1 h at room temperature in the presence of 20 U RiboLock. Streptavidin magnetic beads (50 µl; Pierce, 88816) were added and incubated for an additional 1 h at room temperature with gentle agitation. Beads were collected using a magnetic rack and washed twice with wash buffer (0.6 M NaCl, 20 mM Tris, 10 mM EDTA, 1 mM DTT, 1× PIC, pH 7.4). Mass spectrometry After the final wash, resin-bound proteins were denatured for 10 mins. in 65°C in presence of 0.5% sodium deoxycholate (DOC:Na, Merck). Next, 100 ng of Trypsin (EMS0006, Merck) was added to the samples for an overnight digestion in 37°C. On the following day, the peptide solution was separated from the resin, DOC:Na was removed via acidification and centrifugation, and the supernatants underwent desalting with the STAGE tip protocol. The obtained peptide pellet was resuspended in a 0.1% formic acid (FA) + 3% acetonitrile (ACN) solution. LC-MS was carried out on an M-Class Acquity nanoUPLC system coupled to a Synapt XS HDMS equipped with a nanoESI ion source interface. Mobile phase A consisted of H2O + 0.1% FA, while mobile phase B of ACN + 0.1% FA. Samples were separated on an analytical column (nanoEase HSS T3 C18 100Å, 1.8μm, 75μm x 150mm) kept at 40°C with a 40min linear gradient of 5-35%B at a 300nL/min flow rate. A 3min trapping step was performed prior to gradient separation. Injected volume was 3µL. MS data were collected in DIA mode, scanning at 2Hz through a 50-2000m/z range in ESI+ and analyser mode set to Resolution (25000 FWHM at 785.84m/z). For MS2 of all ions, a collision energy ramp of 25-45V was set on instrument's trap cell. Source conditions were fine-tuned and kept constant during acquisitions. A Glu-Fib (Waters) solution was acquired in the mass reference function, and the correction was applied post-acquisition. Raw processing was performed using Progenesis QI for Proteomics v4.2.7 (Non-linear Dynamics). QI''s autovalues were unchanged for the deconvolution, and the set low and high energy ion thresholds were 200 and 20, respectively. Deconvoluted data were searched via Ion Accounting against a redundant human protein sequence databank (UP000005640, 83526 sequences) appended with the CCP version of cRAP contaminant sequences. The search parameters were as follows; peptide/fragment mass tolerance: 10/20 ppm; min. fragments/peptide: 3; min. fragments/protein: 7; min. peptides/protein: 1; max. protein mass: 1MDa; enzyme: trypsin; max. missed cleavages: 4; variable modification: oxidation of methionine; FDR: 4% (protein-level). Protein grouping was enabled, and relative quantitation was set according to non-conflicted peptides. Lists containing uniquely identified and quantified proteins from two independent experiments (SF1, SF2) were exported for further analysis.
We collected alder swamp soil from North Zealand, Denmark, and used soil-derived suspensions to inoculate organic agricultural soil for the cultivation of Arabidopsis thaliana ecotype Columbia-0 (Col-0). For bacterial 16S rRNA gene amplicon sequencing, the V5-V7 region was amplified using the 799F and 1193R primer pair. Amplicon libraries from 133 samples were subjected to paired-end sequencing (2 × 250 bp) on an Illumina NovaSeq 6000 platform.
TASTANOVA, AIZHAN · Staeger, Ramon · Mitchell, Levesque
This dataset contains processed Seurat (v5) objects generated from single-cell CITE-seq analysis of human skin biopsies collected from patients with actinic keratosis (AK), UV-exposed normal skin (UVES), and non-UV-exposed normal skin (NUVES). The objects contain normalized gene expression data, integrated cell metadata, dimensionality reductions, clustering, and cell-type annotations used in the analyses presented in the accompanying manuscript. The dataset supports the identification of AK-specific keratinocytes (ASK), multimodal characterization of keratinocyte populations, differential gene expression analyses, gene signature scoring, regulatory network inference, copy number inference, and integration with spatial transcriptomics.
Pescador-Dionisio, Sara · García-Robles, Inmaculada
This repository contains the supplementary datasets, computational notebooks and structural prediction files generated for the study "A TSS-aware catalogue of microRNA-encoded peptides in tomato ( Solanum lycopersicum )" . The dataset supports the identification, annotation, structural characterization and comparative analysis of candidate microRNA-encoded peptides (miPEPs) associated with tomato pri-miRNAs. The repository is organised into the following files: TableS1.xlsx: Metadata and sequence features of the curated tomato miPEP catalog. TableS2.xlsx: Predicted ORFs associated with miRNA loci in seven plant species. TableS3.xlsx: Exact amino-acid k-mer matches detected between tomato miPEPs and peptide datasets from the surveyed non-tomato species. TableS5.xlsx: Physicochemical values and statistical summaries. TableS6.xlsx: Strand-aware RNA-seq support for MIR loci screened for miPEP prediction in tomato leaves and roots. description_tables_full.txt: full description of the tables included. eggplant_premiR_candidates_BLAST_EN.ipynb: Homology-based identification of putative eggplant pre-miRNA loci through BLAST alignment of known Solanaceae pre-miRNAs, followed by extraction of candidate genomic regions. generic_pri_miRNA_upstream_from_GFF3_chr_to_NC_FIXED.ipynb: Extraction of strand-aware upstream sequences from annotated pri-miRNA loci using genome and GFF3 files, with automated chromosome-to-genome identifier matching. miPEPs_ORFs_kmer_overlap_colab.ipynb: Detection of exact amino-acid k-mer matches (6-30 aa) between predicted ORF peptides and the curated tomato miPEP dataset for comparative conservation analyses. orfs_from_fasta_ATG_6_200aa_v2_EN.ipynb: Identification and extraction of all ATG-initiated ORFs (6-200 amino acids) from DNA FASTA sequences, including optional reverse-complement strand analysis and peptide FASTA generation. upstream_from_miRBase_hairpin_list_noGFF3_EN.ipynb: Extraction of upstream genomic regions from miRBase hairpin sequences when genomic annotations (GFF3 files) are unavailable, using genome mapping to infer precursor coordinates. mipeps_alphafold_prediction.zip: AlphaFold2 structural predictions for all 107 tomato miPEPs, including structure files, confidence metrics and pLDDT-coloured structural visualizations generated from the AlphaFold2 models. mipepORFS_mappingdata.zip: BAM alignment files and their associated BAI index files generated and used for visual inspection in IGV (Integrative Genomics Viewer) of genomic regions corresponding to putative miPEPs investigated in this study. The dataset includes files from four leaf samples and four root samples.The BAM files contain sequencing reads aligned to the reference genome used in the analysis. These files preserve the information required to visualize: the genomic position of the reads, the coverage across each locus, the orientation of the alignments File use To correctly visualize the alignments from mipepORFS_mappingdata.zip, the files should be opened in IGV together with: • the corresponding reference genome: Solanum lycopersicum reference genome sequence, assembly SL3.0 • the mipep annotation track used in the study: mipeps_prediction_SL3.0.gtf, included in the file • For improved visualization and interpretation in IGV, the gene annotation file Solanum_lycopersicum.SL3.0.57.chr.gtf can also be loaded together with the BAM and BAI files. The tomato reference genome sequence (Solanum lycopersicum assembly SL3.0) and the corresponding gene annotation file can be downloaded from Ensembl Plants.
Dataset of the publication "Single-cell RNA-seq using UltraMarathonRT expands the known transcriptome". This is the updated version of the previous BioArxiv preprint "Single-cell RNA-seq using UltraMarathonRT expands the known transcriptome", https://doi.org/10.1101/2025.10.06.680646 Included is the following data: R scripts with code for every main and supplemental figure Bash scripts for the preprocessing of the data Count tables obtained after featureCounts Prepared RDS objects GENCODE 47 annotation files used in the study GENCODE47 + SIRVOme annotation file used in the study Scripts for TSO concatemer counting Raw fastq files are found here: HeLa S3 : https://doi.org/10.5281/zenodo.20812770 K562 : https://doi.org/10.5281/zenodo.20812231
These data contain records of statutes for each count of conviction for criminal defendants who were sentenced pursuant to provisions of the Sentencing Reform Act (SRA) of 1984 and reported to the United States Sentencing Commission (USSC) during fiscal year 2023. The data are one of two supplementary files that should be used in conjunction with the primary analysis file, which contains records for all defendants sentenced under the guidelines. These data can be linked to the primary analysis file using the unique identifier variable SEQ_NUM. The number of records for a defendant in the current data corresponds to the total number of counts of conviction for that defendant, and that total is recorded in the NOCOUNT variable. As an example, if a defendant has five counts of conviction (NOCOUNT=5), he or she will have five records in the current data. As it is possible for defendants to have multiple statutes applying to a single count of conviction, up to three statutes (STA1-STA3) are recorded for each count of conviction. The data were obtained from the United States Sentencing Commission's Office of Policy Analysis' (OPA) Standardized Research Data File. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
The data contain records of charges filed against defendants whose cases were terminated by United States attorneys in United States district court during fiscal year 2023. The data are charge-level records, and more than one charge may be filed against a single defendant. The data were constructed from the Executive Office for United States Attorneys (EOUSA) Central Charge file. The charge-level data may be linked to defendant-level data (extracted from the EOUSA Central System file) through the CS_SEQ variable, and it should be noted that some defendants may not have any charges other than the lead charge appearing on the defendant-level record. The Central Charge and Central System data contain variables from the original EOUSA files as well as additional analysis variables. Variables containing identifying information (e.g., name, Social Security Number) were either removed, coarsened, or blanked in order to protect the identities of individuals. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
The data contain records of charges filed against defendants whose cases were filed by United States attorneys in United States district court during fiscal year 2023. The data are charge-level records, and more than one charge may be filed against a single defendant. The data were constructed from the Executive Office for United States Attorneys (EOUSA) Central Charge file. The charge-level data may be linked to defendant-level data (extracted from the EOUSA Central System file) through the CS_SEQ variable, and it should be noted that some defendants may not have any charges other than the lead charge appearing on the defendant-level record. The Central Charge and Central System data contain variables from the original EOUSA files as well as additional analysis variables. Variables containing identifying information (e.g., name, Social Security Number) were either removed, coarsened, or blanked in order to protect the identities of individuals. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
These data contain records of guideline computations and adjustments for each count of conviction for criminal defendants who were sentenced pursuant to provisions of the Sentencing Reform Act (SRA) of 1984 and reported to the United States Sentencing Commission (USSC) during fiscal year 2010. The data are one of two supplementary files that should be used in conjunction with the primary analysis file, which contains records for all defendants sentenced under the guidelines. These data can be linked to the primary analysis file using the unique identifier variable SEQ_NUM. The number of records for a defendant in the current data corresponds to the total number of guideline computations, which may or may not equal the total counts of conviction for that defendant, dependent upon the grouping rules of the particular guideline in question (see Section 3D1.2 of the guidelines manual). As an example, a defendant with five counts of drug trafficking will only have one guideline computation because each of the drug weights for each count are simply added together and only one calculation is necessary. However, if a defendant has five counts of bank robbery, he or she will have five separate guideline computations because bank robbery is considered to be a nongroupable offense. The data were obtained from the United States Sentencing Commission's Office of Policy Analysis' (OPA) Standardized Research Data File. These data are part of a series designed by the Urban Institute (Washington, DC) and the Bureau of Justice Statistics. Data and documentation were prepared by the Urban Institute.
These data contain records of statutes for each count of conviction for criminal defendants who were sentenced pursuant to provisions of the Sentencing Reform Act (SRA) of 1984 and reported to the United States Sentencing Commission (USSC) during fiscal year 2022. The data are one of two supplementary files that should be used in conjunction with the primary analysis file, which contains records for all defendants sentenced under the guidelines. These data can be linked to the primary analysis file using the unique identifier variable SEQ_NUM. The number of records for a defendant in the current data corresponds to the total number of counts of conviction for that defendant, and that total is recorded in the NOCOUNT variable. As an example, if a defendant has five counts of conviction (NOCOUNT=5), he or she will have five records in the current data. As it is possible for defendants to have multiple statutes applying to a single count of conviction, up to three statutes (STA1-STA3) are recorded for each count of conviction. The data were obtained from the United States Sentencing Commission's Office of Policy Analysis' (OPA) Standardized Research Data File. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
The data contain records of charges filed against defendants whose cases were terminated by United States attorneys in United States district court during fiscal year 2022. The data are charge-level records, and more than one charge may be filed against a single defendant. The data were constructed from the Executive Office for United States Attorneys (EOUSA) Central Charge file. The charge-level data may be linked to defendant-level data (extracted from the EOUSA Central System file) through the CS_SEQ variable, and it should be noted that some defendants may not have any charges other than the lead charge appearing on the defendant-level record. The Central Charge and Central System data contain variables from the original EOUSA files as well as additional analysis variables. Variables containing identifying information (e.g., name, Social Security Number) were either removed, coarsened, or blanked in order to protect the identities of individuals. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
The data contain records of charges filed against defendants whose cases were filed by United States attorneys in United States district court during fiscal year 2022. The data are charge-level records, and more than one charge may be filed against a single defendant. The data were constructed from the Executive Office for United States Attorneys (EOUSA) Central Charge file. The charge-level data may be linked to defendant-level data (extracted from the EOUSA Central System file) through the CS_SEQ variable, and it should be noted that some defendants may not have any charges other than the lead charge appearing on the defendant-level record. The Central Charge and Central System data contain variables from the original EOUSA files as well as additional analysis variables. Variables containing identifying information (e.g., name, Social Security Number) were either removed, coarsened, or blanked in order to protect the identities of individuals. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
EL Southeast conducted a randomized control trial in the Mississippi Delta to test the impact of a kindergarten vocabulary instruction program on students' expressive vocabulary—the words students understand well enough to use in speaking. The study, The Effectiveness of a Program to Accelerate Vocabulary Development in Kindergarten, found that the 24-week K-PAVE program had a significant positive impact on students' vocabulary development and academic knowledge and on the vocabulary and comprehension support that teachers provided during book read-alouds and other instructional time. K-PAVE is designed to build children's vocabulary and comprehension skills, oral language skills, and enhance teacher-child relationships. K-PAVE is one of only a few kindergarten-age-appropriate vocabulary interventions and the only intervention with teacher training materials. An existing preschool version of K-PAVE had already demonstrated some evidence of positive effects from an impact study. The K-PAVE intervention group included 64 schools, 128 kindergarten classrooms and teachers, and 1,296 kindergarten students (596 treatment and 700 control students). Other findings include: Kindergarteners who received the K-PAVE intervention were one month further ahead in vocabulary development and academic knowledge at the end of kindergarten compared with their peers who did not receive the intervention. However, there were no statistically significant differences between the two groups on listening comprehension. The kindergarten teachers trained in the program were significantly more likely than their peers who did not receive K-PAVE training to provide vocabulary and comprehension support to students during book read alouds and other instructional times (e.g., providing background information; making connections to students' experiences; asking students to analyze, explain, make inferences; introducing vocabulary words). The program did not produce a statistically significant impact on either instructional support or emotional support (promoted during the K-PAVE relationship-building conversations) in the classroom. Additionally, the program did not impact the amount of instructional time spent on literacy in areas other than vocabulary and comprehension.
The 2018 NAEP Oral Reading Fluency (ORF) study was conducted by NCES to examine public-school fourth-grade students’ ability to read passages out loud with sufficient speed, accuracy, and expression, as well as foundational skills, to gauge underlying sources of poor fluency. The restricted-use data file contains oral reading performance variables, student contextual information, responses collected from the NAEP reading student survey questionnaire, NAEP reading scores and achievement level and below NAEP Basic subgroup variables, and student sample weights and replicate weights. The data companion provides background on the 2018 NAEP Oral Reading Fluency (ORF) study and a description of oral reading fluency and the study design. It also describes the 2018 NAEP ORF study instrument development and provides the sampling design and response rates, variables included in the data file, information needed to conduct statistical analyses using the 2018 NAEP ORF data, and control statement files to create system files in SAS and Stata.
These data contain records of criminal defendants who were sentenced pursuant to provisions of the Sentencing Reform Act (SRA) of 1984 and reported to the United States Sentencing Commission (USSC) during fiscal year 2021. These data can be linked to the primary analysis file using the unique identifier variable SEQ_NUM. It is estimated that over 90 percent of felony defendants in the federal criminal justice system are sentenced pursuant to the SRA of 1984. The data were obtained from the United States Sentencing Commission's Office of Policy Analysis' (OPA) Standardized Research Data File. The Standardized Research Data File consists of variables from the Monitoring Department's database, which is limited to those defendants whose records have been furnished to the USSC by United States district courts and United States magistrates, as well as variables created by the OPA specifically for research purposes. The data include variables from the Judgment and Conviction (J and C) order submitted by the court, background and guideline information collected from the Presentencing Report (PSR), and the report on sentencing hearing in the Statement of Reasons (SOR). These data contain detailed information such as the guideline base offense level, offense level adjustments, criminal history, departure status, statement of reasons given for departure, and basic demographic information. These data are the primary analysis file and include only statute, guideline computation, and adjustment variables for the most serious offense of conviction. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
The data contain records of charges filed against defendants whose cases were filed by United States attorneys in United States district court during fiscal year 2021. The data are charge-level records, and more than one charge may be filed against a single defendant. The data were constructed from the Executive Office for United States Attorneys (EOUSA) Central Charge file. The charge-level data may be linked to defendant-level data (extracted from the EOUSA Central System file) through the CS_SEQ variable, and it should be noted that some defendants may not have any charges other than the lead charge appearing on the defendant-level record. The Central Charge and Central System data contain variables from the original EOUSA files as well as additional analysis variables. Variables containing identifying information (e.g., name, Social Security Number) were either removed, coarsened, or blanked in order to protect the identities of individuals. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.
These data contain records of statutes for each count of conviction for criminal defendants who were sentenced pursuant to provisions of the Sentencing Reform Act (SRA) of 1984 and reported to the United States Sentencing Commission (USSC) during fiscal year 2020. The data are one of two supplementary files that should be used in conjunction with the primary analysis file, which contains records for all defendants sentenced under the guidelines. These data can be linked to the primary analysis file using the unique identifier variable SEQ_NUM. The number of records for a defendant in the current data corresponds to the total number of counts of conviction for that defendant, and that total is recorded in the NOCOUNT variable. As an example, if a defendant has five counts of conviction (NOCOUNT=5), he or she will have five records in the current data. As it is possible for defendants to have multiple statutes applying to a single count of conviction, up to three statutes (STA1-STA3) are recorded for each count of conviction. The data were obtained from the United States Sentencing Commission's Office of Policy Analysis' (OPA) Standardized Research Data File. These data are part of a series designed by Abt Associates and the Bureau of Justice Statistics. Data and documentation were prepared by Abt Associates.