-
McDonald Vaughn posted an update 1 year, 7 months ago
Ribosome profiling shows potential for studying the function of long noncoding RNAs (lncRNAs). We introduce a bioinformatics pipeline for detecting ribosome-associated lncRNAs (ribo-lncRNAs) from ribosome profiling data. CC-90001 JNK inhibitor Further, we describe a machine-learning approach for the characterization of ribo-lncRNAs based on their sequence features. Scripts for ribo-lncRNA analysis can be accessed at ( https//ribolnc.hamadalab.com/ ).Single-cell analysis has contributed greatly to gaining a better understanding of human brain function and has implications for neurodegenerative and neuropsychiatric disorders. Long noncoding RNAs (lncRNAs) acting, in part, as epigenetic regulators exist in brain cells in high abundance exhibiting a large diversity that play important roles in neural development, function, and neurodegenerative disease. Due to lncRNA tissue-type and cell-type specific expression characteristics, it is important to analyze lncRNA at single-cell resolution. In this chapter, we highlight a method named scTISA (single-cell transcription in situ with antisense RNA amplification), which is applicable to fixed single cells and can yield polyA+ lncRNAs and mRNAs data at the same time.Metazoan genomes produce thousands of long-noncoding RNAs (lncRNAs), of which just a small fraction have been well characterized. Understanding their biological functions requires accurate annotations, or maps of the precise location and structure of genes and transcripts in the genome. Current lncRNA annotations are limited by compromises between quality and size, with many gene models being fragmentary or uncatalogued. To overcome this, the GENCODE consortium has developed RNA capture long-read sequencing (CLS), an approach combining targeted RNA capture with third-generation long-read sequencing. CLS provides accurate annotations at high-throughput rates. It eliminates the need for noisy transcriptome assembly from short reads, and requires minimal manual curation. The full-length transcript models produced are of quality comparable to present-day manually curated annotations. Here we describe a detailed CLS protocol, from probe design through long-read sequencing to creation of final annotations.While more than a hundred thousand long noncoding RNAs (lncRNAs) have been identified in human genome, their biological functions and regulation are largely elusive. Here we present AnnoLnc, a one-stop online annotation portal for human lncRNAs ( http//annolnc1.gao-lab.org/ ). As the first (and the most comprehensive) Web server to provide on-the-fly annotation for novel human lncRNAs, AnnoLnc exploits more than 700 data sources to annotate inputted lncRNA systematically, spanning genomic location, secondary structure, expression patterns, coexpression-based functional annotation, transcriptional regulation, miRNA interaction, protein interaction, genetic association, and evolution. Moreover, in addition to a user-friendly Web interface, AnnoLnc can also be integrated into existing pipelines by either a set of JSON-based web service APIs or a stand-alone version for Linux server.A number of difficulties exist when studying long noncoding RNAs (lncRNAs) from a biological standpoint. As it is uncertain what percentage of human lncRNAs play meaningful roles in biology or consists of transcriptional artifacts, one prominent challenge is to decide which lncRNAs to study out of a potential 70,000 putative lncRNA genes. Integration of GWAS and eQTL signals has led to the identification of functional genes for disease susceptibility (Barbeira et al., Nat Commun 9(1)1825, 2018). In this chapter we describe a protocol for building bioinformatic evidence for lncRNA and trait/disease association.The INFERNO method provides an integrative computational framework for characterizing the causal variants, tissue contexts, affected regulatory mechanisms, and target genes underlying noncoding genetic variants associated with any phenotype or disease of interest. Here we describe the computational steps required to run the full INFERNO pipeline on any dataset of interest.Long noncoding RNAs are well studied for their regulatory actions through interaction with DNA regulating biological roles of DNA, RNA, or protein. However, direct binding of lncRNA with DNA is rarely demonstrated in experiments. The present protocol explains genome wide computational strategies to choose lncRNAs that can bind directly to the chromatin by forming highly stable DNA-DNA-RNA triplexes. The chapter also focuses on biophysical methods that can be used to validate the computationally derived lncRNA-gene targets in vitro.K-mer based comparisons have emerged as powerful complements to BLAST-like alignment algorithms, particularly when the sequences being compared lack direct evolutionary relationships. In this chapter, we describe methods to compare k-mer content between groups of long noncoding RNAs (lncRNAs), to identify communities of lncRNAs with related k-mer contents, to identify the enrichment of protein-binding motifs in lncRNAs, and to scan for domains of related k-mer contents in lncRNAs. Our step-by-step instructions are complemented by Python code deposited in Github. Though our chapter focuses on lncRNAs, the methods we describe could be applied to any set of nucleic acid sequences.CPAT (Coding-Potential Assessment Tool) is a logistic regression model-based classifier that can accurately and quickly distinguish protein-coding and noncoding RNAs using pure linguistic features calculated from the RNA sequences. CPAT takes as input the nucleotides sequences or genomic coordinates of RNAs and outputs the probabilities p (0 ≤ p ≤ 1), which measure the likelihood of protein coding. Users can run CPAT online ( http//lilab.research.bcm.edu/cpat/ ) or from the local computers after installation. CPAT provides prebuilt logistic models to recognize RNAs originated from human (Homo sapiens), mouse (Mus musculus), zebrafish (Danio rerio), and fly (Drosophila melanogaster) genomes. Instructions on how to train models for other genomes are described in CPAT website ( http//rna-cpat.sourceforge.net/ ) and this chapter.

