0
settings
الوضع الليلي
moon
انماط الصفحة الرئيسية arrow
EN
1
المرجع الالكتروني للمعلوماتية

النبات

مواضيع عامة في علم النبات

الجذور - السيقان - الأوراق

النباتات الوعائية واللاوعائية

البذور (مغطاة البذور - عاريات البذور)

الطحالب

النباتات الطبية

الحيوان

مواضيع عامة في علم الحيوان

علم التشريح

التنوع الإحيائي

البايلوجيا الخلوية

الأحياء المجهرية

البكتيريا

الفطريات

الطفيليات

الفايروسات

علم الأمراض

الاورام

الامراض الوراثية

الامراض المناعية

الامراض المدارية

اضطرابات الدورة الدموية

مواضيع عامة في علم الامراض

الحشرات

التقانة الإحيائية

مواضيع عامة في التقانة الإحيائية

التقنية الحيوية المكروبية

التقنية الحيوية والميكروبات

الفعاليات الحيوية

وراثة الاحياء المجهرية

تصنيف الاحياء المجهرية

الاحياء المجهرية في الطبيعة

أيض الاجهاد

التقنية الحيوية والبيئة

التقنية الحيوية والطب

التقنية الحيوية والزراعة

التقنية الحيوية والصناعة

التقنية الحيوية والطاقة

البحار والطحالب الصغيرة

عزل البروتين

هندسة الجينات

التقنية الحياتية النانوية

مفاهيم التقنية الحيوية النانوية

التراكيب النانوية والمجاهر المستخدمة في رؤيتها

تصنيع وتخليق المواد النانوية

تطبيقات التقنية النانوية والحيوية النانوية

الرقائق والمتحسسات الحيوية

المصفوفات المجهرية وحاسوب الدنا

اللقاحات

البيئة والتلوث

علم الأجنة

اعضاء التكاثر وتشكل الاعراس

الاخصاب

التشطر

العصيبة وتشكل الجسيدات

تشكل اللواحق الجنينية

تكون المعيدة وظهور الطبقات الجنينية

مقدمة لعلم الاجنة

الأحياء الجزيئي

مواضيع عامة في الاحياء الجزيئي

علم وظائف الأعضاء

الغدد

مواضيع عامة في الغدد

الغدد الصم و هرموناتها

الجسم تحت السريري

الغدة النخامية

الغدة الكظرية

الغدة التناسلية

الغدة الدرقية والجار الدرقية

الغدة البنكرياسية

الغدة الصنوبرية

مواضيع عامة في علم وظائف الاعضاء

الخلية الحيوانية

الجهاز العصبي

أعضاء الحس

الجهاز العضلي

السوائل الجسمية

الجهاز الدوري والليمف

الجهاز التنفسي

الجهاز الهضمي

الجهاز البولي

المضادات الميكروبية

مواضيع عامة في المضادات الميكروبية

مضادات البكتيريا

مضادات الفطريات

مضادات الطفيليات

مضادات الفايروسات

علم الخلية

الوراثة

الأحياء العامة

المناعة

التحليلات المرضية

الكيمياء الحيوية

مواضيع متنوعة أخرى

الانزيمات

قم بتسجيل الدخول اولاً لكي يتسنى لك الاعجاب والتعليق.

Bioinformatic approaches to predicting genes and their functions, and gene annotation

المؤلف:  Strachan, T., & Read, A.

المصدر:  Human molecular genetics

الجزء والصفحة:  5th E, P215-218

2026-09-28

87

+

-

20

Most human genes had not been uncovered before the Human Genome Project. As the DNA sequences of chromosomes were being spewed out of the big sequencing centers they were scanned by powerful bioinformatics programs to predict genes that could subsequently be validated. Some programs sought to identify genes ab initio by using hidden Markov models to analyze the human sequence by itself. However, comparisons with genome sequence data from other organisms and from previously studied genes provided extremely valuable support in identifying human genes, and also often provided important functional clues as to what kind of product a predicted human gene might make. As functions became mapped to genes, collaborative efforts sought to annotate genes in a comprehensive, systematic way and the Gene Ontology project was developed to address the need for consistent descriptions of gene products across species and data bases, as described below.

Predicting genes and their functions in silico Genome sequence data can be interrogated by a suite of computer programs that seek novel genes by screening for general gene characteristics. Genes are transcribed into RNA, for example, and so confidence in candidate gene sequences is high when sequence homology searching (Box 7.4) reveals many highly-related expressed sequence tags (ESTs) and/or cDNA sequences in sequence databases. Vertebrate genes also often show altered base composition: in addition to having a significantly higher percentage of GC (%GC) than the genome average, they are frequently associated with CpG islands, regions around 1 kb in length that are both GC-rich and also have a significantly higher frequency of the dinucleotide CpG than the bulk of the genome. (We consider CpG islands in more detail in Chapter 9.) For protein-coding genes, three gene-associated facets have been especially exploited to identify novel genes, as listed below.

• Open reading frames (ORFs). Open reading frames are needed in long coding DNA sequences to make a protein. There are three possible reading frames for each of the two DNA strands. If we make the simplifying assumption that the three types of termination codon occur with equal frequencies, then in human DNA, with an average base composition of 41% GC, one might expect that a stop codon would occur by chance roughly once every 50 nucleotides or so in each of the six possible reading frames (3/64 codons are termination codons, but termination codons are AT-rich: seven out of the nine nucleotides in the three stop codons are A or T). Coding DNA has a significantly higher %GC, and so statistically longer ORFs can generally be expected in coding DNA (making the same assumptions as above, a stop codon might be expected to occur by chance once every 80 nucleotides or so in DNA regions with 50% GC). Added to that is the frequent occurrence of intervening introns: the coding DNA is usually split, and an average internal exon in a human protein-coding gene is about 150 nucleotides in length. Long ORFs (>300 nucleotides) become prioritized for follow-up investigations.

• Exon prediction. Programs such as GENSCAN exploit the observation that there are short conserved sequences at splice junctions and assign high probability to a predicted exon if there is also a large ORF.

 • Evolutionary conservation. Homology searching—comparing sequences at the nucleotide level, or a predicted protein against protein sequence databases or against translated nucleotide sequences—is an especially powerful tool for gene identification (see Box 7.4). That is so because functionally important sequences, such as proteins and coding DNA sequences, have been highly conserved during evolution. A recently discovered human gene would often be found to have previously characterized homologs in model organisms (including other vertebrates, invertebrates such as Drosophila and C. elegans, or even in microbial cells). Because protein sequences are more conserved than the corresponding DNA sequences, predicted translation products of a candidate gene are typically used to search for homology against all known protein sequences using BLASTP, and against the predicted translation products of all known nucleotide sequences using TBLASTN. Identifying related homologs in other species may provide valuable clues to the function of the human gene. Additionally, even if a related homolog is not found, homology searching may identify small subcomponents of a candidate gene that are suggestive of a function—for example, sequence motifs associated with DNA-binding properties, such as zinc fingers.

Integrated gene-finding software packages have been developed that combine programs designed to identify exons and gene-associated motifs with general sequence homology-based database-searching programs. Although the bioinformatics approaches have been of great help in predicting genes (and some aspects of their function), they have also had their limitations. Gene prediction in silico has been especially valuable in the case of protein-coding genes, but programs like GENSCAN do have a tendency toward overprediction.

The Gene Ontology (GO) Project

 Searching across databases (and the broad scientific literature) has been hampered by the sometimes wide variation in terminologies used in different species. The Gene Ontology (GO) Consortium was formed in 1998 to institute a standardized system of gene ontology to represent gene function across genomes and species, and the resulting GO project began to develop a controlled vocabulary to describe the attributes of genes and gene products in any organism. Three separate ontologies—biological process, cellular component, and molecular function—were developed to allow for the annotation of molecular characteristics across species, and they may be broad or more focused. For example, the biological process can be as broad as signal transduction or more restricted, such as alpha-glucoside transport. Each vocabulary is structured so that any term may have more than one parent as well as zero, one, or more children. This makes attempts to describe biology much richer than would be possible with a hierarchical graph. By 2014 the GO project had developed formal ontologies to represent more than 40,000 biological concepts.

اشترك بقناتنا على التلجرام ليصلك كل ما هو جديد