Wednesday, February 8, 2012

Must a Gene Have a Function?

Biology is such a messy subject.1 It's impossible to come up with simple definitions of fundamental concepts in biology because there are exceptions to everything. In the case of "gene," there are so many exceptions that it seems hopeless to propose a general definition of such an important term. Nevertheless, we need some basic ground rules to prevent the situation from getting out-of-hand.

In an earlier posting from 2007 [What Is a Gene?], I suggested the following ...
This essay describes various modern definitions of physical genes (Gene-D). I like to define a gene as “a DNA sequence that’s transcribed” but that’s a bit too brief for a formal definition. We need to include something that restricts the definition of gene to those entities that are biologically significant. Hence,

A gene is a DNA sequence that is transcribed to produce a functional product.

This eliminates those parts of the chromosome that are transcribed by accident or error. These regions are significant in large genomes; in fact, the confusion between accidental transcripts and real transcripts is responsible for the overestimates of gene number in many genome projects. (In technical parlance, most ESTs are artifacts and the sequences they come from are not genes.)
Let's not quibble about all of the exceptions. Most of them are covered in my original article and in the comments there. I want to concentrate here on the idea that a gene has to have a "function" of some sort. As I explained in the comments ....
I don't know if I can come up with a catchy definition of "function." What I mean is that the transcript or it's product has to do some biochemical duty in order to qualify. It doesn't have to be an essential function but it has to make a difference of some sort.
This is important because there's a growing tendency to label all kinds of things as "genes" just because they produce small RNA molecules or, in some cases, a small protein. In most cases the products have no known biological function.

Here's a couple of examples.

De Novo Ptotein-Encoding Genes

It's plainly obvious that new genes must arise from time to time in various lineages. Lot's of people are interested in the evolution of humans and in particular the changes that distinguish us from our closest cousins. Almost all of the changes can be explained by alterations in the timing or location of orthologous gene expression but that doesn't exclude the possibility that entirely new genes might arise de novo in some lineages.

Let's just think about genes that encode proteins. There are three steps required for the de novo creation of a new protein-encoding gene. (1) A part of the ancestral genome must be transcribed. (2) The transcript must contain an open reading frame with a start and stop codon. (3) The new protein must have a function.

That last step needs explaining. If the new protein doesn't have a function then the putative new gene is no different than a pseudogene or a mutant gene that produces a truncated protein because of a premature stop codon. It's also indistinguishable from bits of the genome that are accidentally transcribed and just happen to have an open reading frame.

Wu et al. (2010) looked at the evolution of new genes in the lineage leading to humans. The title of their paper is: "De Novo Origin of Human Protein-Coding Genes." I want to challenge their definition of "gene" by suggesting that what they've really discovered are "potential" or "candidate" genes that don't deserve to be called "genes" until one discovers a biological function for their products.

The authors searched the human genome (build 56) for annotated "genes" with small open reading frames greater that 100 codons long. Then they examined the corresponding loci in the chimpanzee and orangutan genomes looking for case where there was no open reading frame in the other apes. Various expressed RNA databases and two expressed peptide databases were screened to see if the candidate genes were expressed as RNA and protein. They found 27 examples. These are the candidates for de novo genes in humans.

Their collection did not contain some of the de novo "genes" reported by others. As it turns out, those "genes" were annotated in previous versions of the human genome (builds 40-55) but were dropped from the latest versions because there were no homologues in the other ape genomes. By using those older builds, Wu et al. discovered another 33 candidates for a total of 60 putative new protein-encoding genes in the human genome.

Wu et al. concede that the expression levels of these candidate genes are "very low" but unfortunately they don't give us any specific levels. This is important because there's plenty of evidence that the expressed RNA databases contain spurious transcripts [How to Evaluate Genome Level Transcription Papers].

I wonder how many spurious peptides are in the peptide databases? Wu et al report that one of the peptides used to identify an earlier example of a de novo gene (Knowles and McLysaght, 2009) has been removed from the current build of PeptideAtlas. What happened to it?

The authors are aware of the fact that function is important, especially if they want to argue that these new genes conferred some selective advantage on our hominid ancestors. The only "evidence" they offer is that the putative genes are expressed at a low level in testis and brains but at an even lower level in other tissues. This is no evidence at all since we've known for fifty years that the complexity of RNA sequences in brain and testis is much higher than in other tissues. We still don't know whether that's due to elevated spurious transcription in those tissues of whether it is biologically significant.

Are these 60 candidates really new "protein-coding genes"? I don't think so. I don't think they can be called "genes" until it has been demonstrated that the products have a biological function. Guerzoni and McLysaght (2010) seem to agree because they write,
The observation by Wu et al. that some of the candidate de novo genes are expressed at their highest in brain tissues and testis is interesting, but by no means proves they are functional. A major challenge remains to demonstrate functionality of the de novo genes.

Genes that Encode Functional RNAs

The people who annotate the human genome are somewhat skeptical of these new genes and that's why so many putative genes have disappeared from the more recent builds. (But the Ensembl group still lists 434 "novel protein-coding genes.")

However, they don't seem to be as skeptical when it comes to genes that produce small RNAs. The most recent Ensembl build (GRCh37.p5, Feb 2009), for example, lists 12,523 RNA genes [Ensembl: Human Genome].

What are the criteria they use to prove that these are really genes? It can't have anything to do with biological function since it's simply not true that the human genome contains more that twelve thousand genes that produce an RNA whose function has been demonstrated.

Should that be a requirement before declaring that a bit of transcribed DNA is a gene? You're damn right it should because otherwise every bit of DNA that's accidentally transcribed in some tissue at some time during development qualifies as a gene. That makes no sense [What is a gene, post-ENCODE?].


1. That's why it's much more difficult than physics where there's talk about unifying the entire discipline under a single theory of everything. :-)

Guerzoni D, McLysaght A. (2011) De novo origins of human genes. PLoS Genet. 2011 Nov;7(11):e1002381. Epub 2011 Nov 10. [PLoS Genetics]

Knowles, D.G. and McLysaght, A. (2009) Recent de novo origin of human protein-coding genes. Genome Res. 19:1752-1759. PLoS Genet. 2011 Nov;7(11):e1002379. Epub 2011 Nov 10. [doi: 10.1101/gr.095026.109]

Wu, D.D., Irwin, D.M., and Zhang, Y.P. (2010) De novo origin of human protein-coding genes. [PLoS Genetics]

Monday, February 6, 2012

DM Jean Bingen, 26 mars 1920-6 fevrier 2012

How Much of Our Genome Is Sequenced?

I'm getting ready for a class on the size and composition of the human genome so I thought I'd check to see the latest estimate of its size. Recall that in an earlier posting I concluded that the size of the human genome was 3,200,000,000 bp (3,200,000 kb, 3,200 Mb, 3.2 Gb) [How Big Is the Human Genome?].

You might think that all you have to do is check out the human genome websites and look up the exact size. That doesn't work because not all of the human genome has been sequenced and organized into a contiguous assembly of 24 different strands (one for each chromosome). So that prompts the question, how much of the human genome has actually been sequenced?1

The latest assembly is GRCh37 Patch Release 7 (GRCh37.p7), released on Feb. 3, 2012. If you look at the data for this assembly you will see an estimate of the "Total Sequenced Bases in the Assembly." The number is 3,173,036,847 bp or 3.17 Gb. This value is close to estimates of the genome size from the years before the first draft of the genome sequence was published.

I was suspicious of this number since we know that there are many gaps in the human genome sequence. The largest gaps cover highly repetitive parts of the genome—mostly around the centromeres and other heterochromatic regions. There were also gaps at the locations of several gene clusters (e.g. ribosomal RNA genes) where it's impossible to determine the exact number of copies. In the case of ribosomal RNA gene clusters, these gaps have now been closed.

Deanna Church posted a few comments on my earlier posting. She's with the Genome Reference Consortium (GRC). That's the group responsible for updating the human genome. Deanna explained that "Total Sequenced Bases in the Assembly" is not an accurate representation of the truth.2 What it actually means is total sequenced bases plus estimated sizes of the gaps. In other words, it's a good estimate of the size of the genome.

So, how much of the genome is actually sequenced and organized into "scaffolds," or contiguous stretches of DNA? You can see the actual numbers by clicking on Ungapped Lengths on the NCBI website.

The total number of sequenced base pairs that have been organized into scaffolds and placed on a particular chromosome is 2,861,332,606 bp. An additional 6,110,758 bp have been sequenced but the blocks of sequence cannot be placed in the assembly. Most of this unassigned sequence is on chromosomes 1,4,9, and 17 but some of it can't even be associated with a particular chromosome.

If we assume that the true haploid genome size is 3.2 Gb, or 3,200 Mb, then the sequenced and assigned part of the genome represents 89.6% and the unassigned sequenced part is 0.2%.

We can say that only 90% of the human genome has been sequenced and the remaining 10% falls into 357 gaps scattered throughout the genome. (Every chromosome has unsequenced gaps but some have more than others and it doesn't depend on the size of the chromosome.)

The The Wellcome Trust Sanger Institute is part of the Genome Reference Consortium but it maintains its own website on the human genome [Whole Genome]. The data on the e!Ensembl page refers to build CRCh37.p5 from Feb. 2009 but it also says the data was updated in Dec. 2011.

According to the Sanger Institute, the size of the sequenced genome is 3,283,984,159 bp and the "golden path length" is 3,101,804,739 bp. I've tried to find out what these numbers mean but if the information is present on the Ensembl website then it's very well hidden.

Are you interested in the number of genes? Here's the data from Ensembl. It indicates that the human genome contains 33,399 genes! [What Is a Gene?] [What is a gene, post-ENCODE?] This inflated value is calculated by including 12,523 genes that make an RNA product that's not translated. This is almost certainly a highly inflated number.

The data indicates that there are 181,744 gene transcripts or between 5 and 9 transcripts per gene depending on how you count the genes. I don't believe there are this many biologically functional transcripts per gene. I think the actual number is much closer to one (1) [Genes and Straw Men].


1. It ceratinly doesn't "beg the question." That means something else entirely [Begging the Question].

2. That's a euphemism for "It's a lie!"

Monday's Molecule #158

 
This molecule is responsible for one of the distinguishing features of an entire group of species. Sadly, most undergraduates have never heard of this molecule and they never study the fundamental process that it represents. In my experience, about 90% of all introductory biochemistry courses skip the relevant chapter(s) in the textbooks. There's no reasonable excuse for that omission. It's just bad teaching.

Identify the molecule—the common name will do. Post your answer in the comments. I'll hold off releasing any comments for 24 hours. The first one with the correct answer wins. I will only post correct answers to avoid embarrassment.

There could be two winners. If the first correct answer isn't from an undergraduate student then I'll select a second winner from those undergraduates who post the correct answer. You will need to identify yourself as an undergraduate in order to win. (Put "undergraduate" at the bottom of your comment.)

Some past winners are from distant lands so their chances of taking up my offer of a free lunch are slim. (That's why I can afford to do this!)

In order to win you must post your correct name. Anonymous and pseudoanonymous commenters can't win the free lunch.

Winners will have to contact me by email to arrange a lunch date.

UPDATE: The molecule is phycocyanobilin the light absorbing pigment in cyanobacteria (and some other species). This blue pigment is found in large structures called phycobilosomes and it is the reason why cyanobacteria were called blue-green algae. The winners are Thomas Ferraro and Charles Motraghi (undergraduate).

Winners
Nov. 2009: Jason Oakley, Alex Ling
Oct. 17: Bill Chaney, Roger Fan
Oct. 24: DK
Oct. 31: Joseph C. Somody
Nov. 7: Jason Oakley
Nov. 15: Thomas Ferraro, Vipulan Vigneswaran
Nov. 21: Vipulan Vigneswaran (honorary mention to Raul A. Félix de Sousa)
Nov. 28: Philip Rodger
Dec. 5: 凌嘉誠 (Alex Ling)
Dec. 12: Bill Chaney
Dec. 19: Joseph C. Somody
Jan. 9: Dima Klenchin
Jan. 23: David Schuller
Jan. 30: Peter Monaghan

Sunday, February 5, 2012

Fifth International Society for Arabic Papyrology Conference, Carthage, March 28-31, 2012


Fifth International Society for Arabic Papyrology Conference, Carthage, March 28-31, 2012

The fifth ISAP conference will be hosted by the Tunisian Academy of Sciences, Letters and Arts, Beït Al-Hikma(/www.baitelhekma.nat.tn) in Carthage. It will be organized by the International Society for Arabic Papyrology in cooperation with the Institut français d’archéologie orientale (Ifao) in Cairo.
The conference will start on the evening of Wednesday, March 28, and continue through Saturday, March 31. The programme will include 20-minute lectures presenting text editions or studies based on documentary material from the Islamic medieval world, workshops in which unedited Arabic documents will be presented, and evening lectures. There will also be the opportunity to visit the National Library of Tunisia (Tunis) and the National Museum of Islamic Arts of Raqqada (Kairouan), which hosts the only collection of Arabic payri in Tunisia, and important early Islamic manuscripts written on parchment.
Participants are supposed to be or become members of the International Society for Arabic Papyrology.

Summary

The fifth conference of the International Society for Arabic Papyrology (ISAP) will take place at the Tunisian Academy of Sciences, Letters and Arts, Beït Al-Hikma, in Carthage.
It will bring together scholars using documentary evidence to study the history of the early Islamic world, including Arabic, Coptic, and Greek papyri, paper and other documents, as well as epigraphic and numismatic material. Participants may present their research either as 20-minute papers or within the context of workshops on Greek, Coptic, and Arabic papyrology and palaeography.

Conference Format

The Conference will include 1) text workshops and 2) sessions for the presentation of 20-minute papers and 3) evening lectures at local research institutes. Although the "official language" of the conference is English, papers and workshops may be given in English, French, German, or Arabic.

Text Workshops

These workshops will focus on a single text, or group of texts, to be circulated in advance. The texts used may be in any of the languages of the documentary sources relevant to the history of early Islamic Egypt and the wider Mediterranean world (Greek, Coptic, or Arabic). A translation of the text should also be circulated to allow for the widest possible participation. The presenter will have the first thirty minutes to introduce the text and its problems, and then the remaining hour will be spent in discussion.

Paper Sessions

There will be several sessions during which three or four 20-minute papers, followed by questions and discussion, will be read. While the topics addressed need not focus exclusively on documentary evidence, it is expected that documentary sources will be an integral and substantive part of each paper.

Abstracts and Handouts

The deadline for 400-word abstracts is November 1 2011. Please send abstracts to Sobhi Bouderbala (sbouderbala at ifao.egnet.net). If your presentation will require audio-visual equipment of any kind, please describe what is needed. Notification regarding the acceptance of proposals will be made by the end of November 2011.
Also, please send a copy of all texts and translations to be used in the text workshops by 15 February 2011. These will be made available to participants.

Registration

There will be no conference fee charged. Participants who currently have no membership should renew their membership in Carthage on the first day of the conference. Payments have to be made in cash in Euros, dollars or Tunisian dinars. If you are interested in joining ISAP, information can be found at the ISAP sign-up website.
Please send a notice of intent to participate in the Conference to one of the conference organizers, Petra Sijpesteijn (p.m.sijpesteijn at hum.leidenuniv.nl) or Sobhi Bouderbala (sbouderbala at ifao.egnet.net)

Travel Subsidies

It is hoped that the Conference will be able to offer a few awards for scholars not able to get institutional subventions for travel to Carthage.
Please let us know as soon as possible whether you will be in need for such sponsoring.

Conference Organizers

If you have any further questions about the Conference, please contact: Petra Sijpesteijn (p.m.sijpesteijn at hum.leidenuniv.nl) or Sobhi Bouderbala (sbouderbala at ifao.egnet.net).

One last re-issue: G. Zuntz, The Text of the Epistles (Etc.)



The Text of the Epistles: A Disquisition Upon the Corpus Paulinum
The Schweich Lectures of The British Academy, 1946
By G. Zuntz









ETC.:


Light from Ancient Letters
 
Private correspondence in the

 non-literary papyri of Oxyrhynchus 
of the first four centuries and its 
bearing on New Testament language 
and thoughtBy Henry G. Meecham
Retail Price: $18.00
Web Price: $14.40
ISBN 10: 1-59244-473-3;
ISBN 13: 978-1-59244-473-1
Paperback; Published 01/15/2004


ISBN 13: 978-1-55635-370-3
Paperback; Published 04/01/2007









A Literary and Source Analysis
By Charles W. Hedrick
Retail Price: $31.00
Web Price: $24.80
ISBN 10: 1-59752-386-0; ISBN 13: 978-1-59752-386-8
Paperback; Published 09/20/2005


 





  • Here and There 
    Among the Papyri
     By George Milligan
    Retail Price: $20.00
    Web Price: $16.00
    ISBN 10: 1-59244-182-3; 
  • ISBN 13: 978-1-59244-182-2
    Paperback; Published 03/12/2003

Yet another re-issue: P. Comfort, Early Manuscripts and Modern Translations of the New Testament

Early Manuscripts and Modern Translations of the New Testament
By Philip Wesley Comfort