Showing posts with label Genome. Show all posts
Showing posts with label Genome. Show all posts

Wednesday, December 4, 2013

Michael Eisen defends 23andMe against the FDA

The US Food and Drug Administration has asked 23andMe to stop marketing their genetic test product. The company will test your DNA for the presence of several genetic markers that might indicate a predisposition to disease. The FDA is concerned that 23andMe is not doing enough to ensure that its tests are accurate and the advice it gives is medically sound.

Michael Eisen is on the Scientific Advisory Board for 23andme. He responds to the controversy: FDA vs. 23andMe: How do we want genetic testing to be regulated?.

I think some of his points are worth discussing. My position is that the links between certain diseases and certain SNPs are not well-established. The scientific literature on this topic is not all that great and many of the published results have not been repeated. What this means is that private companies like 23andMe are under pressure to be the first to include a new link in their database but may not be exercising the appropriate amount of skepticism.

Read more »

Thursday, November 21, 2013

Claudiu Bandea Shows Why Attacking Dan Graur Is a Very Bad Idea

Claudiu Bandea is a frequent commenter on this blog. Whenever the subject of junk DNA comes up he reminds us that he had a theory over twenty years ago. Now he has published(?) an advertisement at: On the concept of biological function, junk DNA and the gospels of ENCODE and Graur et al.. Here's the abstract ...
In a recent article entitled “On the immortality of television sets: "function" in the human genome according to the evolution-free gospel of ENCODE”, Graur et al. dismantle ENCODE’s evidence and conclusion that 80% of the human genome is functional. However, the article by Graur et al. contains assumptions and statements that are questionable. Primarily, the authors limit their evaluation of DNA’s biological functions to informational roles, sidestepping putative non-informational functions. Here, I bring forward an old hypothesis on the evolution of genome size and on the role of so called ‘junk DNA’ (jDNA), which might explain C-value enigma. According to this hypothesis, the jDNA functions as a defense mechanism against insertion mutagenesis by endogenous and exogenous inserting elements such as retroviruses, thereby protecting informational DNA sequences from inactivation or alteration of their expression. Notably, this model couples the mechanisms and the selective forces responsible for the origin of jDNA with its putative protective biological function, which represents a classic case of ‘fighting fire with fire.’ One of the key tenets of this theory is that in humans and many other species, jDNAs serves as a protective mechanism against insertional oncogenic transformation. As an adaptive defense mechanism, the amount of protective DNA varies from one species to another based on the rate of its origin, insertional mutagenesis activity, and evolutionary constraints on genome size.
It's not a good idea to attack someone who; (a) is an expert in the field, (b) is intelligent and outspoken, and (c) has a blog. But that never stopped Claudiu Bandea before so why should it now?

Here's part of how Dan Graur responds at: A Pre-Refuted Hypothesis on the Subject of “Junk DNA”. There's more, read it all.
The first problem with this hypothesis is that big eukaryotic genomes consist mostly of very few active transposable elements and numerous dead transposable elements. So, big genomes seem to need little protection. Moreover, a positive correlation exists between genome size and number of transposable elements. In 2002, Margaret Kidwell published a paper entitled “Transposable elements and the evolution of genome size in eukaryotes.” In it, she showed that an approximately linear relationship exists between total transposable element DNA and genome size. Copy numbers per family of transposable elements were found to be low and globally constrained in small genomes, but to vary widely in large genomes. Thus, the major characteristic of large genomes is the absence of selective constraint on transposable element copy number.

Given that the vast majority of transposable elements are dead, the most parsimonious explanation is that the continuous accumulation of dead transposable elements is the reason for genomes becoming large. Let me spell it out: the “large” part in “large genomes” is made of transposable elements. Genome do not become large first and then protect genetic information by becoming sinks of transposable elements.

The other problem with the protection-from-mutation hypothesis is that it assumes selection on mutation to be effective. Selection on mutation is referred to in population genetics as second-order selection. The reason is that this type of selection is anticipatory. It protects against a possibility, not an actuality. Second-order selection on mutation (mutability) requires huge effective population sizes, so huge in fact that they are only found in a few bacteria and viruses. Unfortunately for the protection-from-mutation hypothesis, genome size is known to be inversely correlated with effective population size. In other words, huge genomes are found in species that have very small effective population sizes. So small, in fact, that even regular selection (first-order selection) is not very effective.

Thomas Huxley was proven right again: "The great tragedy of Science is the slaying of a beautiful hypothesis by an ugly fact." Several ugly facts in this case.
I can't count the number of people who have tried to explain to Claudiu Bandea that his idea is ridiculous. Hopefully, this last embarrassment will silence him.

Naturally, the Intelligent Design Creationists are all over it [Another response to Darwin’s followers’ attack on the “not-much-junk-DNA” ENCODE findings].


Friday, November 8, 2013

Science Journal Blows It Again

This week's issue of Science contains three separate papers analyzing transcription factor binding sites and chromatin modification sites in the genomes of different individuals. If most of these sites are spurious sites that just happen to contain a consensus sequence, then you would expect a lot of variability since the sites are mostly in junk DNA where the sequences make no difference. That's what all three papers found but, of course, they interpret this to mean that the regulatory sites must be responsible for the variation between individuals.

The papers were summarized in the form of a "press release" called a "Perspective." The complete citation is ...
Furey, T.S. and Sethupathy, P. (2013) Genetics Driving Epigenetics. Science 342:705-706. [doi: 10.1126/science.1246755]
These authors are affiliated with several departments at the University of North Carolina in Chapel Hill but, most significantly, they are part of the Carolina Center for Genome Sciences. This strongly suggests that they know something about genomes.

Read more »

Tuesday, November 5, 2013

Stop Using the Term "Noncoding DNA:" It Doesn't Mean What You Think It Means

Axel Visel is a member of the ENCODE Consortium. He is a Staff Scientist at the Lawrence Berkeley National Laboratory in Berkeley, California (USA). Axel Visel is responsible, in part, for the publicity fiasco of September 2012 where the entire ENCODE Consortium gave the impression that most of our genome is functional.

He is also the senior author on a paper I blogged about last week—the one where some journalists made a big deal about junk DNA when there was nothing in the paper about junk DNA [How to Turn a Simple Paper into a Scientific Breakthrough: Mention Junk DNA].

Dan Graur contacted him by email to see if he had any comment about this misrepresentation of his published work and he defended the journalist. Here's the email response from Axel Visel to Dan Gaur.
Read more »

Friday, November 1, 2013

Vertebrate Complexity Is Explained by the Evolution of Long-Range Interactions that Regulate Transcription?

The Deflated Ego Problem is a very serious problem in molecular biology. It refers to the fact that many molecular biologists were puzzled and upset to learn that humans have about the same number of genes as all other multicellular eukaryotes. The "problem" is often introduced by stating that the experts working on the human genome project expected at least 100,000 genes but were "shocked' when the first draft of the human genome showed only 30,000 genes (now down to about 25,000). This story is a myth as I document in: Facts and Myths Concerning the Historical Estimates of the Number of Genes in the Human Genome. Truth is, most knowledgeable experts expected that humans would have about the same number of genes as other animals. They realized that the differences between fruit flies and humans, for example, didn't depend on a host of new human genes but on the timing and expression of a mostly common set of genes.

This isn't good enough for many human chauvinists. They are still looking for something special that sets human apart from all other animals. I listed seven possibilities in my post on the deflated ego problem:
Read more »

Sunday, October 27, 2013

Trace Dominguez of Discovery News Says 98% of Your Genome Is Junk

Theme Genomes & Junk DNAI happened to stumble on this video where Trace Dominguez (@trace501) promotes the idea of junk DNA based on the C-value Paradox—a version of the Onion Test. It's good that he tells the general public about junk DNA but it's bad that he equates "noncoding DNA" with "junk DNA." It's really silly to tell people that the only important part of your genome is the 2% that codes for proteins.

Just so you know, some of the important known functions of "noncoding DNA" are [What's in Your Genome?] ....
  1. Genes for functional RNAs like ribosomal RNA, tRNA, and a host of others.
  2. Regulatory sequences that control expression of all genes.
  3. Part of intron sequences.
  4. Origins of replication;specific sites where DNA replication begins.
  5. Telomeres.
  6. Centromeres.
  7. SARS or scaffold attachment regions; sites required to organize chromatin.
  8. Functional transposons or "selfish DNA."
  9. Functional DNA and RNA viruses.
Scientists believe that about 2% of our genome encodes proteins and about 8% has other functions. It is not true that all noncoding DNA is junk. No knowledgeable scientist ever said that.

I realize that the kind of presentation shown in this video doesn't lend itself to a detailed description of noncoding DNA functions but surely we can do better than this? Why not say that scientists have determined that genes make up about 2% of our genome and about 8% contains information necessary for the proper functioning of genes and chromosomes? The rest, about 90%, is thought to be junk?

98% of your DNA is junk


Saturday, October 26, 2013

How to Turn a Simple Paper into a Scientific Breakthrough: Mention Junk DNA

Attanasio et al. (2013) published a paper in Science where they identified several thousand possible enhancers that were active in the facial area of developing mouse embryos. About 200 of them appear to be controlling genes that determine the size and shape of the face. (Recall that there are about 20,000 protein-encoding genes in mammals.)

Lynn Yarris of Lawrence Berkeley National Laboratory in California (USA) wrote up the press release [What is it About Your Face?]. It's a really good press release that fairly represents the published work and explains some of the significance. There's no mention of junk DNA in the press release or the published paper.

This is what it looks like when science correspondent Alok Jha published it in The Guardian.
Faces are sculpted by 'junk DNA'

Though everybody's face is unique, the actual differences are relatively subtle. What distinguishes us is the exact size and position of things like the nose, forehead or lips. Scientists know that our DNA contains instructions on how to build our faces, but until now they have not known exactly how it accomplishes this.

Visel's team was particularly interested in the portion of the genome that does not encode for proteins – until recently nicknamed "junk" DNA – but which comprises around 98% of our genomes. In experiments using embryonic tissue from mice, where the structures that make up the face are in active development, Visel's team identified more than 4,300 regions of the genome that regulate the behaviour of the specific genes that code for facial features.
It's pretty clear that science correspondent Alok Jha doesn't understand what he's writing and it's about time we started publicizing the names of those science writers who mislead the public about science. The consensus among knowledgeable scientists is that at least 80-90% of our genome is junk. It's time for science writers to admit that the science favors junk.

Scientists have known for decades that a lot of noncoding DNA is functional. The idea that all noncoding DNA (98%) is junk is false. No knowledgeable scientist ever made such a claim. It is a myth perpetuated, in part, by ignorant science writers; albeit, aided and abetted by ignorant scientists. Scientists have known for fifty (50!!) years that gene expression is controlled by regulatory sequences in noncoding DNA. Scientists have known for at least that length of time that during embryogenesis different genes are turned on and off and that this is due, in part, to binding of transcription factors to those regulatory sequences (enhancers). Scientists have known for one hundred years that the morphological features of mammals, including humans, are controlled by genes.

Move along folks. There's nothing to see here.


Attanasio, C. et al. (2013) Fine Tuning of Craniofacial Morphology by Distant-Acting Enhancers. Science 342: Oct. 25, 2013 [doi: 10.1126/science.1241006]

Monday, October 21, 2013

Jukes to Crick on Junk DNA

Dan Graur discovered that the term "junk DNA" was commonly used in the 1960's—long before Susumu Ohno used "junk" in the title of his 1972 paper. This makes a lot of sense. Apparently the term was quite commonly used in Cambridge by people like Francis Crick and Sydney Brenner. (Perhaps you've heard of them?)

Graur found a 1963 paper that refers to "junk" DNA. This is the earliest known refencee to junk in the scientific literature. Read about his sleuthing at: The Origin of Junk DNA: A Historical Whodunnit.

Meanwhile, a person named "ShadiZl" commented on one my posts and pointed me to a letter from Thomas Jukes to Francis Crick in 1979. Jukes, you might recall, was no Darwinian. He was a proponent of Neutral Theory and random genetic drift. The letter is archived on the National Library of Medicine (USE) site under a section devoted to The Francis Crick Papers: Letter from Thomas H. Jukes to Francis Crick.

The letter is interesting because it reveals how casually the "insiders" talked about junk DNA and about the adaptationist misconception even as far back as 1979. This was when Gould and Lewontin published the "spandrels" paper. It also reveals how misguided the creationists are when it comes to the history of junk DNA. They still think that it was "Darwinists" who "predicted" junk DNA based on their view of natural selection. (Do not read this letter if you are irony-deficient. It will only confuse you.)
December 20, 1979

Dear Francis:

I am sure that you realize how frightfully angry a lot of people will be if you say that much of the DNA is junk. The geneticists will be angry because they think that DNA is sacred. The Darwinian evolutionists will be outraged because they believe every change in DNA that is accepted in evolution is necessarily an adaptive change. To suggest anything else is an insult to the sacred memory of Darwin.

This additive is so pervasive that if no reason can be found for an evolutionary change, it is necessary to invent one. Kimura points out that one author attributed the pink color of flamingos to protective coloration against the setting sun. This type of thinking carries over into people who sequence mRNA. They claim that differences between rabbit and human globin mRNAs are because each species has its own requirements for secondary structure.

Various people have tried to think up possible functions for the regions of DNA that do not code for anything as far as is known. Roy Britten says that such DNA has a regulatory function.

Actually, the scheme proposed by Britten about ten years ago was that occasionally events of saltatory duplication, took place, so that a great many copies of a short piece of DNA were made. As time went by, the composition of a family of identical copies became changed by drift, until the copies no longer closely resemble each other. Figure 55 of the article by Britten shows a diagram of a sort of "junk DNA generating system". I note that he says on page 105 "the rate of increase in DNA content per cell resulting from saltatory replication alone may prove to be embarrassingly large and a mechanism for the loss of DNA may have to be invoked". I gather that you agree with this.

I quoted you on drift in DNA in a talk that I gave at the symposium for Emil Smith (see enclosure). Your concept of "junk DNA" presumably includes this idea. I shall look forward to hearing more about it, and I have been asked by Die Naturwissenschaften to write an article on silent changes, so I hope I can include mention of your new manuscript when I start to write mine.

With best regards,


Thomas H. Jukes



Friday, September 27, 2013

The Extraordinary Human Epigenome

We learned a lot about genes and gene expression in the second half of the 20th century. We learned that genes are transcribed and we have a pretty good understanding of how transcription initiation complexes are formed and how transcription works.

We learned how transcription is regulated through promoter strength, activators, and repressors. Activators and repressors bind to DNA and those binding sites can lie at some distance from the promoter leading to formation of loops of DNA that bring the regulatory proteins into contact with the transcription complex. Much of our basic understanding of this process was derived from detailed studies of bacteriophage and bacterial genes.

THEME:
Transcription

Later on we learned that eukaryotic genes expression was very similar and regulation also required repressors and activators. We discovered that gene expression was associated with chromatin remodeling that opened up regions of the chromosome that were tightly bound to histones in 30nm or higher order structures.

Building on studies in prokaryotes, we learned about temporal gene regulation and differentiation. Much of the work was done in model organisms like Drosophila, yeast, C. elegans, and various mammalian cells in culture.

By the end of the century I was pretty confident that what I wrote in my textbook was a fair representation of the fundamental concepts in gene expression and regulation.

Turns out I was wrong as I just discovered this morning when I read the opening paragraph of a review by Rivera and Ren (2013). Here's what they say ...
More than a decade has passed since the human genome was completely sequenced, but how genomic information directs spatial- and temporal-specific gene expression programs remains to be elucidated (Lander, 2011). The answer to this question is not only essential for understanding the mechanisms of human development, but also key to studying the phenotypic variations among human populations and the etiology of many human diseases. However, a major challenge remains: each of the more than 200 different cell types in the human body contains an identical copy of the genome but expresses a distinct set of genes. How does a genome guide a limited set of genes to be expressed at different levels in distinct cell types?
Wow! The textbooks need to be rewritten! We didn't learn anything in the last century!

It took me the whole first paragraph of this paper to realize that the rest of it was probably going to be worthless unless you were interested in technical details about the field. That's because I'm not as smart as Dan Graur. He only read the title, "Mapping Human Epigenomes" and the abstract before concluding that the authors were speaking in newspeak1 [A “Leading Edge Review” Reminds Me of Orwell (and #ENCODE)].

The Rivera and Ren paper is a "Leading Edge" review in the prestigious journal Cell. It covers all the techniques used to study methylation, histone modification and binding, transcription factor binding, and nucleosome positioning at the genome level. According to the authors, people like me were fooled by studies on individual genes, purified factors, and in vitro binding assays. That didn't really tell us what was going on.

Apparently, the most effective way of learning about the regulation of gene expression in humans is to analyze the entire genome all at once and read off the data from microarrays and computer monitors. (After shoving it through a bunch of code.)
Overwhelming evidence now indicates that the epigenome serves to instruct the unique gene expression program in each cell type together with its genome. The word "epigenetics," coined half a century ago by combining "epigenesis" and "genetics," describes the mechanisms of cell fate commitment and lineage specification during animal development (Holliday, 1990; Waddington, 1959). Today, the "epigenome" is generally used to describe the global, comprehensive view of sequence-independent processes that modulate gene expression patterns in a cell and has been liberally applied in reference to the collection of DNA methylation state and covalent modification of histone proteins along the genome (Bernstein et al., 2007; Bonasio et al., 2010). The epigenome can differ from cell type to cell type, and in each cell it regulates gene expression in a number of ways—by organizing the nuclear architecture of the chromosomes, restricting or facilitating transcription factor access to DNA, and preserving a memory of past transcriptional activities. Thus, the epigenome represents a second dimension of the genomic sequence and is pivotal for maintaining cell-typespecific gene expression patterns.

Not long ago, there were many points of trepidation about the value and utility of mapping epigenomes in human cells (Madhani et al., 2008). At the time, it was suggested that histone modifications simply reflect activities of transcription factors (TFs), so cataloging their patterns would offer little new information. However, some investigators believed in the value of epigenome maps and advocated for concerted efforts to produce such resources (Feinberg, 2007; Henikoff et al., 2008; Jones and Martienssen, 2005). The last five years have shown that epigenome maps can greatly facilitate the identification of potential functional sequences and thereby annotation of the human genome. Now, we appreciate the utility of epigenomic maps in the delineation of thousands of lincRNA genes and hundreds of thousands of cis-regulatory elements (ENCODE Project Consortium et al., 2012; Ernst et al., 2011; Guttman et al., 2009; Heintzman et al., 2009; Xie et al., 2013b; Zhu et al., 2013), all of which were obtained without prior knowledge of cell-type-specific master transcriptional regulators. Interestingly, bioinformatic analysis of tissue-specific cis-regulatory elements has actually uncovered novel TFs regulating specific cellular states.
So, what are all these new discoveries that now elucidate what was previously unknown; namely, "how genomic information directs spatial- and temporal-specific gene expression programs"?

This is a very long review full of technical details so let's skip right to the conclusions.
Six decades ago, Watson and Crick put forward a model of DNA double helix structure to elucidate how genetic information is faithfully copied and propagated during cell division (Watson and Crick, 1953). Several years later, Crick famously proposed the "central dogma" to describe how information in the DNA sequence is relayed to other biomolecules such as RNA and proteins to sustain a cell’s biological activities (Crick, 1970). Now, with the human genome completely mapped, we face the daunting
task to decipher the information contained in this genetic blueprint. Twelve years ago, when the human genome was first sequenced, only 1.5% of the genome could be annotated as protein coding, whereas the rest of the genome was thought to be mostly "junk" (Lander et al., 2001; Venter et al., 2001). Now, with the help of many epigenome maps, nearly half of the genome is predicted to carry specific biochemical activities and potential regulatory functions (ENCODE Project Consortium, et al., 2012). It is conceivable that in the near future the human genome will be completely annotated, with the catalog of transcription units and their transcriptional regulatory sequences fully mapped.
I hope they hurry up. Not only do I have to re-write my description of the Central Dogma2 but I'm going to have to re-write everything I thought I knew about regulation of gene expression and the organization of information in the human genome. That's going to take time so I hope the epigeneticists will publish lots more whole genome studies in the near future so I can understand the new model of gene expression.

Keep in mind that this paper was published in Cell where it was rigorously reviewed by the leading experts in the field. It must be right.


[Image Credit: Moran, L.A., Horton, H.R., Scrimgeour, K.G., and Perry, M.D. (2012) Principles of Biochemistry 5th ed., Pearson Education Inc. page 647 [Pearson: Principles of Biochemistry 5/E] © 2012 Pearson Education Inc.]

1. Newspeak was first described in 1984 proving, once again, that George Orwell (Eric Arthur Blair) was a really smart and prescient guy. For another example see: What Is "Science" According to George Orwell?.

2. Apparently I didn't read the Crick (1970) paper as carefully as they did.

Rivera, C.M. and Ren, B. (2013) Mapping Human Epigenomes. Cell 155:39-55 [doi: 10.1016/j.cell.2013.09.011]

Transcription Initiation Sites: Do You Think This Is Reasonable?

I'm interested in how scientists read the scientific literature and in how they distinguish good science from bad science. I know that when I read a paper I usually make a pretty quick judgement based on my knowledge of the field and my model of how things work. In other words, I look at the conclusions first to see whether they conflict with or agree with my model.

Many of my colleagues do it differently. They focus on the actual experiments and reach a conclusion based on how the perceive the data. If the experiments look good and the data seems reliable then they tentatively accept the conclusions even if they conflict with the model they have in their mind. They are much more likely to revamp their model than I am.

I'm about to give you the conclusions from a recently published paper in Nature. I'd like to hear from all graduate students, postdocs, and scientists on how you react to those conclusions. Do you think the conclusions are reasonable (as long as the experiments are valid) or do you think that the conclusions are unreasonable, indicating that there has to be something wrong somewhere?

The paper is Venters and Pugh (2013). It's title is Genomic organization of human transcription complexes. You don't need to read the paper unless you want to get into a more detailed debate. All I want to hear about is your initial reaction to their final two paragraphs.
Consolidated genomic view of initiation

...The discovery that transcription of the human genome is vastly more pervasive than what produces coding mRNA raises the question as to whether Pol II initiates transcription promiscuously through random collisions with chromatin as biological noise or whether it arises specifically from canonical Pol II initiation complexes in a regulated manner. Our discovery of ~150,000 non-coding promoter initiation complexes in human K562 cells and more in other cell lines suggests that pervasive non-coding transcription is promoter-specific, regulated, and not much different from coding transcription, except that it remains nuclear and non-polyadenylated. An important next question is the extent to which transcription factors regulate production of ncRNA.

We detected promoter transcription initiation complexes at 25% of all ~24,000 human coding genes, and found that there were 18-fold more non-coding complexes than coding. We therefore estimate that the human genome potentially contains as many as 500,000 promoter initiation complexes, corresponding to an average of about one every 3 kilobases (kb) in the non-repetitive portion of the human genome. This number may vary more or less depending on what constitutes a meaningful transcription initiation event. The finding that these initiation complexes are largely limited to locations having well-defined core promoters and measured TSSs indicates that they are functional and specific, but it remains to be determined to what end. Their massive numbers would seem to provide an origin for the so-called dark matter RNA of the genome, and could house a substantial portion of the missing heritability.
Looking forward to hearing from you.

Keep in mind that this is a Nature paper that has been rigorously reviewed by leading experts in the field. Does that influence your opinion?


Venters, B.J. and Pugh, B.F. (2013) Genomic organization of human transcription initiation complexes. Nature Published online 18 September 2013 [doi: 10.1038/nature12535] [PubMed] [Nature]

Wednesday, August 28, 2013

An Example of a Very Bad Press Release

Cornelius Hunter is gloating over another study that disputes the notion of junk DNA [More Functions For “Junk” DNA, and More Functions For “Junk” DNA]. His article sounded interesting so I followed the link to the press release.

There was something about the press release that sounded suspicious and that prompted me to seek out the original published paper. Here it is with the abstract ...
Wong, J.J.-L., Ritchie, W., Ebner, O.A., Selbach, M., Wong, J.W., Huang, Y., Gao, D., Pinello, N., Gonzalez, M. and Baidya, K. (2013) Orchestrated Intron Retention Regulates Normal Granulocyte Differentiation. Cell 154:583-595. [PDF] [doi: 10.1016/j.cell.2013.06.052]

Intron retention (IR) is widely recognized as a consequence of mis-splicing that leads to failed excision of intronic sequences from pre-messenger RNAs. Our bioinformatic analyses of transcriptomic and proteomic data of normal white blood cell differentiation reveal IR as a physiological mechanism of gene expression control. IR regulates the expression of 86 functionally related genes, including those that determine the nuclear shape that is unique to granulocytes. Retention of introns in specific genes is associated with downregulation of splicing factors and higher GC content. IR, conserved between human and mouse, led to reduced mRNA and protein levels by triggering the nonsense-mediated decay (NMD) pathway. In contrast to the prevalent view that NMD is limited to mRNAs encoding aberrant proteins, our data establish that IR coupled with NMD is a conserved mechanism in normal granulopoiesis. Physiological IR may provide an energetically favorable level of dynamic gene expression control prior to sustained gene translation.
The authors found 86 genes expressed in mouse granulocytes where there were at least some transcripts that retained an intron. This could be due to mistakes in splicing but the authors prefer to think that intron retention is part of a regulatory step. The transcripts that retain an intron are degraded and this reduces the level of protein that would have been made if a properly spliced transcript had produced a functional mRNA.

It's an example of down-regulation, according to the authors. In most cases the intron-retaining transcripts make up only a few percent of the total transcripts but this is presumably enough to make a difference. In 25 of the genes, the aberrant transcripts are more that 25% of the total cytoplasmic transcripts.

There's nothing in the paper that mentions junk DNA.

Contrast this with the press release from Centenary Institute, Sydney Australia. I reproduce it below ...
How 'Junk DNA' Can Control Cell Development

Aug. 2, 2013 — Researchers from the Gene and Stem Cell Therapy Program at Sydney's Centenary Institute have confirmed that, far from being "junk," the 97 per cent of human DNA that does not encode instructions for making proteins can play a significant role in controlling cell development.

And in doing so, the researchers have unravelled a previously unknown mechanism for regulating the activity of genes, increasing our understanding of the way cells develop and opening the way to new possibilities for therapy.

Using the latest gene sequencing techniques and sophisticated computer analysis, a research group led by Professor John Rasko AO and including Centenary's Head of Bioinformatics, Dr William Ritchie, has shown how particular white blood cells use non-coding DNA to regulate the activity of a group of genes that determines their shape and function. The work is published today in the scientific journal Cell.

"This discovery, involving what was previously referred to as "junk," opens up a new level of gene expression control that could also play a role in the development of many other tissue types," Rasko says. "Our observations were quite surprising and they open entirely new avenues for potential treatments in diverse diseases including cancers and leukemias."

The researchers reached their conclusions through studying introns -- non-coding sequences which are located inside genes.

As part of the normal process of generating proteins from DNA, the code for constructing a particular protein is printed off as a strip of genetic material known as messenger RNA (mRNA). It is this strip of mRNA which carries the instructions for making the protein from the gene in the nucleus to the protein factories or ribosomes in the body of the cell.

But these mRNA strips need to be processed before they can be used as protein blueprints. Typically, any non-coding introns must be cut out to produce the final sequence for a functional protein. Many of the introns also include a short sequence -- known as the stop codon -- which, if left in, stops protein construction altogether. Retention of the intron can also stimulate a cellular mechanism which breaks up the mRNA containing it.

Dr Ritchie was able to develop a computer program to sort out mRNA strips retaining introns from those which did not. Using this technique the lead molecular biologist of the team, Dr Justin Wong, found that mRNA strips from many dozens of genes involved in white blood cell function were prone to intron retention and consequent break down. This was related to the levels of the enzymes needed to chop out the intron. Unless the intron is excised, functional protein products are never produced from these genes. Dr Jeff Holst in the team went a step further to show how this mechanism works in living bone marrow.

So the researchers propose intron retention as an efficient means of controlling the activity of many genes. "In fact, it takes less energy to break up strips of mRNA, than to control gene activity in other ways," says Rasko. "This may well be a previously-overlooked general mechanism for gene regulation with implications for disease causation and possible therapies in the future."
The published paper has nothing to do with junk DNA. Even if intron retention were a common mechanism of gene regulation (it is not), that would only account for about 100 base pairs per gene of additional sequence-dependant information. That's less than 0.1% of the genome.

This is a bad press release because it highlights information that is not in the published paper. The authors bear responsibility for press releases from their own institute that distort their published work. While they may not have written the press release, they presumably are quoted correctly and they should be aware of what's in the press release.

I wonder if they are willing to defend this press release as an accurate representation of their published work?


Saturday, August 24, 2013

John Mattick vs. Jonathan Wells

John Mattick and Jonathan Wells both believe that most of the DNA in our genome is functional. They do not believe that most of it is junk.

John Mattick and Jonathan Wells use the same arguments in defense of their position and they quote one another. Both of them misrepresent the history of the junk DNA debate and both of them use an incorrect version of the Central Dogma of Molecular Biology to make a case for the stupidity of scientists. Neither of them understand the basic biochemistry of DNA binding proteins leading them to misinterpret low level transcription as functional. Jonathan Wells and John Mattick ignore much of the scientific evidence in favor of junk DNA. They don't understand the significance of the so-called "C-Value Paradox" and they don't understand genetic load. Both of them claim that junk DNA is based on ignorance.

Read more »

Friday, August 23, 2013

Some Questions for IDiots

Here's a short quiz for proponents of Intelligent Design Creationism. Let's see if you have been paying attention to real science. Please try to answer the questions below. Supporters of evolution should refrain from answering for a few days in order to give the creationists a chance to demonstrate their knowledge of biology and of evolution.

The bloggers at Evolution News & Views (sic) are promoting another creationist book [see Biological Information]. This time it's a collection of papers from a gathering of creationists held in 2011. The title of the book, Biological Information: New Perspectives suggests that these creationists have learned something new about biochemistry and molecular biology.

One of the papers is by Jonathan Wells: Not Junk After All: Non-Protein-Coding DNA Carries Extensive Biological Information. Here's part of the opening paragraphs.
James Watson and Francis Crick’s 1953 discovery that DNA consists of two complementary strands suggested a possible copying mechanism for Mendel’s genes [1,2]. In 1958, Crick argued that “the main function of the genetic material” is to control the synthesis of proteins. According to the “ Sequence Hypothesis,” Crick wrote that the specificity of a segment of DNA “is expressed solely by the sequence of bases,” and “this sequence is a (simple) code for the amino acid sequence of a particular protein.” Crick further proposed that DNA controls protein synthesis through the intermediary of RNA, arguing that “the transfer of information from nucleic acid to nucleic acid, or from nucleic acid to protein may be possible, but transfer from protein to protein, or from protein to nucleic acid, is impossible.” Under some circumstances RNA might transfer sequence information to DNA, but the order of causation is normally “DNA makes RNA makes protein.” Crick called this the “ Central Dogma” of molecular biology [3], and it is sometimes stated more generally as “DNA makes RNA makes protein makes us.”

The Sequence Hypothesis and the Central Dogma imply that only protein-coding DNA matters to the organism. Yet by 1970 biologists already knew that much of our DNA does not code for proteins. In fact, less than 2% of human DNA is protein-coding. Although some people suggested that non-protein-coding DNA might help to regulate gene expression, the dominant view was that non-protein-coding regions had no function. In 1972, biologist Susumu Ohno published an article wondering why there is “so much ‘ junk’ DNA in our genome” [4].
  1. Crick published a Nature paper on The Central Dogma of Molecular Biology in 1970. Did he and most other molecular biologists actually believe that "only protein-coding DNA matters to the organism?"
  2. Did Crick really say that "DNA makes RNA makes protein" is the Central Dogma or did he say that this was the Sequence Hypothesis? Read the paper to get the answer—the link is below).
  3. Is it true that, in 1970, the majority of molecular biologists did not believe in repressor and activator binding sites (regulatory DNA)?
  4. Is it true that in 1970 molecular biologists knew nothing about the functional importance of non-transcribed DNA sequences such as centromeres and origins of DNA replication?
  5. It is true that most molecular biologists in 1970 had never heard of genes for ribosomal RNAs and tRNAs (non-protein-coding genes)?
  6. If the answer to any of those questions contradicts what Jonathan Wells is saying then why do you suppose he said it?

Crick, F. (1970) Central Dogma of Molecular Biology. Nature 227:561-563. [PDF]

Wednesday, July 31, 2013

The Dark Matter Rises

John Mattick is a Professor and research scientist at the Garvan Institute of Medical Research at the University of New South Wales (Australia).

John Mattick publishes lots of papers. Most of them are directed toward proving that almost all of the human genome is functional. I want to remind you of some of the things that John Mattick has said in the past so you'll be prepared to appreciate my next post [The Junk DNA Controversy: John Mattick Defends Design].

Mattick believes that the Central Dogma means DNA makes RNA makes protein. He believes that scientists in the past took this very literally and discounted the importance of RNA. According to Mattick, scientists in the past believed that genes were the only functional part of the genome and that all genes encoded proteins.

If that sounds familiar it's because there are many IDiots who make the same false claim. Like Mattick, they don't understand the Central Dogma of Molecular Biology and they don't understand the history that they are distorting.

Mattick believes that there is a correlation between the amount of noncoding DNA in a genome and the complexity of the organism. He thinks that the noncoding DNA is responsible for making tons of regulatory RNAs and for regulating expression of the genes. This belief led him to publish a famous figure (left) in Scientific American.

Mattick has many followers. So many, in fact, that the Human Genome Organization (HUGO) recently gave him an award for his contributions to the study of the human genome. Here's the citation.
Theme
Genomes
& Junk DNA
The Award Reviewing Committee commented that Professor Mattick’s “work on long non-coding RNA has dramatically changed our concept of 95% of our genome”, and that he has been a “true visionary in his field; he has demonstrated an extraordinary degree of perseverance and ingenuity in gradually proving his hypothesis over the course of 18 years.”
Let's see what this "true visionary" is saying this year. The first paper is "The dark matter rises: the expanding world of regulatory RNAs" (Clark et al., 2013). Here's the abstract ...
The ability to sequence genomes and characterize their products has begun to reveal the central role for regulatory RNAs in biology, especially in complex organisms. It is now evident that the human genome contains not only protein-coding genes, but also tens of thousands of non–protein coding genes that express small and long ncRNAs (non-coding RNAs). Rapid progress in characterizing these ncRNAs has identified a diverse range of subclasses, which vary widely in size, sequence and mechanism-of-action, but share a common functional theme of regulating gene expression. ncRNAs play a crucial role in many cellular pathways, including the differentiation and development of cells and organs and, when mis-regulated, in a number of diseases. Increasing evidence suggests that these RNAs are a major area of evolutionary innovation and play an important role in determining phenotypic diversity in animals.
This is his main theme. Mattick believes that a large percentage of the human genome is devoted to making regulatory RNAs that control development. He believes that the evolution of this complex regulatory network is responsible for the creation of complex organisms like humans, which, incidentally, are the pinnicle of evolution according to the figure shown above.

The second paper I want to highlight focuses on a slightly different theme. It's title is "Understanding the regulatory and transcriptional complexity of the genome through structure." (Mercer and Mattick, 2013). In this paper he emphasizes the role of noncoding DNA in creating a complicated three-dimensional chromatin structure within the nucleus. This structure is important in regulating gene expression in complex organisms. Here's the abstract ...
An expansive functionality and complexity has been ascribed to the majority of the human genome that was unanticipated at the outset of the draft sequence and assembly a decade ago. We are now faced with the challenge of integrating and interpreting this complexity in order to achieve a coherent view of genome biology. We argue that the linear representation of the genome exacerbates this complexity and an understanding of its three-dimensional structure is central to interpreting the regulatory and transcriptional architecture of the genome. Chromatin conformation capture techniques and high-resolution microscopy have afforded an emergent global view of genome structure within the nucleus. Chromosomes fold into complex, territorialized three-dimensional domains in concert with specialized subnuclear bodies that harbor concentrations of transcription and splicing machinery. The signature of these folds is retained within the layered regulatory landscapes annotated by chromatin immunoprecipitation, and we propose that genome contacts are reflected in the organization and expression of interweaved networks of overlapping coding and noncoding transcripts. This pervasive impact of genome structure favors a preeminent role for the nucleoskeleton and RNA in regulating gene expression by organizing these folds and contacts. Accordingly, we propose that the local and global three-dimensional structure of the genome provides a consistent, integrated, and intuitive framework for interpreting and understanding the regulatory and transcriptional complexity of the human genome.
Other posts about John Mattick.

How Not to Do Science
John Mattick on the Importance of Non-coding RNA
John Mattick Wins Chen Award for Distinguished Academic Achievement in Human Genetic and Genomic Research
International team cracks mammalian gene control code
Greg Laden Gets Suckered by John Mattick
How Much Junk in the Human Genome?
Genome Size, Complexity, and the C-Value Paradox


Clark, M.B., Choudhary, A., Smith, M.A., Taft, R.J. and Mattick, J.S. (2013) The dark matter rises: the expanding world of regulatory RNAs. Essays in Biochemistry 54:1-16. [doi:10.1042/bse0540001]

Mercer, T.R. and Mattick, J.S. (2013) Understanding the regulatory and transcriptional complexity of the genome through structure. Genome research 23:1081-1088 [doi: 10.1101/gr.156612.113]

Thursday, July 25, 2013

Every non-lethal genome position is variable in the human population

Melissa Wilson Sayres blogs at mathbionerd and Panda's Thumb. A recent post on Panda's Thumb address a tweet from Daniel Wegmann where he said "Every non-lethal genome position is variable in the human population."

She asks "Is this true?" and proceeds to show that it is [How many mutations?]. She assumes that the human mutation rate is 1.2 × 10-8 per sit per generation. Multiply this by 7.16 billion people on the planet and you get an average of 86 mutations at every single base pair in the human genome.1

Many of these mutations will be deleterious and they will be quickly eliminated from the population if they are lethal or cause severe problems. Some moderately and slightly deleterious mutations will be present in the population because they haven't yet been eliminated by negative selection. (Some will have no effect if they are present in only one copy of your diploid genome.)

To a first approximation, the statement is pretty accurate. If it's true that most of our genome is junk then the nucleotide sequence is not important.2 As we sequence more and more genomes we should see heterogeneity at 90% of the base pairs in the genome. We haven't reached this sort of coverage but all available evidence is consistent with the idea that most positions can be variable.


1 I prefer a larger mutation rate of 100 new mutations per generation for a total of 112 mutations at every site.

2. This doesn't rule out functions that are not sequence-specific. Such functions are known to exist but there are no reasonable hypotheses that justify such functions for most of the genome.

Monday, July 15, 2013

How the IDiots View Genome Research

It's safe to say that a majority of knowledgeable scientists now agree that the most of our genome is junk. This is bad news for Intelligent Design Creationism because they have staked their credibility on the idea that if the DNA is present is must be the product of god(s) the intelligent designer and it must be there for a reason.

One of the latest posts on Evolution News & Views (sic) emphasizes this point [More Clues that Intergenic DNA Is Functional]. Its author doesn't identify herself/himself. The point of the post is to cherry-pick a couple of papers from the scientific literature, including the horrible paper by Hangauer et al. (2013) [see How to Make a Scientific Argument ].

Read more »

Sunday, July 14, 2013

What Did Dan Graur Say in Chicago?

Dan Graur gave a fantastic and entertaining talk at SMBE2013 [Powerpoint]. He covered a lot of bases, but unfortunately left some out 'cause he had many slides that he didn't get to because of time limitations. Most of the audience enjoyed the talk very much—there was much laughter and enthusiastic head nodding. (I figure that two thirds of the audience agreed with his stance on junk DNA and ENCODE.)

Apparently Cornelius Hunter was in the audience because he blogged about it on Darwin's God: Dan Graur Gave a Great Talk This Week (copied on Uncommon Descent: Dan Graur Gave a Great Talk This Week). It's a shame he didn't make himself known to me or Reed Cartwright or Nick Matzke.

One of the things that Graur said was that if ENCODE is right then evolution is wrong. Now this may be a bit of an exaggeration, but not by much. If the ENCODE leaders are right that 85% of our genome has a biological function then we really do have to rethink the C-value paradox and genetic load. We also have to re-think a lot of biochemistry. I think that's what Dan meant. Cornelius Hunter says that's the one "glaring error" in Graur's presentation.

Theme

Genomes
& Junk DNA
Graur said that junk DNA is a "known known" and by that he means that there's plenty of evidence for junk DNA. It's not "dark matter" and we don't use the term "junk" to mask our ignorance. Speaking of ignorance, Cornelius Hunter seems to be completely unaware of any evidence for junk DNA. I guess he didn't stay to hear my presentation.


How Not to Do Science

Theme
Genomes
& Junk DNA
Many reputable scientists are convinced that most of our genome is junk. However, there are still a few holdouts and one of the most prominent is John Mattick. He believes that most of our genome is made up of thousand of genes for regulatory noncoding RNA. These RNAs (about 100 of them for every single protein-coding gene) are mostly involved in subtle controls of the levels of protein in human cells. (I'm not making this up. See: John Mattick on the Importance of Non-coding RNA )

It was a reasonable hypothesis at one point in time.

How do you evaluate a hypothesis in science? Well, one of the things you should always try to do is falsify your hypothesis. Let's see how that works ...
  1. The RNAs should be conserved. FALSE
  2. The RNAs should be abundant (>1 copy per cell). FALSE
  3. There should be dozens of well-studied specific examples. FALSE
  4. The hypothesis should account for variations in genome size. FALSE
  5. The hypothesis should be consistent with other data, such as that on genetic load. FALSE
  6. The hypothesis should be consistent with what we already know about the regulation of gene expression. FALSE
  7. You should be able to refute existing hypotheses, such as transcription errors. FALSE
Normally, you would abandon a hypothesis that had such a bad track record but true believers aren't about to do that. So what's next? Maybe these regulatory RNAs don't show sequence conservation but maybe their secondary structures are conserved. In other words, these RNAs originated as functional RNAs with a secondary structure but over the course of time all traces of sequence conservation have been lost and only the "conserved" secondary structure remains.1 The Mattick lab looked at the "conservation" of secondary structure as an indicator of function using the latest algorithms (Smith et al., 2013). Here's how they describe their attempts to prove their hypothesis in light of conflicting data ...
The majority of the human genome is dynamically transcribed into RNA, most of which does not code for proteins (1–4). The once common presumption that most non–protein-coding sequences are nonfunctional for the organism is being adjusted to the increasing evidence that noncoding RNAs (ncRNAs) represent a previously unappreciated layer of gene expression essential for the epigenetic regulation of differentiation and development (5–8). Yet despite an exponential accumulation of transcriptomic data and the recent dissemination of genome-wide data from the ENCODE consortium (9), limited functional data have fuelled discourse on the amount of functionally pertinent genomic sequence in higher eukaryotes (1, 10–12). What is incontrovertible, however, is that evolutionary conservation of structural components over an adequate evolutionary distance is a direct property of purifying (negative) selection and, consequently, a sufficient indicator of biological function The majority of studies investigating the prevalence of purifying selection in mammalian genomes are predicated on measuring nucleotide substitution rates, which are then rated against a statistical threshold trained from a set of genomic loci arguably qualified as neutrally evolving (13, 14). Conversely, lack of conservation does not impute lack of function, as variation underlies natural selection. Given that the molecular function of ncRNA may at least be partially conveyed through secondary or tertiary structures, mining evolutionary data for evidence of such features promises to increase the resolution of functional genomic annotations.
Here's what they found ..
When applied to consistency-based multiple genome alignments of 35 mammals, our approach confidently identifies >4 million evolutionarily constrained RNA structures using a conservative sensitivity threshold that entails historically low false discovery rates for such analyses (5–22%). These predictions comprise 13.6% of the human genome, 88% of which fall outside any known sequence-constrained element, suggesting that a large proportion of the mammalian genome is functional.
Apparently 13.6% of the human genome is a "large proportion." Taken at face value, however, the Mattick lab has now shown that the vast majority of transcribed sequences don't show any of the characteristics of functional RNA, including conservation of secondary structure. Of course, that's not the conclusion they emphasize in their paper.

Why not?

1. I can't imagine how this would happen, can you? You'd almost have to have selection AGAINST sequence conservation.

Smith, M.A., Gese, T., Stadler, P.F. and Mattick, J.S. (2013) Widespread purifying selection on RNA structure in mammals. Nucleic Acid Research advance access July 11, 2013 [doi: 10.1093/nar/gkt596]

Thursday, July 4, 2013

Five Things You Should Know if You Want to Participate in the Junk DNA Debate

Here are five things you should know if you want to engage in a legitimate scientific discussion about the amount of junk DNA in a genome.
  1. Genetic Load
    Every newborn human baby has about 100 mutations not found in either parent. If most of our genome contained functional sequence information, then this would be an intolerable genetic load. Only a small percentage of our genome can contain important sequence information suggesting strongly that most of our genome is junk.
  2. C-Value Paradox
    A comparison of genomes from closely related species shows that genome size can vary by a factor of ten or more. The only reasonable explanation is that most of the DNA in the larger genomes is junk.
  3. Modern Evolutionary Theory
    Nothing in biology makes sense except in the light of population genetics. The modern understanding of evolution is perfectly consistent with the presence of large amounts of junk DNA in a genome.
  4. Pseudogenes and broken genes are junk
    More than half of our genomes consists of pseudogenes, including broken transposons and bits and pieces of transposons. A few may have secondarily acquired a function but, to a first approximation, broken genes are junk.
  5. Most of the genome is not conserved
    Most of the DNA sequences in large genomes is not conserved. These sequences diverge at a rate consistent with fixation of neutral alleles by random genetic drift. This strongly suggests that it does not have a function although one can't rule out some unknown function that doesn't depend on sequence.
If you want to argue against junk DNA then you need to refute or rationalize all five of these observations.


Tuesday, July 2, 2013

A Philosopher Trashes Junk DNA

I am one of those scientists who think that the discipline of "philosophy of science" is catering to some pretty stupid philosophers. Dan Graur found one of them, his name is Max Andrews and he's a graduate student in philosophy at the University of Edinburgh, Scotland ["I’ve Got a Little List" & “Let the Punishment Fit the Crime"].

You can read Max Andrews' blog posting at: Junk DNA Isn’t Junk. Be careful, you might find it very difficult to see the connection between this philosophy student's view of biology and anything you might recognize as real science.

It goes without saying that Max Andrews gets the Central Dogma wrong—many scientists make the same mistake. But here's a taste of what else he gets wrong.
The argument from junk DNA suggests that a designer would be maximally efficient in his use of information. There appears to be some information that does not execute or have any meaningful coding. Darwinism takes this issue and uses it as the result of the prediction that there would be left over information not being used due to natural selection and random mutation. However, it doesn’t appear that all junk DNA is actually junk.
Genome organization is patterned to be maximally informative. The overlapping codes observed are known to be evolutionarily costly, because random mutations will likely have a deteriorating effect, not an instructing role So the complex specified information entailed by any genomic region is orders of magnitude higher than previously suspected by, say, Dembski. Any seemingly random aspect of chromosome sequence arrangement is not. A case in point involves endogenous retroviruses (ERV’s). This implies that the taxonomically-specific formatting, indexing, punctuation, etc., of genomes were precisely written. Morphogenetic information is not reducible to the genotype—though it is strongly dependent upon it. Therefore, changes in DNA do not equal changes in the information that structures the body plan.
I wonder who his supervisor is? Maybe Dan or I could be external reviewer on his Ph.D. oral?