Showing posts with label Genes. Show all posts
Showing posts with label Genes. Show all posts

Friday, December 6, 2013

Do you understand this Nature paper on transcription factor binding in different mouse strains?

I've published a few papers on the regulation of transcription of a mouse gene and students in my lab have done the standard promoter-bashing experiment to define transcription factor binding sites. I did ny Ph.D. in a lab that specialized in DNA binding proteins. I've kept up with the basic ideas in eukaryotic gene expression in order to teach undergraduate courses on that topic and in order to write appropriate information in my textbook.

I've been interested in genome organization for several decades and I've been following the literature on pervasive transcription and transcription factor binding in whole genome studies. I'm reasonably familiar with the techniques although I've never done them myself.

I'm not bragging; I'm just saying that I know a little bit about this stuff so when I saw this paper in one of the latest issues of Nature I decided to look more carefully.
Heinz, S., Romanoski, C., Benner, C., Allison, K., Kaikkonen, M., Orozco, L. and Glass, C. (2013) Effect of natural genetic variation on enhancer selection and function. Nature 503:487-492. [doi: 10.1038/nature12615]
Read more »

Wednesday, December 4, 2013

Michael Eisen defends 23andMe against the FDA

The US Food and Drug Administration has asked 23andMe to stop marketing their genetic test product. The company will test your DNA for the presence of several genetic markers that might indicate a predisposition to disease. The FDA is concerned that 23andMe is not doing enough to ensure that its tests are accurate and the advice it gives is medically sound.

Michael Eisen is on the Scientific Advisory Board for 23andme. He responds to the controversy: FDA vs. 23andMe: How do we want genetic testing to be regulated?.

I think some of his points are worth discussing. My position is that the links between certain diseases and certain SNPs are not well-established. The scientific literature on this topic is not all that great and many of the published results have not been repeated. What this means is that private companies like 23andMe are under pressure to be the first to include a new link in their database but may not be exercising the appropriate amount of skepticism.

Read more »

Monday, September 30, 2013

The Problems With The Selfish Gene

Lots of people fail to understand that the "selfish gene" is a metaphor. They criticize Richard Dawkins for promoting the idea that genes can actually take on the characteristics of selfishness.

Andrew Brown and Mary Midgley are prominent examples of people with this kind of misunderstanding and Jerry Coyne has set them straight in Poor Richard’s Almanac: Andrew Brown and the Pope go after The Selfish Gene and “Selection pressures” are metaphors. So are the “laws of physics.”

However, there are two other problem with the metaphor. The first is rather trivial, it refers to the fact that it's actually alleles, or variants, of a gene that are "selfish." Dawkins knows this. He explains it in his book but I don't think he puts enough emphasis on the concept and in most parts of the book he uses "gene" when he should be saying "allele." I grant that The Selfish Allele is not a catchy title.

Read more »

Friday, September 6, 2013

Darwin's Doubt: The Genes Tell the Story?

The main goal of Intelligent Design Creationism is to cast doubt on modern science, especially evolutionary biology. Most of the IDiot books are devoted to attacks on evolution. The underlying assumption is that if modern science is discredited then "god-did-it" becomes a viable alternative.

The latest book by Stephen Myer is no exception. The theme is that evolutionary biologists cannot explain the Cambrian Explosion; therefore, God must have created all the animals in the space of a few million years back in the Cambrian Era (about 530 million years ago).

Most of the book is about the lack of transitional fossils that document the slow transition from primitive worm-like creatures to modern phyla such as arthropods and chordates. Others have dealt with this and I'm not going to comment because it's outside of my area of expertise.1

There is strong evidence from molecular evolution that the major animal phyla share common ancestors and that these common ancestors predate the Cambrian by millions of years. In other words, there's a "long fuse" of evolution leading up to the Cambrian Explosion. Meyer refers to this as the "deep-divergence" assumption.

There are many versions of these trees. The one shown here is from Erwin et al. (2011). It's the one shown in the book The Cambrain Explosion by Douglas Erwin and James Valentine. It isn't necessarily correct in all details but that's not the point.

The point is that molecular phylogenies demonstrate conclusively that the major groups of animals share common ancestors AND that the overall pattern does not conform to a massive radiation around 530 million years ago. Also, it's very clear that the pattern is consistent with evolution and not with God creating all the animals at once.

Stephen Meyer has to address this evidence because it casts doubt on his main theme (God did it). I suppose I don't need to tell you what he says ... it's typical creationist denial. He claims that the evidence doesn't exist. Here are his reasons ...
  1. There are no fossils to support the earliest branches in the molecular phylogenies.
  2. There are many different molecular trees and they don't all agree with each other in terms of branching order and timing.
  3. Evolutionary biologists cherry-pick the data by only picking molecules that give reasonable trees.
  4. The trees rely on questionable assumptions; namely, that the molecular clock ticks at a constant rate and that there is a universal tree.
  5. The molecules being compared must be homologous but this is what is being tested so the argument is circular.
The conclusion is ....
Comparative genetic analyses do not establish a single deep-divergence point, and thus do not compensate for the lack of fossil evidence for key Cambrian ancestors—such as the ur-bilateran or the ur-metazoan ancestor. The results of different studies diverge too dramatically to be conclusive, or even meaningful; the methods of inferring divergence points are fraught with subjectivity; and the whole enterprise depends on a question-begging logic. Many leading Cambrian paleontologists, and even some leading evolutionary biologists, now express skepticism about both the results and the significance of deep-divergence studies.
I'm hoping to find time to go over each of Meyer's objections since they reveal a lot about IDiot misconceptions of evolution (and science) and a lot about how they employ strawmen, lies, quote-mining, and distortions in order to discredit an entire field (molecular evolution).2


1. Most IDiots are experts in everything. I'm not as smart as they are.

2. It always amazes me to discover that IDiots like Stephen Meyer think they know more than thousands of expert biologists who do this sort of stuff for a living.

Erwin, D.H., Laflamme, M., Tweedt, S.M., Sperling, E.A., Pisani, D. and Peterson, K.J. (2011) The Cambrian conundrum: early divergence and later ecological success in the early history of animals. Science 334:1091-1097. [doi: 10.1126/science.1206375]

Thursday, August 29, 2013

Richard Lenski's Classic Papers: Luria and Delbrück, 1943

Richard Lenski has joined a number of other biologists and blogged about classic "must-read" papers. His first example is Luria and Delbrück (1943)—the Fluctuation Test. It's an excellent description and there's a personal touch.

John Dennehy [The Fluctuation Test and Jonathan Eisen [Luria and Delbrück] also picked the same paper. That means it must really be a "must-read"! (I agree.)

Given that the early history of molecular biology is no longer being taught, I imagine that there are quite a few of you who have never heard of Max Delbrück (1906-1981) or Salvador Luria (1912-1991) in spite of the fact they are Nobel prize winners. Here's some of my posts on them ....

The Velvet Underground of Molecular Biology
Nobel Laureates Max Delbrück, Alfred D. Hershey, Salvador E. Luria


Wednesday, August 28, 2013

An Example of a Very Bad Press Release

Cornelius Hunter is gloating over another study that disputes the notion of junk DNA [More Functions For “Junk” DNA, and More Functions For “Junk” DNA]. His article sounded interesting so I followed the link to the press release.

There was something about the press release that sounded suspicious and that prompted me to seek out the original published paper. Here it is with the abstract ...
Wong, J.J.-L., Ritchie, W., Ebner, O.A., Selbach, M., Wong, J.W., Huang, Y., Gao, D., Pinello, N., Gonzalez, M. and Baidya, K. (2013) Orchestrated Intron Retention Regulates Normal Granulocyte Differentiation. Cell 154:583-595. [PDF] [doi: 10.1016/j.cell.2013.06.052]

Intron retention (IR) is widely recognized as a consequence of mis-splicing that leads to failed excision of intronic sequences from pre-messenger RNAs. Our bioinformatic analyses of transcriptomic and proteomic data of normal white blood cell differentiation reveal IR as a physiological mechanism of gene expression control. IR regulates the expression of 86 functionally related genes, including those that determine the nuclear shape that is unique to granulocytes. Retention of introns in specific genes is associated with downregulation of splicing factors and higher GC content. IR, conserved between human and mouse, led to reduced mRNA and protein levels by triggering the nonsense-mediated decay (NMD) pathway. In contrast to the prevalent view that NMD is limited to mRNAs encoding aberrant proteins, our data establish that IR coupled with NMD is a conserved mechanism in normal granulopoiesis. Physiological IR may provide an energetically favorable level of dynamic gene expression control prior to sustained gene translation.
The authors found 86 genes expressed in mouse granulocytes where there were at least some transcripts that retained an intron. This could be due to mistakes in splicing but the authors prefer to think that intron retention is part of a regulatory step. The transcripts that retain an intron are degraded and this reduces the level of protein that would have been made if a properly spliced transcript had produced a functional mRNA.

It's an example of down-regulation, according to the authors. In most cases the intron-retaining transcripts make up only a few percent of the total transcripts but this is presumably enough to make a difference. In 25 of the genes, the aberrant transcripts are more that 25% of the total cytoplasmic transcripts.

There's nothing in the paper that mentions junk DNA.

Contrast this with the press release from Centenary Institute, Sydney Australia. I reproduce it below ...
How 'Junk DNA' Can Control Cell Development

Aug. 2, 2013 — Researchers from the Gene and Stem Cell Therapy Program at Sydney's Centenary Institute have confirmed that, far from being "junk," the 97 per cent of human DNA that does not encode instructions for making proteins can play a significant role in controlling cell development.

And in doing so, the researchers have unravelled a previously unknown mechanism for regulating the activity of genes, increasing our understanding of the way cells develop and opening the way to new possibilities for therapy.

Using the latest gene sequencing techniques and sophisticated computer analysis, a research group led by Professor John Rasko AO and including Centenary's Head of Bioinformatics, Dr William Ritchie, has shown how particular white blood cells use non-coding DNA to regulate the activity of a group of genes that determines their shape and function. The work is published today in the scientific journal Cell.

"This discovery, involving what was previously referred to as "junk," opens up a new level of gene expression control that could also play a role in the development of many other tissue types," Rasko says. "Our observations were quite surprising and they open entirely new avenues for potential treatments in diverse diseases including cancers and leukemias."

The researchers reached their conclusions through studying introns -- non-coding sequences which are located inside genes.

As part of the normal process of generating proteins from DNA, the code for constructing a particular protein is printed off as a strip of genetic material known as messenger RNA (mRNA). It is this strip of mRNA which carries the instructions for making the protein from the gene in the nucleus to the protein factories or ribosomes in the body of the cell.

But these mRNA strips need to be processed before they can be used as protein blueprints. Typically, any non-coding introns must be cut out to produce the final sequence for a functional protein. Many of the introns also include a short sequence -- known as the stop codon -- which, if left in, stops protein construction altogether. Retention of the intron can also stimulate a cellular mechanism which breaks up the mRNA containing it.

Dr Ritchie was able to develop a computer program to sort out mRNA strips retaining introns from those which did not. Using this technique the lead molecular biologist of the team, Dr Justin Wong, found that mRNA strips from many dozens of genes involved in white blood cell function were prone to intron retention and consequent break down. This was related to the levels of the enzymes needed to chop out the intron. Unless the intron is excised, functional protein products are never produced from these genes. Dr Jeff Holst in the team went a step further to show how this mechanism works in living bone marrow.

So the researchers propose intron retention as an efficient means of controlling the activity of many genes. "In fact, it takes less energy to break up strips of mRNA, than to control gene activity in other ways," says Rasko. "This may well be a previously-overlooked general mechanism for gene regulation with implications for disease causation and possible therapies in the future."
The published paper has nothing to do with junk DNA. Even if intron retention were a common mechanism of gene regulation (it is not), that would only account for about 100 base pairs per gene of additional sequence-dependant information. That's less than 0.1% of the genome.

This is a bad press release because it highlights information that is not in the published paper. The authors bear responsibility for press releases from their own institute that distort their published work. While they may not have written the press release, they presumably are quoted correctly and they should be aware of what's in the press release.

I wonder if they are willing to defend this press release as an accurate representation of their published work?


Friday, August 23, 2013

How IDiots Would Activate the GULOP Pseudogene

The enzyme L-glucono-γ-lactone oxidase is required for the synthesis of vitamin C. Humans cannot make this enzyme because the gene for this enzyme is defective [see Human GULOP Pseudogene]. The GenBank entry for this pseudogene is GeneID=2989. GULOP is located on chromosome 8 at p21.1 in a region that is rich in genes.

Here's a diagram that compares what is left of the human GULOP pseudogene with the functional gene in the rat genome.

Read more »

Some Questions for IDiots

Here's a short quiz for proponents of Intelligent Design Creationism. Let's see if you have been paying attention to real science. Please try to answer the questions below. Supporters of evolution should refrain from answering for a few days in order to give the creationists a chance to demonstrate their knowledge of biology and of evolution.

The bloggers at Evolution News & Views (sic) are promoting another creationist book [see Biological Information]. This time it's a collection of papers from a gathering of creationists held in 2011. The title of the book, Biological Information: New Perspectives suggests that these creationists have learned something new about biochemistry and molecular biology.

One of the papers is by Jonathan Wells: Not Junk After All: Non-Protein-Coding DNA Carries Extensive Biological Information. Here's part of the opening paragraphs.
James Watson and Francis Crick’s 1953 discovery that DNA consists of two complementary strands suggested a possible copying mechanism for Mendel’s genes [1,2]. In 1958, Crick argued that “the main function of the genetic material” is to control the synthesis of proteins. According to the “ Sequence Hypothesis,” Crick wrote that the specificity of a segment of DNA “is expressed solely by the sequence of bases,” and “this sequence is a (simple) code for the amino acid sequence of a particular protein.” Crick further proposed that DNA controls protein synthesis through the intermediary of RNA, arguing that “the transfer of information from nucleic acid to nucleic acid, or from nucleic acid to protein may be possible, but transfer from protein to protein, or from protein to nucleic acid, is impossible.” Under some circumstances RNA might transfer sequence information to DNA, but the order of causation is normally “DNA makes RNA makes protein.” Crick called this the “ Central Dogma” of molecular biology [3], and it is sometimes stated more generally as “DNA makes RNA makes protein makes us.”

The Sequence Hypothesis and the Central Dogma imply that only protein-coding DNA matters to the organism. Yet by 1970 biologists already knew that much of our DNA does not code for proteins. In fact, less than 2% of human DNA is protein-coding. Although some people suggested that non-protein-coding DNA might help to regulate gene expression, the dominant view was that non-protein-coding regions had no function. In 1972, biologist Susumu Ohno published an article wondering why there is “so much ‘ junk’ DNA in our genome” [4].
  1. Crick published a Nature paper on The Central Dogma of Molecular Biology in 1970. Did he and most other molecular biologists actually believe that "only protein-coding DNA matters to the organism?"
  2. Did Crick really say that "DNA makes RNA makes protein" is the Central Dogma or did he say that this was the Sequence Hypothesis? Read the paper to get the answer—the link is below).
  3. Is it true that, in 1970, the majority of molecular biologists did not believe in repressor and activator binding sites (regulatory DNA)?
  4. Is it true that in 1970 molecular biologists knew nothing about the functional importance of non-transcribed DNA sequences such as centromeres and origins of DNA replication?
  5. It is true that most molecular biologists in 1970 had never heard of genes for ribosomal RNAs and tRNAs (non-protein-coding genes)?
  6. If the answer to any of those questions contradicts what Jonathan Wells is saying then why do you suppose he said it?

Crick, F. (1970) Central Dogma of Molecular Biology. Nature 227:561-563. [PDF]

Tuesday, July 2, 2013

On the Misuse of the Term "Genetic Code"

Dan Graur is fed up with journalists who don't know the difference between the "genetic code" and the sequence of a genome. He's not alone. But, unlike the rest of us, Dan has a solution. It may be a little difficult to enforce ...

See: An Artistic Inspiration for Putting an End to the Misuse of the Term “Genetic Code”.


Wednesday, March 27, 2013

The Hardy-Weinberg Equilibrium

It's important to understand modern evolutionary theory and that means it's important to understand the Hardy-Weinberg Equation and what it means.

The significance is explained in all the leading textbooks on genetics and evolution. I've chosen the explanation given by Carl Zimmer and Douglas Emlen because I know that Carl has spent a good deal of time getting it right in his new book Evolution: Making Sense of Life.

Imagine that you have a population with two alleles, A and a, at a single locus. The frequency of the first allele is f(A) to which we assign the value p. The frequency of the second allele is f(a)=q. In a randomly mating sexual population the probability of an A sperm being produced is p and the probability of an a sperm being produced is q. Similarly, the probability of an A egg cell is p and the probability of an a egg cell is q. These probabilities, p and q, do not have to be equal.

We can calculate the probabilities of all possible combinations or sperm and eggs in the population from a the following diagram (Punnett square). This one is from Wikipedia.


Since the total probability has to equal one, we have ....

p² + 2pq + q² = 1
This is the Hardy-Weinberg equation or the Hardy-Weinberg Equilbrium. What does it mean? Let's quote Zimmer and Emlen (page 156).
Hardy and Weinberg demonstrated that in the absence of outside forces (which we describe later), the allele frequencies of the population will not change from one generation to the next. As we'll see below, this theorem is a powerful tool for population geneticists looking for evidence of evolution in populations. But it's important to bear in mind that it rests upon some assumptions.

One assumption of the model is that a population is infinitely large. If a population is finite, allele frequencies can drift randomly from generation to generation simply due to chance variation, in which alleles happen to be passed on to the next generation. (We will explore genetic draft in detail later in the chapter.) While no real population is infinite, of course, very large ones behave quite similarly to the model. That's because variation due to chance is inconsequential, and the allele frequencies will not change very much from generation to generation.

The Hardy-Weinberg theorem also requires all of the genotypes of the locusts are equally likely to survive and reproduce. If individuals with certain genotypes produced twice as many offspring as individuals with other genotypes, for example, then the alleles that these certain individuals carry will comprise a greater proportion of the total in the offspring generation than would be expected given the Hardy-Weinberg theorem. In other words, selection for or against particular genotypes may cause the relative frequencies of alleles to change and results in evolution.

Yet another assumption of the Hardy-Weinberg theorem is that no alleles enter or leave a population through migration. This assumption can be violated in a population if some individuals disperse out of it or if new individuals arrive. The model also assumes that there is no mutation in the population, because it would lead to new alleles

In each of these four cases, the offspring genotype frequencies will differ from the equilibrium predictions of the Hardy-Weinberg theorem. That is, because they alter allele frequencies from one generation to the next, selection, migration, and mutation are all possible mechanisms of evolution.

The Hardy-Weinberg theorem is useful because it provides mathematical proof that evolution will not occur in the absence of selection, drift, migration, or mutation. By explicitly delineating the conditions under which allele frequencies do not change, the theorem serves as a useful null model for studying ways of allele frequencies do change. The Hardy-Weinberg theorem helps us understand explicitly how and why populations evolve. By studying how populations deviate from the Hardy-Weinberg equilibrium, we can learn about the mechanisms of evolution.
There you have it. The Hardy-Weinberg describes the situation where evolution DOES NOT HAPPEN and thus serves as the null hypothesis for testing whether evolution is happening. Every undergraduate knows this.

Let's see if the Intelligent Design Creationists know this. I'm quoting "niwrad" from a post on one of the leading ID websites, Uncommon Descent: The equations of evolution.
For the Darwinists “evolution” by natural selection is what created all the species. Since they are used to say that evolution is well scientifically established as gravity, and given that Newton’s mechanics and Einstein’s relativity theory, which deal with gravitation, are plenty of mathematical equations whose calculations pretty well match with the data, one could wonder how many equations there are in evolutionary theory, and how well they compute the biological data related to the Darwinian creation.

....

The Hardy-Weinberg law mathematically describes how a population is in equilibrium both for the frequency of alleles and for the frequency of genotypes. Indeed because this law is a fundamental principle of genetic equilibrium, it doesn’t support Darwinism, which means exactly the contrary, the breaking of equilibrium toward the increase of organization and creation of entirely new organisms. To claim that the Hardy-Weinberg law explains evolution is as to say that in mechanics a principle of statics (immobility) explains dynamics (movement and the forces causing it).

....

So the initial question, how well math support Darwinian evolution, has the short answer: it doesn’t support evolution at all. Despite of the pretension of evolution to be a scientific theory with the mathematical certitude of the hard sciences, properly the equations of evolution do not exist.
As you can see, the Intelligent Design Creationists interpret the "Hardy-Weinberg law" very differently, I wonder who is right?

Let's check with Joe Felsenstein. He's an expert on population genetics so he should know. Read his decision at: Evolution disproven — by Hardy and Weinberg?.


Tuesday, March 26, 2013

Who Owns Your Genome?

The sequence of your genome contains lots of information about you. It also contains lots of information about your parents, your siblings, and your children. That's why you should not make your genome sequence public without obtaining their permission.

The sequence of Henrietta Lacks' genome was just published (HeLa cells) and nobody bothered to seek permission from her survivors. Jonathan Eisen has a comment and he has also collected all the information on the internet [HeLa genome sequenced w/o obtaining permission/consent from family - some comments and background]. Be sure to read the New York Times article by Rebecca Skloot: The Immortal Life of Henrietta Lacks, the Sequel. She says,
LAST week, scientists sequenced the genome of cells taken without consent from a woman named Henrietta Lacks. She was a black tobacco farmer and mother of five, and though she died in 1951, her cells, code-named HeLa, live on. They were used to help develop our most important vaccines and cancer medications, in vitro fertilization, gene mapping, cloning. Now they may finally help create laws to protect her family’s privacy — and yours.
In my opinion, there is no excuse for publishing this genome sequence without consent.

Razib Khan disagrees. He thinks that he can publish his genome sequence without obtaining consent from anyone else and I assume he feels the same way about the sequence of the HeLa genome [Henrietta Lacks’ genome, and familial consent].


Friday, March 15, 2013

On the Meaning of the Word "Function"

A lot of the debate over ENCODE's publicity campaign concerns the meaning of the word "function." In the summary article published in Nature last September the authors said, "These data enabled us to assign biochemical functions for 80% of the genome ...." (The ENCODE Project Consortium, 2012).

Here's how they describe function.
Operationally, we define a functional element as a discrete genome segment that encodes a defined product (for example, protein or non-coding RNA) or displays a reproducible biochemical signature (for example, protein binding, or a specific chromatin structure).
What, exactly, do the ENCODE scientists mean? Do they think that junk DNA might contain "functional elements"? If so, that doesn't make a lot of sense, does it?

Ewan Birney tried to address this definitional morass on his blog [ENCODE: My own thoughts] where he says ....
It’s clear that 80% of the genome has a specific biochemical activity – whatever that might be. This question hinges on the word “functional” so let’s try to tackle this first. Like many English language words, “functional” is a very useful but context-dependent word. Does a “functional element” in the genome mean something that changes a biochemical property of the cell (i.e., if the sequence was not here, the biochemistry would be different) or is it something that changes a phenotypically observable trait that affects the whole organism? At their limits (considering all the biochemical activities being a phenotype), these two definitions merge. Having spent a long time thinking about and discussing this, not a single definition of “functional” works for all conversations. We have to be precise about the context. Pragmatically, in ENCODE we define our criteria as “specific biochemical activity” – for example, an assay that identifies a series of bases. This is not the entire genome (so, for example, things like “having a phosphodiester bond” would not qualify). We then subset this into different classes of assay; in decreasing order of coverage these are: RNA, “broad” histone modifications, “narrow” histone modifications, DNaseI hypersensitive sites, Transcription Factor ChIP-seq peaks, DNaseI Footprints, Transcription Factor bound motifs, and finally Exons.
That's about as clear as mud.

We all know what the problem is. It's whether all binding sites have a biological function or whether many of them are just noise arising as a property of DNA binding proteins. It's whether all transcripts have a biological function or whether many of those detected by ENCODE are just spurious transcripts or junk RNA. These questions were debated extensively when the ENCODE pilot project was published in 2007. Every ENCODE scientist should know about this problem so you might expect that they would take steps to distinguish between real biological function and nonfunctional noise.

Their definition of "function" is not helpful. In fact, it seems deliberately designed to obfuscate.

Let's see how other scientist interpret the ENCODE results. In a News & Views article published in Nature last September, Joseph R, Ecker (Salk Institute scientist) said ...
One of the more remarkable findings described in the consortium's 'entre&eacute:' paper is that 80% of the genome contains elements linked to biochemical function, dispatching the widely held view that the human genome is mostly 'junk DNA.'
That makes at least one genomics worker who thinks that "biochemical function" and junk DNA are mutually exclusive.

Recently a representative of GENCODE responded to Dan Graur's criticism [On the annotation of functionality in GENCODE (or: our continuing efforts to understand how a television set works)]. This person (JM) says ...
Q1: Does GENCODE believe that 80% of the genome is functional?

As noted, we will only discuss here the portion of the genome that is transcribed. According to the main ENCODE paper, while 80% of the genome appears to have some biological activity, only “62% of genomic bases are reproducibly represented in sequenced long (>200 nucleotides) RNA molecules or GENCODE exons”. In fact, only 5.5% of this transcription overlaps with GENCODE exons. So we have two things here: existing GENCODE models largely based on mRNA / EST evidence, and novel transcripts inferred from RNAseq data. The suggestion, then, is that there is extensive transcription occurring outside of currently annotated GENCODE exons.
There's another scientist who thinks that 80% of the genome has some biological activity in spite of the fact that the ENCODE paper says it has "biochemical function." I don't think "biological activity" is compatible with "junk DNA," but who knows what they think?

Since this person is part of the ENCODE team, we can assume that at least some of the scientists on the team are confused.

The Sanger Institute (Cambridge, UK) was an important player in the ENCODE Consortium. It put out a press release on the day the papers were published [Google Earth of Biomedical Research]. The opening paragraph is ...
The ENCODE Project, today, announces that most of what was previously considered as 'junk DNA' in the human genome is actually functional. The ENCODE Project has found that 80 per cent of the human genome sequence is linked to biological function.
It looks like the Sanger Institute equates "biochemical function" and "biological function" and it looks like neither one is compatible with junk DNA.

I think the ENCODE leaders, including Ewan Birney, knew exactly what they were doing when they defined function. They meant "biological function" even though they equivocated by saying "biochemical function." And they meant for this to be interpreted as "not junk" even though they are attempting to backtrack in the face of criticism.


The ENCODE Project Consortium (2012) An integrated encyclopedia of DNA elements in the human genome. Nature 489: 57-74. (E. Birney, corresponding author)

Thursday, March 14, 2013

Anonymous Nature Editors Respond to ENCODE Criticism

There are now been four papers in the scientific literature criticizing the way ENCODE leaders hyped their data by claiming that most of our genome is functional [see Ford Doolittle's Critique of ENCODE ]. There have been dozens of blog postings on the same topic.

The worst of the papers were published by Nature—this includes the abominable summary that should never have made it past peer review (Encode Consortium, 2012).

The lead editor on the ENCODE story was Brendan Maher and he promoted the idea that the ENCODE results showed that most of our genome has a function [ENCODE: The human encyclopaedia]
The consortium has assigned some sort of function to roughly 80% of the genome, including more than 70,000 ‘promoter’ regions — the sites, just upstream of genes, where proteins bind to control gene expression — and nearly 400,000 ‘enhancer’ regions that regulate expression of distant genes.
Read more »

Wednesday, March 13, 2013

Ford Doolittle's Critique of ENCODE

Ford Doolittle has never been one to shy away from controversy so it's not surprising that he weighs in against the misleading publicity campaign launched by ENCODE leaders last September (Doolittle, 2013). Recall that Ewan Birney and other prominent members of the consortium promoted the idea that our genome contained an extensive array of regulatory elements and that 80% of our genome was functional [Ewan Birney: Genomics' Big Talker] [ENCODE Leader Says that 80% of Our Genome Is Functional] [The ENCODE Data Dump and the Responsibility of Scientists].

This is the fourth paper that's critical of the ENCODE hype. The first was Sean Eddy's paper in Current Biology (Eddy, 2012). The second was a paper by Niu and Jiang (2012), and the third was a paper by Graur et al. (2013). In my experience this is unusual since the critiques are all directed at how the ENCODE Consortium interpreted their data and how they misled the scientific community (and the general public) by exaggerating their results. Those kind of criticisms are common in journal clubs and, certainly, in the blogosphere, but scientific journals generally don't publish them. It's okay to refute the data (as in the arsenic affair) but ideas usually get a free pass no matter how stupid they are.

In this case, the ENCODE Consortium did such a bad job of describing their data that journals had to pay attention. (It helps that much of the criticism is directed at Nature and Science because the other journals want to take down the leaders!)

Read more »

Thursday, January 31, 2013

Theme: Mutation

This is a collection of Sandwalk posts on mutation starting in 2007. The latest ones are at the bottom of the list.

March 27, 2007
Silent Mutations and Neutral Theory
Neutral Theory and random genetic drift explains variation and it also explains molecular evolution and the (approximate) molecular clock. There are no other explanations that make sense and nobody has offered a competing explanation since Motoo Kimura (1968) or Jack King and Thomas Jukes (1969) published their papers almost fifty years ago. (Aside from occasional nitpicks, of course. There are always scientists who like to show that some mutations that were thought to be neutral are actually beneficial or deleterious. None of them have mounted a serious claim that most variation or most of molecular evolution can be explained by natural selection.)

April 19, 2007
Haldane's Dilemma
This is very interesting. Dembski has teamed up with Walter ReMine, demonstrating once again that the old addage "opposites attract" does not apply to kooks.

ReMine has an article on Uncommon Descent where he pushes his usual whine about evil scientists and how their world-wide conspiracy has kept him from revealing the fatal flaw in evolution [Evolutionist withholds evidence on Haldane’s Dilemma]. I can see how similar this is to Intelligent Design Creationism.


Read more »

Wednesday, January 16, 2013

Why Do the IDiots Have So Much Trouble Understanding Introns?

Most eukaryotic genes have introns. Introns make up about 18% of the DNA sequences in our genome. Most of these sequences are junk but introns are functional and up to 80bp of each intron is required for proper splicing. The essential sequences contain the 5′ splice site (~10 bp); the 3′ splice site (~30 bp): the branch site (~10 bp); and enough additional RNA to form a loop (~30 bp). The branch site and the splice sites are where specific proteins bind to the mRNA precursor [Junk in Your Genome: Protein-Encoding Genes]. It turns out that within introns about 0.37% of the genome is essential and about 17% is junk.


Read more »

Thursday, December 6, 2012

James Shapiro Never Learns

One of the remarkable things about kooks is that they are incredibly resistant to learning from their mistakes. James Shapiro gives us a fine example in one of his latest articles on The Huffington Post where he tries to convince us that the old definition of "gene" has outlived its usefulness. According to Shapiro, "DNA and molecular genetics have brought us to a fundamentally new conceptual understanding of genomes, how they are organized and how they function."

Really? While we all can agree that there's no definition of "gene" that doesn't have exceptions, we can surely agree that some definitions work pretty well. I've argued that defining a gene as, "a DNA sequence that is transcribed to produce a functional product" works well in most cases [What Is a Gene?].

Let's see how James Shapiro handles this problem.1 He says,
The identification of DNA as the key molecule of heredity and Crick's Central Dogma of Molecule Biology [Crick 1970] initially seemed to confirm Beadle and Tatum's "one gene -- one enzyme" hypothesis.
I've already explained that Shapiro doesn't understand the Central Dogma of Molecular Biology even though he quotes the Francis Crick papers that explain it correctly [Revisiting the Central Dogma in the 21st Century]. I also made this point in my review of his book: Evolution: A View from the 21st Century.

Read more »

Friday, October 5, 2012

An Online Course for Intelligent Design Creationists

About 99% of all books and posts by Intelligent Design Creationist consists of criticisms of evolution—which they mistakenly refer to as "Darwinism."

What this usually reveals is that the typical IDiot doesn't understand evolution. But there's at least one Intelligent Design Creationist who recognizes that this is a problem. Jonathan McLatchie (Jonathan M) recommends that his colleagues take an free online course in order to learn about evolution [Free Online Course: Introduction to Genetics and Evolution]. He writes,
Critics of modern evolutionary theory have an intellectual responsibility to strive to understand the paradigm that they are critiquing, preferably to a level where they can clearly articulate the key propositions of evolutionary theory and offer a standard defense of them.

Richard Hoppe, at the Panda’s Thumb blog, drew my attention to a free online course on the subject of genetics and evolution. You can, as I have done, sign up for (and read about) the course at this link.

...

I particularly recommend that those among us who don’t have a strong biology background take this course. It is very important that we ID proponents make sure we have a robust grasp of what evolutionary theory is saying and why it says it, so that no one can say we haven’t given it a fair hearing.
Wouldn't it be nice if most IDiots followed Jonathan McLatchie's advice? In just a few months they could learn that modern evolution and genetics includes all sorts of things that Darwin never knew! Imagine what a relief it would be if they stopped referring to us all as "Darwinists" and started to understand that evolution is a fact.

Not holding my breath.


Tuesday, October 2, 2012

Reddit: We are the Encyclopedia of DNA Elements (ENCODE) Consortium.

There's been a lot of talk recently about the discussion on reddit concerning the ENCODE publicity fiasco.

Here's the forum ...
AskScience Special AMA: We are the Encyclopedia of DNA Elements (ENCODE) Consortium. Last week we published more than 30 papers and a giant collection of data on the function of the human genome. Ask us anything!
It's interesting to see how some of the consortium members are responding to criticism. My personal view is that none of them seem to be very knowledgeable about genome biology and the work that has been published over the past 40 years.


Friday, September 14, 2012

Does the Central Dogma Still Stand?

Lots of people don't understand the Central Dogma of Molecular Biology and that's probably why there are so many articles announcing its death. The article and book by James Shapiro is just one example [Revisiting the Central Dogma in the 21st Century].

The correct version of the Central Dogma of Molecular Biology is .... [see Basic Concepts: The Central Dogma of Molecular Biology]
... once (sequential) information has passed into protein it cannot get out again (F.H.C. Crick, 1958)

The central dogma of molecular biology deals with the detailed residue-by-residue transfer of sequential information. It states that such information cannot be transferred from protein to either protein or nucleic acid. (F.H.C. Crick, 1970)
Eugene Koonin has an article in Biology Direct entitled Does the central dogma still stand (Koonin, 2012).

Read more »