Showing posts with label DNA. Show all posts
Showing posts with label DNA. Show all posts

Wednesday, October 05, 2022

The physics of life and living things

 kw: book reviews, nonfiction, biophysics, dna, biomolecules, self-assembly

A friend of mine is a biophysicist. I asked him once what he does. He said he wasn't working in biophysics, but was writing computer code for a government agency. I didn't press further. When I saw that a book about biophysics I decided to read it: So Simple a Beginning: How Four Physical Principles Shape Our Living World by Raghuveer Parthasarathy. The book's title opens a key sentence in the last paragraph of Darwin's On the Origin of Species.

I was a physics major in 1969 and 1970. While physics deals with phenomena on all scales, from the gravitational and electromagnetic fields that can span the universe to the Planck Length, the smallest possible "useful" unit of length, most physicists at that time worked with subatomic particles, smaller than an atom by a factor of about 10,000, but still very large compared to the Planck Length: If a proton were enlarged to span the distance between Hartford, CT and Providence, RI, about 100 km, the Planck Length would become about the size of a proton.

If we move up in the scale of things to the nano-realm, from the size of an atom (an iron atom's diameter is 0.26 nanometers) and that of a DNA molecule (~10 nm in diameter, but very, very long) to the size of bacterial cells (500 nm to 10,000 nm), and further to the cells of animals and plants (10,000 - 100,000 nm), we are in the realm of biophysics.

The author first presents four physical principles that govern living things:

  • Self-assembly – biological things typically "build themselves", such as the "liquid membrane" of a cell or a cell's nucleus, or a soap bubble as seen here. The electrochemical properties of all biomolecules facilitate their roles.
  • Regulatory Circuits – phenomena such as the expression of a gene involve feedback loops with several elements.
  • Predictable Randomness – this is the basis of statistical inference, and underlies Brownian Motion, which is the "motor" of many actions within cells.
  • Scaling – relationships between length, area and volume regulate what is possible at different sizes, and underlie the dramatic difference between the kinds of legs that work for a rhinoceros beetle, compared to those of a rhinoceros, for example.

The book contains many illustrations drawn by the author, such as the ones shown above. 

The author proceeds from basic facts about atoms and molecules to the molecules need to operate a living cell, primarily DNA, RNA, proteins, sugars and lipids (fats). Examples of self-assembly introduce the ways these molecules' properties facilitate the construction of all the organelles in a cell. Certain operations require more specialized machinery; an example is ferrying certain products over longer distances (clear across an animal cell, which is 10-100x as wide as a bacterium, for example), because Brownian motion is too slow. This is carried out by special kinds of molecules that "walk" along the fibers that form an internal skeleton of the cell. Shorter range transport is typically carried out quite efficiently by relying on Brownian movement to jostle molecules around until they latch onto their targets. When a motion of a micron or so is needed, transport time is around a microsecond.

The four principles listed above are emergent properties of biomolecules in an environment warm enough for Brownian motion to help them go where they need to go, at least inside bacterial cells. Over evolutionary time, mechanisms have been developed that facilitate larger-scale things and operations, right up to the size of a blue whale or redwood tree. This was apparently a hard problem. The "boring billion" refers to a billion-year period during which bacteria and archaea, having developed quite a lot of sophistication, including the ability to aggregate into large assemblages such as stromatolites, didn't do much at all. Finally, eukaryotic cells arose, and things got a lot less boring. Animals, plants, fungi and protozoa are composed of eukaryotic cells (the word means "cells with a nucleus"). The largest eukaryotic cells are the neurons that run end-to-end in large animals such as whales or giant squids. The largest bacteria or archaea are 1/100 millimeter long (well, there are a very few species of bacteria that are 10-20 mm long and 3/4 mm diameter. All the rest are microscopic).

The last section of the book deals with the genetic revolution, first in reading ("sequencing") DNA and now writing it, or editing it. The prospect of "designer babies" and "clone armies" emphasizes that these matters have moral aspects. We have to work out "who decides what is moral" (particularly because most genetic scientists are atheists and so have no external moral compass). The author is optimistic that this can be carried out without much drama. 

I am less optimistic. The author discusses Chinese researcher He Jiankui, who announced having used CRISPR/CAS9 to gene-edit twin embryos. The girls were born in 2018. The Chinese government, partially under outside pressure, reacted strongly, shut down He's lab and jailed him. I suspect the next researcher who decides to give it a go won't announce anything. This may have already happened. Not everyone is willing to wait for consensus. The technique of "gene drive", which can rapidly send a species, such as a noxious sort of mosquito, into extinction, is an even scarier prospect. There is no guarantee that a gene drive that works in the Anopheles mosquito only will not mutate into one that crosses into another species, and eventually spreads and spreads. Think of "Ice-Nine" in Cat's Cradle by Kurt Vonnegut.

On another note: In the present technical environment dominated by Big Data, the author presents a good case for understanding—based on hypothesis, experiment, synthesis, and theory—wherever possible. He uses the example of making numerous experiments with a ball, rolled down a ramp and off the table, and measuring where it hits the floor. One could prepare a table based on thousands of such experiments. Then someone could use that table to determine, based on a ball's velocity and height from the floor, to predict where it will land. But a smaller number of experiments can underlie the development of a formula by which one can calculate the landing distance, without needing to interpolate from a table. The formula is based on understanding what gravity does, and experiments to confirm the strength of gravity. It isn't too extreme to say that Big Data is often used blindly. Physics, including biophysics, leads to understanding and removes the blinders.

I probably haven't demonstrated a great deal of my own understanding of biophysics. I have a lot to think over. This book is a marvelous introduction to the subject.

Errata: On p.266, illustrating how gene drive works, the example is a species of mosquito, gray in color. Sometimes a mutation occurs, yielding a black insect. In mid-discussion this sentence occurs, "Suppose just one individual has the gray mutation." It should be "…black mutation", as is clear from the accompanying illustration and the rest of the discussion.

Friday, June 25, 2021

Mechanisms of evolutionary saltation

 kw: book reviews, nonfiction, evolution, development, evo-devo, dna, molecular biology

It took Stephen Jay Gould twenty years to write The Structure of Evolutionary Theory. It will take me longer than that to read it. I bought a copy when it was released in 2002, and I am only one-third of the way through it. I intend to read it all.

You may know that Dr. Gould is one originator of the hypothesis of Punctuated Equilibrium: fossils show that species tend to persist almost unchanged for periods of a million years to tens of millions of years, and then undergo rapid change, during which new species arise quickly. His book discusses this matter, and much more, in a historical context. I probably haven't come to the "good bits" yet. But others have, and scientists continue to discover new aspects of genetics and evolution, so I read widely in the field.

A stellar new volume is Some Assembly Required: Decoding Four Billion Years of Life, from Ancient Fossils to DNA, by Neil Shubin, a researcher and professor of Organismal Biology and Anatomy. His book outlines certain events in the history of evolutionary thought and genetic discovery, with an emphasis on a seminal thought expressed by one of his mentors, "Things didn't start when you think they did."

For example, he discusses wings and flight. Flight arose at least four times, in insects, pterosaurs (reptiles), bats (mammals), and birds. In each case, wings didn't appear all at once, but we find that earlier tissues and structures with different functions were co-opted to become wings, in a rather short time span. Furthermore, by digging into the genetics of wing development, he and others have found that the precursors to wings have similar origins in these very different types of animals. I can't do justice to an explanation of this. The book's discussion is brief yet illuminating. Bottom line: structures that could later become wings were developed long ago, for other purposes, and only millions of years later did the new function of "catching air" arise, requiring comparatively modest further development.

Another example every school child of my generation learned (do they still?): lungs developed from flotation bladders in fish. Whether the bladder developed a connection to the mouth by accident or for another reason, once that occurred, the already-existing practice many fish had of gulping air when the oxygen supply in the water was low, when combined with a new place to put that air, allowed these fish to survive better. Also, fins in some fish species were modified with "lobes", and these precursors of legs were used to move along the bottom of a lake or stream. "Walking" in this way keeps the animal below the worst of currents that it wants to move against; only later were the "legs" used to move onto and across the land, and eventually they were strengthened into legs strong enough to support amphibian bodies.

The pace of evolutionary development was very slow long ago, but has been accelerated over time with various developments. The first living things were like bacteria, or perhaps their cousins, the archaea. These together are called prokaryotes ("before the nucleus"): a prokaryote cell's DNA is a loosely-wound loop that runs throughout the interior of the cell. After a half billion years of gradual proliferation, some prokaryotes developed photosynthesis. Before that all life was chemosynthetic, using processes such as robbing sulfur from metal sulfides for energy. There are several kinds of photosynthesis; only one, initially, used CO2 and water to produce sugar, with oxygen (O2) as a waste product. Today's cyanobacteria (also called blue-green algae) are descended from O2-producing bacteria that arose about 3,500 million years ago.

At first, all the excess oxygen was used up by oxidizing sulfides into oxides and sulfates. This slowed down after another billion years, and oxygen began accumulating into the atmosphere. From 2,500 million to 1,500 million years ago, during the "boring billion", O2 slowly increased to about 2%. Then things began to change more rapidly. About that time, or perhaps a few hundred million years earlier, more complex cells developed. The DNA was encapsulated inside its own membrane, and at least two events of engulfment happened. Most probably the first "guests" invited into a larger cell (or they were invaders that were subdued and enslaved) were cyanobacteria, which were put to work turning air into sugar, while being kept safe inside the cell. Now they are called chloroplasts. Almost immediately, the second event was the capture of certain small, energy-efficient bacteria that probably looked a lot like E. coli. These became mitochondria. These larger, compound types of cell are called eukaryotes ("good nucleus"). A discussion of this process on pp 195-6 seems to imply that plants have chloroplasts but not mitochondria; not so, they have both. They need both!

Single-celled eukaryotes are still with us, most familiarly in the form of protozoa such as Amoeba and Paramecium. Some time before 1,000 million years ago, molecular mechanisms that were being used to attach to a substrate or to food particles before "swallowing" them, were re-purposed to allow cells to cling together. In the book a lovely discussion of choanoflagellates discusses how this works. The earliest multi-cellular creatures, whether they were proto-plants (with chloroplasts) or proto-animals, had a variety of shapes, but mostly looked quilt-like or mat-like. Some time around 600 million years ago an organizing principle arose. To introduce it, we must look into segmentation.

The prototype of segmented animals is the earthworm. You can see the segments, a lot of them. We vertebrates are segmented also. Our spine expresses the segmentation. Not all animals are segmented; in fact most phyla are not, but all have some kind of body plan. The Homeobox, or HOX, genes are controllers of body plan development. Every animal species has them. The HOX genes are organizers, and represent a kind of meta-control. The simple idea that we have "a gene" for this or that is a big distortion. Even in a simple animal such as a 1mm nematode, there are HOX genes that make the difference between front and rear and so forth. The more complicated sets of HOX genes found in more complex animals arose from reduplication.

Reduplication is a big theme in genetics. The added sets of HOX genes we need are an example. Mutation isn't a matter of creating a new, complex function out of whole cloth. It proceeds by various errors of copying, which will usually just kill the animal, but occasionally are at least mostly harmless, and over time, the odd bit can gain a new function. The most common mutations are single-point changes, such as from an A to a G in the genetic code. But whole segments can be duplicated, particularly during the "crossover" that occurs during the production of eggs and sperm. If an extra set of HOX genes is produced, one set can go its merry way, controlling the body's development, while the other set is modified and can lead to an extra function or body part or even whole section. Again, this isn't usually good for the animal, but it can be.

Segmentation arose by reduplication. In some cases, many identical segments were produced (earthworm). In others, the segments became specialized. The HOX genes control all this. The illustration, from this article at Socratic.org, compares the HOM genes (as they are called for insects) with the multiple sets of HOX genes in humans and mice. The segmentation of the insect's body is emphasized in the drawing.

It may seem strange that we share this organizing principle with fruit flies, mice and everything else. From an evolutionary perspective, it makes sense. The system works, and we can see that it works, for it has produced millions of species of animal.

Now to the matter of saltation, as in this review's title. Saltation is a dirty word to most evolutionists. It has come to mean things like a rabbit suddenly "evolving" into a dog or a horse. That's ludicrous.

In a proper sense, saltation means "jumping", and the concept (if not the term) had to be coped with once Barbara McClintock discovered jumping genes in corn. They have since been found in every species, and certain kinds of them form much of the "junk DNA" found between the genes in our genome. But others have been put to use, and HOX may be an example.

Just by the way, there's a lot less "junk" in our DNA than early reports claimed. Just 2% of it codes for proteins. An additional 2% (perhaps much more) consists of regulatory sequences that control when and how the genes make those proteins, a further 8-10% consists of deactivated viruses, which form a "library" of stuff gathered from everywhere, that can be re-purposed. Some is apparently second- and third-level regulatory stuff. About 2/3 is "palindromic repeats" (such as AATTGCACGTTAA) that consist of head-to-toe copies of "stuff", which at the moment, is at least useful for landmarks used by CRISPER/CAS gene editing.

All these things, and many more discussed in the book, are mechanisms for more rapid evolutionary change, compared to waiting for single-letter mutations to accumulate. Even over millions of years, that process is dreadfully slow. The beauty of these mechanisms, still being discovered, is that they allow big changes to occur without disaster.

The Earth would seem quite full of many species, were there only a few tens of thousands of them. It is astonishing that there are millions! I work in the "shell room" of a museum, and every time I open a cabinet I see something new, just among the mollusks! That room contains specimens for more than 20,000 species...of seashell! Nearly 100,000 are known. It seems that life, having figured out how to spin out new kinds of creatures, is still ramping up. While we may be driving thousands of species to extinction, it is likely that new species are arising even faster. If we attain wisdom enough to let nature alone and "live lightly", we may see even more variety in the multiplicity of life in the future.

Monday, August 10, 2020

DNA is part of a feedback loop

kw: book reviews, nonfiction, science, genetics, dna, history, stories

As I mentioned a couple of reviews back I bought an eBook bundle, three by Sam Kean. This is the third, The Violinist's Thumb: And Other Tales of Love, War, and Genius, as Written in Our Genetic Code.

Firstly, a remark on the cover art:

The cover designs for the first two books are by Will Staehle. Some might call them "busy", but I really like them. Keith Hayes designed the third cover, and I think it matches the subject very well. Kudos to some folks who are seldom recognized.

The violinist in question is Niccolò Paganini (1782-1840), who was probably the most accomplished virtuoso ever. His astonishing ability is typically attributed to the freakish flexibility of his hand and finger joints (and all the other joints also). It is worth mentioning that he worked very hard, practicing endlessly. We find that such flexibility is both a blessing and a curse, a product of a handful of rare snippets of DNA, alleles (a better word than "mutations") that together affect the ligaments. The "curse" part is that such loose joints are frequently painful. However, the times being what they were, and the fact that Paganini also contracted both tuberculosis and syphilis, make his case hard to diagnose two centuries after the fact. Furthermore, the thumb of Paganini was notable not mainly for flexibility, but incredible strength. He could hold a saucer in one hand, press with the thumb, and break it.

Sam Kean likes his books to progress from the micro- to the macro-scale (In Disappearing Spoon, about the elements, each chapter had to have its own structure). The micro-scale of DNA is small indeed, and so many other authors have "gone into the weeds" with ribose sugars, bases, and hydrogen bonds, that there is little need to dwell on them yet again. The lower level stuff we'll find in this book is more about transcription, translation, and gene editing (I'd like to have seen more about gene editing, but in 2012 the details of intron removal and the various ways 20,000 "genes" can produce a few million proteins were little known, and much is still opaque on that subject).

As was brought out even more forcibly in Dueling Neurosurgeons, we learn a lot from the ways things go wrong. Paganini was one who turned a genetic handicap into a career. But when we say that a certain disorder is "in the genes", it is not always so clear-cut. Cystic fibrosis results when the cftr gene doesn't work right, allowing salt transport to fail and thin mucus in the lungs to thicken. Other syndromes such as diabetes and cancer do not result from faulty genes, per se, but from mis-regulation: a genetic sequence may be either over-stimulated (most cancers) or under-stimulated (hypoglycemia or certain kinds of diabetes). One tragic case in Chapter 8 showed that the placenta does not prevent all possible genetic transfer between mother and child, for example.

The actual number of human genes is still disputed. So is the definition of "gene". The prior dogma was DNA→RNA→protein. This is, like, dinosaur-level out of date! One article by M. Pertea and others counts 21,306 protein-coding genes and 21,856 non-coding genes. At one time, the only first group would have been called "genes". The non-coding genes carry on regulatory functions, such as directly triggering or halting the activity of a coding gene or producing RNA that does so less directly; or they affect how introns are removed and the bits of RNA are stitched together. Some proteins can only be produced after more than 100 strings of RNA are connected (and in the right order!) from a transcription that may be much, much larger than the final, edited transcript. There are tons of things going on that we don't yet understand. 

Those 43,000+ "genes" still comprise only a few percent of our DNA. About 8% (at least twice as much!) is made up of various broken virus genomes, and perhaps some that aren't so broken. Retroviruses leave workable copies of their genome in every cell they infect. We all carry many such.

How does all this go together to produce a human? or, for that matter, a fruit fly, a tiny worm, or a blue whale? The book takes a step in this direction with Chapter 8: "Love and Atavisms; what makes a mammal a mammal?" Somehow, a line of reptiles developed the placenta, using lots of virus DNA to do so. Many details are found in this chapter. The placenta has a heck of a job. It has to protect a growing baby from the immune system of its mother, and it must also protect the mother from the developing immune system of this "new resident", all the while allowing the baby to conscript a large proportion of the mother's nutrition for its own use. Many viruses are adept at avoiding or even silencing immune system counterattacks, and these capabilities are built into cells that face both ways, outward from the baby/mother boundary. But the placenta is not bullet-proof, to mix a simile. A mother who has many sons may notice that the older ones are typically more "manly" and the younger ones less so, if not effeminate (Bible readers will recall that King David, though called "mighty", was the youngest, and smallest, of eight sons). This may be due to the mother's immune system getting better at influencing the environment of the developing baby within.

And what of our immense brains? In two places, the book discusses two genes, microcephalin and aspm, that are related to brain size. Certain alleles of these genes lead to babies with no brain or a very small one, a tragic circumstance. Still more fascinating, checking the DNA clock on these genes shows that the modern form of microcephalin arose about 37,000 years ago and soon swept through the entire population of humans, and aspm did the same thing about 6,000 years ago (Bible readers who happen to accept evolution will find that intriguing, because the Biblical story of God molding Adam's body and then putting a spirit of life into him is thought to date to just 6,000 years ago).

A big lesson of the book is the author's growing realization that DNA is not destiny. Or, not usually. There is seldom a single-point "thing" that causes a trait. Even blue eyes/brown eyes are more complicated than that, and eye color is a rather simple system. The author had a DNA test done, but initially asked that the gene(s) "for" Parkinson's Disease be hidden from him, because of family history. Months or years later he reports that he realized that DNA is probabilistic, not causative. So he unlocked the locked section, and found that there was apparently no problem, but then a revision a few days later showed a "slight chance" that he might develop Parkinson's. By then he had the mental fortitude to accept the news.

I've had similar worries, because one line of my family carries Alzheimer's Disease, and another line (this is very recent news to us) carries Lewy Body Dementia. Two arrows pointed at my brain. Considering my age and general health, Alzheimer's is a no-show (Mom was afflicted beginning in her fifties), but Lewy Body shows up later, so who knows? Probability is not destiny.

It will take a long time, may be a really, really long time, to know enough about DNA to begin to "take control". From time to time a sci-fi novel gets into this territory, and posits a future of people "engineered" to live on Mars without a spacesuit, or people with gills who can live under the sea, and so forth. We are a long, long way from learning if this is even possible without screwing up something else that underlies our humanity.

The last chapter introduces DNA as a computing mechanism. The stuff is very, very good at pattern matching. Some tests have been made, and the author describes a DNA algorithm to solve the "traveling salesman" problem, something your GPS unit has to do, and it usually does it pretty well. DNA has the potential to solve huge problems, like the salesman routing problem with 500 stops; we're talking age-of-the-universe time scales for today's supercomputers to deal with that one. However, once the "problem" is solved—it takes about a minute—you have a vat of DNA soup with "the answer" and billions or trillions of partial answers, and you are faced with winnowing out the longest chain in the whole bowl from all the others. That may also be an age-of-the-universe sized problem!

A side note: sorting has been studied more than any other kind of computer operation. The most optimized sort method can still take a long time if you need to sort billions of items. By contrast, the "spaghetti sort" is very fast. Just produce strands of spaghetti (or a more robust sort of rigid rod), cut to length, with the key value written on each one. For a modest size sort, a few hundred strands, you can hold it in your hand and stand it on the table. Then remove the strands, longest first, and read off the key numbers on each. Now, making the strands, and reading the results can be time consuming, but the sorting operation takes a fraction of a second. To sort billions or even trillions of "strands" (presumably of very long pieces of welding rod or something), the actual sorting operation would be almost instant, but the construction of the rods, and reading the results, are still incredibly time-consuming! That's my analogy to DNA calculation.

The book is incredibly fun to read. I like Sam Kean's writing. After catching up on books for other subjects, I may just snarf up another triple-pack of his books.

Thursday, March 29, 2018

Can rewilding rescue the permafrost?

kw: book reviews, nonfiction, science, cloning, DNA, woolly mammoths, rewilding

All the people in Woolly: The True Story of the Quest to Revive one of History's Most Iconic Extinct Creatures are real, as are all the events prior to the last two chapters. Author Ben Mezrich used interviews and published materials to produce a narrative that recounts events over the past half century or so—though mostly over the past twenty-odd years—leading into a concerted effort to produce the DNA needed to revive the species Mammuthus primigenius, the Woolly Mammoth.

The main protagonist is Dr. George Church, a very active and productive genetic researcher. If you've heard his name at all, it is probably in connection with the Human Genome Project. Dr. Church is somewhat self-effacing compared to others who "got famous". Famous or not, he is a prime problem-solver, and gathers problem-solvers around him. That's what you need to tackle a project like this.

I was most intrigued by a side theme of the book, the rewilding of Siberia and possibly northern Canada, with the aim of restoring the permafrost. This entails gathering not just extinct pachyderms, but a number of living cold-adapted herbivores such as Musk Oxen. As I understand it, the large mammals of the Pleistocene fauna could churn the upper surface of the ground, which tends to allow the winter chill to make new permafrost in wintertime but blocks solar heating in summertime. The idea is to keep the huge carbon stores of the permafrost from being oxidized and thus adding many-fold to the greenhouse heating being caused by extra carbon dioxide already released by our burning of fossil fuels. That idea alone was enough to push Dr. Church over the threshold from "We can revive the Mammoth, but should we?" to "We can and we should!"

The larger key idea of the book is the concept, not of simply "finding" mammoth DNA, but learning enough from the DNA sequence to determine the key differences between mammoth DNA and Asian elephant DNA, so as to rewrite critical sections of an elephant genome and thus produce a viable mammoth ovum.

The book ends with a scene of the first mammoth returned to Siberia, perhaps as early as about 2020. However, in an epilogue by Dr. Church, he considers a more realistic figure to be 15-20 years from now. Considering the number of breakthroughs already made, a living mammoth might appear sooner than that. Producing a herd of them will take longer, but a herd is needed to have a useful effect on Siberian (or Canadian) permafrost.

Saturday, November 19, 2005

Dinosaur Construction 101

kw: book reviews, nonfiction, dinosaurs, DNA, genetic engineering

About twelve years ago, shortly after Jurassic Park hit the big screen, a colleague told me he was briefly famous for the first recovery of proteins from a fossil. In the '70s, when he got his PhD, he discovered a relationship between the normal body temperature of a mammal and the ratios of certain "structural" amino acids in their proteins.

Brief aside: The structure of a protein shifts with temperature. It won't work outside a certain range. About half the 20 amino acids (AAs) are mainly structural, forming the helices and sheets that form the shape of a protein. Biochemistry is mainly geometry. For a protein to work best at a different temperature, a shift in the proportions of certain AAs is required.

When my colleague published his results, a friend asked him if his method would work on the proteins from an extinct animal. He said, "Why not. But how would you get some?" The friend brought him some bones of Smilodon, the best-known sabre-tooth cat, from the Rancho La Brea tar pits in Los Angeles. When the animals died there, they were quickly dried by the tar, and the proteins in bone cavities were often preserved.

They were able to extract sufficient protein to work the method, and published a letter stating their findings. My colleague was at a conference in England when the letter was published. Suddenly, he got many calls from reporters, and a British paper published a cartoon of him, sneaking up on a huge Smilodon, and carrying a spear-sized rectal thermometer!

Now, the tar pits contain bones aged between 40,000 and 10,000 years. Hardly dinosaur-age stuff, which is between 65 million and 200 million years old. But impressive for 1970 or thereabouts.

It is a long way from body temperature to a dinosaur clone. A small measure of the difficulty is presented in Jurassic Park, both the movie and the book by Michael Crichton. A recent book makes it clear how much harder it actually is. Rob DeSalle and David Lindley, a working scientist and a highly expert science writer, in 1997 published The Science of Jurassic Park and the Lost World, subtitled, Or, How to Build a Dinosaur.

Dr. DeSalle isolated the first dinosaur-age bit of DNA in 1992, from an insect in amber. It was insect DNA, though, not dinosaur DNA. Older bits have been found since, as old as 135 million years. So when he outlines how one might (just barely, maybe) retrieve dinosaur DNA and eventually produce a living dinosaur, he has it right.

He agrees that amber is a good place to begin looking, but he prefers amber from New Jersey, which is the right age, to Dominican amber, which is only 30 million years old. But what guarantee do we have, if we find a biting critter with a belly full of blood, that it was a dinosaur's blood?

I have recently read of the recovery of soft tissue from deep inside a Tyrannosaur hip bone. Perhaps we ought to be looking there, instead. Otherwise, you're more likely to find the blood of a proto-possum than a dinosaur, which is quite a bit harder to bite...we do have samples of dinosaur skin, so we know.

The authors go through, step by step, what is needed to do the task. They make clear the uncertainties at every step. For example, the DNA sequencing method called "shotgun sequencing" is probably most amenable to this, but it cannot tell you how many chromosomes there were. We only learn this when we sequence, say, a chicken, because we can look at living chicken cells and sequence them one chromosome at a time. If you have a DNA soup with the entire genome in little bits (say from 200 to 1000 bases per chunk, each broken out of a 2- to 3-billion base sequence), you can't really tell where the chromosomes ended. Telomeres (repeated sequences at the ends) are too variable from one animal to the next to prove anything; one critter's telomere might be another's internal repeat sequence.

Suffice it to say, the undertaking is too expensive for an ordinary billionaire. Given the rate that DNA work's price is dropping, however, I expect it might be possible in another decade or two, making initially one assumption: that we can actually recover large enough bits of 80-million-year-old DNA, in sufficient quantity, in the first place. That may be the biggest hurdle of all.

Thursday, October 20, 2005

Alien Hybrids?

kw: space aliens, DNA, central dogma of genetics, genetic code

Do you think Spock could really exist? Is a hybrid between humans (or any Earth species) and an interstellar alien possible? Alien-human hybrids figure prominently in some writers' stories...and they are the core fear driving the "alien abduction" folks. Could it happen???

The Central Dogma of genetics is that DNA is transcribed onto RNA, and the transcribed RNA is used to make Proteins. Most of these proteins are Enzymes, that is, peptide catalysts. That isn't all there is to it, because some RNAs are Rybozymes, behaving as Enzymes though they aren't proteins. It turns out that biochemistry is nearly all geometry, and you can form a desired shape from amino acids (proteins) or from RNA. Also, in trying to determine if "junk DNA" is really junk, we are finding that some DNA has a regulatory function within the nucleus, plus there are at least five or six levels of regulatory activity by various protein families.

So, DNA is Transcribed to RNA, and most RNA is Translated to Protein. The translation function of genetics is based on the Genetic Code, which determines how the 64 3-Base Codons are distributed among the 20 Amino Acids and the Start and Stop functions. As it happens, in most Earth creatures, some amino acids correspond to as many as six RNA codons, some to four, on to just three that correspond to a single codon each. There is a minor difference between the code for Prokaryotes and Eukaryotes, but 62 of the codons are used identically in all cells...on Earth.

It appears that many of the amino acids used in active sites have larger codon sets, while "backbone" AAs have smaller sets. AAs that are "important" have more codes than those that are have more "scaffolding" functions. This is not totally consistent, but it explains one source of robustness; many point mutations make no difference in the shape of the protein encoded.

Recently, we find that all parts of this system are malleable. It is possible to synthesize bases that could be used in place of any of the A, C, G, and T (or U in RNA) bases used in Earth DNA. It is also possible to synthesize a great variety of "extra" amino acids. In fact, there are a couple of "extra" AAs that seem to be used by certain bacteria, and some researchers claim that there are a dozen or so variants of the "Earth" genetic code, found among bacteria only.

Researchers have also created the pieces to get certain bacteria to synthesize and use, in proteins, an amino acid not normally found in nature. So, while nearly all Earth life uses the familiar four bases and twenty amino acids, there is no guarantee that life forms that arose on a distant planet do so.

Since we know that all 64 Codons are used, we can think of a few simple variations that would work equally well. If you have a string of three symbols, there are six ways you can arrange them. That means, if you were to perform the same rearrangement on every codon in the set, there are six "genetic codes" that work exactly as the Earth one does.

If instead you determine how many different ways 64 Codons can be distributed among 20 AAs, the number isn't just in the billions, or in the trillions: it is an immense number with more than seventy digits. If we just look at a near-even distribution, with 60 codons parceled out three at a time to 20 items (e.g., the 20 AAs), with the two remaining items getting two codons each, the value is 60!/[(2!)2(3!)20], which comes to 8.7x1072. There are many, many other ways to partition how many codons each AA gets.

Many of these will be less favorable, because certain 'sensitive' AAs don't get sufficient protection from mutation, but there are huge numbers of relatively 'good' codes that could be used, more than enough that each planet in the visible universe (assuming a billion or so per galaxy, and perhaps a trillion galaxies) can choose among billions of possibilities.

As a result, even if every star has a planet bearing life, the likelihood that two of them will have the same genetic code, or even remotely similar genetic codes, is effectively zero. We can determine how close to zero, by means of the "birthday paradox". You may know that, if you get 23 or more people in a room, there is a better-than-even chance that two of them will have the same birthday.

Similarly, if you assume that blondes have at most 100,000 hairs on their head, in any city with more than 100,000 blondes, there will be quite a few pairs with exactly the same number of hairs on their head. What is interesting is, if you have any group of at least 373 blondes, there is a better-than-50% chance that two of them will have the same number of hairs. Which to? You gotta count a lotta follicles!

Now, if the universe of possible DNA codes is of the order of 1072, what are the chances that there is an exact match between two codes, if the number of inhabited planets in our galaxy is ten billion...or in the universe, perhaps a quintillion (billion trillion)? Both are wild guesses, but plausible.

In the first case, the probability of a single match is 1-exp(-10-52), which has 52 zeroes before you get a nonzero digit. In the second case, the match probability is 1-exp(-10-30), which has 30 zeroes ahead of the first nonzero digit. Either way, it's way, way smaller than a chance in a million. Turning the question around, how many "viable codes" might there be, to assure that in the universe, or in our galaxy, the chance of at least one match is at least 50-50? For the galaxy, there must be fewer than 1021 actual DNA codes in use, and for the universe, fewer than 1043.

So, we start with a 72-digit number, and cut it down by 29 digits, just to make some chance that two life forms, that arose in different planets, could hybridize. Not a good bet.

Added note January 2007: There are currently seventeen "genetic codes" known, in use by Earth organisms (see the Wikipedia article Translation (genetics) ). The vast majority of cells use the "standard code", but there are many minor variants used by mitochondria, a couple used by plastids, and a few others used by certain eukaryotic microbes.