Monday, 23 March 2020

Covid-19 and the road to a vaccine

Image result for edward jennerIn my last post, I explained the basic properties of Covid-19 and how it relates to other well known viruses. Of perhaps the greatest concern to everyone is how science and medicine can be harnessed to beat this virus. There are currently a limited number of pharmacological ways of combating the pathological effect of a virus. The first one I am going to discuss is vaccination and in a separate post I shall focus on antiviral drugs. Antidotes represent a specialist category of medicines that are administered to reverse the effects of a toxin and I won't discuss them here (maybe later). One small point, after a few suggestions from readers, I have tried to explain key terms as I go along. If there is no explanation (eg endosomes) I am assuming these are not essential for general understanding and can easily be found via Google, for the aficionados!

The word vaccine is itself a little unusual, if you don't already know, it is derived from the Latin word vaccinus, pertaining to cows, which makes sense, since Jenner's work was based around cowpox. The World Health Authority's statement on vaccines captures the essence of their use in medicine:

"Vaccination is one of the most effective ways to prevent diseases. A vaccine helps the body’s immune system to recognize and fight pathogens like viruses or bacteria, which then keeps us safe from the diseases they cause. Vaccines protect against more than 25 debilitating or life-threatening diseases, including measles, polio, tetanus, diphtheria, meningitis, influenza, tetanus, typhoid and cervical cancer".

The small glass vials in the above image (typically sealed with a rubber stopper) are in universal use for storing vaccines before injection. But what is inside a typical vaccine? To the pharmacist, the way in which any medicine is packaged and prepared for administration is referred to as formulation. And in the case of vaccines, you may be surprised at the formulation. In addition to the recombinant mixture of antigens (it is directed at 4 molecular variants, or quadravalent), a vial of Afluria vaccine from Seqirus for example, also contains:

sodium chloride, monobasic sodium phosphate, dibasic sodium phosphate, monobasic potassium phosphate, potassium chloride, calcium chloride, sodium taurodeoxycholate, ovalbumin, sucrose, neomycin sulfate, polymyxin B, betapropiolactone, hydrocortisone thimerosal

...all of which are collectively referred to as adjuvants, or substances that enhance the immunogenicity of the vaccine's principle component. (To be precise, some components are adjuvants and others are stabilizers, that ensure the vaccine maintains its potency under the recommended storage conditions).
PDB 1hgd EBI.jpg
Three dimensional structure of haemagglutinin

Of course the key component of the vaccine is the immunogen, which is defined as an antigen that elicits an immune response. The word antigen itself is defined as a foreign molecule that specifically interacts with an antibody (you can read about antibodies at an earlier post). Antigens can be small molecules (where they are sometimes called haptens), proteins or whole cells. In the case of Afluria, the immunogens are given below by the manufacturer: each 0.5ml dose contains 15µg haemagglutinin (HA), total 60µg, from four influenza types and subtypes: A/H1N1, A/H3N2, B/Yamagata, and B/Victoria. The multi-dose vial also contains thimerosal (24.5µg mercury per 0.5ml dose). The image above on the LHS is a representation of haemagglutinin, one of the main immunogens used in influenza vaccines.

H(a)emagglutinin (UK/US spellings), as shown above projects outwards from the surface  of the virus particle. The four variants of HA (above in red) are mutants found associated with different viral strains. The quadravalent vaccine aims to eliminate all four major types of influenza in circulation. Why target haemagglutinin? To answer this  we need to understand how influenza virus acts on us. The influenza virus is shown schematically on the RHS. Covid-19 has a single major spike protein, but the flu virus has two: neuraminidase (N) and haemagglutinin. (This why you may hear of flu viruses called H1N1, which is a shorthand for a specific combination of haemagglutinin and neuraminidase sequences). 

Influenza virus HA first binds to sialic acid residues on glycoproteins or glycolilipid receptors on the surface of the host cell, in response, the cell then engulfs (or endocytoses) the virus. In the acidic environment of the endosomes, the virus changes shape and fuses its envelope with the endosomal membrane. This is followed by a signal to release the virus nucleocapsid into the host cytoplasm. From there, the nucleocapsid travels to the host nucleus and a train of events has now been triggered that leads to virus replication. Here is a nice video simulation of the process. Unlike the virus itself, a vaccine will stimulate the production of antibodies that will in turn block this sequence of events by masking the HA or N proteins.

Unfortunately, these proteins are susceptible to mutation (as discussed in the previous post) and, as a result, completely new vaccines must be prepared each year. The design of each new vaccine is determined following a twice yearly international consultation and evaluation of the epidemiology of viral infection and the determination of the genome sequences of the most common viruses. The time taken for a flu vaccine to be produced just fits into the "window" between the February (northern hemisphere) decision meeting and the surge in cases typical of flu in October as can be seen from the 1918 Spanish Flu pandemic (top LHS): we are unsure yet whether Covid-19 is seasonal. 

Image result for flu vaccine administration methodsSo here's what happens when you are injected with a flu vaccine (or any other vaccine for that matter). The formulated preparation of antigens stimulates the production (in this case ) of HA/N specific antibodies. Shortly after the injection some people experience mild flu-like symptoms, but importantly, within around 14 days, you will have produced a reservoir of antibodies that can be mobilized rapidly if you become infected with the virus in the future. You are now immune to the virus. [I shall come back to the issue of the level and the longevity of immunity later in the post.] I have given the example of influenza vaccination which is usually injected intra-muscularly, but a vaccine can be administered orally, subcutaneously (the needle penetrates the fatty tissue beneath the skin), intra-nasally (though the nose) or intra-venously. Again, each method will be associated with a specific formulation. 

Last week in the journal Science, a US structural biology group used cryo-electron microscopy to determine the structure of the Covid-19 spike protein, This is/will be the major vaccination target. The spike protein makes an interaction with a protein called ACE-2 (angiotensin converting enzyme-2: this enzyme is displayed on the membrane of cells from a number of tissue types including the lung, where it plays a role in lowering blood pressure). I hope you can see from one of the images taken from the paper that the spike protein (on the right) changes shape on making contact with the target cell. The green coloured domain adopts the "up position" revealing a surface that makes a strong interaction with ACE-2. At this point, the virus is engulfed and the viral RNA makes its way into the cytoplasm where a combination of transcription (producing the mRNA needed to manufacture new viral components) and replication takes place, as the virus overwhelms the cell. The race is on to produce a vaccine against the Covid-19 virus.

The time taken to produce a new vaccine is at least 18 months (allowing for design, manufacturing, safety testing and trials) to many years. The best flu vaccines available in 2020 offer at best around 50% protection against hospitalisation as a result of flu infection. Similarly, 40 years on from the emergence of HIV and AIDS, there is currently no vaccine that will prevent HIV infection, or treat those who have it. You may have heard about the Moderna  vaccine that is currently undergoing safety testing in volunteers (Phase I Cliical Trial). The innovation here is to bypass the need to purify the protein-based immunogen (or whole virus), by direct injection of the mRNA encoding the immunogen (the spike protein). The use of mRNA as the immunogen, which in some ways mimics the way in which Covid-19 operates, was first suggested nearly 30 years ago, but only in the last 10 years has technology been available to translate this concept into clinical practice. I will be watching the outcome of this work with great interest, but the likely availability of an effective vaccine is still at least a year away: and possibly longer.

As promised earlier, I said I would mention the durability of vaccines and vaccination. In the case of seasonal flu, those who are perceived to be most vulnerable are a priority for vaccination. [You will also have no doubt heard reports of deaths resulting from Covid-19 and their "underlying conditions". 
This is a catch-all phrase to emphasise that those most likely to experience life-threatening  consequences of infection,are those with an already challenged immune system. In addition, since Covid-19  gains entry via the respiratory system, among those at high risk will be chronic asthmatics and cystic fibrosis sufferers]. The durability of an individual's immune response seems to be variable. In addition each type of vaccine appears to show differences. There is a nice article here on the factors known to influence the longevity of a vaccine. Suffice to say though at this stage, in the absence of a Covid-19 vaccine, only time will tell.

I shall leave you with a link to a recent editorial in the journal Science, which  in my view makes some powerful and important observations on the need to ensure that we understand the underpinning Science and respect the fundamental laws of Nature that will eventually enable us to develop a vaccine. As Richard Feynman famously once said: 

"For a successful technology, reality must take precedence over public relations, for Nature cannot be fooled." 

Wednesday, 18 March 2020

Understanding Covid-19 . I Definitions and Origins


This is a short series of science oriented posts on the current Coronavirus pandemic. I intend to release them in bite-sized narratives and so there will be 3-4 to follow. All questions and comments welcome, but I should make it clear I am not a virology expert.

Image result for corona virus
The subject on everyone's lips (hopefully metaphorically and not physically!) is the Corona virus pandemic. With a tsunami of information everywhere, I thought I would provide some background to the issues that have cropped up in my conversations with colleagues, students and friends. My emphasis will be scientific and not behavioural, but hopefully it will clarify some misconceptions and provide the facts as they relate to other historic viral epidemics. At the end of each post I shall provide a glossary to ensure you have a good working definition of key words (which I shall highlight in bold at the first mention). Ideally, these posts will serve as a resource and please feel free to post questions, which I will do my best to answer in a timely way. I shall begin with some generic properties of all known viruses and then I shall focus on some of the well known examples.

Let's begin with a the word virus itself, originating from Latin, where it was used to describe a poison from an animal such as a snake's venom. It was then in 1900 that medical researchers recognized that some infectious agents or particles could be selectively removed by simple filtering procedures. This provides the dictionary definition below with the reference to virus dimensions (20-300nm). Interestingly Google's Ngram analysis identifies the 1980-1995 period as the period of most frequent printed use of the word virus (I am betting this will be overtaken by 2015-2025, when it gets analysed!). My favourite definition is given below and was obtained from www.dictionary.com


Image result for coronavirus electron microscope
A virus is an ultramicroscopic (20 to 300 nm in diameter), metabolically inert, infectious agent that replicates only within the cells of living hosts, mainly bacteria, plants, and animals: composed of an RNA or DNA core, a protein coat, and, in more complex types, a surrounding envelope.

Let's unpack some of these terms. Ultra-microscopic implies that the virus cannot be "seen" using a conventional "light microscope", used routinely to look at bacteria or human blood cells (for example). An electron microscope is required for direct visualization, and the image on the right was taken by researchers at the University of Hong Kong. The scale bar shows the virus (looking like a crown or corona from the Latin again, or ancient Greek for a wreath (classical scholars among you will be familiar with the Olympic victor's wreath, made from olive branches) entering a human cell prior to its reproduction. The white scale bar shows the diameter of the virus as approximately 500nm (slightly larger than the dictionary definition). The virus genome will be discussed later, but the term metabolically inert refers to the fact that viruses of this kind are absolutely dependent on the host for providing energy (in the form of ATP) for their reproduction (often called replication): they cannot produce energy from food: they "steal" it from the host. The terms RNA, DNA, protein (coat and envelope) describe the classes of macromolecules associated with information (RNA and not DNA in the case of covid-19), structural and catalytic components (proteins) and the "protective shell" or envelope surrounding the virus particle.

In short then, covid-19 is an RNA virus capable of infecting humans primarily through a respiratory route (nose and mouth). It is around 500nm in diameter (a typical lung cell has a diameter of 8000nm: assuming viruses and cells are perfect spheres, what would be their respective volumes and what would be the capacity of each cell for virus particles?). It is an enveloped virus surrounded by a lipid membrane, into which a number of major "spike" proteins are embedded. It is these spike proteins that establish contact with respiratory cells prior to invasion of the cell itself (as shown above). The spike protein will assume importance in subsequent discussions of vaccination strategies.


Related image
The leaf on the left is healthy, but the one on the
right has been overwhelmed by a TMV infection
The first detailed analysis of any RNA virus was the subject of the 1946 Nobel Prize awarded to three protein scientists, including Wendell Stanley for his work on Tobacco Mosaic Virus (TMV). In a seminal paper published in 1935 (yes 85 years ago!). Stanley purified and crystallized this cylindrical virus and demonstrated that the pure preparation could elicit the biologically common infection of tobacco plant leaves as shown in the image on the left. In 1935, proof that DNA was the genetic material was unavailable, and moreover, methods did not exist to easily identify and characterize RNA molecules. Today we know that some of the major pathogenic viruses are encoded by RNA and not DNA genomes.

The diagram on the right, compares the structural features of covid-19 (left) and TMV (middle)  alongside another well known pathogen, the norovirus. These three viruses share some things in common, but covid-19 (like HIV) is an enveloped virus: the other two have capsids not envelopes and importantly, alcohol gel is ineffective in eradicating norovirus, but effective at inactivating covid-19. In general hand-washing with soap is the best way to reduce viral transmission, since the viruses we are likely to encounter could be of either type. All RNA viruses that infect eukaryotic host cells (plant or animal) must either express their genes from the injected RNA by direct transcription, or by first converting the RNA genome into DNA, through the action of the enzyme Reverse Transcriptase. With the viral genome now in the form of DNA, the host cell machinery that converts its own structural genes into first mRNA and then proteins, is hijacked by the virus genome. The result is a rapid accumulation of the building blocks needed to make more virus particles. When the cell capacity is exceeded, the infectious particles are released and the host begins to mount an immune defence. (I shall discuss this process in the light of the recent work published by the Australian virology group, but I do like this NYT graphic summary).

To end this first installment, one question I have been asked is "how does covid-19 compare with small-pox virus"? By the end of the eighteenth century, smallpox was responsible for killing around 10% of the world population..... and then along came Edward Jenner (here is one link that looks at the historical eradication of smallpox). The origins of variola virus (the cause of small pox and cow pox) are unknown, but certainly date back to the third century BCE. Variola viruses are unlike covid-19 in one important respect: they are DNA viruses. Why is this important? Viruses encoded by DNA genomes are less likely to mutate than RNA viruses, since the process of viral genome replication is less "error-prone" in DNA viruses and therefore the equivalents of the "spike proteins" in variola viruses present an easier target for vaccine production. For this reason, unlike flu vaccines, which are produced seasonally, smallpox vaccines can be stockpiled through the agency of the World Health Organisation (a valuable resource in all ways). There are over 30m shots of vaccine in deep storage, just in case this disease, which is in fact the only disease to have been completely eradicated globally, should ever re-surface. Clearly the successful vaccination against covid-19 is a priority and... 

I shall discuss vaccines in the next installment...


Glossary of Terms (in order of appearance in the text)


Epidemic and pandemic (from the US center for disease control)


Occasionally, the amount of disease in a community rises above the expected level. Epidemic refers to an increase, often sudden, in the number of cases of a disease above what is normally expected (this called endemic)  in that population in that area. Outbreak carries the same definition of epidemic, but is often used for a more limited geographic area. Cluster refers to an aggregation of cases grouped in place and time that are suspected to be greater than the number expected, even though the expected number may not be known. Pandemic refers to an epidemic that has spread over several countries or continents, usually affecting a large number of people.


Filtering is simply a process by which small and large particles are separated. It is a slightly more sophisticated form of sieving in which a liquid is passed through a barrier (such as paper or a plastic mesh). The holes in the filter can be manufactured to allow (in this case) particles of less than 1000nm diameter through (the viruses), leaving the cells on the filter itself. 


Related imageRNA, DNA protein (these are discussed above), but formally, deoxyribonucleic acids differ chemically  from ribonucleic acids by virtue of a single oxygen atom per sugar (see RHS). One of the consequences of this difference is that DNA molecules form Watson and Crick based paired double helices, whereas RNA molecules form heterogeneous mixtures of helical regions and single-strands. In both bacterial and mammalian cells Francis Crick proposed the central dogma of molecular biology which states that DNA makes RNA (the process called transcription) makes protein (the process of translation). In the case of retroviruses, by definition, the genomic RNA must first be turned into DNA (through the action of the enzyme Reverse Transcriptase) before the proteins that make up the viral coat can be expressed in the host cell. 

Related imageReplication is the term used to describe the duplication of a genome. In the case of DNA genomes, the enzyme DNA polymerase (which varies considerably in its complexity, from bacteria to man, but is renowned for making very few mistakes) catalyses the copying of DNA to generate new chromosomes. In the case of viral replication, many copies of the viral genome must be replicated to be packaged into the viral capsid (protein shell) or envelope (protein and lipid coat). The original concept of replication was suggested by Watson and Crick in 1953 and through teh pioneering work of (Arthur) Kornberg and Meselson and Stahl, we now have a good understanding of the molecular basis of DNA replication as shown diagrammatically on the left. If the genome is made from RNA instead of DNA, replication is much more like transcription and since he transcribing enzyme, RNA polymerase is more careless than DNA polymerase, mutations arise more frequently in RNA viruses.

Thanks for the request from Anudhi regarding the specifics of replication in corona viruses, here is a summary of the properties of the Replicase gene. Like many RNA viruses, the proteins encoded in the RNA are expressed as a poly-protein which requires processing by a protease prior to assembly of the functional protein: in this case the replicase. The two replicase proteins combine to catalyse both transcription of the viral genome and replication in order that the genome can be packaged following assembly of new virus particles. The genome is just less than 30 000 nucleotides and an overview of the related SARS corona-virus can be found here for those of you who want more details.

Tuesday, 19 November 2019

Nour's story

Nour A_T is a Y12 student at the Liverpool Life Sciences UTC. Read her personal journey at the UTC as she embarks on her Extended Project Qualification.


About me. My name is Nour and I’m a year 12 student in the UTC who is currently studying Biology, Chemistry, Maths and Arabic A- level. My interests include sport psychology and chemistry. I’m close to submitting my Crest Award application that aims to investigate how sports affect the emotions of the performer. I’ve already published an article about my personal experiences and a bit of my background as part of the young changemakers program.

Why I joined the UTCI joined Liverpool Life Sciences UTC in year 9 as part of the accelerated programme. I sat Business, Spanish and PE GCSE in year 10 and then I sat my core subjects and 3 further options, including Business AS in year 11. My experiences in the UTC  so far have been very valuable and informative as they have helped build my background knowledge in various areas and I have also seen my interest in chemistry and related fields increase over time due to the exposure to the labs and the insight I have received into the different areas of science and the wider world; PBL projects and science related events the school has held, along with events I attended due to the schools connections, i.e.: 100,000 genomes project-transforming health care. As well as attending a gender equality conference and a conference related to East Asia in UWC Atlantic College. As I’m keen on developing the working environment I'm in, I’m part of the Junior Leadership Team as I believe it can help contribute to the progress and development of the UTC.

Current future career ideas. I’ve always been interested in medicine and medical related careers. I’m currently looking into studying pharmacy and developing drugs in the lab as well as travelling to conferences around the world in order to gain more understanding about other drugs as well as presenting my findings. I’ve expressed an interest in medicine for a long time, but I haven't really looked into the numerous pathways it opens up. I’ve done a week work experience in a pharmacy in order to gain a more in depth understanding of what a pharmacist’s role entitles and this allowed me to have a more vivid idea of what I want my future career to look like.



Role in the Baltic Research Institute (BRI). I'm currently one of the three co-directors of the BRI as well as being head of the technical centre. This will allow me to develop my leadership skills as well as improving my networking skills and being able to understand other people's perspectives on how they would approach problem solving and collectively work on overcoming a problem.
I’m also head of technical centres as I’m very passionate about this sector as it will help improve my technical skills and improve my knowledge on aspects such as molecular biology and biochemistry. This is due to the fact that I will be able to oversee how my colleagues’ projects and roles are developing and I will be able to help them overcome any difficulties they might face, and this would inevitably improve my lab and technical skills.

EPQ idea. My proposed project is titled: “How are the bacteria, campylobacter and salmonella, affected by the antibiotics: Ciprofloxacin, azithromycin and Ceftriaxone, within antibiotic cocktails?” The overall aim is to find out how do these antibiotics affect the growth and colonisation of the bacteria and whether or not antibiotic cocktails are more effective than using one type of antibiotic to inhibit the growth of the bacteria. This may be done by comparing the zones of inhibition around each antibiotic.

I’ll use secondary research to gain a better understanding of the structure and properties of both the antibiotics and the bacteria and how they interact with each other. Secondary research may also include getting test strips from the company ‘Mast Diagnostics’ that they’ve made and comparing them with my own. Analysis involves Comparing the zones of inhibition caused by the antibiotics in the agar plates and see whether antibiotic cocktails were more effective than one type of antibiotic. Moreover, I’ll have two agar plates with just the bacteria on them as they are a controlled experiment that can be used as a baseline for comparison. Overall, this would help build the main structure of my research and the final conclusion/argument.

Steph's story

Steph R is a Y12 student at the Liverpool Life Sciences UTC. Read her personal journey at the UTC as she embarks on her Extended Project Qualification.

Hi, I’m Steph, I’m a student at Life Sciences UTC and I’m currently studying Chemistry, Physics and Maths. I find physical and inorganic chemistry most interesting and I’m planning on studying either chemistry or chemical engineering at uni, although I’ve no idea where. Places such as Imperial, Brunel, Durham and Loughborough are on my list, but the latter is mainly down to the sports facilities as I do athletics in my spare time. This is something which I find complements science really well but has enough differences to be a way to switch off from academics.

To help me find out more about chemistry and the roles available within it I’ve chosen to do my EPQ about the applications of graphene and am being aided by the Manchester Graphene Institute with this. I think graphene is a really interesting material and it has the potential to create many new technologies which are at the forefront of research. So far I have done some research on the properties and applications of graphene and have found it has great potential in many different areas. These cover a wide variety of topics so gives me plenty of choice when choosing which specific area I focus on. Currently I’m leaning towards exploring the use of graphene mats or electrodes as this is something I don’t currently know much about but feel like there’s lots to learn and have found the electrodes could be used to create efficient, low energy solar cells which may be crucial if we want a sustainable future. Another area which I think may be exciting is the potential to create lightweight but very strong composites, something which I could link to my athletics as we are constantly looking for ways to improve our times by small margins while still producing large forces through the ground. I understand I may not be able to test any graphene samples myself however there is plenty of data out there for me to do a more theory based EPQ and I hope to be able to talk to researchers at the institute about what they are investigating and testing.

So far I’ve not had much involvement in the chemistry industry but I hope to get some work experience in both labs and industry to find which area I find most fascinating and could have an exciting and rewarding career in. For me I’d enjoy a career where I can create technologies which can help improve either the science industry or people’s daily lives. One way I’m beginning to develop my chemical skills is through the student led Baltic Research Institute (BRI) for which I am head of chemistry, a new department I’m looking forward to setting up. The main purpose will be to ensure chemistry has a bigger role in the school as the current focus of the BRI is more biology based. This will involve aiding students with experiments or their EPQs as well as developing my own lab skills and techniques which will massively help me in the future when I go into a role in chemistry. I will be working with a team of Y12 students in the department and I’m sure between us we’ll have some great ideas and give others the opportunity to find out more about chemistry.

Tuesday, 3 October 2017

I got rhythm: and so have you?

The Nobel Committee in Stockholm, yesterday awarded the Prize in Physiology or Medicine to three Biologists who discovered the genetic basis of our own personal "cellular clocks". You may recall in 2014, the Prize was awarded for work on the brain's very own "satnav" system, which makes sure we know where we are going and where we have been. The work on "Circadian Rhythms" by Jeffrey Hall, Michael Rosbash (Brandeis University, Boston) and Michael Young (Rockefeller University, New York) provides an insight into our connection with night and day. Through their work, we can now appreciate how life on Earth is related to the movements of the sun, the moon and the Universe in general. 

Image result for night and dayThere are some words you will need to be familiar with in understanding this post. Diurnal is derived from the Latin meaning of the day or daily and comes from the word dyeu to shine (diamond?). Nocturnal, I think you will know refers to the night and finally, circadian (combined here with rhythm), comes from circa (about or around) and day and is an adjective that describes a process that occurs on a 24 hour cycle. Some species are awake in the day and sleep at night, like us (and I will come back to the problems that shift work and long plane journeys can cause) and others are nocturnal, like bats: sleeping all day and foraging at night. I hope you can see here how evolutionary adaptation is linked to the motion of the planets. Let's look at what they discovered, before I return to the planets!


In 1971, the great US geneticist Seymour Benzer and his colleague Robert Konopka, isolated a mutation in the fruit fly (drosophila), that had an altered pattern of behaviour. More specifically, the mutant flies seemed to have an elongated circadian rhythm of 29 hours instead of 24. They named this gene period, or per for short. This landmark discovery paved the way for the work honoured this week, by the three Nobel Laureates above. The per gene turns out to encode a protein that regulates its own production and destruction: levels of the messenger RNA (mRNA) that encode the PER protein peak during the night and drop down in the daytime (as you can see graphically, top left: the x-axis is in hours). However, the PER protein is produced in the cytoplam of the cell and of course the genes are in the nucleus. It was Young who, in the early 1990s identified the timeless gene, encoding the protein TIM. TIM binds specifically to PER and transports it to the nucleus, where it can now act accordingly. A further protein, identified by Young and this time called doubletime fine tunes the whole process, allowing for occasional adjustments to the 24 hour fixed period. In humans, the network of genes/proteins involved in regulation the body's circadian rhythms is a little more complex than I have described, with the usual collection of protein phosphorylating molecules (kinases and their counterparts phosphatases) ensuring the levels of PER proteins are exquisitely balanced. Finally, the "steady state" levels of the PER proteins are subject to controlled proteolysis in a manner reminiscent of the control of the cell cycle itself by cyclins, which Sir Tim Hunt told us all about when he visited the UTC.

I hope you can appreciate that the genes and proteins involved in maintaining the correct 24 hour clock in cells has now been firmly established by an elegant combination of genetics and biochemistry. Seymour Benzer, who was originally a physicist with interests similar to the young Einstein, set the Nobel Laureates on a journey that has linked Biological Evolution to our place in the Universe. The typical day on Earth is 24 hours, which is determined by the rotation of the Earth during its orbit around the sun, but night and day is opposite to us, if you live in Australia. Given that Mars has a solar cycle that is just over 39 minutes longer than our own: we might expect that life forms on other planets have evolved to match their own solar cycles.

Finally, beyond the fundamental importance of molecular basis of Circadian Rhythms, why may they be important in respect of our health and well-being? As I mentioned above, by flying in the face of our diurnal nature, the regulation of the PER system needs to adapt: this happened when you fly a long distance or you adopt the working habits of a badger and work shifts. Many people choose to work permanent nights, and must therefore re-configure their body clock. Try it once when you are not forced to! [I should mention that the PER network of regulation interacts with light sensors (cryptochromes) thereby providing the cell (and the body) with valuable cues which fine tune the body clock. It is also becoming clear that there is a correlation between the action of certain drugs and the body clock. If you are interested you can read these "rapid response" articles that appeared just after the announcement.

The Biochemical Society
The New York Times
The Nobel Foundation

Thursday, 28 September 2017

Crambin: my molecule for September 2017

This month I have picked a protein I first heard about from a crystallographer, visiting Sheffield around 1981. What interested me was that the structure had been refined to a very high resolution (less than one angstrom, in the early 1980s!). What also interested me was that this was a protein with no known function. So what was the point of all that effort! The molecule itself is pretty unremarkable (see the image left). It contains two alpha helices oriented into a Y-shape, linked by a short, constrained loop. The N-and C-termini seem to have considerable freedom, with a small amount of beta sheet, but surprisingly perhaps, the protein crystallises very easily.


The molecular envelope shown right, shows crambin to be a globular molecule with a well defined shape. Without any knowledge of its function, the surface of the long helix is presented for interaction and the N and C termini could re-fold around a small ligand or another macromolecule. But there is no evidence of a metal ion or any significant space for a small ligand or substrate. After all, it contains less than 50 amino acids, which makes it an ideal candidate for NMR spectroscopy.

The protein was initially isolated from an Abyssinian cabbage (or kale) and is now known to belong to the family of toxins called thionins.  [The source of proteins used by biochemists would make a nice Blog post for the future!] The key to the stability of the terminal segments is (as the name thionin suggests) the disulphide bond. The sequence of crambin is shown below, with the Cys residues highlighted. If you look at the representation shown left, you can see the yellow sulphurs and the small network of disulphide bonds that contribute to  the stability of the structure. This is a feature of many extracellular proteins, including immunoglobulins. The analysis of structures at such high resolution provides a molecular framework for defining the precise geometry of such bonding phenomena and provide a nice 

TTCCPSIVARSNFNVCRLPGTPEALCATYTGCIIIPGATCPGDYAN

experimental opportunity to address the role of disulphides in stabilisation and in protein folding pathways of proteins in general. Try mapping the bonds from the structure onto the sequence. You will immediately appreciate that primary structures must be considered in three dimensions in order to fully appreciate the significance of sequence conservation! 
Image result for crambin nmr structure

The possible application of this plant product in cancer treatment is being investigated, but remains at an early stage to date. One other point I would like to draw your attention to is a comparison between methods of structure determination. X-ray crystallography "prefers" proteins of several hundred or more amino acids in each polypeptide, whereas Nuclear Magnetic Resonance (NMR) spectroscopy "prefers" proteins with molecular weights below 20 000 (<200 amino acids). These rules aren't hard and fast, but they do significantly improve the probability of obtaining a high resolution data set (required to fix the position of side-chain atoms). The structural representation on the right was obtained by NMR. NMR structure determination generates an "ensemble" of structures that are consistent with the spectral data. (See here for an introduction to protein NMR). The first thing you realise is that some parts of the protein are better defined than others. In X-ray crystallography, any significant "flexibility" in a protein structure usually prevents the assignment of electron density in that region and this may mean that this section of a protein is not included in the deposited structural file (it is usually pointed out in the publication). In the case of crambin, the NMR structure suggests that the N and C termini are pretty rigid. A consequence of the disulphide bonding, explaining why the protein is so compact and probably explains why it such an amenable molecule for obtaining high resolution atomic data.

So finally, why has so much effort been invested by structural biologists in a molecule of such poorly defined function? This is an important issue in Science in general: what should we (as tax payers, versus say drug companies) spend our money on? Molecules like crambin can help establish the fundamental principles of protein structure. Some medically important molecules may be difficult to purify, may be unstable or may yield poor diffraction (in the case of X ray crystallography) or may be difficult to solubilise and show poor spectral resolution (for NMR). The insight we gain from "well-behaved" proteins can help us fill in the gaps with molecules that we can easily recognise as being of societal value. Moreover, workhorses like crambin can help us push the envelope of techniques like X-ray crystallography and NMR, which may then make it more likely that we can interpret the data from proteins that are less well-behaved! One final point is that structure determination alone can rarely determine the function of a molecule. It may be that we need to solve the structures of the entire proteome of humans before we are ready to comprehensively link structure and function in Biology!

Wednesday, 23 August 2017

The challenges of making RNA in bacteria. A late summer selection of molecules

Related imageThis month is a little later than I had hoped, mainly due to my choice of molecule. I decided I wanted to get more familiar with the (expanding number of) regulatory proteins involved in the control of gene transcription in bacteria; partly out of a research interest, and partly because it allows me to combine my interest in language (or more specifically alphabets) and Science. The down-side is that I will have to replace all of my as with αand my bs with βs etc. which is always a little clunky on my free Blogging software! The molecule I am focusing on is RNA Polymerase (Pol), which I have covered earlier. However this time I am going to take a look at the transcription factors, or "known associates" of this multi-subunit protein complex that bridges the gap between information and function. The genomes of all prokaryotes contain a set of between 2 000 and 5 000 protein coding genes together with a few hundred genes that encode functional RNAs. This is all information; but in order to "translate" from nucleic acid speak (Nu-speak: sorry George!) to the language of amino acids and proteins (Pep-speak?), the ribosome is required. However, a limited number of RNA species combine an information mode with function, such as the hammerhead ribozyme, that can catalyse specific RNA cleavage in the absence of any proteins (a good future molecular candidate perhaps?). The nucleotide sequence of a ribozyme is no different than that in the genome (apart from an additional oxygen atom per sugar),and it also determines its three-dimensional fold. And therefore its biological function. 

You may be interested to know the source of the images used in this post. I have chosen, where possible, to include the beautiful models created and exhibited at the Pingry Biomolecular Modelling Project web site which is just one of the incredibly impressive Pingry School initiatives at the school: more information on this ground-breaking collaboration between the Milwaukee School of Engineering (MSOE University) staff and the students and teachers at the Pingray School can be found here. On the right is an image of the components of RNA Pol in the early stages of transcriptional initiation. I hope you will agree with me that these models capture both structure and function in a beautiful and informative way.

The Basics RNA Pols, in their simplest forms (let's leave bacteriophage enzymes on the side for now), comprise two α subunits, a β and a variation of β, called β-prime (written β')(there are some enzymes in which the β-type subunits are fused, but these are only occasional exceptions). This hetero-tetrameric "apo-enzyme" then associates with a number of "regulators" to form the "holo-enzyme", the most important being the σ subunit, which is critical for determining the DNA sequence specificity associated with the choice of the promoter to be transcriptionally active. The image below the reaction scheme shows the promoter sequences recognised by the RNA Pol holoenzyme, with the -10 and -35 elements (recognised by the sigma factor) highlighted. As we we shall see below, the σ subunit comes in a number of different "flavours".The reaction catalysed by RNA Pols is shown below: it is important to remember that while I am discussing sequence specific DNA binding, RNA Pols are catalysts and DNA and RNA represent substrates and products respectively. 

The prefixes apo and holo are derived from the Greek: meaning away from and complete, respectively, and are used frequently by Biochemists to describe proteins without (apo) a key component, such as a co-factor compared with the fully functional molecule (holo): apo-haemoglobin lacks the haem, for example. Which brings me to the inevitable glossary: an essential set of definitions of terms, symbols and concepts needed to understand gene transcription and for those of you are unfamiliar with the idiosyncrasies of the Greek alphabet, I have included my suggested (phonetic) pronunciations: remember when discussing Science, it really helps if you feel confident about the pronunciation of some of the rather ludicrous terms!

The Greek alphabet and my advice on pronunciation! [A "hard" consonant, eg the first and last G in gang is written gg, while the soft G in German is written as a single j. Where there is no ambiguity, e.g. the letter D, it is shown as a single d. If the vowel is drawn out, like the two Es in meet, it is again doubled].

α (alff-a)
β (bee-ta (UK), bayta (USA))
γ (ggamm-a)
δ (delt-a)
ε (ep-ssee-lon)
ζ (zee-ta)
η (new)
θ (thee-ta (UK), sometimes tha-yta (USA))
ι (eye-oh-ta)
κ (kapp-a)
λ (lamm-da)
μ (mew)
ν (new)
ξ (k-ss-eye)
ο (oh-mee-kron)
π (p-eye, or for English readers pie!)
ρ (row)
σ (ssigg-ma)
τ (torr)
υ (up-ssee-lon)
φ (ff-eye)
χ (kai, or k-eye [not kee])
ψ (p-ss-eye, as in psychology)
ω (oh-mee-ga (UK) or oh-may-ga (USA))

A short glossary

A

Apoenzyme: an incomplete molecule, usually requires a coenzyme (such as FAD, an additional protein (such as σ) or an RNA molecule for full function
H
Holoenzyme: an complete molecule, usually incorporating an essential coenzyme (such as FAD, an additional protein (such as σ) or an RNA molecule and expressing full biological function
O

Operator is the term given to a promoter that is flanked by a repressor (or an activator) binding site. The sequence of the promoter is extended in either direction (or possibly both

Promoter: a stretch of double-stranded DNA sequence to which an RNA Pol binds and, through a series of orchestrated molecular interactions, marks the initiation point for the transcription of a particular gene or group of genes. In bacteria, the DNA sequence comes in two sections: the -10 box comprises around 10 base pairs which are recognised by a σ factor (which is itself associated with the apo-enzyme for of RNA Pol). The -35 "box" provides contacts for the αβ subunits. The negative sign indicates the distance between the two "boxes" and the nucleotide that forms the 5' end of the transcript. The diagram below should help explain these concepts.
R
Ribosomes: a multi-component molecular machine comprising rRNA and polypeptides in the form of two "subunits" referred to by their sedimentation properties in an analytical ultracentrifuge. The 30S (small) and 50S (large) subunits co-assemble during the initiation of protein synthesis in the presence of initiation factors aminoacylated tRNAs, mRNA and an energy supply. You can read more here
T
Transcription: the catalytic, template mediated synthesis of RNA from double stranded DNA. The products are a range of RNAs, including messenger, transfer etc and the enzyme may be a single species such as bacterial RNA Polymerase, or a dedicated one such as RNA PolII in eukaryotes that catalyses mRNA biosynthesis.
Translation: the biosynthesis of polypeptide chains from mRNA templates via the ribosome. Each ribosome can accommodate virtually any mRNA and in higher organisms, aggregates of ribosomes are called polysomes

Sigma factors One of the many returns on our collective investment in genome sequencing, has been the insights gained into those genes that are essential for cell growth and reproduction. Not surprisingly, the genes encoding the polypeptides that make up RNA Pols are essential for cell viability. However, while all prokaryotes possess the genes encoding the α (rpoA)β/β'(rpoB and C) and the major σ factor,  σ70 (rpoD or sigA), there are some other regulatory factors that seem to confer advantages in regulating gene expression, that are likely to add to the physiological versatility of the organisms in which they are expressed. In the well-studied prokaryote E.coli, in addition to σ70 , we find  the following σ factors:

σ19 (fecI) - regulates the fec gene for iron transport
σ24 (rpoE) - the extreme heat stress factor
σ28 (rpoF) - the flagellar factor
σ32 (rpoH) - the heat shock factor, that is turned on when the bacteria are exposed to heat.  Some of the enzymes that are expressed upon activation of σ32 are chaperones, proteases and DNA-repair enzymes.
σ38 (rpoS) - the starvation/stationary phase sigma factor
σ54 (rpo
N) - the nitrogen-limitation factor.

Before (L) and After (R)

σ factors interact with the RNA Pol apoenzyme to generate the holoenzyme and in doing so, provide the enzyme with the capacity to recognise the -10 and -35 elements of a promoter (see figure and scheme above). The "before and after" images (LHS) show the location of the (orange) σ factor in the complex, and how its elongated shape facilitates recognition of the -10 and -35 elements (the promoter is the blue and pale green duplex above the RNA Pol). The initiation of transcription of all constitutive genes only requires the RNA Pol holoenzyme as in the "before" image. As soon as the transcriptional start site is exposed and a supply of NTPs is made available, the σ factor dissociates (the "after" image) and the elongation phase of transcription gets underway. The role of the σ factor is primarily to "target" the catalytic apparatus: by replacing the house-keeping σ factor with any of the above sigma variants, selective sets of genes can be expressed in response to one or more environmental cues. Pretty straight forward I think you'll agree. This principle of combining a core function, in this case RNA synthesis, with a variety of targeting polypeptides (in this case sigma subunits), is a common strategy used in Biology, with antibodies being a well known example. 

Anti-sigma factors The potency of σ factors has led to the evolution of antagonistic molecules, called anti-sigma factors. In some organisms, σ factors need to be attenuated [slowed] (or even abrogated [stopped]): this can be achieved by the expression of anti-sigmas. Again, the logic is pretty simple. A σ factor can be maintained in complex with an anti-sigma, until an environmental queue is triggered. Through an induced conformational switch, such as a pH transition, or the binding of a small molecule to the anti-sigma component, the two components (see the image of the T4 phage anti-sigma-σ complex, RHS) are able to dissociate and the σ factor is free to promote targeted transcription.

File:Lambda repressor.jpgRepressors These molecules have a special place in the history of Molecular Genetics. The work of Jacob and Monod (see an earlier post on RNA Pol) in the early 1960s laid the foundations for our understanding of gene regulation in prokaryotes and higher organisms. At the centre of their logic was the concept of the repressor, which was later defined in molecular terms as a protein molecule (although it can also be an RNA molecule) that interferes with transcription. The mode of action of repressors can be simply described as creating a road block in the path of a promoter bound RNA Pol, but since this simple concept was proposed, genetic, structural and kinetic studies have shown that repressors can inhibit RNA Pol progress by a variety of mechanisms which do not always arise from simply blocking the path of the RNA Pol, or by competing for a specific sequence in at or around the promoter. In fact, some repressors (including the lambda repressor shown left) are able to act as both repressors and activators of RNA Pol mediated transcription, and this forms the basis of the "plot" of the remarkable work from Mark Ptashne's laboratory, whose short book on this topic is a "must read" for all Molecular Biology students. Since most repressors do not form stable interactions with RNA Pols (although this is not meant to be a dogmatic statement), I will not discuss them further in this post.

Termination factor ρ , which is shown on the right, is responsible for terminating RNA Pol mediated transcription, but once again ρ acts like a classical repressor in recognising a specific RNA termination sequence of around 70 nucleotides, signalling the end of the road for RNA Pol: the ρ protein does not form a stable complex with RNA Pol. Bacteria like E.coli invest significant energy in synthesising this hexameric homo-polymeric protein and it is essential for viability in most prokaryotes. In fact the transcription of about half of the genes in E.coli are terminated via ρ while the remainder are said to be ρ-independent, or alyternatively utilise the proteins τ or nusA

The ω and δ factors. These are both bona fide components of the RNA Pol holoenzyme.  ω seems to be involved in chaperoning and stabilising the interactions of the β' subunit. Unlike ρ it is dispensable, in that ω knockouts survive; but it does seem to improve the net efficiency of transcription: I expect growth rates in ω knockouts are lower than wild-type strains. δ is also formally enshrined in the RNA Pol holoenzyme, and like ω, the gene encoding this not factor is not essential, but its removal from a genome, does give rise to some strange morphological changes in growing cells (abnormal elongation in particular). A complete understanding of the roles of these two factors in transcription remains to be elucidated, but both primary structures are highly conserved amongst prokaryotes and a number of groups are currently looking at the functions of these accessory factors during infections in pathogenic bacteria.

I want to close with a mention of a growing number of regulators of transcription that seem to modulate transcription and bind to DNA, or indirectly via σ factors, and thereby RNA Pol transcription through a redox signal mediated by an iron-sulphur (Fe-S) cluster, buried in the heart of the protein, or sometimes in a flexible subdomain. One such regulator is SoxR (containing an 2Fe-2S cluster, shown in red-yellow on the LHS). I think you can see how the distortion of DNA might be induced by SoxR and this can modulate transcription initiation. Environmental signals such as reactive oxygen species and NO, trigger gene expression events that ultimately lead to the elaboration of processes that defend the cell against this metabolic challenge. The wbl proteins are a class of gene regulators (originally identified in Streptomyces strains), but which are also found amongst the Mycobacteria (think TB). The main reason for including them is however that they have become a major area of interest of one of my colleagues at Sheffield, Professor Jeff Green. And because I really like the story emerging from his lab (see a review here ), that connects redox sensing and the control of gene expression, which may have wider implications for a number of prokaryotes and may possibly modulate the mode of action of some antibiotics. 

In summary, the extraordinary focal point for gene regulation is RNA Pol in bacteria and we are learning every day about the plethora of polypeptides and RNAs that influence its activity. I hope this has given you a flavour of the structure and function of this area of Molecular Biology. At some point, when I am brave enough, I'll look at the eukaryotic RNA Pols!