Tuesday, May 20, 2014

Final Poster Session & Bicentennial!

My internship has finally come to an end. On April 30, 2014, we presented our poster during the assembly block to faculties and students. I am glad that many people stopped by and asked me questions. After analyzing my data, we concluded that most birds acquired the infection independently, while some acquired int through bird-to-bird transmission (not shown in the tree). What I was more happy about was that my mentor  mentor was able to come to my presentation, and he also enjoyed other students' presentations.


Fortunately, I was also able to share my experience in one of the Academy class over the bicentennial weekend. The group asked a lot of questions and I was really happy to share my experience with STEM program. I think the fact that Emma Willard is developing this Signature Program is incredible, and I am glad to be one of the earliest participant. After 2 years of STEM intern, I'm more certain of what I want to do in the future, and I hope the program can help more students to fins their own passions!

Dr. Miller and I at the end of the year presentation.


Monday, April 21, 2014

18 Samples + ATCC Reference Strain

I didn't meet with my mentor last week because I was on a college visit. Nonetheless, the week before (April 11th, 2014), I had attempted to make several BEAST tress with 6 different BEAUti settings (no time included). I would totally include some visualizations since I've got some cool pictures except that my laptop decided to crash down and all the memory is lost (All the data run is probably too demanding for the computer oops). So, I will try to explain what I did in words.. please refer to my previous posts if needed!

I mainly focused on the effect of clock model and site model on the tree. Although I didn't actually look at the output tree visualization, I examine my the goodness of my output data on Tracer (see my previous posts). Out of 6, only one looked decent (approx. Normal), with the setting of Strict Clock Model and the GTR site model. So, I went on including the time of the samples in BEAUti and ran it in BEAST, yet the resulting tree didn't look so good..

However, my mentor gave me a new input file where our previous reference strain MAV104 was substituted by ATCC strain because MAV104, which was isolated from an AID patient was too distant from our M.avium samples, and that ATCC was originated in birds as well. Therefore, I ran the new data on BEAST (time included) as well as on GARLI and RAxML, where the trees were made based exclusively on the sequences. Something promising happened! The BEAST tree looks almost exactly like the RAxML tree with a little variations, which is reasonable since BEAST builds a tree based on both sequences and time. I will go ahead and analyze the GARLI tree to see if I get the resemblance.

In the meantime, I am starting to work on my final poster! Whoo the time has gone by so quick! I just got my new emergency laptop and am trying to reinstall many software. Thankfully that I ran most of my data online so I can still download them to my computer. The thing that will take a while is the actual pictures of the trees, in which I will try to rebuild them these days.

Finally, since there is no visualization in this post, I would like to share a picture from my college visit at LA! This is me standing next to my future school mascot - the Bruin! Can you guess where will I be spending my next 4 years at? Yes - UCLA!!!! I am so glad that the application process is finally over and I can't wait for my college life!!!


Sunday, April 6, 2014

Spring:)

Last Friday (April 4, 2014), I finally met with my mentor again after so long. However, we have a shorter meeting because I have a flight to catch that evening to Atlanta, GA for college visit. Also, since I have been missing several posts (I have been travelling a lot this month!), I will include what has been going on in this post. He has forwarded me several emails over the break to keep me updated about his conversation with the zoo people. So, over the break the zoo people have reassemble many of their mycobacterium sample sequence to increase the number of signals as well as eliminating sequence error. On the other hand, my mentor has been assessing the sequence alignment of the samples, and he did found some assembling errors such as one sequence contains an extra 1000bp and another contains extra 99bp. Additionally, he found 3 pairs of identical sequences, so we will later eliminate the duplicates.

Before Friday, I edited the 21-sample file Dr. Miller sent to me and ran it on BEAST with default setting. However, the graph on the Tracer did not look so good - it had 3 bumps, skewing overall to the left in stead of the fine Normal shape we want, so I did not proceed to the tree making. I showed my result to my mentor and we will just keep testing out with different settings.

Meanwhile, he showed me what has he been doing over the break - comparing each sample sequence to a reference sequence, in this case MAV.104, using a software called mauve to assess the validity of sample sequence and how similar / different are they to / from the reference. So to align the sequence, I clicked Align with progressiveMauve and added two sequences MAV 104 (reference) and myc01.


The top one is the reference sequence and the bottom one is our sample (myc01). Each color segment represent a contig, which ia set of overlapping DNA segments that together represent a consensus region of DNA. The bottom sequence is 2-sided simply because when we sequence the sample, there are some pieces copied from the positive strand and others copied from negative strands. Thus, what can be useful is if we reorder the sequences and convert them all to the positive strand, and you can do this by selecting Tools --> move contigs.


After the reordering, we can really see how similar is our sample sequence to that of the reference. The white gaps simply mean certain places do not match. However, this image does not tell the exact order of contigs of our sample since we've reordered it. One sequence may be preserved in bacteria / viruses though without being at the exact same site since small organisms have great ability to reorder DNA sequences.

This week my assignment is to compare several more contigs with the reference to get myself familiar with Mauve while trying to reinforce our tree. I will continue working on our tree with different BEAUti settings including using relax clock and UPGMA starting trees.

Thursday, March 20, 2014

Tree Comparison

Right before the break (early March), Dr. Miller and I had been working on constructing the 26 samples tree using various programs (see previous posts). After a lot of trial and error, we finally got decent trees from each program. However, since the outcome trees looked very different in each program, we used a software TreeGraph to standardize and compare the trees. When loading the files into the software, you would have to convert the file into nexus file simply by adding .nex in the file name. Below is the comparison between 4 trees:

In order of RAxMl, GARLI, MrBayes, BEAST
Fortunately, trees from GARLI and MrBayes looked exactly the same while the other two resembled them without great differences. This gave us a pretty good picture of the time-included tree in which we will later be working on. This is a rather short post because the process was lengthy and complicated that I don't think I will be able to explain it comprehensively here.

Time is going fast and it'll be April when we returned from the break! Hopefully we will be able to achieve some work before the year ends!

Sunday, February 23, 2014

More Trees with GARLI and RAxML

Last Friday (February 21, 2014), I didn't meet with my mentor because he was out of town. Nonetheless, I continued making trees with GARLI (using the right one this time) and RAxML.

GARLI_1

The numbers on the branches are the bootstrap result. In this trial, I set a bootstrap repetition of 50. Basically, the bigger the number is, the greater support we have for that particular branch, with the greatest number of 50. Though overall the tree looks pretty decent, notice that in this tree we have only little support fomyc 1,2,5,23,16, and 30.

Next, I moved on to run the samples with RAxML. The branch lengths varied significantly and I am still understanding the implication of it. The maximum bootstrap number is 100. While most of them were pretty big, the numbers for were still very small (splitto the extreme). This lack of support was due to the same sequences. myc01 is identical to myc02 where as myc16 is identical to myc23. 



Therefore, the next thing I did was to run both GARLI and RAxML again with identical sequences eliminated so that they wouldn't confuse the program.

GARLI_3 with identical samples eliminated.
RAxML_2 with identical samples eliminated.
This time the resulting trees all had branches with very high bootstrap numbers, suggesting that our trees were very strong. 

The next thing I am going to do it is to run the samples with MrBayes and compare the result topologies. I am finally meeting with my mentor this coming Friday, and hopefully we'll discuss more about the results!

Sunday, February 16, 2014

Examining Different Trees Using Different Programs

Last Friday (February 14, 2014), my mentor wasn't able to come because his flight was cancelled due to the snow storm. However, we did chat through Skype and accomplished some work via email. In continuing our tree making, after reading several publications, we decided to run the larger data (26 samples) at once so it might be more accurate. Wayne, one of the zoo people, sent us an updated genomic sequence of the samples this time with more identified SNPs. We also want to examine the overall topology before taking time into consideration, so that we can first get a sense of how our tree would look like. Thus, we decide to look at different tree visualization tools and software for describing the difference between any two phylogenetic trees.

The programs I will be exploring in addition to BEAST are RAxML, GARLI, and MrBayes. I ran GARLI first, and did it on my mentor's website CIPRES. The whole process took a while and was quite complicated to describe it here. Basically our goal was to find a way to get a majority rule consensus tree form GARLI output. I later found out that I chose the wrong tool to run my data (there were several GARLI choices), but I still went ahead and analyze the data. I converted the output to nexus file so it can be read on Archaeopterix, a powerful tree visualization tool that supports many file formats. My final tree look like this:


I won't know how good was this tree until I make more with other tools so that I can compare them. I will also run GARLI again, using the right tool this time!

Tuesday, February 4, 2014

A Week Off

Last week I didn't meet with my mentor because I was under the weather:( We will try to catch up our work this Friday. However, it was also the Chinese New Year, so Happy Year of Horse everyone!!

Retrieved from http://eastweek.my-magazine.me/?aid=30771

Friday, January 24, 2014

Small Myco bacterium Trees!

January 24, 2014. Last Friday I didn't meet with my mentor because he was out of town. Yet, over the week I was able to successfully run the BEAST on my mentor's website (finally!!). Today we made our first attempt to generate a tree for the small mycobacterium data!

The two main outputs of BEAST are the .log and the .tree files. First we analyzed our data by using a program called Tracer. After I imported the .log file...


One of the most important columns is the effective sample size (ESS) on the left. ESS is the number of independent samples that the trace is equivalent to, and it can help identify autocorrelation in our samples that might result from poor mixing. The ideal is to have the number >200. To do that, I eventually went back to BEAUti and changed the chain length to 5,000,000 with sampling step of 2000 (that increases our sample size in trace to 2500). It is also ideal for our graph to look Normal (this one is pretty good:)).

After confirming that our data converged to a stable posterior distribution, I used TreeAnnotator to summarize the information from a sample of trees (we have 2500) produced by BEAST onto a single “target” tree. The output of TreeAnnotator is a .nex file, which was to be loaded on FigTree program for visualization.


Voila my first tree! Everything looked good except that part in the red. In our data, myco16 and 23 were genetically identical. However, the sampling time of 23 and 30 were closer together, making closer together on the tree. Therefore, we wondered how much does date vs. DNA weighed in Bayesian trees. I went back to BEAUti and excluded myco16 from the taxa so that one sequence only correspond to one sample. The order of the samples on my second tree was good except this time the "time" (in days) was way too large!


I eventually spent the rest of the time going back and forth to see what setting in BEAUti has what effects on my tree. Again, I have generated at least 10 files in this process, including some that failed in BEAST run (I had to go over ALL the steps from BEAUti to FigTree). However, at least we were finally able to get a sense of what our tree would look like:) More trees next week!

Saturday, January 11, 2014

First Meet in Second Semester!

January 10, 2014. Happy New Year everyone!! After enjoying some home time, I had my first meet after so long! However, I felt that I wasn't in my best condition since the jetlag made me a bit dizzy throughout the meeting. Nonetheless, since we have been lagging off for a month, we decide to start running BEAST with the smaller data we got from the zoo people.

We started making our input file with BEAUti. The data was consisted of 6 samples. I had previously converted them into NEXUS file, which was the only format accepted by BEAUti. Dr. Miller sent me a paper (estimating divergence time of viruses is close to estimating that of bacteria) and suggested us to compare / discuss our work after working separately for a while. 


So, for my part, I entered the date (the day each was sampled) of my samples in months (since Jan, 2001). Then, I moved on to setting the substitution model as was suggested in the paper - HKY. I would like to test out different models later as I have read several other combinations that would fit our data. But for our first run I would just stick with the tutorial.

Tips Lane - sampling date

As for Clock lane, I set strict clock (constant rate) because of the low diversity data we are analyzing. Next is the Tree Prior. The Priors panel allows the user to specify informative priors for all the parameters in the model, which can be helpful or burdensome, especially if no obvious prior distribution for a parameter exists (like in our case). Thus, we would need to try out different settings later. My initial setting shown below. 

Exponential because bacteria usually grow exponentially
When all was ready, I clicked generate BEAST file. However, when I uploaded the file to run BEAST, it repeatedly said error! For some reason it said that my setting resulted to 0 prior distribution. One possible reason, in which I read from the BEAST user group online, was that "BEAST starts with a randomly generated tree and if you set tight priors for the parameters, the starting tree may have an extremely low probability which makes it impossible for BEAST to proceed." I went back to the Tree Prior panel and changed random starting tree to UPGMA starting tree. While it ran fine on my computer initially, it terminated to error the next time. After going back and forth, I in total generated 7 files!! Still not working. We decided to make our own starting tree using my mentor's website next time. Yet, the whole process kind of gave me a headache for staring at the screen for 3 hours. To be honest, I am quite intimidated by the computer program as we couldn't really understand the math behind those parameters to fully understand why we were doing. Yet again, those tools are indispensable to many scientists though they would not have invented them themselves. I could only hope that they would start making more sense along the way. Trials and errors - GO SCIENCE!!

Friday, December 27, 2013

Models of Lineage-Specific Rate Variation

December 27, 2013. Unfortunately, because of the conjunction of Revels and winter break, I will not be meeting with Dr. Miller for almost a month! However, most of our work now involves in reading papers online and using computer programs, so the long break time won't be too big of a problem. For a little recap, our goal is to find the divergence time of mycobacterium strains identified in our samples. We are going to construct a tree and calculate the divergence time by using compute program BEAUTi (Bayesian Evolutionary Analysis Utility) and BEAST (Bayesian Evolutionary Analysis Sampling Trees) (very creative names). BEAUti is a program with a graphical user interface for creating an input file for BEAST, which muct be written in XML language. The part that really takes time is deciding which models to use in estimating the divergence time of bacteria. Therefore, what we are doing now is reading journals and previous studies on similar experiments to help us decide our models.

The first type of models introduced here is lineage-specific rate variation. A tutorial mentioned five common models:
  • global molecular clock
  • local molecular clock
  • compound poisson process
  • autocorrelated rate
  • uncorrelated rate
Sebastián Duchêne's  power point provides simplified explanation of those molecular clock, in which I found very helpful.
Different color represents different rate of substitution. Relaxed clocks include auto- and uncorrelated rates. Credit to S.  Duchene (page 8).
After reading several papers and websites, I decided to try global and local clocks because bacteria evolve quickly in a short period of time. Therefore, the lineage-specific rates variation of different strains are more likely to be same or similar to one and another.

While I am still absorbing more info, I wish everyone happy holidays! I can't believe another year has passed and that I will be graduating in less than 6 months! Nonetheless, I hope to have significant contribution to the bird project before I graduate:)

Friday, November 22, 2013

Happy Holiday!

Today (November 22, 2013), we had a short talk on a phone with the zoo people whom we will be working with from San Diego. They are a group of very vibrant and intelligent people even without seeing their faces. Although most of the time I was listening to the conversation as there was a lot to digest, I was nonetheless very excited for this project! After several follow-ups emails, my goal was narrowed down to a single, clear task: What is the divergence time for these avium micrbacterium strains?

Dr. Miller and I looking at data on our computer. Most of our work can be achieved by computer programs, so no fancy gears or species to show B-)


We won't be meeting next Friday because it's Thanksgiving break (Black Friday shopping!!)! Happy Holiday everyone!:)

Tuesday, November 12, 2013

Bayesian Inference vs. Maximum Likelihood

Last Friday (November 8, 2013), my mentor and I discussed a bit about the Bayesian Inference of Trees and the program BEAST.

Like Maximum Likelihood (ML), Bayesian Inference is a character-based tree method, and they both generate several trees and use some criterion to decide which tree is the best. However, BI differs from ML in that BI "seeks the tree that is most likely given the data and the chosen substitution model" whereas ML "seeks the tree that makes the data the most likely" (Hall, 140). Sounds the same, I know. However, after a short period of intense research and seeing a lot of alien equations, I interpreted the major distinction like this (please correct me if I am wrong, clarification needed!!):

  • ML - make  different trees and calculate the likelihood of each tree
  • BI -  finds the tree that constitute the X (observation) with given knowledge (i.e. the descendants or a substitution model).
Mathematically, according to Bayes' Theorem:

P(A|B) = \frac{P(B | A)\, P(A)}{P(B)}\cdot \,

In Bayesian inference, the event B is fixed (the "given knowledge") in the discussion, and we wish to consider the impact of its having been observed on our belief in various possible events A (how the tree split). In such a situation the denominator of the last expression, the probability of the given evidence B, is fixed; what we want to vary is A. Thus, the posterior probabilities are proportional to the numerator: 


P(A|B) \propto  P(A) \cdot P(B|A) \

  • P(B|A) is the prior probability, also the likelihood. It can be interpreted as P(data | hypothesis). Prior probability indicates our state of knowledge about the truth of a hypothesis before we have observed the data. 
  • P(A|B) is the posterior probability. We can rewrite as (hypothesis | data). It shows how well our models agree with the observed data --> Bayesian
I drew a picture and hopefully it helps with visualization. Suppose we know there are species A, B, C, E, and F. For ML, we are finding in what way are the species located on the tree that P(A)xP(B|A) is the largest. For BI, we know from data that A, D, E, B are in the order they are. Given this knowledge, we are finding where C and F locate so that P(A|B) is the greatest.

This is at least what I've got so far. Nonetheless, BI can be burdensome if no obvious prior distribution for a parameter exists. The researcher to ensure that the prior selected is not inadvertently influencing the posterior distribution of parameters of interest. I clarified my thoughts every time I wrote it out. Before we get into the bird data, we try to ask ourselves some fundamental questions about the phylogenetic trees. Next Friday, I will be joining a talk with the data holders to discuss our participation in the Avian mycobacteria project. 

Tuesday, November 5, 2013

More Phylo Talks!

Last Friday (November 1st, 2013), My mentor and I discussed more about the bootstrapping methods in terms of how it work. He gave me a paper, in which I will be reading this week. We also talked about the derivation of the Jukes Cantor model, which is the simplest evolutionary model used to predict the rate at which nucleotide substitutions occur during evolution. It has two main assumptions:

  • Equal frequencies of the four bases
  • The probability of changing from one state to a different state is always equal, i.e. A-->G is as likely to happen as G-->C

Q Matrix of Jukes Cantor Model
f(t) represents the mutation rate, which is equal for all if a nucleotide is substituted with a different nucleotide. The probability of changing from A to A is 1- f(t) since the sum of the probabilities should add to 1.

However, there are many modifications for this model later published, including Felenstein 81 (F81), Kimura 2-parameter (K2P), HKY85, TN93, GTR, etc.

While we are still waiting for the bird data for my project, Dr, Miller sent me a practice data, and he would like me to explore his website on my own first (http://www.phylo.org/). So I looked at the demo, uploaded the data, and let the data run by accepting the default setting. I found the website pretty user-friendly, yet I don't really know what those result means. Hopefully I will learn more about what I can get from these results this Friday.


Friday, October 25, 2013

Neighbor Joining Trees

October 25, 2013. For the past two weeks, I had been reading about the introduction of several major methods for estimating phylogenetic trees. There are primarily two approaches to tree estimation:

  • Algorithmic - uses an algorithm to estimate a tree from the data. Advantages of using this method include fast speed and yielding only a single tree. Algorithmic method includes Neighbor Joining, UPGMA
  • Tree-searching - estimates many trees and then uses some criterion to decide which is the best tree of all.
Another way to categorize those methods are distance vs. character-based methods.
  • distance - Neighbor Joining, UPGMA
  • character-based - Parsimony, Maximum Likelihood, Bayesian Interference
Our discussion today focused on Neighbor Joining Method (NJ). Nonetheless, it is very important to acknowledge that the 'right' tree does not exist since we cannot know what exactly happened in the past. All of these methods only allow us to deduce the the order in which existing taxa (sequences) diverged from a hypothetical common ancestor, and to calculate the amount of changes along the branches between the diverging events.

NJ is one of the most popular distance algorithmic method. It produces a single, strictly bifurcating tree (meaning that each internal node has exactly two branched descending from it). I downloaded another file for practice, in which I am going to show you here.

First I opened up the file LargeData.meg from MEGA. The window shows DNA sequence alignment. The program only show a base when it is different from that of the first sequence. Otherwise it'll just show "."


Then we calculate the average Jukes-Cantor (JC) distance for our data. The data are not suitable for NJ if JC >1.0 In this case, I got 0.811 for this data.

Next, I constructed the tree by choosing Construct/Test Neighbor Joining Tree from Phylogeny menu. A window of Option Summary will pop up. While it is good to stay with default setting (Maximum Composite Likelihood), my mentor and I discussed several other models for the Model/Method section (for simplicity, I won't get into detail here). 
  • Jukes-Cantor model
  • Kimura 2-Parameter model (K2P)
  • Tamura-Nei model
In order to test the reliability of our tree, we also entered "bootstrap method" for Phylogeny Test with a replication of 1000 times. Bootstrapping is, in essence, a method to test how likely a certain "split" is going to occur by using computational simulation.

As we hit compute...
The number assigned to each internal node is the bootstrap value (bootstrap percentage). The high number shows that the split is not random. We can resize the tree a little bit to get a better view by clicking Display Only Topology.
Each internal node only splits into two branches (strictly bifurcating)
The bootstrap value less than 50% means we really have no idea what the branching order is. Scientists usually collapse those branches into polytomies - nodes from which more than two branches descend.
Majority rule tree
Finally, we can do various things to change the appearance of the tree to make it best represent our data...
A circular tree is an unrooted tree
This chapter was quite hard for me because it involved in a lot of equations and concept. Nonetheless, I am happy that I constructed a descent-looking phylogenetic tree in the end! Next week I will discuss with my mentor more about the bootstrapping method and hopefully it will make more sense to me!:)

Tuesday, October 8, 2013

More MEGA 5 with Protein Sequences

Last Friday (October 4th, 2013) I learned more about using MEGA 5 and Blast for constructing phylogenetic trees. This time, however, we used protein blast instead of nucleotide blast because searching protein sequences can detect much more distantly related homologs than searching nucleotide sequences. If two sequences are "homologous," we assume that they descended from a common ancestor (need to be distinguished from "similarity"). Recall that an amino acid is coded by a codon, and that the same amino acid can be coded by several codons, a silent substitution will not cause the amino acid sequence to change. For DNA sequences, there are only 4 possible states (A,T,C,G) of each characters, so when the sequences diverge greatly that there are only 25% identical, the program would classified them as not closely related even though they may have very similar amino acid sequences if you translate them. Thus, the solution id to use protein sequence as query. Proteins have 20 possible states so the lower limit of detectable homology drops to about 5% (instead of 25%). 

For this exercise we used EbgC protein sequence as a query and use blastp.


Those protein sequences are identified as closely related to our query sequences. If you click on the first entry...

The first protein sequences actually include 1087 DNA sequences. Those DNA sequences all code for the same protein sequences even though their DNA sequences vary (silent mutation). Therefore, all those 1087 sequences are very closely related. If we originally did a blastn instead of blastp, the result would only show 100 of these 1087 sequences, and we wouldn't even know other distantly related sequences. The second entry (evolved beta-galactosidase subunit beta [Escherichia coli]) is an example of distantly related sequences (even though they are still pretty close).

Later I chose several protein sequences and translated them back to DNA sequences to align them on MEGA and established a tree. The whole process took so long! Nonetheless, the main idea here is that protein sequences give us a bigger picture regarding the relatedness between species. We start from DNA sequence --> protein sequence --> protein structure --> functions. The bigger we look at , the more distantly related species we are incorporating.

I will be learning some computer code for the command line for the next few weeks. However, I will not meet with Dr. Miller for the next two weeks because of the schedule conflicts. In the meantime, I will continue reading the book and get myself more familiar with blast and MEGA:)

Monday, September 30, 2013

Tutorial: A Phylogenetic Tree with MEGA 5

Last Friday (September 27, 2013) instead of having an internship, I read the tutorial book my mentor gave me and made my first phylogenetic tree!


First of all, I had to download a program called MEGA 5, which allows us to align and compare related sequences with our sequence of interest. For this tutorial, we use a sequence of an alpha-glucuronidase gene from the bacterium Thermotoga petrophilia. We can obtained the sequence by clicking Do BLAST Search from the Align menu, and it will lead us to the following window:
MEGA 5 is linked to BLAST.
I entered the "accession numbers" and the "query subrange", which are given by the book, chose "Neucleotide collection" for the Database, and blastn for Program Selection.

The results showed up few seconds after I hit search. It gave all the sequences that produce similar alignment to our query sequence. The first entry was eventually my query sequence (because that's what I searched for!) Each entry consists of 8 elements: accession number, description, Max score, Total score, query coverage, E-value, and Max identity. For the sake of time, I will just briefly explained Max score, Total score, query coverage, and E-value. Based on the definitions given by the book,

  • Max score - the score for the highest scoring segments of the subject sequence
  • Total score - the sum of the scores for all the segments that aligned (including noncontiguous alignments)
  • Query coverage (%) - shows how much of the query sequence aligns with the subject sequence.
  • E-value - describes the number of hits with scores this high that one would "expect" to see by chance when searching a database of a particular size. The closer the E-value to 0 the better.

 The book asked us to include those subject sequences that have E-value <10^-3 and query coverage >60%. For that, I chose 9 sequences, and I clicked "Add to Alignment" for each (some sequenced require to be reversed before added to alignment).

Sequences of the nine samples
Then I clicked "Align DNA" and get:

This is the end portion of the sequences after being aligned
Finally, I construct a phylogenetic tree from my DNA alignment by choosing "Construct / Test Neighbor Joining Tree" and tada~
My first phylogenetic tree!
This Friday, I will discuss with Dr. Miller more about the assumptions of making phylogenetic tree and do more practices with MEGA 5. I can't wait!:)

Friday, September 20, 2013

Cladisticules

(September 20, 2013) I began my internship today!!! This year I am very lucky to be able to do intern on campus, which cut off all the travelling time:) Dr. Miller, who is my mentor, currently works in UCSD as a part of the CIPRES project (http://www.phylo.org/). He has been developing software and information technologies to assist biomedical research and now works with a supercomputer located in San Diego Supercomputer Center that maintains a database for phylogeneticic tree inference.

Last year, I wrote down for one of the reasons to apply for internship was to "widen my horizon," and here is the opportunity! All of my past experiences have more to do with molecular biology, such as growing cells and running PCRs in a lab. However, this year I get to learn biology in an evolutionary perspective. Dr. Miller raised a good point that studying phylogenetic trees is very critical for it allows us to take a step back and look at the organisms in the most fundamental view, to see a bigger picture of the biological world. I also learned many applications of phylogenetic trees, including predicting the properties of new pathogens, fighting against diseases that destroy food sources, and maintaining biodiversity. Nonetheless, I actually had had some experiences with some database such as BLAST and CLUSTAL before, which would help me through this course.
Different trees have different patterns, which I think are pretty impressive and very beautiful.
http://explorebio.wikispaces.com/The+Art+of+Phylogeny

After the introduction, we began a short exercise - to establish a simple cladogram.

  1. There are 8 different species. We described each using different characters. The variants within a character is called character states. For example, the color of the abdomen is a character where as white / black abdomen are the character states.
  2. Then we made a data matrix, with 0 meaning primary states and 1 being the developed states.
  3. Based on the matrix, we constructed a Venn diagram.
  4. Finally, we drew a cladogram based on the Venn diagram.
This exercise was actually really tricky because black abdomen rise up in two different places, which is called a homoplastic character. Although I had some basic background from AP Biology, this exercise was harder than expected. Luckily, we now have technologies that help us do this (which is what I will be learning!), and this was just to give me a taste of how long scientists used to spend on sorting out the evolutionary tree.

I definitely learned a lot, including many vocabs, during today's meeting. While Dr. Miller won't be here next week, he gave me a book Phylogenetic Trees Made Easy: A How-To Manual to read. My assignment is to learn to establish a phylogenetic trees on MEGA. Today was a great start and I am ready to explore the beauty of a phylogenetic tree!


Here is the link for the video Dr. Miller showed me during the introduction, and I thought it was pretty cool and well explained: http://archive.peabody.yale.edu/exhibits/treeoflife/film_discovering.html

Sunday, September 8, 2013

A Brand New Year - STEM Intern 2013-2014!

September 8th, 2013 Time went by so fast and it has already been a week since the school start! This year, I am really glad that I got the opportunity to participate in EWS STEM Internship program again. However, it feels different doing an internship this year, especially after the intern experiences from last year and from this past summer. Last year, I focused more on learning lab techniques and the principles under systematic and molecular biology, and this past summer I practiced a lot of these techniques such as PCR, digestion, and gel electrophoresis. Thus, this year I hope I can learn more advanced techniques such as western blot.  If possible, I wish to have my own project to work on, which would allow me to be familiar with the process - question, hypothesis, experiments, and results - that a lab project is involved in. Last but not least, I would like to explore my topic of interest (molecular / cellular biology, genetics, oncology) even more in depth. I am very exciting for the program to start, and I believe this year will be another wonderful year!

The image on the left shows a cell undergoing extocytosis, and the image on the right shows the vesicles on dock located in the synapse in the cerebellar cortex. My work for this past summer involved in a protein that participated in intracellular trafficking in yeast. The image is retrieved from Collins Lab at Cornell Univeristy. http://blogs.cornell.edu/collinslab/


Wednesday, May 1, 2013

End of the Year Poster Section

Today (May, 1st 2013) was my last day of internship of the year! Time passed by so fast and I couldn't believe I have come so far! I made aposter and we interns presented our research during lunch today in Kellas Commons. Many people came and we explained our researches to them. Not until I repeated myself over and over again did I realize how much I actually learned. What seemed so incomprehensible back in November now made perfect sense to me:)
A screenshot of my final presentation poster.
That afternoon I went to RPI for the one last time. Eun Ji gave me a general view of some cool stuff she was working on in other labs. Guess what are those?
Frozen bovine (cow) cornea!
On the project on malaria prevention in fetus during pregnancy, Eun Ji was trying to extract and modify a compound from bovine cornea that could bind to specific receptors on placenta to block parasite-encoded variant surface antigens to bind to those receptors and transmit malaria to fetus. Later, I also helped prepare some dye for silver staining that is used widely in protein detection in gel.

In the end, I said goodbye to both Eun Ji and Namita. I really had a wonderful year doing STEM internship at RPI! From the people I worked with, to lab environment, to the actual lab work, it gave me a taste of the life of a scientist. The uppermost thing I would take away from this experience in addition to a plethora of laboratory techniques is the “qualities of a scientist” – precise, robust, patient, inquisitive, inventive, and inspirational. Not everything will yield the result we want, and we just have to keep trying. This experience makes me more certain about my goals. In the future, I would like to participate in more research focusing on molecular biology to enrich my experience and widen my horizon! Last but not the least, I will be working in another lab at Cornell this summer on intracellular communication, which I am very excited about! I will keep posting interesting things and events, so stay tuned!:)

When I grow up...

Saturday, April 27, 2013

Happy National DNA Day!

This Thursday April 25, 2013 was the National DNA Day!!! This year particularly is the 60th anniversary of Watson & Cricks' discovery of DNA structure and the 10th anniversary of Human Genome Project (HGP)!
Retrieved from ASHG
I am especially interested in genetics because of its intricacy and roles in diseases. HGP provides scientists an unprecedented opportunity to better understand the role of genetics in human health and gradually reveal the mystery of genetic diseases. Today, genetics becomes increasingly important in diagnosis, drug development, and new treatments. Bellow is an excerpt of an essay I wrote on HGP:
HGP, though it does not boost the speed, increases the “success rate” of drug development. Identifying a specific mutation allows scientists to develop a targeted drug that directly tackles that mutation. Many studies on targeted drugs are done in cancers. Targeted cancer therapies are drugs or substances that inhibit the uncontrollable growth of cancer cells by blocking growth factors or inducing cell death (apoptosis). HGP facilitates the identification of these “targets”, which are usually defective genes that encode for proteins involved in cell signaling pathways. For instance, in chronic myeloid leukemia (CML), researchers had identified gene BCR-ABL – a result of translocation between chromosome 9 and 22. This gene produces a hyperactive protein that keeps Abl signaling pathway active and causes continuous proliferation of CML cells. Researchers can then develop a drug that represses this defective gene and treat the deadly disease (National Cancer Institute, 2011).
Here is the speech delivered by Francis Collins, the director of NIH, on National DNA day: http://directorsblog.nih.gov/dnas-double-anniversary/#more-1194 I really I could attend the annual ASHG meeting someday! Anyways, Happy National DNA Day!!

Because I had proctor training this Wednesday, I wasn't able to go to RPI. My internship is soon coming to an end, and I have working on my poster for my presentation:) I hope it'll all go well!