Wednesday, April 18, 2007

A Pro-Linux Cartoon

I see this picture at http://www.whylinuxisbetter.net/. Creative and Interesting!








Monday, March 12, 2007

Recommended Firefox Extensions

1. Tab Mix Plus

Tab Mix Plus enhances Firefox's tab browsing capabilities. It includes such features as duplicating tabs, controlling tab focus, tab clicking options, undo closed tabs and windows, plus much more. I strongly recommend the options of Select the tab pointed (Tool->Tab Mix Plus Option->Mouse->Mouse Gestures). It greatly decrease the amount of clicks when browsing.
2. Session Manager
Session Manager saves and restores the state of all windows - either when you want it or automatically at startup and after crashes. For example, if you start you work everyday by opening some webpages , you can store them in a session. Additionally it offers you to reopen (accidentally) closed windows and tabs. If you're afraid of losing data while browsing - this extension allows you to relax...
3. Firefox Showcase
If you habitually find yourself awash in open tabs (It seems that one of my classmates usually runs into such situation), clicking around looking for the page you need, Firefox Showcase will save you a lot of aggravation. Once you install the extension, you'll have a new Showcase submenu under the View menu. From here you can choose to show thumbnails of all tabs in the current window or all tabs in all windows. Firefox has lots of options and keyboard shortcuts, however I will never dive into those complex options. Simply click F12 to get the thumbnail view, uparrow and downarrow to select the intended tab and enter to see the full view the select tab. I can also exit the thumbnail view by clicking Esc. That's all. For more, see View->Showcase and for a sidebar view, follow View->Sidebar->Showcase Sidebar.
4. Download Statusbar
If you're tired with that sometimes-pesky Downloads window that pops up whenever you download a file in Firefox. Download Statusbar suppresses that window from popping up, and instead provides you the same information in the status bar at the bottom of the browser window. You can roll your mouse over the filename and get a pop-up tool tip with some extra information about your download, too
5. DownThemAll
DownThemAll is is download manager and accelerator. It lets you download all the links or images contained in a webpage and much more: you can refine your downloads by fully customizable criteria to get only what you really want! Simply, it saves you them time to open a shell and use wget.
6. Fotofox
Using Fotofox, you are able to grab a picture from any pages to the Fotofox sidebar, title it, tag it and upload them to your flickr album with a simple click. It is also compitable with other picture hosting site such as Tabblo, 23hq, Smugmug, Marela, and Kodak EasyShare Gallery.
7. Fasterfox
Exactly I don't know its performance.
8. Google Browser Sync
Google Browser Sync for Firefox is an extension that continuously synchronizes your browser settings – including bookmarks, history, persistent cookies, and saved passwords when you are using firefox on several computers. It also allows you to restore open tabs and windows across different machines and browser sessions. Alternatively You can choose to sync cookies, or not to sync cookies, but you can't make the decision based on individual cookies. Suppose that you are now using a public computer. At first you can install this extension, sign in with your google account. The Google Browser Sync begins to synchronize your firefox settings on the server and thus you get a familiar firefox in the public computer. When you are going to leave, choose Tools->Google Browser Sync->Stop Syncing and click Tools->Clear Private Data... (Ctrl+Shilt+Delete) to clean your personal leaved on the public computer.
9. Google Notebook
In my opinion Google Notebook is a simple but great product though most people ignorate it. This is possibly because that there is not a convinient on-one-click client interface and that the usability is poor in the webpage user interface. For example, I have been long hoping for the tag feature and sharing ability based on a note. Google Notebook it the client interface. It simple the process you make notes. Just select any part of the webpages, whether be it text or picture, click Note This from pop menu activated by right mouse button. I believe Google notebook will outrun Clipmarks or other similar services.
10. Firefox Google Bookmarks
Firefox Google Bookmarks (GBookmarks) creates a menu to access your google bookmarks from any computer. ( Your google bookmarks resides in your search history). Additionally it server a backup mechanism for all bookmarks. This extension may overlap with the Google Browser Sync. Personally, I consider the Bookmark in firefox a lightweight bookmarks that store only the everyday used links and use GBookmarks as a heavyweight repository of all links that may be useful some day.
11. StumbleUpon
StumbleUpon is also a bookmark service that incorporate social network elements. It resembles delicious. StumbleUpon lets you "channelsurf" the best-reviewed sites on the web. It is a collaborative surfing tool for browsing, reviewing and sharing great sites with like-minded people. This helps you find interesting webpages you wouldn't think to search for. You can also share pages of interest within a community. Particularly, it will insert into your google search results which shows you other people's rating and reviews of the search result.
12. Greasemonkey
Greasemonkey basically allows you to add JavaScript to any Web page, which implies infinite control over the behavior of web pages. Greasemonkey is not for the faint of heart. The good news is that there are many generous souls out there who share the scripts they create. Check out userscripts.org for a script repository. If you want to write your own scripts, try diveintogreasemonkey.org or pick up Mark Pilgrim's Greasemonkey Hacks from O'Reilly Media. Personally I strongly recommend Gmail Macros. After installation, it empower gmail with striking and convinient keyboard shortcuts like Google Reader.
13. Adblock
Adblock is a content filtering plug-in for the Mozilla and Firefox browsers. It allows the user to specify filters, which remove unwanted content based on the source-address. Adblock supports two types of filters: simple, and Regular Expression. Adblock will also provide a default list of filters, which is enough for a lazy person like me.
14. ScrapBook
ScrapBook is a Firefox extension, which helps you to save Web pages and easily manage collections. This enable you to surf web pages off line. You can also directly copy the ScrapBook folder from your firefox option folder to other computers and see what you saved there.
References:
20 must-have Firefox extensions
Firefox Add-ons Recommended Add-ons
千种风情千种树: My Firefox Extensions

Friday, February 09, 2007

Alon Halevy and Peter Norvig


Alon Halevy and Peter Norvig, two Googlers, have been selected for the 2006 class of ACM Fellows.

Peter, who was Google's first director of search quality and is currently director of Google Research, has been recognized for his many contributions to the disciplines of artificial intelligence and information retrieval. His personal websites is http://norvig.com/.

Alon, who leads one of our structured data initiatives, has been honored for his contributions in data integration and knowledge representation. His old personal website: http://www.cs.washington.edu/homes/alon/, new personal website: http://alonhalevy.googlepages.com/, his blog is http://www.alonhalevy.blogspot.com/

Saturday, February 03, 2007

A Elegent Blog: Designer's Block


The blog Designer's Block is elegant blog. The writer is from UK. I really like the design (see left), grace and mysterious style. I will choose it as my blog's background.








update: In fact, the painter of above paintings are Melissa Mossart. In her website there are also many paintings of similar style.





































See more paintings at http://www.melissamossart.com/paint.htm
Melissa Mossart and nine other women artists' gallery: http://www.tenwomen.org/venicegallery.html

Thursday, February 01, 2007

Research Projects on Microarray Analysis


BioConductor


Bioconductor is an open source and open development software project for the analysis and comprehension of genomic data.

Bioconductor is primarily based on the R programming language but we do accept contributions in any programming language. Although initial efforts focused primarily on DNA microarray data analysis, many of the software tools are general and can be used broadly for the analysis of genomic data, such as SAGE, sequence, or SNP data.

The broad goals of the projects are to

  • provide access to a wide range of powerful statistical and graphical methods for the analysis of genomic data;
  • facilitate the integration of biological metadata in the analysis of experimental data: e.g. literature data from PubMed, annotation data from LocusLink;
  • allow the rapid development of extensible, scalable, and interoperable software;
  • promote high-quality documentation and reproducible research;
  • provide training in computational and statistical methods for the analysis of genomic data.
If you are new to Bioconductor you might consider buying Bioinformatics and Computational Biology Solutions Using R and Bioconductor


Gene Ontology


The Gene Ontology (GO) project is a collaborative effort to address the need for consistent descriptions of gene products in different databases. The GO project has developed three structured controlled vocabularies (ontologies) that describe gene products in terms of their associated biological processes, cellular components and molecular functions in a species-independent manner. There are three separate aspects to this effort: first, the development and maintenance of the ontologies themselves; second, the annotation of gene products, which entails making associations between the ontologies and the genes and gene products in the collaborating databases; and third, development of tools that facilitate the creation, maintenance and use of ontologies.

The use of GO terms by collaborating databases facilitates uniform queries across them. The controlled vocabularies are structured so that they can be queried at different levels: for example, you can use GO to find all the gene products in the mouse genome that are involved in signal transduction, or you can zoom in on all the receptor tyrosine kinases. This structure also allows annotators to assign properties to genes or gene products at different levels, depending on the depth of knowledge about that entity.

International HapMap Project


The HapMap is a catalog of common genetic variants that occur in human beings. It describes what these variants are, where they occur in our DNA, and how they are distributed among people within populations and among populations in different parts of the world. The International HapMap Project is not using the information in the HapMap to establish connections between particular genetic variants and diseases. Rather, the Project is designed to provide information that other researchers can use to link genetic variants to the risk for specific illnesses, which will lead to new methods of preventing, diagnosing, and treating diseases.

Microarray Gene Expression Data Society - MGED Society

The Microarray Gene Expression Data (MGED) Society is an international organisation of biologists, computer scientists, and data analysts that aims to facilitate the sharing of microarray data generated by functional genomics and proteomics experiments. The current focus is on establishing standards for microarray data annotation and exchange, facilitating the creation of microarray databases and related software implementing these standards, and promoting the sharing of high quality, well annotated data within the life sciences community. A long-term goal for the future is to extend the mission to other functional genomics and proteomics high throughput technologies


The MolTools consortium



The MolTools consortium started on January 1st 2004, as a joint research programme bringing together 12 leading European academic groups, four biotech SMEs and one US laboratory working in the area of postgenomic technology development. The partners have pioneered a series of important molecular techniques and will now work together to establish next-generation tools for molecular analysis. Its scientific aims are to establish genome analysis technologies set to monitor extensive molecular repertoires, and with the capacity to investigate even single molecules. Its current research projects include:

Tuesday, January 30, 2007

Resources of Genomics and Microarray Analysis

Genomics and Microarrays:
Stanford Microarray Database:
storing lots of raw and normalized data from microarray experiments

Genomics tutorial at Genome Canada
http://www.genomecanada.ca/xpublic/dnaBasics/index.asp?l=e

Introductions to microarray at NCBI
http://www.ncbi.nlm.nih.gov/About/primer/microarrays.html

Microarray (movie)
http://www.broad.harvard.edu/chembio/lab_schreiber/anims/videos/microarray.html

other resourses for microarray
http://www.learner.org/channel/courses/biology/units/genom/images.html

image analysis for microarray
http://www.maths.usyd.edu.au/u/jeany/ (publication)
http://www.stat.berkeley.edu/users/terry/zarray/Talks/image/jpegindex.html http://cmm.ensmp.fr/~angulo/research/dnamicro.htm

Bioconductor
The richest source of freely available packages for genomic data analysis

Nature article
the perspective of biologists facing heaps of noisy genomic data including their urgent need for better methods and computationally and statistically skilled support.

dChip Software
http://biosun1.harvard.edu/complab/dchip/

dChip Software: Analysis and visualization of gene expression and SNP microarrays



Biology:

Retroviruses
http://www.whfreeman.com/kuby/content/anm/kb03an01.htm (FLASH)

human genome project(movies)
http://www.genome.gov/Pages/EducationKit/download.html

the central dogma of molecular biology (wonderful movie):
http://www.genome.gov/Pages/EducationKit/video/qt/3D.mov

EBM & Clinical Research Workstation menu
http://www.shdem.com/ebm/default.asp

Biochemistry & Epidemiology useful link:
http://www.med-ed.virginia.edu/menu/otherMedEd.cfm
 
Statistics:

online textbook for statistics
http://www.stat.berkeley.edu/~stark/SticiGui/Text/toc.htm

Terry Speed's Microarray homepage
statistical challenges related to microarray data
new webpage: http://www.stat.berkeley.edu/~terry/Group/home.html

Statistics
http://www.bettycjung.net/statsiteS.htm

Computing technology

Introduction to R for biologists (by Natalie Roberts, WEHI, Melbourne)
R manuals under link "Manuals" (left column)
manuals: http://cran.r-project.org/manuals.html
R tutorial: http://www.personality-project.org/r/
R package: Statistics for Microarray Analysis

Directionary of Blogs About Microarray Analysis: Draft

Biodefense Bioinformatics Blog http://ai59694.blogspot.com/
Rotten bananas - http://heathermaughan.blogspot.com/index.html
Synthetic Biology and Gene Synthesis - http://syntheticbio.blogspot.com
Genomics Online - http://genomics-info.blogspot.com/index.html
formerscienceguy - http://formerscienceguy.blogspot.com/index.html

Friday, January 26, 2007

Alan Perelson

Dr. Perelson received his B.S. degrees in Life Science and Electrical Engineering from MIT in 1967, and a Ph.D. in Biophysics, under the supervision of Aharon Katchalsky-Katzir, from UC Berkeley in 1972. He was Acting Assistant Professor, Division of Medical Physics, Berkeley, in 1973 and a postdoctoral fellow at the Department of Chemical Engineering, University of Minnesota, in 1974. He was a staff member in the Theoretical Biology and Biophysics Group at Los Alamos National Laboratory from 1974 - 1991, a Laboratory Fellow from 1991 - 2002, head of the Theoretical Biology and Biophysics Group between 1995 - 2001, and is currently a Los Alamos National Laboratory Senior Fellow. He spent the 1978 and 1979 academic years at Brown University as an Assistant Professor of Medical Sciences in the Division of Biology and Medicine and the Lefschetz Center for Dynamical Systems, was a visiting scientist at the Mathematical Institute, Oxford University in 1986 and a visiting professor of Physics at Ecole Normale Superieure, Paris in 1990, and the University of Paris VII in 1992. He is also a member of the Science Board and head of the Theoretical Immunology Program at the Santa Fe Institute. He is also an adjunct professor of Bioinformatics at Boston University and an adjunct professor of biology at the University of New Mexico.

Research Interests

Mathematical and theoretical biology, with an emphasis on problems in immunology, virology,
and cell and molecular biology.

Time Zone of United States

PST: Washington, Oregon, Neveda, California

MT: Montana, Wyoming, Idaho, Utah, Colorado, Arizona, New Mexico and parts of
North Dakota, South Dakota and Nebraska

CT: Parts of North Dakota, South Dakota and Nebraska, Kansas, Oklahoma, Texas,
Minnesota, Iowa, Missouri, Arkansas, Louisiana Wisconsin, Illinois,
Tennessee, Mississippi, Alabama

EST: Michigan, Indiana, Ohio, Kentucky, Georgia, New York, Pennsylvania, West
Virginia, Virginia, North Carolina, South Carolina, Florida, Washington DC
New Jersey, Connecticut, Ehode Island, Massachusetts, New Hampshire,
Vermont, Maine

Thursday, January 25, 2007

Friday, January 19, 2007

A Haplotype Map of the Human Genome

A Haplotype Map of the Human Genome

David Altshuler
Harvard Medical School, Massachusetts General Hospital, Whitehead Institute

Eric Lander
Whitehead Institute and MIT

Goal

The next key step of the Human Genome Project (HGP) (following the creation of the genetic, physical, sequence and SNP maps) is the generation of a "haplotype" map of the human genome. Such a "haplotype" map consists of a high density of SNPs defining the small number of ancestral haplotypes (blocks of tightly correlated genetic variants) in each region of the human genome. Knowledge of these haplotypes will allow comprehensive and efficient testing of the association of human genes with human diseases. The haplotype map can and should be generated rapidly and should be made freely available to researchers worldwide.

Background

A haplotype map of the human genome has become both justified and practical due to significant advances over the last two years.

Specifically, these advances include:

  • Genomic Sequence: The development of a complete genome sequence - integrated with human genes and annotations - providing a reference framework on which to layer knowledge about allelic variation.

  • Genetic Variants: The development of a dense (and rapidly growing) map of 1.4 million human SNPs provides a genome-wide resource of genetic variation adequate to uniquely tag the vast majority of human haplotypes.

  • Genotyping Technology: The development of high-throughput methods, allowing a rapid, efficient and cost-effective experimental approach to a project of the required scale.

  • Long-range LD: The discovery that human SNPs display strong linkage disequilibrium (LD or allelic association) over large distances. LD is detectable over distances in the range of 100kb and is extremely strong over regions spanning several tens of kb (the size of typical genes). For such regions, the vast majority of chromosomes in the population carry one of a handful of highly conserved haplotypes. As a result, genetic diversity in the region can be represented by a small number of well-chosen SNPs.

Impact on biomedical research

The availability of a haplotype map of the human genome will have a substantial impact on human genetic studies.

Specifically, these studies include:

  • Comprehensive association studies of individual genes. The association of genes with disease has traditionally been probed by testing individuals SNPs one-at-a-time. The drawback to this approach is that the task is never-ending: one can exclude particular SNPs as playing a role, but one cannot exclude a gene. Once the haplotype structure of the genome is defined, one can (1) comprehensively test all significant haplotypes in the gene, and (2) decrease the number of SNPs needed by selecting a subset that defines the population variability. This will allow haplotype studies of individual genomic loci in an unbiased manner, without assumption about the locations of causal mutations in coding regions, promoters or regulatory sites at significant distance away. And, it will greatly decrease the technical and financial barriers faced by laboratories in undertaking such work

  • Genome-wide association studies. A genome-wide haplotype map will make possible whole-genome scans for association in the population. Rather than focusing only on 'candidate' genes, it will become possible to search the genome in an unbiased manner for genes whose common variation contributes to disease in the population. Routine use of genome-wide association studies will also require further decreases in genotyping costs, but such decreases are likely to be driven by the development of the haplotype map.

  • Human population structure and history. Knowledge of haplotypes will transform our understanding of human population structure and history. The LD pattern turns out to be an extremely sensitive indicator of population history, because the multi-allelic nature of haplotypes provides rich detail and because the breakdown of haplotypes follows a predictable clock set by recombination rates. In particular, LD patterns are more powerful than traditional studies of allele frequencies per se. Information about human population history is interesting in its own right, but is also very valuable in the design of medical studies (such as admixture mapping).

Technical Issues

Generating a haplotype map would involve the following components:

  • Population Samples. Development of appropriate population samples, consisting of parent-offspring trios (to allow inference of haplotypes). We estimate that a total of about 300 samples will be needed, representing major ethnic groups in a manner appropriate for generating a map that can be used for medical studies in all populations. The population samples should be a renewable resource (i.e., immortalized cell lines).

  • Sample and Data Availability. The samples should be made freely available so that any interested scientific group can contribute data (in the manner of the CEPH panel and the DNA Polymorphism Discovery Resource). Conversely, all data generated by the project should be immediately released into the public domain without restrictions of any kind.

  • Numbers of SNPs to be genotyped. It is estimated that generating the haplotype map will require successful genotyping of 450,000 SNPs, which will in turn require initial testing of some 800,000 to 900,000 SNPs. The required scale is now well within reach: the Whitehead and Sanger Centre are each currently engaged in pilot projects involving 25,000 SNPs using automated genotyping setup and MALDI-TOF-based detection. Given the required scale and efficiencies, it is likely that the bulk of the work should be performed by a few large groups, but all groups should be encouraged to participate in the project by analyzing genes and regions of interest.

  • Analytical Tools. The project will require various analytical tools to readily define haplotype blocks from genotype data, software systems to aid in the hierarchical selection of SNPs to fill in blocks, and databases to make the information maximally useful to the community. Prototype systems have been developed, but focused effort will be needed to develop mature systems.

Wednesday, December 20, 2006

Chris Sander


Chris Sander
Chris Sander

Trained as a theoretical physicist, Chris Sander, Director of Memorial Sloan-Kettering Cancer Center's Computational Biology Center and Chairman of the Sloan-Kettering Institute's Computational Biology Program, knew early in his career he wanted to do more than compute the behavior of elementary particles. After a bold move from physics to biology, he helped to develop the field of computational biology, which aims to use mathematical algorithms and information systems to simulate the behavior of molecules, cells, and organisms, using these simulations to make useful diagnostic and therapeutic predictions.

While an undergraduate at the University of Berlin, I was intensely engaged in the study of theoretical physics and mathematics. Yet I felt that analyzing life, the living system on this planet, was a more fascinating and challenging scientific problem than studying the world of elementary particles. To explore this potential change in career direction, I sought the advice of Max Delbrück at the California Institute of Technology, a physicist originally from Berlin who was one of the founders of molecular genetics.

After visiting friends in Texas, I hopped on a Greyhound bus to Los Angeles, made my way to Caltech, found Dr. Delbrück's office, knocked on the door, and with a dash of chutzpah, asked the Nobel laureate if he had a moment to chat with an aspiring graduate student. Dr. Delbrück described the areas of theoretical physics that might be relevant to biology in the future -- information that was in the forefront of my mind as I entered the graduate physics program at the University of California, at Berkeley, in 1967.

Wanting to move from the theoretical physics of my PhD thesis to the theoretical biology I had dreamt of, I made my second pilgrimage, this time to see Manfred Eigen, a Nobel Prize-winning chemist who was studying biological evolution in Göttingen, Germany. Dr. Eigen surprised me by explaining that the field barely existed, but he did point me to three mathematical biology research problems: a theory of the immune response, neuronal mapping in the brain, and protein folding. So I packed my bags and I moved from the University of Heidelberg to the Weizmann Institute of Science, in Israel, where I began work with Shneior Lifson on the prediction of three-dimensional protein structures.

Enter the third motivator of my career: the first completely sequenced genome -- no, not in 2000, but in 1977! In that year, I saw an amazing paper in the journal Nature from Fred Sanger's group in Cambridge, United Kingdom, which included two entire pages filled with 5,375 letters, all As, Ts, Gs and Cs, representing the genetic blueprint of a small virus. I walked down the hall to ask my friend Georg Schulz, "With this kind of cryptic information coming from genomes, won't biology need computational science to decipher it?" His answer was yes, and I spent the next 23 years of my professional life preparing for the day, in the year 2001, when the 3.5 billion letters of the human genome finally became available. In the process, I helped to develop the field now known as computational biology.

The real value of computational science, when applied to any system, is to predict what's going to happen next. Weather forecasting is an example. There's an enormous amount of data collected about the weather, but the data, by themselves, are unintelligible. What's required is the application of the appropriate mathematical equations embodied in a software system, which, using the data, allows one to compute tomorrow's weather.

Applied to cancer biology, we want to be able to predict, for example, if a cancer will go from a nonaggressive to an aggressive form, or more importantly, to predict accurately the consequences of possible therapeutic interventions. The goal is to have an impact on human disease, and to do this you have to work in collaboration with physicians. In 2002, Harold Varmus presented his vision of Memorial Sloan-Kettering as the perfect environment for this -- a place with open doors between basic and clinical research, where close collaboration is encouraged.

My first action at Memorial Sloan-Kettering Cancer Center was to start the Computational Biology Center (CBC) and its Bioinformatics Core Facility. The CBC's researchers and engineers are devoted both to basic science and to the goal of developing diagnostic and therapeutic tools that help improve the lives of people affected by cancer. We often collaborate with researchers in the lab and in the clinic to translate data -- data such as the molecular profiles of cells and tissues, the billions of letters of genome sequences, and the functions and structures of key genes -- into biological insights and prediction tools. And the Bioinformatics Core, ably led by Alex E. Lash, provides internal bioinformatics training, collaboration, and infrastructure support.

One concrete example of the practical uses of computational biology is the work we have been doing with Howard I. Scher, Chief of Memorial Sloan-Kettering Cancer Center's Genitourinary Oncology Service, and Francis M. Sirotnak, Member Emeritus and Head of Sloan-Kettering Institute's Laboratory of Molecular Therapeutics. The idea is that cancer cells, like any system that recovers from a round of major damage, might be especially sensitive after a first round of therapy. With this concept in mind, we are aiming to prevent the development of aggressive prostate cancer by looking at the molecular profile of prostate cancer cells after androgen removal in mice, using DNA chips provided by our Genomics Core Laboratory. We use computer software to find needles in a haystack -- the perhaps tens of genes, out of tens of thousands, that may be a characteristic signature of how prostate cancer reacts to such therapy. We hope this will lead us to an Achilles' heel to target to avoid recurrence. It's a long-term effort but the idea is to arrive computationally at the best therapeutic intervention.

Overall, what's been most rewarding for me during my short time here is the opportunity not just to predict the behavior of biological systems, but hopefully to help improve the quality of people's lives. The dream that started with Max Delbrück's advice is now within reach.

Saturday, December 09, 2006

Super Computing

 Logo
The TOP500 project was started in 1993 to provide a reliable basis for tracking and detecting trends in high-performance computing. Twice a year, a list of the sites operating the 500 most powerful computer systems is assembled and released. The best performance on the Linpack benchmark is used as performance measure for ranking the computer systems. The list contains a variety of information including the system specifications and its major application areas.


NCSA Home
The National Center for Supercomputing Applications (NCSA), one of the five original centers in the National Science Foundation's Supercomputer Centers Program, opened its doors in January 1986. Since then, NCSA has contributed significantly to the birth and growth of the worldwide cyberinfrastructure for science and engineering, operating some of the world's most powerful supercomputers and developing the software infrastructure needed to efficiently use these systems (for example, NCSA Telnet and, in 1993, NCSA Mosaic™, the first readily available graphical Web browser). Today the center is recognized as an international leader in deploying robust high-performance computing resources and in working with research communities to develop new computing and software technologies

Blue Gene




Blue Gene is an IBM Research project dedicated to exploring the
frontiers in supercomputing: in computer architecture, in the software required to program and control massively parallel systems, and in the use of computation to advance our understanding of important biological processes such as protein folding.

The full Blue Gene/L machine was designed and built in collaboration with the Department of Energy's NNSA/Lawrence Livermore National Laboratory in California, and has a peak speed of 360 Teraflops. Blue Gene systems occupy the #1 (Blue Gene/L) and #2 (Blue Gene Watson) positions in the TOP500 supercomputer list announced in November 2005, as well as 17 more of the top 100.

IBM now offers a Blue Gene Solution. IBM and its collaborators are currently exploring a growing list of applications including hydrodynamics, quantum chemistry, molecular dynamics, climate modeling and financial modeling.

SDSC - San Diego Super Computer Center

Founded in 1985, the San Diego Supercomputer Center (SDSC) enables international science and engineering discoveries through advances in computational science and high performance computing. Continuing this legacy into the era of cyberinfrastructure, SDSC is a strategic resource to science, industry and academia, offering leadership in the areas of data management, grid computing, bioinformatics, geoinformatics, high-end computing as well as other science and engineering disciplines. The mission of SDSC is to extend the reach of scientific accomplishments by providing tools such as high-performance hardware technologies, integrative software technologies and deep inter-disciplinary expertise, to the community.

SDSC was founded with a $170 million grant from the National Science Foundation's (NSF) Supercomputer Centers program. From 1997 to 2004, SDSC extended its leadership in computational science and engineering to form the National Partnership for Advanced Computational Infrastructure (NPACI), teaming with approximately 40 university partners around the country. Today, SDSC is an organized research unit of the University of California, San Diego primarily funded by NSF with a staff of talented scientists, software developers and support personnel.





The National Resource for Biomedical Supercomputing (NRBSC) pursues leading edge research in high performance computing and the life sciences, and fosters exchange between PSC expertise in computational science and biomedical researchers nationwide.

Our focus is two-fold: computational biomedical research and outreach to the national biomedical research community through education and publications.

Research at NRBSC is centered in three areas: microphysiology; volumetric visualization and analysis; and computational structural biology.

NRBSC's education arm includes not only user training, but also software distribution, publications, and other outreach activities such as online courses and workshop webcasts.

The National Resource for Biomedical Supercomputing, formerly the Biomedical Initiative, was established at the Pittsburgh Supercomputing Center in 1987 as the first extramural biomedical supercomputing program in the country funded by the National Institutes of Health.

Tuesday, December 05, 2006

Python Resources Collection

Website



Python® is a dynamic object-oriented programming language that can be used for many kinds of software development. It offers strong support for integration with other languages and tools, comes with extensive standard libraries, and can be learned in a few days. Many Python programmers report substantial productivity gains and feel the language encourages the development of higher quality, more maintainable code.

Jython
Python is an implementation of the high-level, dynamic, object-oriented language Python written in 100% Pure Java, and seamlessly integrated with the Java platform. It thus allows you to run Python on any Java platform.




Stored in these dark caverns you may find rich veins of Python code, collected caches of Python information, and all manner of sundry Python passageways to explore. With candle or torch in hand, good hunting this night to all.Those not familiar with Python perhaps might start your quest at a brighter place.

NumPy

The fundamental package needed for scientific computing with Python is called NumPy. This package contains: a powerful N-dimensional array objectsophisticated (broadcasting) functionsbasic linear algebra functions basic Fourier transformssophisticated random number capabilitiestools for integrating Fortran code.

SciPy.org
SciPy (pronounced "Sigh Pie") is open-source software for mathematics, science, and engineering. It is also the name of a very popular conference on scientific programming with Python. The core library is NumPy which provides convenient and fast N-dimensional array manipulation. The SciPy library is built to work with NumPy arrays, and provides many user-friendly and efficient numerical routines such as routines for numerical integration and optimization. Together, they run on all popular operating systems, are quick to install, and are free of charge. NumPy and SciPy are easy to use, but powerful enough to be depended upon by some of the world's leading scientists and engineers. If you need to manipulate numbers on a computer and display or publish the results, give SciPy a try!


DISLIN Homepage
DISLIN is a high-level plotting library for displaying data as curves, polar plots, bar graphs, pie charts, 3D-color plots, surfaces, contours and maps.


wxPython is a GUI toolkit for the Python programming language. It allows Python programmers to create programs with a robust, highly functional graphical user interface, simply and easily.


VPython is a package that includes: the Python programming languagethe IDLE interactive development environment "Visual", a Python module that offers real-time 3D output, and is easily usable by novice programmers"Numeric", a Python module for fast processing of arrays

PyOpenGL Logo

PyOpenGL is the cross platform Python binding to OpenGL and related APIs. The binding is created using the SWIG wrapper generator, and is provided under an extremely liberal BSD-style Open-Source license.


The Biopython Project is an international association of developers of freely available Python tools for computational molecular biology.It is a distributed collaborative effort to develop Python libraries and applications which address the needs of current and future work in bioinformatics. The source code is made available under the Biopython License, which is extremely liberal and compatible with almost every license in the world. We work along with the Open Bioinformatics Foundation, who generously provide web and CVS space for the project

PyZine
The journal of Python Language



PyLucene is a GCJ-compiled version of Java Lucene integrated with Python. Its goal is to allow you to use Lucene's text indexing and searching capabilities from Python. It is designed to be API compatible with the latest version of Java Lucene.

Courses with an emphasis on scientific computing

Python course in Bioinformatics
Introduction to Python and Biopython with biological examples.



Monday, December 04, 2006

LaTeX Resources Collections

To start with Latex, the most convenient way is to read a short book with a strange name( sorry, I can not recall it now).
And Professor Schneider provides a page to introduce Latex for Biologist, here it is
LaTeX Style and BiBTeX Bibliography Formats for Biologists: TeX and LaTeX Resources

Monday, November 27, 2006

The International Society for Computational Biology (ISCB)


The International Society for Computational Biology (ISCB) is incorporated in the United States as a 501(c)(3) non-profit corporation, and registered in the state of California as a Charitable Trust. Now hosted at the San Diego Supercomputer Center at University of California, San Diego, the Society was officially formed in 1997 as an outgrowth of the conference on Intelligent Systems for Molecular Biology (ISMB). From humble beginnings, both ISCB's membership and ISMB's annual attendance have kept pace with the overall explosive growth experienced in the field of bioinformatics/computational biology.

For the complete story from conception to present day please visit the History link below. Additional links provide copies of documents that detail the legal structure of ISCB, which may prove interesting to our current and prospective members, as well as be of some use to regional groups around the world wanting to form national or regional societies and not knowing how to begin. And finally, ISCB's mission, vision and values are detailed in the 2003 Strategic Plan, and activities over the years are documented in the Newsletter Archives. We encourage you to peruse both of these links for a perspective on where we've been and where we may be going in the years ahead.

Ryan Songer

Sunday, November 26, 2006

ThinkWiki: Install Linux on ThinkPad

kkk recommended this site.

ThinkWiki

From ThinkWiki

Jump to: navigation, search

This is ThinkWiki, the Wiki Web for ThinkPad users. Here you find anything you need to install your favourite Linux distribution on your ThinkPad. Windows users shouldn't run away, there's a lot of useful information for them as well.

Please support us and help to extend this wiki. Thank you!



Ryan Songer

Wednesday, October 11, 2006

Zooomr and Its Free Pro Account

bird01bird01 Hosted on Zooomr

Frankly speaking, I have been a flickr fan for a long time thought it has a monthly upload limit, because I didn't expect I would use up the 50M limits(about 30+ pics) per month. However, I was frustrated when I returned from a journey with more than 50+ pics. This experience forced me to find a new photo hosting site. For a poor students like me, it should be free, free of upload limit, large enough, stable and comfortable. Under this criteria, picasa web album is poorly 250MB and still too simple. Fotkit is not free and has a messy UI. ... I have been looking for it until I came across a blog post Do We Love Bloggers? Yes We Do!
I signed up and write this post in hope a free pro account.

Thanks Kris. Yes, now I get pro account after some errors occurred though. Its seems that Zooomr is not reliably stable at this time, but I would like to try it.

Sunday, September 24, 2006

Can a Biologist Fix a Radio? -- or, What I Learned while Studying Apoptosis

Y. Lazebnik

Cold Spring Harbor Laboratory, Cold Spring Harbor

As a freshly minted Assistant Professor, I feared that everything in my field would be discovered before I even had a chance to set up my laboratory. Indeed, the field of apoptosis, which I had recently joined, was developing at a mind-boggling speed. Components of the previously mysterious process were being discovered almost weekly, frequent scientific meetings had little overlap in their contents, and it seemed that every issue of Cell, Nature, or Science had to have at least one paper on apoptosis. My fear led me to seek advice from David Papermaster (currently at the University of Connecticut), who I knew to be a person with pronounced common sense and extensive experience. David listened to my outpouring of primal fear and explained why I should not worry.

David said that every field he witnessed during his decades in biological research developed quite similarly. At the first stage, a small number of scientists would somewhat leisurely discuss a problem that would appear esoteric to others, such as whether cell cycle is controlled by an oscillator or whether cells can commit suicide. At this stage, the understanding of the problem increases slowly, and scientists are generally nice to each other, a few personal antipathies notwithstanding. Then, an unexpected observation, such as the discovery of cyclins or the finding that apoptosis failure can contribute to cancer, makes many realize that the previously mysterious process can be dissected with available tools and, importantly, that this effort may result in a miracle drug. At once, the field is converted into a Klondike gold rush with all the characteristic dynamics, mentality, and morals. A major driving force becomes the desire to find the nugget that will secure a place in textbooks, guarantee an unrelenting envy of peers, and, at last, solve all financial problems. The assumed proximity of this imaginary nugget easily attracts both financial and human resources, which results in a rapid expansion of the field. The understanding of the biological process increases accordingly and results in crystal clear models that often explain everything and point at targets for future miracle drugs. People at this stage are not necessarily nice, though, as anyone who has read about a gold rush can expect. This description fit the then current state of the apoptosis field rather well, which made me wonder why David was smiling so reassuringly. He took his time to explain.

At some point, David said, the field reaches a stage at which models, that seemed so complete, fall apart, predictions that were considered so obvious are found to be wrong, and attempts to develop wonder drugs largely fail. This stage is characterized by a sense of frustration at the complexity of the process, and by a sinking feeling that despite all that intense digging the promised cure-all may not materialize. In other words, the field hits the wall, even though the intensity of research remains unabated for a while, resulting in thousands of publications, many of which are contradictory or largely descriptive. The flood of publications is explained, in part, by the sheer amount of accumulated information (about 10,000 papers on apoptosis were published yearly over the last few years), which makes reviewers of the manuscripts as confused and overwhelmed as their authors. This stage can be summarized by the paradox that the more facts we learn the less we understand the process we study.

It becomes slowly apparent that even if the anticipated gold deposits exist, finding them is not guaranteed. At this stage, the Chinese saying that it is difficult to find a black cat in a dark room, especially if there is no cat, comes to mind too often. If you want to continue meaningful research at this time of widespread desperation, David said, learn how to make good tools and how to keep your mind clear under adverse circumstances. I am grateful to David for his advice, which gave me hope and, eventually, helped me to enjoy my research even after my field did reach the state he predicted.

At some point, I began to realize that David's paradox has a meaning that is deeper than a survival advice. Indeed, it was puzzling to me why this paradox manifested itself not only in studies of fundamental processes, such as apoptosis or cell cycle, but even in studies of individual proteins. For example, the mystery of what the tumor suppressor p53 actually does seems only to deepen as the number of publications about this protein rises above 23,000.

The notion that your work will create more confusion is not particularly stimulating, which made me look for guidance again. Joe Gall at the Carnegie Institution, who started to publish before I was born, and is an author of an excellent series of essays on history of biology [1], relieved my mental suffering by pointing out that a period of stagnation is eventually interrupted by a new development. As an example, he referred to the studies of cell death that took place in the 19th century, faded into oblivion, and re-emerged a century later with about 60,000 studies on the subject published during a single decade. Even though a prospect of a possible surge in activity in my field was relieving, I started to wonder whether anything could be done to expedite this event, which brought me to think about the nature of David's paradox. The generality of the paradox suggested some common fundamental flaw of how biologists approach problems.

To understand what this flaw is, I decided to follow the advice of my high school mathematics teacher, who recommended testing an approach by applying it to a problem that has a known solution. To abstract from peculiarities of biological experimental systems, I looked for a problem that would involve a reasonably complex but well understood system. Eventually, I thought of the old broken transistor radio that my wife brought from Russia (Fig. 1, see color insert). Conceptually, a radio functions similarly to a signal transduction pathway in that both convert a signal from one form into another (a radio converts electromagnetic waves into sound waves). My radio has about a hundred various components, such as resistors, capacitors, and transistors, which is comparable to the number of molecules in a reasonably complex signal transduction pathway. I started to contemplate how biologists would determine why my radio does not work and how they would attempt to repair it. Because a majority of biologists pay little attention to physics, I had to assume that all we would know about the radio is that it is a box that is supposed to play music.

Figure 1

Fig. 1. The radio that has been used in this study.

How would we begin? First, we would secure funds to obtain a large supply of identical functioning radios in order to dissect and compare them to the one that is broken. We would eventually find how to open the radios and will find objects of various shape, color, and size (Fig. 2, see color insert). We would describe and classify them into families according to their appearance. We would describe a family of square metal objects, a family of round brightly colored objects with two legs, round-shaped objects with three legs and so on. Because the objects would vary in color, we will investigate whether changing the colors affects the radio's performance. Although changing the colors would have only attenuating effects (the music is still playing but a trained ear of some people can discern some distortion), this approach will produce many publications and result in a lively debate.

Figure 2

Fig. 2. The insides of the radio. See text for description of the indicated components. The inset is an enlarged portion of the radio. The horizontal arrows indicate tunable components.

A more successful approach will be to remove components one at a time or to use a variation of the method, in which a radio is shot at a close range with metal particles. In the latter case, radios that malfunction (have a “phenotype”) are selected to identify the component whose damage causes the phenotype. Although removing some components will have only an attenuating effect, a lucky postdoc will accidentally find a wire whose deficiency will stop the music completely. The jubilant fellow will name the wire Serendipitously Recovered Component (SRC) and then find that SRC is required because it is the only link between a long extendable object and the rest of the radio. The object will be appropriately named the Most Important Component (MIC) of the radio. A series of studies will definitively establish that MIC should be made of metal and the longer the object is the better, which would provide an evolutionary explanation for the finding that the object is extendable.

However, a persistent graduate student from another laboratory will discover another object that is required for the radio to work. To the delight of the discoverer, and the incredulity of the flourishing MIC field, the object will be made of graphite and changing its length will not affect the quality of the sound significantly. Moreover, the graduate student would convincingly demonstrate that MIC is not required for the radio to work, and will suitably name his object the Really Important Component (RIC). The heated controversy, as to whether MIC or RIC is more important, will be fueled by the accumulating evidence that some radios require MIC while other, apparently identical ones, need RIC. The fight will continue until a smart postdoctoral fellow will discover a switch, whose state determines whether MIC or RIC is required for playing music. Naturally, the switch will become the Undoubtedly Most Important Component (U-MIC). Inspired by these findings, an army of biologists will apply the pull-it-out approach to investigate the role of each and every component. Another army will crush the radios into small pieces to identify components that are on each of the pieces, thus providing evidence for interaction between these components. The idea that one can investigate a component by cutting its connections to other components one at a time or in a combination (“alanine scan mutagenesis”) will produce a wealth of information on the role of the connections.

Eventually, all components will be catalogued, connections between them will be described, and the consequences of removing each component or their combinations will be documented. This will be the time when the question, previously obscured by the excitement of productive research, would have to be asked: Can the information that we accumulated help us to repair the radio? It will turn out that sometimes it can, such as if a cylindrical object that is red in a working radio is black and smells like burnt paint in the broken radio (Fig. 2, inset, a component indicated as a target). Replacing the burned object with a red object will likely repair the radio.

The success of this approach explains the pharmaceutical industry's mantra: “Give me a target!”. This mantra reflects the belief in a miracle drug and assumes that there is a miracle target whose malfunction is solely responsible for the disease that needs to be cured.

However, if the radio has tunable components, such as those found in my old radio (indicated by yellow arrows in Fig. 2, inset) and in all live cells and organisms, the outcome will not be so promising. Indeed, the radio may not work because several components are not tuned properly, which is not reflected in their appearance or their connections. What is the probability that this radio will be fixed by our biologists? I might be overly pessimistic, but a textbook example of the monkey that can, in principle, type a Burns poem comes to mind. In other words, the radio will not play music unless that lucky chance meets a prepared mind.

Yet, we know with near certainty that an engineer, or even a trained repairman could fix the radio. What makes the difference? I think the languages that these two groups use (Fig. 3, see color insert). Biologists summarize their results with the help of all-too-well recognizable diagrams, in which a favorite protein is placed in the middle and connected to everything else with two-way arrows. Even if a diagram makes overall sense (Fig. 3a), it is usually useless for a quantitative analysis, which limits its predictive or investigative value to a very narrow range. The language used by biologists for verbal communications is not better and is not unlike that used by stock market analysts. Both are vague (e.g., “a balance between pro- and anti-apoptotic bcl-2 proteins appears to control the cell viability, and seems to correlate in long-term with the ability to form tumors”) and avoid clear predictions.

Figure 3

Fig. 3. The tools used by biologists and engineers to describe processes of interest: a) the biologist view of a radio. See Fig. 2 and text for description of the indicated components; b) the engineer view of a radio (please note that the circuit diagram presented is not that of the radio used in the study; the diagram of the radio was lost, which, in part, explains why the radio remains broken).

These description and communication tools are in a glaring contrast with the language that has been used by engineers (compare Figs. 3a and 3b). Because the language (Fig. 3b) is standard (the elements and their connections are described according to invariable rules), any engineer trained in electronics would unambiguously understand a diagram describing the radio or any other electronic device. As a consequence, engineers can discuss the radio using terms that are understood unambiguously by the parties involved. Moreover, the commonality of the language allows engineers to identify familiar patterns or modules (a trigger, an amplifier, etc.) in a diagram of an unfamiliar device. Because the language is quantitative (a description of the radio includes the key parameters of each component, such as the capacity of a capacitor, and not necessarily its color, shape or size) it is suitable for a quantitative analysis, including modeling.

I would like to argue that the absence of such language is the flaw of biological research that causes David's paradox. Indeed, even though the impotence of purely experimental approaches might be a bit exaggerated in my radio metaphor, it is common knowledge that the human brain can keep track of only so many variables. It is also common experience that once the number of components in a system reaches a certain threshold, understanding the system without formal analytical tools requires geniuses, who are so rare even outside biology. In engineering, the scarcity of geniuses is compensated, at least in part, by a formal language that successfully unites the efforts of many individuals, thus achieving a desired effect, be that design of a new aircraft or of a computer program. In biology, we use several arguments to convince ourselves that problems that require calculus can be solved with arithmetic if one tries hard enough and does another series of experiments.

One of these arguments postulates that the cell is too complex to use engineering approaches. I disagree with this argument for two reasons. First, the radio analogy suggests that an approach that is inefficient in analyzing a simple system is unlikely to be more useful if the system is more complex. Second, the complexity is a term that is inversely related to the degree of understanding. Indeed, the insides of even my simple radio would overwhelm an average biologist (this notion has been proven experimentally), but would be an open book to an engineer. The engineers seem to be undeterred by the complexity of the problems they face and solve them by systematically applying formal approaches that take advantage of the ever-expanding computer power. As a result, such complex systems as an aircraft can be designed and tested completely in silico, and computer-simulated characters in movies and video games can be made so eerily life-like. Perhaps, if the effort spent on formalizing description of biological processes would be close to that spent on designing video games, the cells would appear less complex and more accessible to therapeutic intervention.

A related argument is that engineering approaches are not applicable to cells because these little wonders are fundamentally different from objects studied by engineers. What is so special about cells is not usually specified, but it is implied that real biologists feel the difference. I consider this argument as a sign of what I call the urea syndrome because of the shock that the scientific community had two hundred years ago after learning that urea can be synthesized by a chemist from inorganic materials. It was assumed that organic chemicals could only be produced by a vital force present in living organisms. Perhaps, when we describe signal transduction pathways properly, we would realize that their similarity to the radio is not superficial. In fact, engineers already see deep similarities between the systems they design and live organisms [2].

Another argument is that we know too little to analyze cells in the way engineers analyze their systems. But, the question is whether we would be able to understand what we need to learn if we do not use a formal description. The biochemists would measure rates and concentrations to understand how biochemical processes work. A discrepancy between the measured and calculated values would indicate a missing link and lead to the discovery of a new enzyme, and a better understanding of the subject of investigation. Do we know what to measure to understand a signal transduction pathway? Are we even convinced that we need to measure something? As Sydney Brenner noted, it seems that biochemistry disappeared in the same year as communism [3]. I think that a formal description would make the need to measure system's parameters obvious and would help to understand what these parameters are.

An argument that is usually raised privately is why to bother with all these formal languages if one can make a living by continuing with purely experimental research that took years to learn. There are at least two reasons. One is that formal approaches would make our research more meaningful, more productive and might indeed lead to miracle drugs. A more immediate reason is that formal approaches may become a basic part of biology sooner than we, experimental biologists, expect. This transition may be as rapid as that from slides to PowerPoint presentations, a change that forced some graphics designers to learn how to use a computer and put others out of work.

Of course, a plea for a formal approach in biology is not new. The general systems theory, developed by Ludwig von Bertalanffy because of his fascination with the complexity of live organisms, was formulated 60 years ago, as well as his concept of organisms as physical systems [4]. Bertalanffy's fundamental studies have been followed by several attempts to approach cells as systems, the latest of which, system biology, has been rapidly developing into an active field [5-11]. Available computer power and advances in analysis of complex systems raise hope that this time the system approach will change from an esoteric tool that is considered useless by many experimental biologists, to a basic and indispensable approach of biology.

The question is how to facilitate this change, which is not exactly welcomed by many experimental biologists, to put it mildly [12]. Learning computer programming was greatly facilitated by BASIC, a language that was not very useful to solve complex problems, but was very efficient in making one comfortable with using a computer language and demonstrating its analytical power. Similarly, a simple language that experimental scientists can use to introduce themselves to formal descriptions of biological processes would be very helpful in overcoming a fear of long-forgotten mathematical symbols. Several such languages have been suggested [13, 14] but they are not quantitative, which limits their value. Others are designed with modeling in mind but are too new to judge as to whether they are user-friendly [15]. However, I hope that it is only a question of time before a user-friendly and flexible formal language will be taught to biology students, as it is taught to engineers, as a basic requirement for their future studies. My advice to experimental biologists is to be prepared.

Friday, September 08, 2006

Mathematics is a game of life

source http://www.hno.harvard.edu/gazette/2001/02.01/03-mathematics.html

Ahha, another guy in my dreaming research field. However, I think his path to computational biology is at a way of probability. It takes too much time for him to find the right target of his research and he lacks a broad understanding biology in perspective of others. However, his road is promising to me, proving the era of computational biology. --by Ryan

Jun Liu uses statistics to understand genes

By William J. Cromie
Gazette Staff

Jun Liu remembers being interested in mathematics as early as age 12. It was a hard interest to pursue in the waning years of the Cultural Revolution in China. Computers were not available to him. He didn't own a calculator. Mathematics books were difficult to find.

Jun's parents, both teachers, scrounged books wherever they could. They borrowed books from older professors who had hidden them away. His father, a professor at a technical university in Beijing, copied one book entirely by hand.

"I couldn't tell high school from college texts, so I read everything," Liu recalls. "Doing math was like a game you could play with only a piece of paper and pencil. On Sundays, I rode my bike for an hour to meet friends and do problems."

Sitting in his office at the Science Center, Liu, now 35, still shows a youthful enthusiasm for math problems. He wants to find answers to fundamental questions about genes and how they control life. "Every cell in your body contains a complete set of genes; each cell could become a part of your eye, your hand, or your brain. The question that challenges many scientists is how cells decide to be part of one organ or the other."

Liu thinks he may be able to get some of the answers with statistics more quickly than biologists can with experiments.

Commenting on Liu's recent tenure appointment, fellow professor of statistics Donald Rubin called him "a great asset to the University both as a teacher and a colleague. Harvard needs his strength in computational biology. In addition, he's a wonderful warm guy, soft-spoken but with a fabulous sense of humor."

Fooling around

Liu attended Beijing University, where he admits he spent most of his time playing bridge, and "fooling around." Nevertheless, after graduation in 1985, he was among a group of top math students who came to the United States on a scholarship program supported by the Society for Industrial and Applied Mathematics.

Sent to Rutgers University in New Jersey, Liu experienced cultural shock. "It was like a movie," he recalls. "It didn't feel real."

Language turned out to be the highest barrier. "Fortunately, I could get by with understanding formulas and equations," he says. "I got straight As without knowing much of what the teachers were saying. After a year, my English was worse than when I first came."

In 1986, Liu transferred to the University of Chicago. While there, he became interested in human rights issues and spent a lot of time participating in student protests. That activism brought his adviser to ask him if he wanted to be "a politician or a mathematician."

That was a turning point for Liu. He decided he had to apply himself to a career in statistics. "I didn't want to solve problems just because they hadn't been solved by others," he says. "I wanted to connect to reality. Although I didn't know exactly what statistics is, the field appealed to me for that reason."

"Liu has very original ideas, exceptional ability, and amazing computational skills," says Wing Wong, who supervised his Ph.D research in Chicago and who is now a professor of statistics at the Harvard School of Public Health. "He's also a clear thinker who communicates his ideas well, and is an easygoing, warm, and helpful person."

Tenure struggle

Liu earned a Ph.D. in statistics in 1991, then came to Harvard. During his first year, he met Wei Zhang, a graduate student studying immunology. They fell in love and married in 1994.

The same year, Stanford offered him a position in its statistics department. "I knew the chance of getting tenure at Stanford was nonexistent," Liu says. "But the chance of getting it at Harvard was one magnitude less than that, so I went to the West Coast."

His wife earned her Ph.D. in 1995, and followed him west. She got a job as a consultant in Los Angles and they commuted between apartments in the two different cities.

Despite his pessimism about tenure, Liu was offered the honor last year at both Stanford and Harvard.

The choice was not a difficult one for him. "I liked teaching undergraduate courses at Harvard; the students seemed more interested in the work than students at Stanford," he says.

Liu is also impressed by "the many great biologists and chemists here." Also of high interest to him is the new Bauer Center of Genomics Research where biologists, mathematicians, chemists, and others will work together to find the general principles that underlie life.

Liu now pursues the mystery of how genes are turned on and off. Using various statistical and computer techniques, he studies repetitive patterns in the DNA that lies between genes. This material contains instructions for regulating the expression of genes, and it is involved in whether the proteins produced by genes will become part of a brain or a big toe.

These on/off switches can be found by doing difficult, time-consuming experiments that require copying and mutating genes. If a region close to a gene is mutated and the gene stops producing a certain protein, that region must be part of a genetic switch. Liu believes he can locate such switches by statistical analysis of the genetic sequence patterns that occur between the actual genes.

Liu has done some of this work with collaborators at Harvard Medical School and the New York State Department of Health. For example, he has made about 2,000 predictions of where switches are located in the bacterium e-coli. In cases where these switches have actually been found by experiments, his predictions are 80 percent correct.

"What's nice (or not so nice) about this field," he says, "is that there's always a judgment day. You are making predictions of actual happenings, so you'll always know how right (or wrong) you are."

Liu is pleased with his results so far. But humans have at least 10 times more genes than bacteria, so he knows that things will get a lot more complicated.

That's OK with him. Talking with Liu in his office on a bright winter afternoon, it's easy to see that he is glad to be here. He misses playing bridge, fooling around, and traveling as a student on a few dollars a day. And his wife, now a consultant, makes more money than he does. But Liu has tenure now, and it's good to be back among attentive students and good colleagues. Best of all, he feels he can strengthen the contributions that Harvard makes to understanding what biological life is all about.