Sunday, June 03, 2007

写好英语科技论文的诀窍zz

写好英语科技论文的诀窍:
主动迎合读者期望,预先回答专家可能质疑

周耀旗
印地安那大学信息学院
印地安那大学医学院计算生物学和生物信息中心

以此文献给母校中国科技大学五十周年校庆

前言
我 的第一篇英语科技论文写作是把在科大的学士毕业论文翻译成英文。当我一九九零年从纽约州立大学博士毕业时,发表了20多篇英语论文。 但是,我对怎样写高 质量科技论文的理解仍旧处于初级阶段,仅知道尽量减少语法错误。之所以如此,是因为大多数时间我都欣然接受我的博士指导老师 Dr. George Stell和Dr. Harold Friedman的修改,而不知道为什么要那样改,也没有主动去问。这种情况一直持续到我去北 卡州立大学做博士后。我的博士后指导老师Dr. Carol Hall建议我到邻近的杜克大学去参加一个为期两天的写作短训班。这堂由Gopen教授主办 的短训班真使我茅塞顿开。第一次,我知道了读者在阅读中有他们的期望,要想写好科技论文,最有效的方法是要迎合他们的期望。这堂写作课帮我成功地完成了我 的第一个博士后基金申请,有机会进入哈佛大学Dr. Martin Karplus组。在哈佛大学的五年期间,在Karplus教授的指导下,我认识到一 篇好的论文需要从深度广度进行里里外外自我审查。目前,我自己当了教授,有了自己的科研组,也常常审稿。我觉得有必要让我的博士生和博士后学好写作。 我 不认为我自己是写作专家。我的论文也常常因为这样或那样的原因被退稿。但是我认为和大家共享我对写作的理解和我写作的经验教训,也许大家会少走一些我走过 的弯路。由于多年未用中文写作,请大家多多指正。来信请寄:
yqzhou@iupui.edu。 欢迎访问我的网站:http://sparks.informatics.iupui.edu

相关连接
1.周耀旗文章下载:http://csbl.bmb.uga.edu/~ffzhou/how-to-write.pdf
2.周耀旗参考的英文文章 The Science of Scientific Writing by G. D. Gopen and J. A. Swan, Scientific American, 78, 550-558, 1990. http://www.americanscientist.org/template/AssetDetail/assetid/23947?fulltext=true&print=yes

Wednesday, April 18, 2007

A Pro-Linux Cartoon

I see this picture at http://www.whylinuxisbetter.net/. Creative and Interesting!








Monday, March 12, 2007

Recommended Firefox Extensions

1. Tab Mix Plus

Tab Mix Plus enhances Firefox's tab browsing capabilities. It includes such features as duplicating tabs, controlling tab focus, tab clicking options, undo closed tabs and windows, plus much more. I strongly recommend the options of Select the tab pointed (Tool->Tab Mix Plus Option->Mouse->Mouse Gestures). It greatly decrease the amount of clicks when browsing.
2. Session Manager
Session Manager saves and restores the state of all windows - either when you want it or automatically at startup and after crashes. For example, if you start you work everyday by opening some webpages , you can store them in a session. Additionally it offers you to reopen (accidentally) closed windows and tabs. If you're afraid of losing data while browsing - this extension allows you to relax...
3. Firefox Showcase
If you habitually find yourself awash in open tabs (It seems that one of my classmates usually runs into such situation), clicking around looking for the page you need, Firefox Showcase will save you a lot of aggravation. Once you install the extension, you'll have a new Showcase submenu under the View menu. From here you can choose to show thumbnails of all tabs in the current window or all tabs in all windows. Firefox has lots of options and keyboard shortcuts, however I will never dive into those complex options. Simply click F12 to get the thumbnail view, uparrow and downarrow to select the intended tab and enter to see the full view the select tab. I can also exit the thumbnail view by clicking Esc. That's all. For more, see View->Showcase and for a sidebar view, follow View->Sidebar->Showcase Sidebar.
4. Download Statusbar
If you're tired with that sometimes-pesky Downloads window that pops up whenever you download a file in Firefox. Download Statusbar suppresses that window from popping up, and instead provides you the same information in the status bar at the bottom of the browser window. You can roll your mouse over the filename and get a pop-up tool tip with some extra information about your download, too
5. DownThemAll
DownThemAll is is download manager and accelerator. It lets you download all the links or images contained in a webpage and much more: you can refine your downloads by fully customizable criteria to get only what you really want! Simply, it saves you them time to open a shell and use wget.
6. Fotofox
Using Fotofox, you are able to grab a picture from any pages to the Fotofox sidebar, title it, tag it and upload them to your flickr album with a simple click. It is also compitable with other picture hosting site such as Tabblo, 23hq, Smugmug, Marela, and Kodak EasyShare Gallery.
7. Fasterfox
Exactly I don't know its performance.
8. Google Browser Sync
Google Browser Sync for Firefox is an extension that continuously synchronizes your browser settings – including bookmarks, history, persistent cookies, and saved passwords when you are using firefox on several computers. It also allows you to restore open tabs and windows across different machines and browser sessions. Alternatively You can choose to sync cookies, or not to sync cookies, but you can't make the decision based on individual cookies. Suppose that you are now using a public computer. At first you can install this extension, sign in with your google account. The Google Browser Sync begins to synchronize your firefox settings on the server and thus you get a familiar firefox in the public computer. When you are going to leave, choose Tools->Google Browser Sync->Stop Syncing and click Tools->Clear Private Data... (Ctrl+Shilt+Delete) to clean your personal leaved on the public computer.
9. Google Notebook
In my opinion Google Notebook is a simple but great product though most people ignorate it. This is possibly because that there is not a convinient on-one-click client interface and that the usability is poor in the webpage user interface. For example, I have been long hoping for the tag feature and sharing ability based on a note. Google Notebook it the client interface. It simple the process you make notes. Just select any part of the webpages, whether be it text or picture, click Note This from pop menu activated by right mouse button. I believe Google notebook will outrun Clipmarks or other similar services.
10. Firefox Google Bookmarks
Firefox Google Bookmarks (GBookmarks) creates a menu to access your google bookmarks from any computer. ( Your google bookmarks resides in your search history). Additionally it server a backup mechanism for all bookmarks. This extension may overlap with the Google Browser Sync. Personally, I consider the Bookmark in firefox a lightweight bookmarks that store only the everyday used links and use GBookmarks as a heavyweight repository of all links that may be useful some day.
11. StumbleUpon
StumbleUpon is also a bookmark service that incorporate social network elements. It resembles delicious. StumbleUpon lets you "channelsurf" the best-reviewed sites on the web. It is a collaborative surfing tool for browsing, reviewing and sharing great sites with like-minded people. This helps you find interesting webpages you wouldn't think to search for. You can also share pages of interest within a community. Particularly, it will insert into your google search results which shows you other people's rating and reviews of the search result.
12. Greasemonkey
Greasemonkey basically allows you to add JavaScript to any Web page, which implies infinite control over the behavior of web pages. Greasemonkey is not for the faint of heart. The good news is that there are many generous souls out there who share the scripts they create. Check out userscripts.org for a script repository. If you want to write your own scripts, try diveintogreasemonkey.org or pick up Mark Pilgrim's Greasemonkey Hacks from O'Reilly Media. Personally I strongly recommend Gmail Macros. After installation, it empower gmail with striking and convinient keyboard shortcuts like Google Reader.
13. Adblock
Adblock is a content filtering plug-in for the Mozilla and Firefox browsers. It allows the user to specify filters, which remove unwanted content based on the source-address. Adblock supports two types of filters: simple, and Regular Expression. Adblock will also provide a default list of filters, which is enough for a lazy person like me.
14. ScrapBook
ScrapBook is a Firefox extension, which helps you to save Web pages and easily manage collections. This enable you to surf web pages off line. You can also directly copy the ScrapBook folder from your firefox option folder to other computers and see what you saved there.
References:
20 must-have Firefox extensions
Firefox Add-ons Recommended Add-ons
千种风情千种树: My Firefox Extensions

Friday, February 09, 2007

Alon Halevy and Peter Norvig


Alon Halevy and Peter Norvig, two Googlers, have been selected for the 2006 class of ACM Fellows.

Peter, who was Google's first director of search quality and is currently director of Google Research, has been recognized for his many contributions to the disciplines of artificial intelligence and information retrieval. His personal websites is http://norvig.com/.

Alon, who leads one of our structured data initiatives, has been honored for his contributions in data integration and knowledge representation. His old personal website: http://www.cs.washington.edu/homes/alon/, new personal website: http://alonhalevy.googlepages.com/, his blog is http://www.alonhalevy.blogspot.com/

Saturday, February 03, 2007

A Elegent Blog: Designer's Block


The blog Designer's Block is elegant blog. The writer is from UK. I really like the design (see left), grace and mysterious style. I will choose it as my blog's background.








update: In fact, the painter of above paintings are Melissa Mossart. In her website there are also many paintings of similar style.





































See more paintings at http://www.melissamossart.com/paint.htm
Melissa Mossart and nine other women artists' gallery: http://www.tenwomen.org/venicegallery.html

Thursday, February 01, 2007

Research Projects on Microarray Analysis


BioConductor


Bioconductor is an open source and open development software project for the analysis and comprehension of genomic data.

Bioconductor is primarily based on the R programming language but we do accept contributions in any programming language. Although initial efforts focused primarily on DNA microarray data analysis, many of the software tools are general and can be used broadly for the analysis of genomic data, such as SAGE, sequence, or SNP data.

The broad goals of the projects are to

  • provide access to a wide range of powerful statistical and graphical methods for the analysis of genomic data;
  • facilitate the integration of biological metadata in the analysis of experimental data: e.g. literature data from PubMed, annotation data from LocusLink;
  • allow the rapid development of extensible, scalable, and interoperable software;
  • promote high-quality documentation and reproducible research;
  • provide training in computational and statistical methods for the analysis of genomic data.
If you are new to Bioconductor you might consider buying Bioinformatics and Computational Biology Solutions Using R and Bioconductor


Gene Ontology


The Gene Ontology (GO) project is a collaborative effort to address the need for consistent descriptions of gene products in different databases. The GO project has developed three structured controlled vocabularies (ontologies) that describe gene products in terms of their associated biological processes, cellular components and molecular functions in a species-independent manner. There are three separate aspects to this effort: first, the development and maintenance of the ontologies themselves; second, the annotation of gene products, which entails making associations between the ontologies and the genes and gene products in the collaborating databases; and third, development of tools that facilitate the creation, maintenance and use of ontologies.

The use of GO terms by collaborating databases facilitates uniform queries across them. The controlled vocabularies are structured so that they can be queried at different levels: for example, you can use GO to find all the gene products in the mouse genome that are involved in signal transduction, or you can zoom in on all the receptor tyrosine kinases. This structure also allows annotators to assign properties to genes or gene products at different levels, depending on the depth of knowledge about that entity.

International HapMap Project


The HapMap is a catalog of common genetic variants that occur in human beings. It describes what these variants are, where they occur in our DNA, and how they are distributed among people within populations and among populations in different parts of the world. The International HapMap Project is not using the information in the HapMap to establish connections between particular genetic variants and diseases. Rather, the Project is designed to provide information that other researchers can use to link genetic variants to the risk for specific illnesses, which will lead to new methods of preventing, diagnosing, and treating diseases.

Microarray Gene Expression Data Society - MGED Society

The Microarray Gene Expression Data (MGED) Society is an international organisation of biologists, computer scientists, and data analysts that aims to facilitate the sharing of microarray data generated by functional genomics and proteomics experiments. The current focus is on establishing standards for microarray data annotation and exchange, facilitating the creation of microarray databases and related software implementing these standards, and promoting the sharing of high quality, well annotated data within the life sciences community. A long-term goal for the future is to extend the mission to other functional genomics and proteomics high throughput technologies


The MolTools consortium



The MolTools consortium started on January 1st 2004, as a joint research programme bringing together 12 leading European academic groups, four biotech SMEs and one US laboratory working in the area of postgenomic technology development. The partners have pioneered a series of important molecular techniques and will now work together to establish next-generation tools for molecular analysis. Its scientific aims are to establish genome analysis technologies set to monitor extensive molecular repertoires, and with the capacity to investigate even single molecules. Its current research projects include:

Tuesday, January 30, 2007

Resources of Genomics and Microarray Analysis

Genomics and Microarrays:
Stanford Microarray Database:
storing lots of raw and normalized data from microarray experiments

Genomics tutorial at Genome Canada
http://www.genomecanada.ca/xpublic/dnaBasics/index.asp?l=e

Introductions to microarray at NCBI
http://www.ncbi.nlm.nih.gov/About/primer/microarrays.html

Microarray (movie)
http://www.broad.harvard.edu/chembio/lab_schreiber/anims/videos/microarray.html

other resourses for microarray
http://www.learner.org/channel/courses/biology/units/genom/images.html

image analysis for microarray
http://www.maths.usyd.edu.au/u/jeany/ (publication)
http://www.stat.berkeley.edu/users/terry/zarray/Talks/image/jpegindex.html http://cmm.ensmp.fr/~angulo/research/dnamicro.htm

Bioconductor
The richest source of freely available packages for genomic data analysis

Nature article
the perspective of biologists facing heaps of noisy genomic data including their urgent need for better methods and computationally and statistically skilled support.

dChip Software
http://biosun1.harvard.edu/complab/dchip/

dChip Software: Analysis and visualization of gene expression and SNP microarrays



Biology:

Retroviruses
http://www.whfreeman.com/kuby/content/anm/kb03an01.htm (FLASH)

human genome project(movies)
http://www.genome.gov/Pages/EducationKit/download.html

the central dogma of molecular biology (wonderful movie):
http://www.genome.gov/Pages/EducationKit/video/qt/3D.mov

EBM & Clinical Research Workstation menu
http://www.shdem.com/ebm/default.asp

Biochemistry & Epidemiology useful link:
http://www.med-ed.virginia.edu/menu/otherMedEd.cfm
 
Statistics:

online textbook for statistics
http://www.stat.berkeley.edu/~stark/SticiGui/Text/toc.htm

Terry Speed's Microarray homepage
statistical challenges related to microarray data
new webpage: http://www.stat.berkeley.edu/~terry/Group/home.html

Statistics
http://www.bettycjung.net/statsiteS.htm

Computing technology

Introduction to R for biologists (by Natalie Roberts, WEHI, Melbourne)
R manuals under link "Manuals" (left column)
manuals: http://cran.r-project.org/manuals.html
R tutorial: http://www.personality-project.org/r/
R package: Statistics for Microarray Analysis

Directionary of Blogs About Microarray Analysis: Draft

Biodefense Bioinformatics Blog http://ai59694.blogspot.com/
Rotten bananas - http://heathermaughan.blogspot.com/index.html
Synthetic Biology and Gene Synthesis - http://syntheticbio.blogspot.com
Genomics Online - http://genomics-info.blogspot.com/index.html
formerscienceguy - http://formerscienceguy.blogspot.com/index.html

Friday, January 26, 2007

Alan Perelson

Dr. Perelson received his B.S. degrees in Life Science and Electrical Engineering from MIT in 1967, and a Ph.D. in Biophysics, under the supervision of Aharon Katchalsky-Katzir, from UC Berkeley in 1972. He was Acting Assistant Professor, Division of Medical Physics, Berkeley, in 1973 and a postdoctoral fellow at the Department of Chemical Engineering, University of Minnesota, in 1974. He was a staff member in the Theoretical Biology and Biophysics Group at Los Alamos National Laboratory from 1974 - 1991, a Laboratory Fellow from 1991 - 2002, head of the Theoretical Biology and Biophysics Group between 1995 - 2001, and is currently a Los Alamos National Laboratory Senior Fellow. He spent the 1978 and 1979 academic years at Brown University as an Assistant Professor of Medical Sciences in the Division of Biology and Medicine and the Lefschetz Center for Dynamical Systems, was a visiting scientist at the Mathematical Institute, Oxford University in 1986 and a visiting professor of Physics at Ecole Normale Superieure, Paris in 1990, and the University of Paris VII in 1992. He is also a member of the Science Board and head of the Theoretical Immunology Program at the Santa Fe Institute. He is also an adjunct professor of Bioinformatics at Boston University and an adjunct professor of biology at the University of New Mexico.

Research Interests

Mathematical and theoretical biology, with an emphasis on problems in immunology, virology,
and cell and molecular biology.

Time Zone of United States

PST: Washington, Oregon, Neveda, California

MT: Montana, Wyoming, Idaho, Utah, Colorado, Arizona, New Mexico and parts of
North Dakota, South Dakota and Nebraska

CT: Parts of North Dakota, South Dakota and Nebraska, Kansas, Oklahoma, Texas,
Minnesota, Iowa, Missouri, Arkansas, Louisiana Wisconsin, Illinois,
Tennessee, Mississippi, Alabama

EST: Michigan, Indiana, Ohio, Kentucky, Georgia, New York, Pennsylvania, West
Virginia, Virginia, North Carolina, South Carolina, Florida, Washington DC
New Jersey, Connecticut, Ehode Island, Massachusetts, New Hampshire,
Vermont, Maine

Thursday, January 25, 2007

Friday, January 19, 2007

A Haplotype Map of the Human Genome

A Haplotype Map of the Human Genome

David Altshuler
Harvard Medical School, Massachusetts General Hospital, Whitehead Institute

Eric Lander
Whitehead Institute and MIT

Goal

The next key step of the Human Genome Project (HGP) (following the creation of the genetic, physical, sequence and SNP maps) is the generation of a "haplotype" map of the human genome. Such a "haplotype" map consists of a high density of SNPs defining the small number of ancestral haplotypes (blocks of tightly correlated genetic variants) in each region of the human genome. Knowledge of these haplotypes will allow comprehensive and efficient testing of the association of human genes with human diseases. The haplotype map can and should be generated rapidly and should be made freely available to researchers worldwide.

Background

A haplotype map of the human genome has become both justified and practical due to significant advances over the last two years.

Specifically, these advances include:

  • Genomic Sequence: The development of a complete genome sequence - integrated with human genes and annotations - providing a reference framework on which to layer knowledge about allelic variation.

  • Genetic Variants: The development of a dense (and rapidly growing) map of 1.4 million human SNPs provides a genome-wide resource of genetic variation adequate to uniquely tag the vast majority of human haplotypes.

  • Genotyping Technology: The development of high-throughput methods, allowing a rapid, efficient and cost-effective experimental approach to a project of the required scale.

  • Long-range LD: The discovery that human SNPs display strong linkage disequilibrium (LD or allelic association) over large distances. LD is detectable over distances in the range of 100kb and is extremely strong over regions spanning several tens of kb (the size of typical genes). For such regions, the vast majority of chromosomes in the population carry one of a handful of highly conserved haplotypes. As a result, genetic diversity in the region can be represented by a small number of well-chosen SNPs.

Impact on biomedical research

The availability of a haplotype map of the human genome will have a substantial impact on human genetic studies.

Specifically, these studies include:

  • Comprehensive association studies of individual genes. The association of genes with disease has traditionally been probed by testing individuals SNPs one-at-a-time. The drawback to this approach is that the task is never-ending: one can exclude particular SNPs as playing a role, but one cannot exclude a gene. Once the haplotype structure of the genome is defined, one can (1) comprehensively test all significant haplotypes in the gene, and (2) decrease the number of SNPs needed by selecting a subset that defines the population variability. This will allow haplotype studies of individual genomic loci in an unbiased manner, without assumption about the locations of causal mutations in coding regions, promoters or regulatory sites at significant distance away. And, it will greatly decrease the technical and financial barriers faced by laboratories in undertaking such work

  • Genome-wide association studies. A genome-wide haplotype map will make possible whole-genome scans for association in the population. Rather than focusing only on 'candidate' genes, it will become possible to search the genome in an unbiased manner for genes whose common variation contributes to disease in the population. Routine use of genome-wide association studies will also require further decreases in genotyping costs, but such decreases are likely to be driven by the development of the haplotype map.

  • Human population structure and history. Knowledge of haplotypes will transform our understanding of human population structure and history. The LD pattern turns out to be an extremely sensitive indicator of population history, because the multi-allelic nature of haplotypes provides rich detail and because the breakdown of haplotypes follows a predictable clock set by recombination rates. In particular, LD patterns are more powerful than traditional studies of allele frequencies per se. Information about human population history is interesting in its own right, but is also very valuable in the design of medical studies (such as admixture mapping).

Technical Issues

Generating a haplotype map would involve the following components:

  • Population Samples. Development of appropriate population samples, consisting of parent-offspring trios (to allow inference of haplotypes). We estimate that a total of about 300 samples will be needed, representing major ethnic groups in a manner appropriate for generating a map that can be used for medical studies in all populations. The population samples should be a renewable resource (i.e., immortalized cell lines).

  • Sample and Data Availability. The samples should be made freely available so that any interested scientific group can contribute data (in the manner of the CEPH panel and the DNA Polymorphism Discovery Resource). Conversely, all data generated by the project should be immediately released into the public domain without restrictions of any kind.

  • Numbers of SNPs to be genotyped. It is estimated that generating the haplotype map will require successful genotyping of 450,000 SNPs, which will in turn require initial testing of some 800,000 to 900,000 SNPs. The required scale is now well within reach: the Whitehead and Sanger Centre are each currently engaged in pilot projects involving 25,000 SNPs using automated genotyping setup and MALDI-TOF-based detection. Given the required scale and efficiencies, it is likely that the bulk of the work should be performed by a few large groups, but all groups should be encouraged to participate in the project by analyzing genes and regions of interest.

  • Analytical Tools. The project will require various analytical tools to readily define haplotype blocks from genotype data, software systems to aid in the hierarchical selection of SNPs to fill in blocks, and databases to make the information maximally useful to the community. Prototype systems have been developed, but focused effort will be needed to develop mature systems.

Wednesday, December 20, 2006

Chris Sander


Chris Sander
Chris Sander

Trained as a theoretical physicist, Chris Sander, Director of Memorial Sloan-Kettering Cancer Center's Computational Biology Center and Chairman of the Sloan-Kettering Institute's Computational Biology Program, knew early in his career he wanted to do more than compute the behavior of elementary particles. After a bold move from physics to biology, he helped to develop the field of computational biology, which aims to use mathematical algorithms and information systems to simulate the behavior of molecules, cells, and organisms, using these simulations to make useful diagnostic and therapeutic predictions.

While an undergraduate at the University of Berlin, I was intensely engaged in the study of theoretical physics and mathematics. Yet I felt that analyzing life, the living system on this planet, was a more fascinating and challenging scientific problem than studying the world of elementary particles. To explore this potential change in career direction, I sought the advice of Max Delbrück at the California Institute of Technology, a physicist originally from Berlin who was one of the founders of molecular genetics.

After visiting friends in Texas, I hopped on a Greyhound bus to Los Angeles, made my way to Caltech, found Dr. Delbrück's office, knocked on the door, and with a dash of chutzpah, asked the Nobel laureate if he had a moment to chat with an aspiring graduate student. Dr. Delbrück described the areas of theoretical physics that might be relevant to biology in the future -- information that was in the forefront of my mind as I entered the graduate physics program at the University of California, at Berkeley, in 1967.

Wanting to move from the theoretical physics of my PhD thesis to the theoretical biology I had dreamt of, I made my second pilgrimage, this time to see Manfred Eigen, a Nobel Prize-winning chemist who was studying biological evolution in Göttingen, Germany. Dr. Eigen surprised me by explaining that the field barely existed, but he did point me to three mathematical biology research problems: a theory of the immune response, neuronal mapping in the brain, and protein folding. So I packed my bags and I moved from the University of Heidelberg to the Weizmann Institute of Science, in Israel, where I began work with Shneior Lifson on the prediction of three-dimensional protein structures.

Enter the third motivator of my career: the first completely sequenced genome -- no, not in 2000, but in 1977! In that year, I saw an amazing paper in the journal Nature from Fred Sanger's group in Cambridge, United Kingdom, which included two entire pages filled with 5,375 letters, all As, Ts, Gs and Cs, representing the genetic blueprint of a small virus. I walked down the hall to ask my friend Georg Schulz, "With this kind of cryptic information coming from genomes, won't biology need computational science to decipher it?" His answer was yes, and I spent the next 23 years of my professional life preparing for the day, in the year 2001, when the 3.5 billion letters of the human genome finally became available. In the process, I helped to develop the field now known as computational biology.

The real value of computational science, when applied to any system, is to predict what's going to happen next. Weather forecasting is an example. There's an enormous amount of data collected about the weather, but the data, by themselves, are unintelligible. What's required is the application of the appropriate mathematical equations embodied in a software system, which, using the data, allows one to compute tomorrow's weather.

Applied to cancer biology, we want to be able to predict, for example, if a cancer will go from a nonaggressive to an aggressive form, or more importantly, to predict accurately the consequences of possible therapeutic interventions. The goal is to have an impact on human disease, and to do this you have to work in collaboration with physicians. In 2002, Harold Varmus presented his vision of Memorial Sloan-Kettering as the perfect environment for this -- a place with open doors between basic and clinical research, where close collaboration is encouraged.

My first action at Memorial Sloan-Kettering Cancer Center was to start the Computational Biology Center (CBC) and its Bioinformatics Core Facility. The CBC's researchers and engineers are devoted both to basic science and to the goal of developing diagnostic and therapeutic tools that help improve the lives of people affected by cancer. We often collaborate with researchers in the lab and in the clinic to translate data -- data such as the molecular profiles of cells and tissues, the billions of letters of genome sequences, and the functions and structures of key genes -- into biological insights and prediction tools. And the Bioinformatics Core, ably led by Alex E. Lash, provides internal bioinformatics training, collaboration, and infrastructure support.

One concrete example of the practical uses of computational biology is the work we have been doing with Howard I. Scher, Chief of Memorial Sloan-Kettering Cancer Center's Genitourinary Oncology Service, and Francis M. Sirotnak, Member Emeritus and Head of Sloan-Kettering Institute's Laboratory of Molecular Therapeutics. The idea is that cancer cells, like any system that recovers from a round of major damage, might be especially sensitive after a first round of therapy. With this concept in mind, we are aiming to prevent the development of aggressive prostate cancer by looking at the molecular profile of prostate cancer cells after androgen removal in mice, using DNA chips provided by our Genomics Core Laboratory. We use computer software to find needles in a haystack -- the perhaps tens of genes, out of tens of thousands, that may be a characteristic signature of how prostate cancer reacts to such therapy. We hope this will lead us to an Achilles' heel to target to avoid recurrence. It's a long-term effort but the idea is to arrive computationally at the best therapeutic intervention.

Overall, what's been most rewarding for me during my short time here is the opportunity not just to predict the behavior of biological systems, but hopefully to help improve the quality of people's lives. The dream that started with Max Delbrück's advice is now within reach.

Saturday, December 09, 2006

Super Computing

 Logo
The TOP500 project was started in 1993 to provide a reliable basis for tracking and detecting trends in high-performance computing. Twice a year, a list of the sites operating the 500 most powerful computer systems is assembled and released. The best performance on the Linpack benchmark is used as performance measure for ranking the computer systems. The list contains a variety of information including the system specifications and its major application areas.


NCSA Home
The National Center for Supercomputing Applications (NCSA), one of the five original centers in the National Science Foundation's Supercomputer Centers Program, opened its doors in January 1986. Since then, NCSA has contributed significantly to the birth and growth of the worldwide cyberinfrastructure for science and engineering, operating some of the world's most powerful supercomputers and developing the software infrastructure needed to efficiently use these systems (for example, NCSA Telnet and, in 1993, NCSA Mosaic™, the first readily available graphical Web browser). Today the center is recognized as an international leader in deploying robust high-performance computing resources and in working with research communities to develop new computing and software technologies

Blue Gene




Blue Gene is an IBM Research project dedicated to exploring the
frontiers in supercomputing: in computer architecture, in the software required to program and control massively parallel systems, and in the use of computation to advance our understanding of important biological processes such as protein folding.

The full Blue Gene/L machine was designed and built in collaboration with the Department of Energy's NNSA/Lawrence Livermore National Laboratory in California, and has a peak speed of 360 Teraflops. Blue Gene systems occupy the #1 (Blue Gene/L) and #2 (Blue Gene Watson) positions in the TOP500 supercomputer list announced in November 2005, as well as 17 more of the top 100.

IBM now offers a Blue Gene Solution. IBM and its collaborators are currently exploring a growing list of applications including hydrodynamics, quantum chemistry, molecular dynamics, climate modeling and financial modeling.

SDSC - San Diego Super Computer Center

Founded in 1985, the San Diego Supercomputer Center (SDSC) enables international science and engineering discoveries through advances in computational science and high performance computing. Continuing this legacy into the era of cyberinfrastructure, SDSC is a strategic resource to science, industry and academia, offering leadership in the areas of data management, grid computing, bioinformatics, geoinformatics, high-end computing as well as other science and engineering disciplines. The mission of SDSC is to extend the reach of scientific accomplishments by providing tools such as high-performance hardware technologies, integrative software technologies and deep inter-disciplinary expertise, to the community.

SDSC was founded with a $170 million grant from the National Science Foundation's (NSF) Supercomputer Centers program. From 1997 to 2004, SDSC extended its leadership in computational science and engineering to form the National Partnership for Advanced Computational Infrastructure (NPACI), teaming with approximately 40 university partners around the country. Today, SDSC is an organized research unit of the University of California, San Diego primarily funded by NSF with a staff of talented scientists, software developers and support personnel.





The National Resource for Biomedical Supercomputing (NRBSC) pursues leading edge research in high performance computing and the life sciences, and fosters exchange between PSC expertise in computational science and biomedical researchers nationwide.

Our focus is two-fold: computational biomedical research and outreach to the national biomedical research community through education and publications.

Research at NRBSC is centered in three areas: microphysiology; volumetric visualization and analysis; and computational structural biology.

NRBSC's education arm includes not only user training, but also software distribution, publications, and other outreach activities such as online courses and workshop webcasts.

The National Resource for Biomedical Supercomputing, formerly the Biomedical Initiative, was established at the Pittsburgh Supercomputing Center in 1987 as the first extramural biomedical supercomputing program in the country funded by the National Institutes of Health.

Tuesday, December 05, 2006

Python Resources Collection

Website



Python® is a dynamic object-oriented programming language that can be used for many kinds of software development. It offers strong support for integration with other languages and tools, comes with extensive standard libraries, and can be learned in a few days. Many Python programmers report substantial productivity gains and feel the language encourages the development of higher quality, more maintainable code.

Jython
Python is an implementation of the high-level, dynamic, object-oriented language Python written in 100% Pure Java, and seamlessly integrated with the Java platform. It thus allows you to run Python on any Java platform.




Stored in these dark caverns you may find rich veins of Python code, collected caches of Python information, and all manner of sundry Python passageways to explore. With candle or torch in hand, good hunting this night to all.Those not familiar with Python perhaps might start your quest at a brighter place.

NumPy

The fundamental package needed for scientific computing with Python is called NumPy. This package contains: a powerful N-dimensional array objectsophisticated (broadcasting) functionsbasic linear algebra functions basic Fourier transformssophisticated random number capabilitiestools for integrating Fortran code.

SciPy.org
SciPy (pronounced "Sigh Pie") is open-source software for mathematics, science, and engineering. It is also the name of a very popular conference on scientific programming with Python. The core library is NumPy which provides convenient and fast N-dimensional array manipulation. The SciPy library is built to work with NumPy arrays, and provides many user-friendly and efficient numerical routines such as routines for numerical integration and optimization. Together, they run on all popular operating systems, are quick to install, and are free of charge. NumPy and SciPy are easy to use, but powerful enough to be depended upon by some of the world's leading scientists and engineers. If you need to manipulate numbers on a computer and display or publish the results, give SciPy a try!


DISLIN Homepage
DISLIN is a high-level plotting library for displaying data as curves, polar plots, bar graphs, pie charts, 3D-color plots, surfaces, contours and maps.


wxPython is a GUI toolkit for the Python programming language. It allows Python programmers to create programs with a robust, highly functional graphical user interface, simply and easily.


VPython is a package that includes: the Python programming languagethe IDLE interactive development environment "Visual", a Python module that offers real-time 3D output, and is easily usable by novice programmers"Numeric", a Python module for fast processing of arrays

PyOpenGL Logo

PyOpenGL is the cross platform Python binding to OpenGL and related APIs. The binding is created using the SWIG wrapper generator, and is provided under an extremely liberal BSD-style Open-Source license.


The Biopython Project is an international association of developers of freely available Python tools for computational molecular biology.It is a distributed collaborative effort to develop Python libraries and applications which address the needs of current and future work in bioinformatics. The source code is made available under the Biopython License, which is extremely liberal and compatible with almost every license in the world. We work along with the Open Bioinformatics Foundation, who generously provide web and CVS space for the project

PyZine
The journal of Python Language



PyLucene is a GCJ-compiled version of Java Lucene integrated with Python. Its goal is to allow you to use Lucene's text indexing and searching capabilities from Python. It is designed to be API compatible with the latest version of Java Lucene.

Courses with an emphasis on scientific computing

Python course in Bioinformatics
Introduction to Python and Biopython with biological examples.



Monday, December 04, 2006

LaTeX Resources Collections

To start with Latex, the most convenient way is to read a short book with a strange name( sorry, I can not recall it now).
And Professor Schneider provides a page to introduce Latex for Biologist, here it is
LaTeX Style and BiBTeX Bibliography Formats for Biologists: TeX and LaTeX Resources

Monday, November 27, 2006

The International Society for Computational Biology (ISCB)


The International Society for Computational Biology (ISCB) is incorporated in the United States as a 501(c)(3) non-profit corporation, and registered in the state of California as a Charitable Trust. Now hosted at the San Diego Supercomputer Center at University of California, San Diego, the Society was officially formed in 1997 as an outgrowth of the conference on Intelligent Systems for Molecular Biology (ISMB). From humble beginnings, both ISCB's membership and ISMB's annual attendance have kept pace with the overall explosive growth experienced in the field of bioinformatics/computational biology.

For the complete story from conception to present day please visit the History link below. Additional links provide copies of documents that detail the legal structure of ISCB, which may prove interesting to our current and prospective members, as well as be of some use to regional groups around the world wanting to form national or regional societies and not knowing how to begin. And finally, ISCB's mission, vision and values are detailed in the 2003 Strategic Plan, and activities over the years are documented in the Newsletter Archives. We encourage you to peruse both of these links for a perspective on where we've been and where we may be going in the years ahead.

Ryan Songer

Sunday, November 26, 2006

ThinkWiki: Install Linux on ThinkPad

kkk recommended this site.

ThinkWiki

From ThinkWiki

Jump to: navigation, search

This is ThinkWiki, the Wiki Web for ThinkPad users. Here you find anything you need to install your favourite Linux distribution on your ThinkPad. Windows users shouldn't run away, there's a lot of useful information for them as well.

Please support us and help to extend this wiki. Thank you!



Ryan Songer

Wednesday, October 11, 2006

Zooomr and Its Free Pro Account

bird01bird01 Hosted on Zooomr

Frankly speaking, I have been a flickr fan for a long time thought it has a monthly upload limit, because I didn't expect I would use up the 50M limits(about 30+ pics) per month. However, I was frustrated when I returned from a journey with more than 50+ pics. This experience forced me to find a new photo hosting site. For a poor students like me, it should be free, free of upload limit, large enough, stable and comfortable. Under this criteria, picasa web album is poorly 250MB and still too simple. Fotkit is not free and has a messy UI. ... I have been looking for it until I came across a blog post Do We Love Bloggers? Yes We Do!
I signed up and write this post in hope a free pro account.

Thanks Kris. Yes, now I get pro account after some errors occurred though. Its seems that Zooomr is not reliably stable at this time, but I would like to try it.