O'Really?

August 4, 2007

Scifoo day 1: Turn up, tune in, drop out

Filed under: google — Duncan Hull @ 9:38 pm
Tags: , , , , , ,

Scifoo campersMy boss, Douglas Kell, who has kindly allowed and paid for me to attend Science Foo Camp (scifoo), says to me “tell me what you get up to”. So here goes. Scifoo day 1, A chance to meet and around 250 engineers, scientists, philosophers and other odd people from all over the world.

Shortly after arriving at the Googleplex, California and being fed by gourmet chefs, it all starts . There is a quick round of introductions from everyone in the room, the conference schedule gets put up on a big board, and interactively edited like a wiki. Sounds chaotic, but it actually works.

The introductions are followed by some lightning talks by selected people, chaired by Tim O’Reilly and Timo Hannay.

  1. Drew Endy from OpenWetWare talked about biotechnology. He drew analogies between civil engineering and bio-engineering. Today we can build wonderful bridges like Viaduc Millau in France. But it hasn’t always been that way. In the stone age, we used rocks as they were to build the likes of Stone Henge. Then we moved to to quarrying rock more systematically, so we can build simple bridges. For biotechnology to succeed in the same way as civil engineering, we need to synthesize DNA in the same way as we synthesis concrete to make bridges. But currently, biotechnology is still in its stone age.
  2. Charles Simonyi gave a talk about his recent trip as a Space tourist. I’ve never met an astronaut before, and never wondered what it smells like or what the quality of your sleep is like in space. You can find out more about Charles in Space</.
  3. Felice Frankel: Visualisation, visualisation, visualisation! (although she doesn’t like that word)

After all this, theres some time for “corridor conversations” with other delegates, which is where most of the interesting stuff goes on. Its difficult to pull out a narrative, because theres all kinds of people here: some people I managed to speak to (note form, sorry!):

In his introduction, Tim O’Reilly described scifoo as “making new synapses in the global brain”. You take a load of people from different disciplines, stick them together, and they find all sorts of interesting connections that they might not otherwise have found. It might sound pretentious, but I think its true. Unlike larger conferences, scifoo is small and intimate enough to be able to talk to lots of different people which is one thing that makes it special. This year, they’ve lifted the blogging ban, so everything is public unless stated otherwise. Which means you’ll be hearing lots more about it from bloggers like me at the conference.

Day two will be fun, theres lots of demos, and more people to meet: Martin Rees, how do we survive the twenty first century given that we’re all going to die?…Must try and pluck up the courage to talk to Sergey but I’m completely starstruck. Brian Cox, Hello, I’ve seen you on the telly…Esther “always make new mistakes” Dyson, Anne Wojcicki, George Church, Eric Lander, Paul Z. Myers Theres a tonne of bio-people here….So many people, so little time!

[this post originally published on nodalpoint]

Creative Commons License

This work is licensed under a

Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.

May 31, 2007

Google Metabolic Maps

Google in the Palm of my HandThese days, new Google products and code seem to appear on a weekly basis. Take, for example, Google Gears which takes advantage of SQLite, mentioned on nodalpoint recently. They certainly don’t hang about at the Googleplex in Mountain View, California. Wouldn’t it be great if Google applied some of that engineering expertise and agility to science and bioinformatics? Just imagine: we could have Google Metabolic Maps, a virtual globe of the cell for scientists everywhere…

Scientists have been drawing metabolic maps for a very long time, but unfortunately when it comes to charting and understanding metabolic pathways, we’re still at the “here be dragons” stage of bio-cartography. I’m obviously not the first person to dream of this, but imagine maps of metabolic pathways looked more like Google Earth or Google Maps, than the old fashioned style maps, many life scientists will be familiar with. Now imagine just a little more, that these maps weren’t just available on conventional screens, but we’re given the Minority Report treatment, courtesy of Mr Bill Gates and his wizzy surface magic at Microsoft. Wouldn’t that be great? Metabolic maps on an interactive tabletop computer. Just like Tom Cruise in the movies, we’d be able to effortlessly swish around metabolism (or the metabolome / proteome / genome / [insert-your-favourite]ome). Imagine if it was all open-source too, no boundaries, no passports…

Now, you may say that I’m a dreamer, but I’m not the only one [1,2,3].

References

  1. Zhenjun Hu, Joe Mellor, Jie Wu, Minoru Kanehisa, Joshua M. Stuart and Charles DeLisi (2007) Towards zoomable multidimensional maps of the cell Nature biotechnology 25 (5), 547-54. DOI:10.1038/nbt1304
  2. Hiroaki Kitano, Akira Funahashi, Yukiko Matuoka and Kanae Oda (2005) Using process diagrams for the graphical representation of biological networks Nature biotechnology 23 (8), 961-6. DOI:10.1038/nbt1111
  3. John Lennon and Yoko Ono (1971) Imagine
  4. this post originally published on nodalpoint with comments

Creative Commons License

This work is licensed under a

Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.


April 13, 2007

Collaboration, collaboration, collaboration!

Geldof Blair collaborationWhat should your three main priorities be as a Scientist? Collaboration, collaboration, collaboration. Quentin Vicens and Phil Bourne have just published Ten Simple Rules for a Successful Collaboration [1] to help you do just that, as part of a continuing series [2,3,4,5].

Tony Bliar once said “Ask me my three main priorities for government, and I tell you: education, education, education.” In Science, its not so much about education as collaboration, collaboration, collaboration. The advice in Ten Simple Rules is all useful stuff, but what caught my eye is the fact that collaboration is on the rise, at least according to the number of co-authors on papers published in PNAS. The average number of co-authors has risen from 3.9 in 1981 to 8.4 in 2001. So before you publish or perish, it seems likely that you’ll also need to collaborate or commiserate… less laboratory, more collaboratory!

Photo credit Garret Keogh

References

  1. Quentin Vicens and Phillip Bourne (2007) Ten Simple Rules for a Successful Collaboration PLOS Computational Biology
  2. Phillip Bourne (2006) Ten Simple Rules for Getting Published PLOS Computational Biology
  3. Philip Bourne and Iddo Friedberg (2006) Ten Simple Rules for Selecting a Postdoctoral Position PLOS Computational Biology
  4. Phillip Bourne and Leo Chalupa (2006) Ten Simple Rules for Getting Grants PLOS Computational Biology
  5. Phillip Bourne and Alon Korngreen (2006) Ten Simple Rules for Reviewers PLOS Computational Biology
  6. This post originally published on nodalpoint with comments

Creative Commons License

This work is licensed under a

Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.

March 30, 2007

This month’s molecule is…

Filed under: biotech — Duncan Hull @ 10:10 pm
Tags: , , , , , ,

Space-filling and backbone model of 1HRYThere are a number of “Molecule of the Month” style mini-reviews on the web, which highlight one particular molecule (usually a protein) every month, in an accessible style. Two of my personal favourites are protein spotlight: one month, one protein written by Vivienne Baillie Gerritsen of the Swiss-Prot team and Molecule of the Month at the Protein Databank PDB edited by David Goodsell. Both these features are worth a quick read because they can help bio-literate and bio-curious users to increase and reinforce their knowledge relatively quickly.

Part of what makes the PDB one worth reading is the colourful visualisations and short descriptions that go with it. For March 2007, PDBs molecule of the month is Zinc Fingers. Meanwhile, over at swissprot, the molecule is Sex-determining region Y protein (Sry), used to illustrate the tenuous nature of sex.

[This post originally published on nodalpoint with comments]

February 22, 2007

NSPNAS: Nature, Science or PNAS?

Filed under: publishing,Uncategorized — Duncan Hull @ 10:19 pm
Tags: , ,

A crude score for benchmarking scientists

TIM Have you ever wanted to compare different scientists by their publication record? It’s not always an easy task, but here is a crude and handy way to benchmark people by their journal publications in Nature, Science or PNAS using PubMed. Let’s call it the NSPNAS score, it’s not the h-index and it’s far from perfect, but it can be useful.

Imagine these scenarios:

  1. You’re a young scientist comtemplating who to do an undergraduate project, Masters degree or PhD with.
  2. You’ve finished your PhD and are wondering which lab could be your Stairway to PostDoc Heaven [1].
  3. You’re lucky enough to have landed a faculty position and you want to check the credibility of your new colleagues.
  4. You want to do some industrial espionage on your competitors in different labs around the world.
  5. You’re a Scientist dammit, and naturally you’re a curious person who just likes to measure things.

In any of these situations, you’ll probably want to look up the people concerned using Google Scholar which will give you a good idea of their research history. But you’re not interested in publications in the Journal of Few Subscribers or the Proceedings of the Boring Incomprehensible Nonsense Society (BINS), even if Google Scholar lists hundreds of their citations. Instead, you care about counting the Big Bang impact publications they have in the über-journals: Nature, Science and PNAS. You can find these publications in PubMed with this simple query:

Surname +Initials[au]+(nature[journal] or science[journal] or Proc Natl Acad Sci U S A[journal])

…and you can obviously modify this query to include popular journals from your own field as appropriate.

Where NSPNAS works

Note, NSPNAS scores were correct at the time of writring in 2007, but will change over time.

When you substitute an authors name and initials into the beginning of that query, you get your NSPNAS score. So Systems Biologist Douglas Kell for example, surname and initials “Kell+D[au]”, has an NSPNAS score of 6.

If the person in question has a unique or unusual surname and initials, its fairly easy to find their score: Nodalpointer Chris Mungall has an NSPNAS score of two while nodalpointer Jason Stajich has an NSPNAS score of three. These results suggest a positive correlation between Californian sunshine and NSPNAS. Meanwhile, back in rainy old Britain, Ensemblian Ewan Birney scores a formidable sixteen, which is just scary for a bloke in his thirties.

Where NSPNAS doesn’t work

Unfortunately, authors with common names like John Smith (who has more than 340 hits) can’t be easily benchmarked with this type of query, without trawling through hundreds of false positives. More importantly, some influential scientists score very low or zero, despite the fact that their work has been important in the world of biomedical science an beyond. This is especially true for Computer Scientists, Mathematicians and Informaticians, for example:

Many important members of the Dead Scientists Society also have low NSPNAS scores…

Conclusions

All these statistics remind us that many important ideas, techniques and results are not published in Nature, Science or PNAS and others are excluded from the PubMed index completely. It also confirms what we already know about peer-reviewed Journal publications not being the be-all and end-all of Engineering, Science or Medicine [3]. But NSPNAS still has its uses, provided the people you’re benchmarking have a rare name and didn’t snuff it before the PubMed index starts.

What is your NSPNAS score? If like me, you score a spectacular “nul points”, console yourself with the fact that you’re in good company with that score and given time, maybe you can change it.

References

  1. Jimmy Page and Robert Plant (1971) Stairway to Heaven
  2. Most of the Clay Mathematics Institute Millenium Prizes are still up for grabs if you get disillusioned with bioinformatics, fancy some fame and winning a million dollar fortune!
  3. Michael Seringhaus and Mark Gerstein (2007) Publishing perishing? Towards tomorrow’s information architecture BMC Bioinformatics 2007, 8:17 DOI:10.1186/1471-2105-8-17
  4. This post originally on nodalpoint, with comments

Creative Commons License

This work is licensed under a

Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.


January 22, 2007

DNA mania

Filed under: bio — Duncan Hull @ 10:29 pm
Tags: , , , ,

What does DNA do when it’s not being transcribed into RNA? It causes DNA mania…

Quote of the Day

“DNA, you know, is Midas’ gold. Everyone who touches it goes mad.”

Maurice Wilkins

Read the rest in [1,2]

Do you or your colleagues ever suffer from DNA mania [3,4]? A biochemist friend of mine once semi-jokingly remarked that people’s manic obsession with DNA is a bit like buying some food and being more interested in the bar-code on the packaging, than the food inside. In his particular area of research, DNA is about as exciting as bar-codes, because it doesn’t even leave the nucleus of the cell, at least in Eukaryotes. I wonder what readers of nodalpoint think of this analogy? Anyway, as a result of this philosophy, most of his community have developed an unhealthy and manic interest in proteins rather than DNA. You could call this particular obsessive-compulsive disorder “protein mania”.

Depending on the scientific obsession(s) of your particular community, you might need to substitute Protein or RNA for DNA in the above quote, as appropriate. And if that is all too molecular for you, substitute any other of your favourite bioinformatics buzzwords.

References

  1. Horace Freeland Judson (1996) The Eighth Day of Creation: Makers of the Revolution in Biology
  2. John Sulston (2006) Won for All: How the Drosophila Genome was sequenced: a book by Michael Ashburner
  3. André Pichot (1999) Histoire de la notion de gène (one of the first documented uses of the phrase “DNA mania”)
  4. Denis Noble (2006) The Music of Life: Biology Beyond the Genome (an antidote to DNA mania and the Dawkinian gene-centric view of Life)
  5. DNA Photograph taken by Unapersona in Ciutat de les Arts i les Ciències, Calatrava building, Valencia, Spain.

January 5, 2007

NAR Database Issue 2007: Not Waving But Drowning?

The 14th annual Nucleic Acids Research (NAR) database issue 2007 has just been published, open-access. This year is the largest yet (again) with 968 molecular biology databases listed, 110 more than the previous one (see figure below). In the world of biological databases, are we waving or drowning?

NAR Database Growth 2007

Nine hundred and sixty eight is a lot of databases, and even that mind-boggling number is not an exhaustive or comprehensive tally. But is counting all these databases waving or drowning [1]? Will we ever stop stamp-collecting the databases and tools we have in molecular biology? What prompted this is, an employee of the The Boeing Company once told me they have given up counting their databases because there were just too many. Just think of all the databases of design and technical documentation that accompanies the myriad of different aircraft that Boeing manufacture, like the iconic 747 jumbo jet. Now, combine that with all the supply chain, customer and employee information and you can begin to imagine the data deluge that a large multi-national corporation has to handle.

Like Boeing, in Biology we’ve clearly got more data than we know what to do with [2,3]. It won’t be news to bioinformaticians and its been said many times before but its worth repeating again here:

  • We know how many databases we have but we don’t know what a lot of the data in these databases means, think of all those mystery proteins of unknown function. It will obviously take time until we understand it all…
  • Most of the data only begins to make sense when it is integrated or mashed-up with other data. However, we still don’t know how to integrate all these databases, or as Lincoln Stein puts it “so far their integration has proved problematic” [4], a bit of an understatement. Many grandiose schemes for the “integration” of biological databases have been proposed over the years, but unfortunately none have been practical to the point of implementation [5]


IMGP4592
Despite this, it is still useful to know how many molecular biology databases there are. At least we know how many databases we are drowning in. Thankfully, unlike Boeing, most biological data, algorithms and tools are open-source and more literature is becoming open access which will hopefully make progress more rapid. But biology is more complicated than a Boeing 747, so we’ve got a long-haul flight ahead of us. OK, I’ve managed to completely overstretch that aerospace analogy now so I’ll stop there.

Whatever databases you’ll be using in 2007, have a Happy New Year mining, exploring and understanding the data they contain, not drowning in it.

References

  1. Stevie Smith (1957) Not waving but drowning
  2. Michael Galperin (2007) The Molecular Biology Database Collection: 2007 update Nucleic Acids Research, Vol. 35, Database issue. DOI:10.1093/nar/gkl1008
  3. Alex Bateman (2007) Editorial: What makes a good database? Nucleic Acids Research, Vol. 35, Database issue. DOI:10.1093/nar/gkl1051
  4. Lincoln Stein (2003) Biological Database Integration Nature Reviews Genetics. 4 (5), 337-45. DOI:10.1038/nrg1065
  5. Michael Ashburner (2006) Keynote at the Pacific Symposium on Biocomputing (PSB2006) in Hawaii seeAlso Aloha: Biocomputing in Hawaii
  6. This post originally published on nodalpoint with comments

Creative Commons License
This work is licensed under a

Creative Commons Attribution-Noncommercial-Share Alike 3.0 License.


December 19, 2006

Taverna 1.5.0

Filed under: Uncategorized — Duncan Hull @ 8:26 pm
Tags: , , , , , ,

Happy Christmas from the myGrid team, who are pleased to announce the release of version 1.5.0 of the Open Source Taverna bioinformatics workflow toolkit [1]. This is now available for download on the Sourceforge site and includes some substantial changes to version 1.4.

IMGP4570Taverna 1.5.0 is a small download, but when first run it will then download and install the required packages which can take some time on slow networks. In the near future there will be a mechanism for downloading a bundle of core packages. There are some significant changes in the underlying architecture of Taverna and how it handles core packages and optional plugins, using a system called Raven, see release notes below.

The documentation is currently being updated and the user documentation should be complete very soon, with the technical documentation following shortly afterwards. The reason for this is to allow the software to be released with some time to spare before the Christmas holidays.

Release notes:

There have been a number of substantial changes in the underlying architecture of Taverna since the previous release. These include:

  • An overhaul of the User Interface (UI), replacing the unpopular Multiple Document Interface with a cleaner and simpler single document UI which can be customised using Perspectives. There are built in perspectives to allow the design and enactment of workflows, and plugins can integrate with the UI by providing perspectives of their own. Together with this, users are able to create their own layouts built from individual components.
  • Taverna now allows for multiple workflows to be open and enacted at the same time.
  • Support for the new BioMart data management system version 0.5, together with backward compatibility for old workflows that used Biomart 0.4.
  • Better provenance generation and browsing support, through a plugin now known as LogBook.
  • Better support for semantic service discovery through the Feta plugin [2].
  • Modulularisation of the Taverna code base.
  • Development and integration of an underlying architecture know as Raven. This allows for Apache Maven like declaration of dependencies which are discovered and incorporated into the Taverna system at runtime. Together with the modularisation of the Taverna code base, Raven gives the benefit that updates can be provided dynamically and incrementally, without the need for monolithic releases as in the past. This allows the provision of updates to bugs, and new features, within a very short timescale if necessary. It also provides plugin developers with a greater degree of autonomy and independance from the core Taverna code base.
  • Improved and more advanced plugin management with the ability to provide immediate updates, and for plugin providers to publish their plugins via xml descriptions.
  • Numerous bug-fixes including the removal of a number of memory leaks.

JIRA generated release notes and bug status reports can be found here and here

References

  1. Peer-reviewed publications about the Taverna workbench in PubMed
  2. Feta: A Light-Weight Architecture for User Oriented Semantic Service Discovery
  3. BioMoby extensions to the Taverna workflow management and enactment software

December 12, 2006

Semantic Web for Life Sciences Book

Filed under: semweb — Duncan Hull @ 4:57 pm
Tags: ,

Revolutionizing Knowledge Discovery in the Life Sciences
All I want for Christmas is a book about the semantic web, written by people who are actually building and using it, rather than “visionaries” who don’t have to. Maybe this year I’ll be lucky…

A group of semantic webheads (aka HCLSIG the Health Care and Life Sciences Interest Group) led by Christopher J. Baker and Kei-Hoi Cheung and gathered together on public-semweb-lifesci@w3.org have written a book about the semantic web for life sciences.

I haven’t seen the final printed version of this book yet, but if you want to add it to your christmas amazon wishlist, its called Semantic Web: Revolutionizing Knowledge Discovery in the Life Sciences (ISBN:0387484361). The table of contents for the book (DOI:10.1007/978-0-387-48438-9) has more details if you are interested.

So what about other readers, what bioinformatics presents (not just books) would they like to find under the Christmas tree this year? If you don’t celebrate Christmas, what Solstice wishes do you have?

(see original post at nodalpoint for comments)

Buggotea: Redundant Links in Connotea

IMGP4570Dear Santa, all I want for Christmas* is a better version of Connotea, please can you sort out it’s duplicated redundant links? In my book this particular bug is “buggotea” number one. Here is the problem… [update: buggotea is partially fixed, see comments from Ian Mulvany at the nodalpoint link in the references below]

There is this handy bioinformatics web application called Connotea which I like to use, built by those nice people in the web team at Nature Publishing Group. Most readers of nodalpoint probably already know about it, but because you’re Santa and you’ve been busy lately, let me explain. Connotea can help scientists (not just bioinformaticians) to organise and share their bibliographic references, whilst discovering what other people with similar interests are reading. It’s good, but it has some bugs in it. Since it’s open-source software, anyone with the time, inclination and skills can get hold of the connotea source code and improve it. There is, however, one particularly nasty redundancy bug in Connotea that is bugging me [1]. I think it should be fixable, and that doing so would make Connotea a significantly better application than it already is. Let’s illustrate this bug with a little story…

(more…)

« Previous PageNext Page »

Blog at WordPress.com.