O'Really?

June 10, 2009

Kenjiro Taura on Parallel Workflows

Kenjiro TauraKenjiro Taura is visting Manchester next week from the Department of Information and Communication Engineering at the University of Tokyo. He will be doing a seminar, the details of which are below:

Title: Large scale text processing made simple by GXP make: A Unixish way to parallel workflow processing

Date-time: Monday, 15 June 2009 at 11:00 AM

Location: Room MLG.001, mib.ac.uk

In the first part of this talk, I will introduce a simple tool called GXP make. GXP is a general purpose parallel shell (a process launcher) for multicore machines, unmanaged clusters accessed via SSH, clusters or supercomputers managed by batch scheduler, distributed machines, or any mixture thereof. GXP make is a ‘make‘ execution engine that executes regular UNIX makefiles in parallel. Make, though typically used for software builds, is in fact a general framework to concisely describe workflows constituting sequential commands. Installation of GXP requires no root privileges and needs to be done only on the user’s home machine. GXP make easily scales to more than 1,000 CPU cores. The net result is that GXP make allows an easy migration of workflows from serial environments to clusters and to distributed environments. In the second part, I will talk about our experiences on running a complex text processing workflow developed by Natural Language Processing (NLP) experts. It is an entire workflow that processes MEDLINE abstracts with deep NLP tools (e.g., Enju parser [1]) to generate search indices of MEDIE, a semantic retrieval engine for MEDLINE. It was originally described in Makefile without a particular provision to parallel processing, yet GXP make was able to run it on clusters with almost no changes to the original Makefile. Time for processing abstracts published in a single day was reduced from approximately eight hours (with a single machine) to twenty minutes with a trivial amount of efforts. A larger scale experiment of processing all abstracts published so far and remaining challenges will also be presented.

References

  1. Miyao, Y., Sagae, K., Saetre, R., Matsuzaki, T., & Tsujii, J. (2008). Evaluating contributions of natural language parsers to protein-protein interaction extraction Bioinformatics, 25 (3), 394-400 DOI: 10.1093/bioinformatics/btn631

June 4, 2009

Improving the OBO Foundry Principles

The Old Smithy Pub by loop ohThe Open Biomedical Ontologies (OBO) are a set of reference ontologies for describing all kinds of biomedical data, see [1-5] for examples. Every year, users and developers of these ontologies gather from around the globe for a workshop at the EBI near Cambridge, UK. Following on from the first workshop last year, the 2nd OBO workshop 2009 is fast approaching.

In preparation, I’ve been revisiting the OBO Foundry documentation, part of which establishes a set of principles for ontology development. I’m wondering how they could be improved because these principles are fundamental to the whole effort. We’ve been using one of the OBO ontologies (called Chemical Entities of Biological Interest (ChEBI)) in the REFINE project to mine data from the PubMed database. OBO Ontologies like ChEBI and the Gene Ontology are really crucial to making sense of the massive data which are now common in biology and medicine – so this is stuff that matters.

The OBO Foundry Principles, a sort of Ten Commandments of Ontology (or Obology if you prefer) currently look something like this (copied directly from obofoundry.org/crit.shtml):

  1. The ontology must be open and available to be used by all without any constraint other than (a) its origin must be acknowledged and (b) it is not to be altered and subsequently redistributed under the original name or with the same identifiers.The OBO ontologies are for sharing and are resources for the entire community. For this reason, they must be available to all without any constraint or license on their use or redistribution. However, it is proper that their original source is always credited and that after any external alterations, they must never be redistributed under the same name or with the same identifiers.
  2. The ontology is in, or can be expressed in, a common shared syntax. This may be either the OBO syntax, extensions of this syntax, or OWL. The reason for this is that the same tools can then be usefully applied. This facilitates shared software implementations. This criterion is not met in all of the ontologies currently listed, but we are working with the ontology developers to have them available in a common OBO syntax.
  3. The ontologies possesses a unique identifier space within the OBO Foundry. The source of a term (i.e. class) from any ontology can be immediately identified by the prefix of the identifier of each term. It is, therefore, important that this prefix be unique.
  4. The ontology provider has procedures for identifying distinct successive versions.
  5. The ontology has a clearly specified and clearly delineated content. The ontology must be orthogonal to other ontologies already lodged within OBO. The major reason for this principle is to allow two different ontologies, for example anatomy and process, to be combined through additional relationships. These relationships could then be used to constrain when terms could be jointly applied to describe complementary (but distinguishable) perspectives on the same biological or medical entity. As a corollary to this, we would strive for community acceptance of a single ontology for one domain, rather than encouraging rivalry between ontologies.
  6. The ontologies include textual definitions for all terms. Many biological and medical terms may be ambiguous, so terms should be defined so that their precise meaning within the context of a particular ontology is clear to a human reader.
  7. The ontology uses relations which are unambiguously defined following the pattern of definitions laid down in the OBO Relation Ontology.
  8. The ontology is well documented.
  9. The ontology has a plurality of independent users.
  10. The ontology will be developed collaboratively with other OBO Foundry members.

ResearchBlogging.orgI’ve been asking all my frolleagues what they think of these principles and have got some lively responses, including some here from Allyson Lister, Mélanie Courtot, Michel Dumontier and Frank Gibson. So what do you think? How could these guidelines be improved? Do you have any specific (and preferably constructive) criticisms of these ambitious (and worthy) goals? Be bold, be brave and be polite. Anything controversial or “off the record” you can email it to me… I’m all ears.

CC-licensed picture above of the Old Smithy (pub) by Loop Oh. Inspired by Michael Ashburner‘s standing OBO joke (Ontolojoke) which goes something like this: Because Barry Smith is one of the leaders of OBO, should the project be called the OBO Smithy or the OBO Foundry? 🙂

References

  1. Noy, N., Shah, N., Whetzel, P., Dai, B., Dorf, M., Griffith, N., Jonquet, C., Rubin, D., Storey, M., Chute, C., & Musen, M. (2009). BioPortal: ontologies and integrated data resources at the click of a mouse Nucleic Acids Research DOI: 10.1093/nar/gkp440
  2. Côté, R., Jones, P., Apweiler, R., & Hermjakob, H. (2006). The Ontology Lookup Service, a lightweight cross-platform tool for controlled vocabulary queries BMC Bioinformatics, 7 (1) DOI: 10.1186/1471-2105-7-97
  3. Smith, B., Ashburner, M., Rosse, C., Bard, J., Bug, W., Ceusters, W., Goldberg, L., Eilbeck, K., Ireland, A., Mungall, C., Leontis, N., Rocca-Serra, P., Ruttenberg, A., Sansone, S., Scheuermann, R., Shah, N., Whetzel, P., & Lewis, S. (2007). The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration Nature Biotechnology, 25 (11), 1251-1255 DOI: 10.1038/nbt1346
  4. Smith, B., Ceusters, W., Klagges, B., Köhler, J., Kumar, A., Lomax, J., Mungall, C., Neuhaus, F., Rector, A., & Rosse, C. (2005). Relations in biomedical ontologies Genome Biology, 6 (5) DOI: 10.1186/gb-2005-6-5-r46
  5. Bada, M., & Hunter, L. (2008). Identification of OBO nonalignments and its implications for OBO enrichment Bioinformatics, 24 (12), 1448-1455 DOI: 10.1093/bioinformatics/btn194

June 2, 2009

Who Are You? Digital Identity in Science

The Who by The WhoThe organisers of the Science Online London 2009 conference are asking people to propose their own session ideas (see some examples here), so here is a proposal:

Title: Who Are You? Digital Identity in Science

Many important decisions in Science are based on identifying scientists and their contributions. From selecting reviewers for grants and publications, to attributing published data and deciding who is funded, hired or promoted, digital identity is at the heart of Science on the Web.

Despite the importance of digital identity, identifying scientists online is an unsolved problem [1]. Consequently, a significant amount of scientific and scholarly work is not easily cited or credited, especially digital contributions: from blogs and wikis, to source code, databases and traditional peer-reviewed publications on the Web. This (proposed) session will look at current mechanisms for identifying scientists digitally including contributor-id (CrossRef), researcher-id (Thomson), Scopus Author ID (Elsevier), OpenID, Google Scholar [2], Single Sign On, PubMed, Google Scholar [2], FOAF+SSL, LinkedIn, Shared Identifiers (URIs) and the rest. We will introduce and discuss each via a SWOT analysis (Strengths, Weaknesses, Opportunities and Threats). Is digital identity even possible and ethical? Beside the obvious benefits of persistent, reliable and unique identifiers, what are the privacy and security issues with personal digital identity?

If this is a successful proposal, I’ll need some help. Any offers? If you are interested in joining in the fun, more details are at scienceonlinelondon.org

References

  1. Bourne, P., & Fink, J. (2008). I Am Not a Scientist, I Am a Number PLoS Computational Biology, 4 (12) DOI: 10.1371/journal.pcbi.1000247
  2. Various Publications about unique author identifiers bookmarked in citeulike
  3. Yours Truly (2009) Google thinks I’m Maurice Wilkins
  4. The Who (1978) Who Are You? Who, who, who, who? (Thanks to Jan Aerts for the reference!)

Michael Ley on Digital Bibliographies

Michael Ley

Michael Ley is visiting Manchester this week, he will be doing a seminar on Wednesday 3rd June, here are some details for anyone who is interested in attending:

Date: 3rd Jun 2009

Title: DBLP: How the data get in

Speaker: Dr Michael Ley. University of Trier, Germany

Time & Location: 14:15, Lecture Theatre 1.4, Kilburn Building

Abstract: The DBLP (Digital Bibliography & Library Project) Computer Science Bibliography now includes more than 1.2 million bibliographic records. For Computer Science researchers the DBLP web site now is a popular tool to trace the work of colleagues and to retrieve bibliographic details when composing the lists of references for new papers. Ranking and profiling of persons, institutions, journals, or conferences is another usage of DBLP. Many scientists are aware of this and want their publications being listed as complete as possible.

The talk focuses on the data acquisition workflow for DBLP. To get ‘clean’ basic bibliographic information for scientific publications remains a chaotic puzzle.

Large publishers are either not interested to cooperate with open services like DBLP, or their policy is very inconsistent. In most cases they are not able or not willing to deliver basic data required for DBLP in a direct way, but they encourage us to crawl their Web sites. This indirection has two main problems:

  1. The organisation and appearance of Web sites changes from time to time, this forces a reimplementation of information extraction scripts. [1]
  2. In many cases manual steps are necessary to get ‘complete’ bibliographic information.

For many small information sources it is not worthwhile to develop information extraction scripts. Data acquisition is done manually. There is an amazing variety of small but interesting journals, conferences and workshops in Computer Science which are not under the umbrella of ACM, IEEE, Springer, Elsevier etc. How they get it often is decided very pragmatically.

The goal of the talk and my visit to Manchester is to start a discussion process: The EasyChair conference management system developed by Andrei Voronkov and DBLP are parts of scientific publication workflow. They should be connected for mutual benefit?

References

  1. Lincoln Stein (2002). Creating a bioinformatics nation: screen scraping is torture Nature, 417 (6885), 119-120 DOI: 10.1038/417119a

Blogging For Profit: Costs and Benefits


Business Graph by nDevilTV
The organisers of the Science Online London 2009 conference are asking people to propose their own session ideas (see some examples here), so here is proposal:

Title: Blogging For Profit: Costs and Benefits

What are the costs and benefits of blogging and how can you make sure the latter justifies the former?

This (proposed) session will look at two kinds of profit, and the costs associated with each.

  1. Research profit (in science and academia), building digital reputations on the Web. Can blogging help your next grant proposal for research funding and if so, how? How can blogging be used to increase the visibility and impact of published research via the likes of ResearchBlogging.org, blogs.nature.com and other aggregators?
  2. Financial profit (in business), making blogging pay the bills. What business models (and infrastructure) exist to support blogging? Including, but not limited to: Nature Network, ScienceBlogs, Google AdSense, “20% time“, “free” tools (WordPress, Blogger, OpenWetWare etc). Going solo vs. joining a club – which business models and tools are right for you?

This could be followed by a general discussion on these benefits. When do they justify their costs (and risks) and make for profitable blogging?

If this is a successful proposal, I’ll need some help. Any offers? If you are interested in joining in the fun, details are at scienceonlinelondon.org

[CC-licensed Business Graph picture by nDevilTV]

June 1, 2009

Scott Marshall on Interoperability

M. Scott MarshallScott Marshall is visiting Manchester this week, he will be doing a seminar on Friday 5th June, here are some details for anyone who is interested in attending:

Speaker: Dr. M. Scott Marshall, The University of Amsterdam

Date/Time: 5th June 2009, 11:00

Location: Room MLG.001 (Lecture Theatre), MIB building, (number 16 on campus map)

Title: Standards Enabled Interoperability: W3C Semantic Web for Health Care and Life Sciences Interest Group

Abstract: The W3C Semantic Web for Health Care and Life Sciences Interest Group (HCLS) has the mission of developing, advocating for, and supporting the use of Semantic Web technologies for biological science, translational medicine and health care. HCLS covers hot topics including data integration and federation, bridging commonly used domain standards such as CDISC and HL7, and the applications of medical terminologies. This talk will introduce the HCLS, as well as provide an overview of the activities that are currently ongoing within the task forces, as well as new developments and the recent Face2Face meeting. The role of information extraction and the current interest in Shared Identifiers will also be discussed.

References

  1. Ruttenberg, A., Rees, J., Samwald, M., & Marshall, M. (2009). Life sciences on the Semantic Web: the Neurocommons and beyond Briefings in Bioinformatics, 10 (2), 193-204 DOI: 10.1093/bib/bbp004

May 26, 2009

Subscribing to O’Really?

Filed under: technology — Duncan Hull @ 4:53 pm
Tags: , , , ,

Feed the WorldJust a quick note about subscribing: if you are a regular reader of this O’Really blog and you don’t already subscribe, there are two ways you can receive automatic notifications when new posts are published here:

  1. Point your feed reader at http://feeds2.feedburner.com/oreally, (the preferred method) or …
  2. Point your feed reader at https://duncan.hull.name/feed/ (the WordPress method) which unfortunately gives unreliable subscriber stats. This  feed is linked to by the blue or orange feed icon (pictured right) you should be able to see in your web browsers address bar.

The first feed, is just the second feed re-routed through the magic of FeedBurner, which gives more useful viewing statistics.

Grants on the Web: Transparent Scientific Funding?

Lord Drayson by DIUSGOVUKAll over the Britain, politicians are getting ready to publish their expenses on the Interweb. Why? Because they are trying to regain their lost credibility, after making some incredibly dodgy and embarrassing expense claims [1-7]. Scandals aside, this is all well and good since this money has come from the UK taxpayers pocket, and politicians are public servants, doing public work which is supposedly in the public good.

Scientists, like politicians, also provide a public service, spending public money, for the public good. Science is public knowledge after all and scientists spend quite a lot of public money. At least £3 billion was spent on scientific research in the UK during 2008 (see Who Funds Science in Britain?) and that was just research, not teaching. Wouldn’t it be great if anyone who was interested could see what all this money had been spent on, who spent it and what the outcomes were?

Thankfully you can already do this for some areas of research. The Engineering and Physical Sciences Research Council (EPSRC) which currently spends around £740 million a year on everything from “mathematics to materials science, and from information technology to structural engineering” has a system called Grants on the Web @ gow.epsrc.ac.uk. You can find out who spent the money and how much money was spent since the system was set up, see an EPSRC example here. Some of the original grant proposals are there too, which can be enlightening. The Biotechnology and Biological Sciences Research Council (BBSRC) also has a similar system (called oasis), though it is not as easy to use and link to – see a BBSRC example here. The trouble is, if you can’t easily link to it, it doesn’t get indexed by search engines. If it doesn’t get indexed by search engines, then it’s almost invisible. Fortunately, the BBSRC are working on improving this, with a new system due for release in the autumn of 2009.

Other organisations are putting grant information on the web too. Recently, thanks to the UK’s PubMed Central database you can also see the published results of publicly funded biomedical research. The funders pages at ukpmc.ac.uk/funders give a breakdown of published results from different funding bodies, as described in this article by Robert Kiley of the Wellcome Trust and this one by Alison Henning.

Now not all the research councils seem to publish their grants on the Web in a transparent manner*, and some of those that do, leave lots of room for improvement. But it is still useful to be able to see where some of that public money went and what the outcomes of the research were. More transparent spending of public money like this isn’t just a desirable extra, it should come as standard.

* (It is difficult to find the details of grants awarded by JISC, NERC, MRC and STFC, but please leave a comment below if you know where this information is published. More commentary on this post over at friendfeed.)

[Creative Commons licensed picture of Baron Paul Drayson, currently UK Science Minister from DIUSGOVUK.]

References

  1. The Daily Telegraph (2009) MPs’ expenses: all the gory details from the Daily Telegraph
  2. The Guardian (2009) Grauniad datablog: MP’s expenses as spreadsheet and Free Our Data: Make taxpayers’ data available to them
  3. Wikipedia (2009) MPs’ expenses in wikipedia
  4. BBC News (2009) MPs’ expenses: A triumph of journalism? A week after its opening salvo, the Daily Telegraph is still reaping great benefit from its exclusive expose of MPs’ expenses.
  5. BBC News (2009) Q&A: MP expenses row explained: Revelations in the Daily Telegraph about exactly what MPs have been claiming on expenses has prompted a public outcry and a pledge to reform the “gentlemen’s club” at Westminster
  6. BBC Newsnight (2009) Stephen Fry dismisses MPs’ expenses row, accusing journalists of hypocrisy
  7. The Guardian (2009) Censored version of MPs’ expenses will break the law, Hugh Tomlinson QC warns

May 21, 2009

Upcoming Gig: The Italian Job at NETTAB

NETTAB: Network Tools and Applications in BiologyNetwork Tools and Applications in Biology (NETTAB) is a series of workshops in Bioinformatics. It focuses on the most promising and innovative ICT tools and their utility in Bioinformatics. These workshops aim to introduce participants to the evolving network standards and technologies that are being applied to the field of biology.

Since 2001, the NETTAB workshops have being doing a Giro d’Italia or  Grand Tour of Italy; Genova, Bologna, Naples, Sardinia, Lake Como and Pisa have all played host to the workshop. This year, NETTAB 2009 is in Catania at the Università degli Studi di Catania in Sicily close to Mount Etna.

There is special theme for this years workshop, held on June 10-13, on Technologies, Tools and Applications for Collaborative and Social Bioinformatics Research and Development. So I’m very pleased that Paolo Romano asked me to do a keynote presentation (w00t!) on the work we have been doing in the REFINE project and myExperiment. Grazie Paolo, grazie. And thanks Carole Goble too for the recommendation.

If you’re going to NETTAB this year, see you there. If you’d like to come, today is the last day for the early bird discount, sign up at the registration page. The scientific programme looks interesting, it will be good to meet Alex Bateman and Tim Clark and the rest of this years speakers.

Now, if my keynote presentation is going to (as Michael Caine once famously said [1]) “blow the bl**dy doors off” [2], it needs loads more work. So I’d better get back to it. Ciao!

[Update: See reports from day one, day two and day three of NETTAB 2009.]

References

  1. Peter Collinson and Troy Kennedy-Martin (1969) The Italian Job
  2. Michael Caine (1969) “You’re only supposed to blow the bl**dy doors off!”
  3. Cannata, N., Schröder, M., Marangoni, R., & Romano, P. (2008). A Semantic Web for bioinformatics: goals, tools, systems, applications BMC Bioinformatics, 9 (Suppl 4) DOI: 10.1186/1471-2105-9-S4-S1

May 19, 2009

Defrosting the John Rylands University Library

Filed under: seminars — Duncan Hull @ 4:14 pm
Tags: , , , , , , , , , , , ,

http://www.flickr.com/photos/dpicker/3107856991/For anyone who missed the original bioinformatics seminar I’ll be doing a repeat of the “Defrosting the Digital Library” talk, this time for the staff in the John Rylands University Library (JRUL) . This is the main academic library in Manchester with (quote) “more than 4 million printed books and manuscripts, over 41,000 electronic journals and 500,000 electronic books, as well as several hundred databases, the John Rylands University Library is one of the best-resourced academic libraries in the country.” The journal subscription budget of the library is currently around £4 million per year, that’s before they’ve even bought any books! Here is the abstract for the talk:

After centuries with little change, scientific libraries have recently experienced massive upheaval. From being almost entirely paper-based, most libraries are now almost completely digital. This information revolution has all happened in less than 20 years and has created many novel opportunities and threats for scientists, publishers and libraries.

Today, we are struggling with an embarrassing wealth of digital knowledge on the Web. Most scientists access this knowledge through some kind of digital library, however these places can be cold, impersonal, isolated, and inaccessible places. Many libraries are still clinging to obsolete models of identity, attribution, contribution, citation and publication.

Based on a review published in PLoS Computational Biology, pubmed.gov/18974831 this talk will discuss the current chilly state of digital libraries for biologists, chemists and informaticians, including PubMed and Google Scholar. We highlight problems and solutions to the coupling and decoupling of publication data and metadata, with a tool called citeulike.org. This software tool (and many other tools just like it) exploit the Web to make digital libraries “warmer”: more personal, sociable, integrated, and accessible places.

Finally issues that will help or hinder the continued warming of libraries in the future, particularly the accurate identity of authors and their publications, are briefly introduced. These are discussed in the context of the BBSRC funded REFINE project, at the National Centre for Text Mining (NaCTeM.ac.uk), which is linking biochemical pathway data with evidence for pathways from the PubMed database.

Date: Thursday 21st May 2009, Time: 13.00, Location: John Rylands University (Main) Library Oxford Road, Parkinson Room (inside main entrance, first on right) University of Manchester (number 55 on google map of the Manchester campus). Please come along if you are interested…

References

  1. Hull, D., Pettifer, S., & Kell, D. (2008). Defrosting the Digital Library: Bibliographic Tools for the Next Generation Web PLoS Computational Biology, 4 (10) DOI: 10.1371/journal.pcbi.1000204

[CC licensed picture above, the John Rylands Library on Deansgate by dpicker: David Picker]

« Previous PageNext Page »

Blog at WordPress.com.