September 9, 2014

Punning with the Pub in PubMed: Are there any decent NCBI puns left? #PubMedPuns

PubMedication: do you get your best ideas in the Pub?

Many people claim they get all their best ideas in the pub, but for lots of scientists their best ideas probably come from PubMed.gov – the NCBI’s monster database of biomedical literature. Consequently, the database has spawned a whole slew of tools that riff off the PubMed name, with many puns and portmanteaus (aka “PubManteaus”), and the pub-based wordplays are very common. [1,2]

All of this might make you wonder, are there any decent PubMed puns left? Here’s an incomplete collection:

  • PubCrawler pubcrawler.ie “goes to the library while you go to the pub…” [3,4]
  • PubChase pubchase.com is a “life sciences and medical literature recommendations engine. Search smarter, organize, and discover the articles most important to you.” [5]
  • PubCast scivee.tv/pubcasts allow users to “enliven articles and help drive more views” (to PubMed) [6]
  • PubFig nothing to do with PubMed, but research done on face and image recognition that happens to be indexed by PubMed. [7]
  • PubGet pubget.com is a “comprehensive source for science PDFs, including everything you’d find in Medline.” [8]
  • PubMine “supports intelligent knowledge discovery” [9]
  • PubNet pubnet.gersteinlab.org is a “web-based tool that extracts several types of relationships returned by PubMed queries and maps them into networks” aka a publication network graph utility. [10]
  • GastroPub repackages and re-sells ordinary PubMed content disguised as high-end luxury data at a higher premium, similar to a Gastropub.
  • PubQuiz is either the new name for NCBI database search www.ncbi.nlm.nih.gov/gquery or a quiz where you’re only allowed to use PubMed to answer questions.
  • PubSearch & PubFetch allows users to “store literature, keyword, and gene information in a relational database, index the literature with keywords and gene names, and provide a Web user interface for annotating the genes from experimental data found in the associated literature” [11]
  • PubScience is either “peer-reviewed drinking” courtesy of pubsci.co.uk or an ambitious publishing project tragically axed by the U.S. Department of Energy (DoE). [12,13]
  • PubSub is anything that makes use of the publish–subscribe pattern, such as NCBI feeds. [14]
  • PubLick as far as I can see, hasn’t been used yet, unless you count this @publick on twitter. If anyone was launching a startup, working in the area of “licking” the tastiest data out of PubMed, that could be a great name for their data-mining business. Alternatively, it could be a catchy new nickname for PubMedCentral (PMC) or Europe PubMedCentral (EuropePMC) [15] – names which don’t exactly trip off the tongue. Since PMC is a free digital archive of publicly accessible full-text scholarly articles, PubLick seems like a appropriate moniker.

PubLick Cat got all the PubMed cream.

There’s probably lots more PubMed puns and portmanteaus out there just waiting to be used. Pubby, Pubsy, PubLican, Pubble, Pubbit, Publy, PubSoft, PubSort, PubBrawl, PubMatch, PubGames, PubGuide, PubWisdom, PubTalk, PubChat, PubShare, PubGrub, PubSnacks and PubLunch could all work. If you’ve know of any other decent (or dodgy) PubMed puns, leave them in the comments below and go and build a scientific twitterbot or cool tool using the same name — if you haven’t already.


April 1, 2014

The Serene Scientists Serenity Prayer via Jon Butterworth



The Church of Banksy

Whatever your religous preferences, the Serenity Prayer by Reinhold Niebuhr captures a certain wisdom about life in general. So it is good to see that physicist Jon Butterworth at UCL has adapted it [1] for scientists:

“Give me grace to accept with serenity the things that cannot be understood,

Data to investigate the things which can be understood,

And the Wisdom to know the difference.”



August 3, 2012

April 2, 2012

Open Data Manchester: Twenty Four Hour Data People

Sean Ryder at the Hacienda

Sean Ryder, the original twenty-four hour Manchester party person of the Happy Mondays, spins the discs at the Wickerman festival in 2008.

According to Francis Maude, Open Data is the raw material for “next industrial revolution”. Now you should obviously take everything politicians say with a large pinch of salt (especially Maude) but despite the political hyperbole, when it comes to data he is onto something.

According to wikipedia, which is considerably more reliable than politicians, Open Data is:

“the idea that certain data should be freely available to everyone to use and republish as they wish, without restrictions from copyright, patents or other mechanisms of control.”

Open Data is slowly having an impact in the world of science [1] and also in wider society. Initiatives like data.gov in the U.S. and data.gov.uk in England, also known as e-government or government 2.0, have put huge amounts of data in the public domain and there is plenty more data in the pipeline. All of this data makes novel applications possible, like cycling injury maps showing accident black spots, and many others just like it.

To discuss the current status of Open Data in Greater Manchester there were two events last week:

  1. The Open Data Manchester meetup “24 hour data people” [2] at the the Manchester Digital Laboratory (“madlab”), which recently made BBC headlines with the DIY bio project
  2. The Discover Open Data event at the Cornerhouse cinema
Here is a brief and incomplete summary of what went on at these events:


February 15, 2012

The Open Access Irony Awards: Naming and shaming them

Ask me about open access by mollyaliOpen Access (OA) publishing aims to make the results of scientific research available to the widest possible audience. Scientific papers that are published in Open Access journals are freely available for crucial data mining and for anyone or anything to read, wherever they may be.

In the last ten years, the Open Access movement has made huge progress in allowing:

“any users to read, download, copy, distribute, print, search, or link to the full texts of these articles, crawl them for indexing, pass them as data to software, or use them for any other lawful purpose, without financial, legal, or technical barriers.”

But there is still a long way to go yet, as much of the world’s scientific knowledge remains locked up behind publisher’s paywalls, unavailable for re-use by text-mining software and inaccessible to the public, who often funded the research through taxation.

Openly ironic?

ironicIronically, some of the papers that are inaccessible discuss or even champion the very Open Access movement itself. Sometimes the lack of access is deliberate, other times accidental – but the consequences are serious. Whether deliberate or accidental, restricted access to public scientific knowledge is slowing scientific progress [1]. Sometimes the best way to make a serious point is to have a laugh and joke about it. This is what the Open Access Irony Awards do, by gathering all the offenders in one place, we can laugh and make a serious point at the same time by naming and shaming the papers in question.

To get the ball rolling, here is are some examples:

  • The Lancet owned by Evilseviersorry I mean Elsevier, recently  published a paper on “the case for open data” [2] (please login to access article). Login?! Not very open…
  • Serial offender and über-journal Science has an article by Elias Zerhouni on the NIH public access policy [3] (Subscribe/Join AAAS to View Full Text), another on “making data maximally available” [4] (Subscribe/Join AAAS to View Full Text) and another on a high profile advocate of open science [5] (Buy Access to This Article to View Full Text) Irony of ironies.
  • From Nature Publishing Group comes a fascinating paper about harnessing the wisdom of the crowds to predict protein structures [6]. Not only have members of the tax-paying public funded this work, they actually did some of the work too! But unfortunately they have to pay to see the paper describing their results. Ironic? Also, another published in Nature Medicine proclaims the “delay in sharing research data is costing lives” [1] (instant access only $32!)
  • From the British Medical Journal (BMJ) comes the worrying news of dodgy American laws that will lock up valuable scientific data behind paywalls [7] (please subscribe or pay below). Ironic? *
  • The “green” road to Open Access publishing involves authors uploading their manuscript to self-archive the data in some kind of  public repository. But there are many social, political and technical barriers to this, and they have been well documented [8]. You could find out about them in this paper [8], but it appears that the author hasn’t self-archived the paper or taken the “gold” road and pulished in an Open Access journal. Ironic?
  • Last, but not least, it would be interesting to know what commercial publishers make of all this text-mining magic in Science [9], but we would have to pay $24 to find out. Ironic?

These are just a small selection from amongst many. If you would like to nominate a paper for an Open Access Irony Award, simply post it to the group on Citeulike or group on Mendeley. Please feel free to start your own group elsewhere if you’re not on Citeulike or Mendeley. The name of this award probably originated from an idea Jonathan Eisen, picked up by Joe Dunckley and Matthew Cockerill at BioMed Central (see tweet below). So thanks to them for the inspiration.

For added ironic amusement, take a screenshot of the offending article and post it to the Flickr group. Sometimes the shame is too much, and articles are retrospectively made open access so a screenshot will preserve the irony.

Join us in poking fun at the crazy business of academic publishing, while making a serious point about the lack of Open Access to scientific data.




* Please note, some research articles in BMJ are available by Open Access, but news articles like [7] are not. Thanks to Trish Groves at BMJ for bringing this to my attention after this blog post was published. Also, some “articles” here are in a grey area for open access, particularly “journalistic” stuff like news, editorials and correspondence, as pointed out by Becky Furlong. See tweets below…

December 17, 2010

Planet Facebook

Whatever your views on Facebook [1], you can’t deny that from space, “Planet Facebook” looks rather intriguing. The wonderful diagram below of Facebook connections has been made by Paul Butler. Even miserable Facebook refuseniks (like me) can’t help but go “ooh that’s pretty” while marvelling at the masterful use of the R language to construct this beautiful map…

Planet Facebook / Planet Earth by Paul Butler


September 1, 2010

How many unique papers are there in Mendeley?

Lex Macho Inc. by Dan DeChiaro on Flickr, How many people in this picture?Mendeley is a handy piece of desktop and web software for managing and sharing research papers [1]. This popular tool has been getting a lot of attention lately, and with some impressive statistics it’s not difficult to see why. At the time of writing Mendeley claims to have over 36 million papers, added by just under half a million users working at more than 10,000 research institutions around the world. That’s impressive considering the startup company behind it have only been going for a few years. The major established commercial players in the field of bibliographic databases (WoK and Scopus) currently have around 40 million documents, so if Mendeley continues to grow at this rate, they’ll be more popular than Jesus (and Elsevier and Thomson) before you can say “bibliography”. But to get a real handle on how big Mendeley is we need to know how many of those 36 million documents are unique because if there are lots of duplicated documents then it will affect the overall head count. (more…)

July 27, 2010

Twenty million papers in PubMed: a triumph or a tragedy?

pubmed.govA quick search on pubmed.gov today reveals that the freely available American database of biomedical literature has just passed the 20 million citations mark*. Should we celebrate or commiserate passing this landmark figure? Is it a triumph or a tragedy that PubMed® is the size it is? (more…)

June 22, 2010

Impact Factor Boxing 2010

[This post is part of an ongoing series about impact factors. See this post for the latest impact factors published in 2012.]

Roll up, roll up, ladies and gentlemen, Impact Factor Boxing is here again. As with last year (2009), the metrics used in this combat sport are already a year out of date. But this doesn’t stop many people from writing about impact factors and it’s been an interesting year [1] for the metrics used by many to judge the relative value of scientific work. The Public Library of Science (PLoS) launched their article level metrics within the last year following the example of BioMedCentral’s “most viewed” articles feature. Next to these new style metrics, the traditional impact factors live on, despite their limitations. Critics like Harold Varmus have recently pointed out that (quote):

“The impact factor is a completely flawed metric and it’s a source of a lot of unhappiness in the scientific community. Evaluating someone’s scientific productivity by looking at the number of papers they published in journals with impact factors over a certain level is poisonous to the system. A couple of folks are acting as gatekeepers to the distribution of information, and this is a very bad system. It really slows progress by keeping ideas and experiments out of the public domain until reviewers have been satisfied and authors are allowed to get their paper into the journal that they feel will advance their career.”

To be fair though, it’s not the metric that is flawed, more the way it is used (and abused) – a subject covered in much detail in a special issue of Nature at http://nature.com/metrics [2,3,4,5]. It’s much harder than it should be to get hold of these metrics, so I’ve reproduced some data below (fair use? I don’t know I am not a lawyer…) to minimise the considerable frustrations of using Journal Citation Reports (JCR).

Love them, loathe them, use them, abuse them, ignore them or obsess over them … here’s a small selection of the 7347 journals that are tracked in JCR  ordered by increasing impact.

Journal Title 2009 data from isiknowledge.com/JCR Eigenfactor™ Metrics
Total Cites Impact Factor 5-Year Impact Factor Immediacy Index Articles Cited Half-life Eigenfactor™  Score Article Influence™ Score
RSC Integrative Biology 34 0.596 57 0.00000
Communications of the ACM 13853 2.346 3.050 0.350 177 >10.0 0.01411 0.866
IEEE Intelligent Systems 2214 3.144 3.594 0.333 33 6.5 0.00447 0.763
Journal of Web Semantics 651 3.412 0.107 28 4.6 0.00222
BMC Bionformatics 10850 3.428 4.108 0.581 651 3.4 0.07335 1.516
Journal of Molecular Biology 69710 3.871 4.303 0.993 916 9.2 0.21679 2.051
Journal of Chemical Information and Modeling 8973 3.882 3.631 0.695 266 5.9 0.01943 0.772
Journal of the American Medical Informatics Association (JAMIA) 4183 3.974 5.199 0.705 105 5.7 0.01366 1.585
PLoS ONE 20466 4.351 4.383 0.582 4263 1.7 0.16373 1.918
OUP Bioinformatics 36932 4.926 6.271 0.733 677 5.2 0.16661 2.370
Biochemical Journal 50632 5.155 4.365 1.262 455 >10.0 0.10896 1.787
BMC Biology 1152 5.636 0.702 84 2.7 0.00997
PLoS Computational Biology 4674 5.759 6.429 0.786 365 2.5 0.04369 3.080
Genome Biology 12688 6.626 7.593 1.075 186 4.8 0.08005 3.586
Trends in Biotechnology 8118 6.909 8.588 1.407 81 6.4 0.02402 2.665
Briefings in Bioinformatics 2898 7.329 16.146 1.109 55 5.3 0.01928 5.887
Nucleic Acids Research 95799 7.479 7.279 1.635 1070 6.5 0.37108 2.963
PNAS 451386 9.432 10.312 1.805 3765 7.6 1.68111 4.857
PLoS Biology 15699 12.916 14.798 2.692 195 3.5 0.17630 8.623
Nature Biotechnology 31564 29.495 27.620 5.408 103 5.7 0.14503 11.803
Science 444643 29.747 31.052 6.531 897 8.8 1.52580 16.570
Cell 153972 31.152 32.628 6.825 359 8.7 0.70117 20.150
Nature 483039 34.480 32.906 8.209 866 8.9 1.74951 18.054
New England Journal of Medicine 216752 47.050 51.410 14.557 352 7.5 0.67401 19.870

Maybe next year Thomson Reuters, who publish this data, could start attaching large government health warnings (like on cigarette packets) and long disclaimers to this data? WARNING: Abusing these figures can seriously damage your Science – you have been warned!




April 30, 2010

Daniel Cohen on The Social Life of Digital Libraries

Daniel Cohen is giving a talk in Cambridge today on The Social Life of Digital Libraries, abstract below:

The digitization of libraries had a clear initial goal: to permit anyone to read the contents of collections anywhere and anytime. But universal access is only the beginning of what may happen to libraries and researchers in the digital age. Because machines as well as humans have access to the same online collections, a complex web of interactions is emerging. Digital libraries are now engaging in online relationships with other libraries, with scholars, and with software, often without the knowledge of those who maintain the libraries, and in unexpected ways. These digital relationships open new avenues for discovery, analysis, and collaboration.

Daniel J. Cohen is an Associate Professor at George Mason University and has been involved in the development of the Zotero extension for the Firefox browser that enables users to manage bibliographic data while doing online research. Zotero [1] is one of many new tools [2] that are attempting to add a social dimension to scholarly information on the Web, so this should be an interesting talk.

If you’d like to come, the talk starts at 6pm in Clare College, Cambridge and you need to RSVP by email via the talks.cam.ac.uk page


