Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Monday, March 03, 2008

Buzzwordomics

I see Lars Juhl Jensen has come up with a fun tag cloud of recently popular buzzwords in the biosciences. He calls it a BuzzCloud. The buzzword from the cloud I've noticed most lately is "Quantative Proteomics" ... quantitation is a good goal for the field of proteomics to aim for, since IMHO it doesn't really deserve the -omics prefix. "Omics" tends to imply the possibility of global proteome coverage, which proteomic studies rarely, if ever, achieve. But enough of the side-rants.

The way Lars' BuzzCloud is constructed by extracting phrases ending in -ics, -ology, -omy, -phy, -chemistry, -medicine, or -sciences etc reminded me of a stupid little CGI application I wrote a few years back ... the Biotech company name generator. When you take common prefixes like "Gene-", "Pept-" or "Chemi-" and suffixes like "-omics" or "-agen" etc, it's amazing how often Googling the name turns up a real honest-to-goodness biotech company.

Feel free to comment on any "biotechie" suffixes and prefixes that I should add ... the hardcoded list in the script isn't that long.

Wednesday, September 19, 2007

If pipettes could talk, oh the tales they could tell !

Occasionally impoverished University labs and early career researchers go looking for bargain priced lab equipment on Ebay ... sometimes hard to get but still very useful equipment also comes up. A colleague of mine found this entertaining auction for a P20 Gilson Pipette, much of it written from the pipettes perspective. Here's a quote:

For the purposes of full disclosure, this pipette has NOT resulted in data that has made it into Nature or Science (and in hindsight then the Cell paper may be considered to be a fluke). However, we choose to believe that that is the fault of both the editorial staff of these journals, and many a short-sighted peer-reviewer, rather than the pipette itself. Nonetheless, you may want to calibrate the pipette upon its arrival.
Fluke or not, a sentient pipette that produces Cell papers has got to be worth more than the mere $48 it's currently sitting at ! Sadly (or maybe happily for them), it looks like the seller is leaving bench science and selling up their gear. I've no idea who it is, but they live in the same suburb as me, so there's a good chance this is a fairly well-published senior scientist that I've crossed paths with at some stage. Also, don't miss the (*) footnote at the bottom of the auction info about impact factors ...

Wednesday, September 12, 2007

ARIA verson 2.2 released

I don't usually post about NMR (Nuclear Magnetic Resonance) and structural biology related stuff, but I've always intended to. In this post I'm pulling out all the stops on specialist lingo and assumed background knowledge, so hopefully it isn't too incomprehensible to the non-structural biology crowd :).

ARIA version 2.2 has been released in the last few weeks. ARIA is an automated NOE assignment and structure calculation package, which (in theory) takes some of the pain and slowness out of producing protein (and DNA and/or RNA) structures from Nuclear Magnetic Resonance data. I'll say up front; I haven't tried this version yet, but some of the improvements look exciting.

Here are two new features worth noting ... followed by what I think it all means:


  • The assignment method has been improved with the introduction of a network-anchoring analysis (Herrmann et al., 2002) for filtering of the initial assignments.
  • The integration of the CCPN has been completed. The imported CCPN distance constraints lists can enter the ARIA process for calibration, violation analysis and network-anchoring analysis. The final constraint lists can be exported as well.

In the past I have done some quick and dirty tests comparing the quality of protein structures produced using Aria 2.1 vs. Peter Gunterts CYANA 1.07 and 2.1, using the exact same NMR peak input lists (with slightly noisy data containing a number of incorrectly picked peaks). CYANA always won hands down, assigning more NOE crosspeaks correctly and producing an ensemble of model structures with much lower RMSD and generally better protein structure quality scores (ie using pretty much any decent pairwise pseudo-energy potential, and Procheck). Also, ARIA produced 'knotted' structures which were almost certainly incorrect, while CYANA did not. Other postdocs and students in my former lab had done similar independent tests with ARIA 1.2 vs. CYANA 1.0.7, and had come to similar conclusions.

The disclaimer: It should be noted here that assessment of the quality of an ensemble of NMR structure coordinates can be problematic, and is really the topic of another long post (and probably tens if not hundreds of peer-reviewed journal articles). So saying "CYANA version X is better then ARIA version X" based on the RMSD of the final calculated ensemble is a bit unfair ... in fact using RMSD of the ensemble to gauge structure quality is just plain wrong in this context. In my (unpublished, non-peer reviewed) tests, it is possible that ARIA could be producing high RMSD but essentially 'correct' structures, while CYANA could be producing tightly defined but 'incorrect' structures, but I doubt it. The gap between the output of each program was wide enough to suggest that under real-world conditions where the input peak list contained a number of 'noise' peaks, ARIA was failing to give a set of consistent solutions (probably due to lack of NOE assignments), while CYANA was giving a set of tightly defined structures (which may or may not have represented the 'correct' solution). Other evaluations (protein structure quality measures, Procheck, comparison to known structures of similar proteins) indicated that the CYANA structures were not grossly 'incorrect', so I'd say CYANA was just giving a better defined (ie lower ensemble RMSD) set of plausible solutions.

My gut feeling is that ARIA 2.2 will perform much better than past versions, due to one key feature that has been 'borrowed' from CYANA; the introduction of a network-anchoring analysis. In a nutshell, network-anchoring scores essentially weight distance constraints (or NOE assignments) based on how 'connected' that constraint is within the graph formed by other constraints. This means that in effect a single, isolated constraint pulling two residues on opposite sides of a protein together is down-weighted, while if multiple constraints link those residues (or their neighboring residues) then those constraints are considered more trusted and hence weighted heavier. For better or worse (usually better), this score simulates what the human NMR spectroscopist would do when assigning NOE crosspeaks manually ... usually two residues in contact will show multiple NOE crosspeaks connecting them and involve multiple different nuclei, however a single lonely NOE between two nuclei which are distant from eachother in the primary protein sequence is heavily scrutinized and regarded with suspicion since it is likely to be mis-assigned. I'm very keen to test ARIA 2.2 on my old data set and see if I'm actually right (I may be able to try it with network anchoring turned on, and off, and see just what sort of contribution that score is making).

Another completed feature, the integration between ARIA and the CCPN libraries/analysis package should also be a big plus. I haven't used the CCPN analysis software yet, but a few years ago I wrote some code to help make CYANA and the Sparky NMR assignment program work together better. The result was functional, but very hackish (and I'm probably the only person in the world who understands how it was intended to be used, since I still haven't got around to writing any documentation. Naughty, naughty). CCPN + ARIA may turn out to be the better option for spectral analysis and structure calculation in the future, as opposed to my currently preferred Sparky + CYANA combination.

I'm really itching to find a good reason to do an NMR structure project now ... back to work !!

Friday, May 25, 2007

Cleaning up the cesspool that is the PDB

Well .. maybe cesspool is a little strong ... there's a lot of great data in the Protein Data Bank, it's just that in the early days it was allowed to grow very large without enforcing better standardization of the data. Things that are being fixed include updating citations for structures from "To be published" to the actual publication if it exists (with PubMed ID), linking to sequence databases (ie UniProt), bringing atom names to standard IUPAC nomenclature (Hooray!!) and loads of other things I haven't mentioned. Don't fret ... none of the raw experimental data or coordinates are going to be changed :)

From the PDB remediation overview document (pdf):

When the RCSB PDB first addressed the remediation issues in 1998, it was with the intention of providing a uniform and consistent content across all formats. It was surprising and very disappointing to find that many PDB users at the time strongly objected to any changes in the released PDB entries, even if these changes addressed serious but correctable errors (e.g., consistency between chemical and coordinate sequence). As a result of this prevailing attitude toward changes in PDB format entries, the RCSB PDB released its corrections in a new set of mmCIF format data files and left the data in PDB file format unchanged. Since that initial release of mmCIF data, new data items and uniformity corrections have been added to the released mmCIF data files.

I've used coordinates from PDB format files for a lot of things over the years, but I've got to admit, I've never used an mmCIF file. The PDB file format is almost always supported by all legacy (and recent) structural biology analysis software, while using mmCIF is rarely an option (unless it's converted to PDB format first). If I'd known the mmCIF versions in the database have been 'remediated' I may have been more inclined to use them (or the somewhat equivalent XML/PDBML files) for some tasks, since the non-uniformity in atom naming in legacy PDB files can become a royal pain in the butt ....

Anyhow, everyone has until July 2007 to check out the new remediated files before the 'mainline' PDB changes over and provides these by default. All new structure releases will follow the remediated format after July. The old versions will still remain available ... but who would want them ... we are getting standardized goodness !!

Thursday, December 21, 2006

I submitted my PhD thesis, and all I got was this crappy balloon



Well, it's not really all that crappy ... the balloon is a nice happy gesture to mark the occasion. I even got to pick the colour. It took me far too long to write and submit this thing, it's a relief to not have to look at it for a few months. My thesis, entitled "The structure of outer mitochondrial protein import receptors", may well be the first Creative Commons Licensed thesis submitted in Australia (although I doubt it) . Once it's been examined (hopefully I pass), I'll release it online and allow everyone to poke holes and rip it to shreds (or they can poke at the associated peer reviewed publication instead .. unfortunately it's probably not Open Access).

Afterthought: One thing that slowed down the final submission was the bloody Latex typesetting. I'm a Latex novice, and while I really like the final result, Latex is an abomination (much like Perl).



Update, 15th October, 2007.

I've finally got around to submitting the final post-examination version of my thesis to the University of Melbourne ePrints server. You can get a PDF copy of my thesis here. I used the xmpincl Latex macro to embed XMP Creative Commons licensing data into the final PDF version generated by pdflatex. I probably didn't get the format of the licensing XML exactly right, but I'm sure it will be good enough that search engines can (or will one day) determine the correct licensing for the work.

Wednesday, November 29, 2006

First Online EMBL PhD Symposium

This looks interesting ... the First Online EMBL PhD Symposium, a sort of 'online' conference for the life sciences. Everybody with a scientific background is invited to participate. Registration is free.

The programme (Career Development Session, Omics Session / Systems Biology, Scientific Communication 2.0 and Participant's Contributions) and speakers list makes it look sort of like a "Biology 2.0" conference.

Apart from the (possible) IRC sessions, hopefully the fact that everything is stored as video/audio + comments on their content managment system means the 'inconvenient' timezone in Australia won't limit my participation too much.

(via the worldwide bioinformatics cabal :), Neil via Pedro, Roland and Stew)

Wednesday, November 22, 2006

International Genetically Engineered Machine competition videos

The 2006 iGEM Jamboree (International Genetically Engineered Machine competition) happened at the start of this month. This is a synthetic biology 'competition' where teams of talented undergraduates from around the world engineer an organism for a specific purpose ... like E. coli that produce mint or banana smell, or form simple logic gates the could potentially be used to make a 'biological computer'.

They are encouraged to use BioBricks from the Registry of Standard Biological Parts, which at the moment is essentially comprised of series many well-characterized DNA constructs (promoters, repressors, selection markers, lots of fluorescence protein coding sequences, etc) with standardized restriction site that can be mixed and matched to produce new and interesting behaviours in bacteria, yeast or mammalian cells. BioBricks are sent out to teams in in 96-well format, so everyone has a good basic set of starting components.

Videos of the student presentations have finally turned up on Google Video. (Unfortunately, the videos only show the speakers, not the slides for the presentation ... which makes some parts pretty hard to follow).

I watched the presentation by the University of Arizona team. They printed bacteria onto paper using a stock-standard inkjet printer, with the ink simply removed from the cartridges and replaced with a solution of bacteria. They could then tranfer this to agar plates to grow in whatever pattern they printed. Very simple, but inkjet hardware hacking crossed with molecular biology is just plain cool. As a side discovery, they noticed some weird fractal patterns in colonies under the confocal microscope, apparently based on variation in the fluorescent protein expression level of cells in a single colony.

I wonder how much interest there would be from undergrads (and their supervising acedemics) to start an Australian iGEM team for 2007 ? Funding would also be a tricky issue, as always.

Tuesday, November 14, 2006

Protein structure sculpture

Check out these amazing protein structure sculptures by Julian Voss-Andreae. The GFP (Green Fluorescent Protein) in shiny steel is particularly striking.

He has even provided instructions on how to construct your own [pdf] ... Before the advent of molecular graphics on computers, making physical models similar to this was what crystallographers (and Linus Pauling) did to build protein models.

I've gotta find time to make one of these ... the question is, do I make something of personal significance, or a choose a structure that is actually a little more challenging ?


Thursday, November 02, 2006

The SDS-PAGE Hall of Shame

For the uninitiated .. SDS-PAGE is a method that biochemists use to separate mixtures of proteins (and sometimes other biomolecules, like short pieces of DNA). It's a really useful technique, and most of the time it works perfectly, giving a nice little 'ladder' of bands with large proteins at the top and the smallest ones at the bottom.

Occasionally, something goes wrong ... enter the SDS-PAGE "Hall of Shame".

This is a hilarious gallery of botched SDS-PAGE gels, which doubles as a useful trouble-shooting guide.

Over the years, despite trying my best to avoid it, I've occasionally run gels which have suffered from most of these problems. This one resembles the gel I ran yesterday ... I was in a hurry and turned the voltage up too high. Normally I'd get away with it, but this time the cooling wasn't adequate enough. It doesn't pay to rush these things.

(For the non-scientists: SDS-PAGE is an acronym for for sodium dodecyl sulphate polyacrylamide gel electrophoresis ... sorry you asked ?)

Friday, October 06, 2006

Ways of seeing the world ...

As blogged by several others ... the Science Magazine Science and Engineering Visualization Challenge winners have been announced.

I'm a big fan of innovative visualization ... sometimes it can the difference between seeing something meaningful in data, or just seeing noise. I was initially disappointed at the large number of finalists that are purely 'educational' in nature, rather than providing novel representations of 'raw' data. But, the National Science Foundation site explains it: "The spirit of the competition is for communicating science, engineering and technology for education and journalistic purposes.". It's important to have these types of events pitched so that the 'general public' (i.e. non-scientists, or scientists of vastly different fields) can get something out of it .. after all, they are often indirectly providing the funds for a lot of the research, and in some cases a cool image is all they get for their tax dollars. Nonetheless, I'd like to see a competition dedicated to innovative visualization of new experimental or statistical results, with no opening for purely 'textbook' style educational compositions (I bet there's one or two out there ... comments anyone ?).

Also, congratulations goes to one of the (tied) 1st place winners in the non-interactive multimedia section, Drew Berry and François Tétaz at The Walter and Eliza Hall Institute (WEHI) and Jeremy Pickett-Heaps at the University of Melbourne. It's nice to see some local Aussies getting some recognition.

Friday, September 29, 2006

Combio 2006, last day roundup

Yesterday I breezed into Brisbane for the last day of the Combio 2006 meeting, to catch some talks, and make a showing to accept an award from the ASBMB.

Neal Saunders has been posting summaries of this meeting in Brisbane on his blog, so I thought I'd give my take on the last day too.

Here's are my highlights:

In David Claphams talk on transient receptor potential (TRP) ion channels, I learnt that menthol feels cold because it binds to an activates a TRP channel involved in cold sensing. Think about that next time you taste that cool minty freshness. (I woke up a 4 am to fly to Brisbane. The brain wasn't really kicking over just yet).

In the "Molecular Basis of Disease and Drug Design" session, K. Krause gave a very honest and entertaining talk on what he termed his "Night Science". ("Day Science" is the stuff that works out nicely, shows logical progression with no nasty inconsistencies or loose ends and gets talked about at plenary lectures. "Night Science" is the stuff that doesn't work out as well as we'd like .. it's confusing, there are loose ends and inconsistencies, despite carefully doing all appropriate controls. Not to be confused with "Bad Science"). Krause and his group were unlucky enough to find that a lead compound discovered through an in silico screen, which initially appeared to be a great inhibitor of alanine racemase, turned out to in fact be a potent inhibitor of another enzyme in their coupled assay. I wasn't inhibiting their target well at all (doh!).

There were actually a few examples of some somewhat disturbing results from in silico screens in this session, which I've seen similar examples of a few times before. Researchers do an in silico screen, and find some top-ranking hits, one or two of which are also good inhibitors in an assay. The co-crystal structure is solved, and reveals that the compound is not actually binding in anything like the conformation that the computational docking predicted (sometimes not even the same site). What is going on here ? Is it just the fact that in twenty random compounds one will turn out to be a weak inhibitor ? Unlikely, since then high-throughput real-world screens would have a much higher hit rate. Is it that the computational docking is half right, fitting one fragment of the compound which has high affinity well, and the other non-binding or weak binding half doesn't matter ? Probably more likely, but it still doesn't explain the cases where the compound binds in a completely unpredicted site. Food for thought: maybe many docking scoring functions for small molecules are good at selecting generally sticky molecules ...... (I don't do this kind of work directly, so I'm really an ignoramus on the issue).

I also went to the "Cancer - Emerging Drug Targets" session. Andrew Scott from the Ludwig Institute for Cancer Research presented some really encouraging results of early clinical trails for an EGFR antibody, and Michelle Haber of the Children's Cancer Institute Australia presented some results from two cell based assays, where 'high-throughput' screens have identified some inhibitors of the N-myc oncogene, and a drug efflux pump (MRP) inhibitor. I'd never really thought about it, but apparently those pesky cancer cells up regulate this efflux channel and actively pump out anti-cancer drugs, in a similar way to some parasites that become multi-drug resistant.

In the final plenary lecture, Nick Proudfoot told us about his work on transcriptional termination. It's still too early for the textbooks, but it looks like transcriptional terminators bind at the termination site and near the promoter regions in a lot of cases, turning genes into physical 'loops'. Whether this helps the RNA polymerase jump from the end of a gene straight back to the start to make the next mRNA transcript is still not proven, but it's an attractive model.

Combio is always a bit of an eclectic mix, but if you take it in the right frame of mind it can be good fun, and a nice way to broaden the scientific horizons a little. Needless to say, I slept like a log after all that.