Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Tuesday, August 04, 2015

Hacking Zebrafish thoughts

The last lab from Scalable Machine Learning with Spark features a guest lecture by Jeremy Freeman, a professor of neuroscience at Janelia Farm Research Campus.

His group produced this gorgeous video of a living zebrafish brain. Little fish thoughts sparkle away, made visible by a technique called light-sheet flourescent microscopy in which engineered proteins that light up when the neurons fire are engineered into the fish.

The lab covers principal component analysis in a lively way. Principal components are extracted from time-series data and mapped onto an HSV color wheel and used to color an image of the zebrafish brain. In the process, we use some fun matrix manipulation to aggregate the time series data in two different ways - by time relative to the start of a visual stimulus and by the directionality of the stimulus (shown below).

The whole series of labs from the Spark classes was nicely done, but this was an especially fun way to finish it out.

Check out the Freeman Lab's papers:

Monday, January 12, 2015

Brave Genius

Brave Genius is an unlikely dual biography of a biologist and a writer who shared a friendship and a common philosophy. Both were active in the French resistance to the German Occupation and both would later receive a Nobel prize. Sean B. Carroll forges an inspiring story from seemingly incongruous elements: the desperate defiance of a few in an occupied country, the exhilarating pursuit of an open scientific question, and a lonely stand on the moral high ground.

In 1940, Jacques Monod was a newly married father of twins and a researcher at the Sorbonne. Albert Camus, having already published a couple of books of essays, departed his native Algeria for France in March of that year to find work.

On May 10 1940, German troops crossed into Holland and Belgium. Panzers raced towards the Atlantic coast severing Allied lines and stranding French and British troops in the low countries. French defenses collapsed and Germans arrived in an undefended Paris on June 14. The armistice signed on June 22nd marked the beginning of four years of occupation.

During those years, Camus edited and wrote for the underground newspaper Combat urging resistance to the occupation. As the tide of the war turned, Monod organized sabotage attacks and armed resistance ahead of the approaching liberators.

“I have always believed that if people who placed their hopes in the human condition were mad, those who despaired of events were cowards. Henceforth, there will be only one honorable choice: to wager everything on the belief that in the end words will prove stronger than bullets.” Camus, Combat (November 30, 1946)

François Jacob, André Lwoff and Jacques Monod were awarded a Nobel prize in 1965 for their work on the control of gene expression, elucidating the regulation of the lac operon by which bacteria switch on metabolism of the sugar lactose.

In his writing, Camus confronts the absurdity of the human search for clarity and meaning in a world that offers only indifference. The attempt to derive meaning and morality without resort to mysticism links Camus's philosophy to Monod's scientific work, which provided some of the first direct evidence that life is mechanistic rather than the result of some magical "vital force" and that its workings could be understood.

“The scientific approach reveals to Man that he is an accident, almost a stranger in the universe.” Monod, in On Values in the Age of Science (1969)

“One of the great problems of philosophy, is the relationship between the realm of knowledge and the realm of values. Knowledge is what is; values are what ought to be. I would say that all traditional philosophies up to and including Marxism have tried to derive the 'ought' from the 'is.' My point of view is that this is impossible.” Monod

Carroll, a biologist himself, embeds philosophy and science into the personal lives of his protagonists and the geopolitical events unfolding around them. Both men did brilliant work in the darkest of times, and did so not by retreating but by fully engaging at great risk with the struggles that faced them. The book serves as a warning of what happens when good people overlook the malfeasance of their leaders, but also as confirmation of the resilience of intellect, creativity and humanity.

More

Tuesday, June 17, 2014

Ada's Technical Books

If you're in Seattle, you owe it to yourself to spend some time in Ada's Technical Books in my old neighborhood. It's on the top of Capitol Hill on 15th between Republican and Harrison.

The shop features books on all sorts of techy topics - science, math, engineering and, of course, computers along with various nifty maker-type gadgets and a cafe full of tasty treats. Geek heaven!

Tuesday, April 22, 2014

The damaging effects of hypercompetition

In Rescuing US biomedical research from its systemic flaws (PNAS 2014), authors Bruce Alberts, Marc Kirschner, Shirley Tilghman and Harold Varmus recount the ways in which research funding is broken, in particular for life sciences.

PI's pour vast amounts of time and energy into competing for a shrinking pool of funds. Pressure for results causes a shift to short-term thinking and engenders a conservatism poorly suited to producing breakthroughs. In much the same way as the corporate sector abandoned fundamental research in favor of product development, the current funding climate overvalues "translational" research over basic science. The situation is especially dire for young scientists facing the prospect of years as an underpaid and overworked post-doc with faint hopes of ever landing a faculty position. The internet is littered with goodbye-academia letters. [1 2 3 4 5 6]

The Alberts paper proposes a few solutions, some sensible and one in particular that doesn't make much sense to me.

Broadening career paths for scientists is a sound idea. In this respect, life science can borrow from information technology. Computer science departments and the tech industry have a long history of exchange. The barriers in biotech are higher, but the flow of people and ideas between academics and industry can only be a good thing. Probably the biggest factor in diversifying away from the shrinking tenure track, is matching expectation with reality.

Managing expectations seems more sensible than trying to match supply and demand by restricting entry into PhD programs based on the reasoning that the system is training too many PhDs. Isn't too many PhD's a good problem to have? Is educating people a bad thing? Scientifically trained people are a tremendous asset. If we're not using a valuable resource effectively, having less of that resource doesn't seem like an ideal solution. With some creative thinking, productive uses for that talent can and are being found.

I'm not convinced that the pyramid shaped structure of science is the problem. Harnessing the idealism, curiosity and naive overconfidence of youth works. What doesn't work in absence of perpetual growth is the expectation that everyone at the bottom of the pyramid will eventually be at the top. But, what's wrong with a system in which people are supported for a time on research grants then continue onto many diverse career paths?

Getting the mix of competition and camaraderie right can be the difference between a situation that's nurturing rather than toxic. Economics teaches that incentives matter. But, re-engineering science will require a nuanced understanding of what those incentives are, both in crass terms of money and prestige but also the more subtle psychology around freedom, inner drive and higher purpose.

Since nobody asked me

If I was in charge of funding science, I'd place the majority of my bets on longer term grants for small focused labs where the PI does hands-on science and training of new scientists. Longer term grants of 5 years or so would give researchers time to do actual science in teams of 2-8 people. The article recognizes the success of the Howard Hughes Institutes model of explicitly encouraging scientists to pursue risky ideas - "the free play of free intellects".

I'd be reluctant to fund 40 person labs because neither mentoring nor creative thinking scales up to that size particularly well. I'd avoid large multi-center grants. Funding agencies seem to favor this sort of thing perhaps as a safer bet, they're a recipe for infighting and politicking. They risk getting all the disadvantages of both collaboration and competition with little of the benefits.

Research has a great track record of paying off in the long run. But, Bill Janeway, author of Doing Capitalism in the Innovation Economy says that there has to be a surplus in the funding system to absorb the unavoidable costs of uncertainty, both in terms of when the payoff comes and who captures the gains. In the past, this surplus has come from monopolistic corporations like Bell or Xerox, from the taxpayer, or from flush venture capital. In The Great Stagnation, Tyler Cowen posits that we've picked the low-hanging fruit of the current wave of technology - that we're in a period of stagnation while the last wave is fully digested and the stage can be set for the next wave. To that end, maybe some hedge-fund wizard can cook up some financial innovation for reliably financing risky, long-term projects with unevenly distributed payouts.

[1] Goodbye academia, I get a life
[2] Goodbye Academia
[3] Goodbye academia? Hello, academia
[4] The Big Data Brain Drain
[5] Why So Many Academics Quit and Tell
[6] On Leaving Academe
[7] Science: The Endless Frontier, Vannevar Bush, 1945
[8] The Decline of Unfettered Research, Andrew Odlyzko, 1995

Thursday, January 09, 2014

Guide to Open Science

Open science is the idea that scientific research should be openly and immediately shared.

…when the journal system was developed in the 17th and 18th centuries it was an excellent example of open science. The journals are perhaps the most open system for the dissemination of knowledge that can be constructed — if you’re working with 17th century technology. But, of course, today we can do a lot better.

Refactoring science to take advantage of digital technology is a what Michael Nielsen, quoted above, calls, “Building a better collective memory.”

The problems with the current system reinforce the case for change. The reproducibility crisis, the prevalence of unreliable findings in published research described by Ioannidis, is giving science a credibility problem. Negative results and replications go unpublished. Peer review is uncompensated and sometimes ineffective. The publication process is slow. Over time, data gets lost. Artificial barriers impede potential synergy among researchers, or between research and industry, students, or patients.

The scientific paper has become something of a “choke point”. Unbundling the functions of a paper might allow more degrees of freedom for progress and innovation.

As science and technology progress, the amount of accumulated knowledge that must be mastered to get to the frontier increases. "If one is to stand on the shoulders of giants, one must first climb up their backs, and the greater the body of knowledge, the harder this climb becomes." This "burden of knowledge" changes the effective organization of innovative activity. As the low hanging fruit is depleted, research becomes more specialized and team oriented.

In this new regime, the strategy that comes into play is leveraging the “scale of the communication” of the Internet. The experience of the Polymath project inspired Gowers and Nielsen to write “mass collaboration will extend the limits of human problem-solving ability.”

Not everything digital is necessarily open, but the interesting developments are concentrated at the intersection of open and digital science - in the interaction between technology and the redesign of centuries old institutions.

Overview

Here are two great places to start:

Open Access

In the primordial days of the web, known to some as 1991, physicist Paul Ginsparg started the arXiv, an on-line preprint library, superseding a mailing list. Today, arXiv offers, "Open access to 901,072 e-prints in Physics, Mathematics, Computer Science, Quantitative Biology, Quantitative Finance and Statistics". Researchers enjoy the convenience and increased exposure enough to negotiate exceptions to journal copyright agreements or to ignore them. The arXiv is currently maintained by the Cornell University Library and supporting member institutions at a cost of $826,000 per year.

The economics of information delivery have changed. But, there's also an ethical argument for making the fruits of publicly funded research public.

In 2000, the National Library of Medicine launched PubMed Central a digital archive of NIH funded biomedical and life sciences journal literature. Open access is particularly strong at the NIH where PLoS Computational Biology founding editor Philip E. Bourne was recently named Associate Director for Data Science. The stories of PubMed Central and PLoS are tightly bound. Harold Varmus, now NCI director and formerly director of the NIH in the late 1990's, co-founded PLoS along with Patrick Brown and Michael Eisen partly out of frustration with the resistance met by Pubmed Central.

There's a struggle going on in the US congress over whether to extend the NIH's open access policies to the rest of federally funded research. Publishers, meanwhile, lobby to block or limit open access mandates.

Open access journals give away content for free, typically recouping costs by charging publication fees and ads. This may incentivize weaker peer review. Science magazine's open access sting, in which a bogus paper was accepted by many open access journals, has some hallmarks of a FUD (fear uncertainty and doubt) campaign, but also shines a light on a real problem. The limitations of peer review, it is argued, might be addressed by a combination of post-publication review and algorithmic filtering, ensuring that quality research rises to the top.

eLife aims for the highest standards. Its founder, Nobel laureate Randy Schekman, wrote recently of the “inappropriate incentives” created by the “luxury” journals.

Open access publishers
  • BioMed Central, the largest open access publisher; founded in 2000; now owned by Springer
  • Elementa Science of the Anthropocene; Publishing original research on the Earth's physical, chemical, and biological systems
  • eLife, a high-impact open access journal formed in a joint initiative of the Wellcome Trust, the Howard Hughes Medical Institute and the Max Planck Society.
  • F1000 Research, a post-publication reviewed journal from Faculty of 1000
  • Frontiers, community-driven journals with it's own peer-review system
  • PeerJ, a peer reviewed journal with a membership model
  • PLOS ,launched October 2003

The DOAJ lists almost 10 thousand open access journals.

Priem and Hemminger propose Decoupling the scholarly journal - unbundling the functions that have been tightly coupled within traditional academic journals since the days of Henry Oldenburg in the 17th century, and freeing up the parts for experimentation.

Traditional functions of a paper
  • Archival: storing scholarship for posterity.
  • Registration: time-stamping discoveries to establish precedence.
  • Dissemination: distribution of scholarly products.
  • Certification: assessing rigor and significance of contributions

Re-engineering the journal makes new things publishable, negative results, for example, which science calls “the neglected stepchild of scientific publishing”. There are further possibilities yet, outside to format of the paper.

Data and Code

Potentially valuable products of scholarship include: papers, data, software, figures, posters, reviews, talks, slides, instructional material.

Data and source code, in particular, are indispensable for replication and should, in theory, be highly re-usable. Data archives like GenBank and GEO exist for this reason and more scientific software is finding its way into GitHub or other public source code repositories.

Data are a classic example of a public good, in that shared data do not diminish in value. To the contrary, shared data can serve as a benchmark that allows others to study and refine methods of analysis, and once collected, they can be creatively repurposed by many hands and in many ways, indefinitely.
Permanent archives for published research data would allow us to write an amendment to the centuries-old social contract governing scientific publishing and give data their due.

Open Data and the Social Contract of Scientific Publishing, BioScience 2010 Todd J. Vision

Longer term, lots of work is still needed in the ongoing project of creating machine readable data standards, open APIs, rich metadata, and semantic integration.

The exact meaning of a “GitHub for science” has many interpretations. There are a handful of platforms looking to implement some version of that idea.

Platforms
  • Arvados spun out of work done in George Church's Lab at Harvard Medical School to support the Personal Genome Project. According to developer Alexander Wait Zaranek, Arvados is "an open-source platform for data-management, analysis and sharing of large biomedical data sets spanning millions of individual humans across numerous organizations and eventually encompassing exabytes of data.".

    Technology: Ruby. SDK libraries in Perl, Python and Ruby. "Keep", a content-addressable distributed file system developed by the Arvados team. Docker.

  • DataDryad is a curated general-purpose repository that makes the data underlying scientific publications discoverable, freely reusable, and citable. Dryad has integrated data submission for a growing list of journals; submission of data from other publications is also welcome.

    Leadership: Todd J. Vision at the Department of Biology at UNC and a board with deep publishing background.

    Technology: Java. Dryad is built on the open source DSpace software for building open digital repositories.

    Funding: NSF and data publishing charges.

  • Figshare is an open science platform for publishing figures, data sets and other research outputs such that they are discoverable, sharable and citable.

    Technology: source not available.

    Support: FigShare is now supported by Digital Science, employer of FigShare founder Mark Hahnel, and a division of Macmillan Publishers which also owns the Nature Pulishing Group.

  • The Open Science Framework is part network of research materials, part version control system, and part collaboration software. The purpose of the software is to support scientific workflow and help increase the alignment between scientific values and scientific practices.” The OSF is the flagship product of the Center for Open Science founded by Brian Nosek and Jeffrey Spies, who presented OSF at SciPy2013.

    Technology: Python; Git on the back end. Hosted on Linode; Source on GitHub promised soon.

    Funding: The Laura and John Arnold Foundation, Alfred P. Sloan Foundation, the Templeton Foundation, and an anonymous donor

  • RunMyCode enables scientists to create companion sites for papers holding code or data. Aimed at reproducibility and compatibility with journals.

    Leadership: Victoria Stodden, Christophe Hurlin, and Christophe Perignon are co-authors on the paper.

    Funding: Sloan, French universities and research agencies.

  • Synapse by Sage Bionetworks is a platform for sharing data, code, and narrative description linked together with provenance. Synapse has been used in conjunction with the DREAM challenges in computational biology and the in The Cancer Genome Atlas (TCGA) Pan-Cancer project.

    Technology: Java back end, javascript web front end, Python client and R client for programmatic access.

    Leadership: Developed by Sage Bionetworks, led by Stephen Friend, the sage team has expertise in drug development, machine learning, oncology, open access and data governance. Engineering group led by Michael Kellen.

    Funding: Sloan, Washington State Life Sciences Discovery Fund, NCI, NIH

By offering collaborative features or the ability to execute code, these platforms aspire to be more than just data repositories, of which there are reckoned to be hundreds by the Registry of Research Data Repositories (re3data.org).

Journals for data

Burying data in supplementary material is a better archival strategy than keeping data on a post-doc's laptop archival strategy, backed up by USB key in post-doc's shirt pocket. A couple big journals are experimenting with publication of data sets in a more prominent way.

Nature Publishing Group's Scientific Data an open-access, online-only publication for descriptions of datasets to launch in May 2014. Data stored in separate repositories. Descriptions will have a narrative part and a structured part, the Data Descriptor metadata, consisting of ISA-Tab formatted information about samples, methods and data processing steps.

GigaScience, a new on-line publication from from BioMed Central, links manuscripts with associated data sets archived in GigaDB, a database hosted by BGI. (src: github.com/gigascience)

The GigaDB paper, starts with a nice quote from Tim Berners-Lee: "Data is a precious thing and will last longer than the systems themselves”. It describes several use cases: an epigenomics pipeline, the mouse methylome dataset, 15 Tb of hepatocellular carcinoma tumor/normal sequence data. GigaDB aims to "accept any large-scale data including proteomic, environmental, and imaging data" and states the goal of "working with authors to make the computational tools and data processing pipelines described in their papers available and, where possible, executable."

Computing infrastructure

If the functions of the journal along with new functions are going to be performed by a loosely connected set of services on the web, a lot of cyberinfrastructure will be needed to connect it all together. One example is ORCID, a service that assigns globally unique IDs for researchers. Similarly, DataCite issues DOIs (Digital Object Identifiers) for data sets providing a convenient and unambiguous way of citing data.

Another bit of programmatic glue, rOpenSci is a set of R packages that wrap APIs to a variety of scientific data repositories and journals. These packages interface between R and several of the tools mentioned here: Dryad, Figshare, Mendeley, NCBI, DataCite, altmetric.com, PLoS and Pubmed, etc.

Consent

Extracting maximal scientific value from data isn't entirely a technical problem. There are sticky legal and ethical issues around the privacy and consent of research subjects, licensing, publication embargoes, the interface between academic and commercial.

Altmetrics

I was once told by a successful P.I. in reference to a scientific database project, "I don't want to be a librarian." Data curation, though undoubtedly valuable, is labor intensive and doesn't always generate much career momentum for a scientist.

Incentives matter. If publishing is the currency of the realm, how does a scientist get credit for curating data or supporting well engineered software? Altmetrics is a catch-phrase for a set of alternate ways of measuring and incentivizing academic progress, including creation of artifacts such as data sets and software.

  • ImpactStory is a project by Jason Priem and Heather Piwowar to compile and display a broad range of scientific contributions including everything from papers to source code to twitter feeds.

    Here are a handful of ImpactStory profiles:

    Technology: Python and javascript source on GitHub

    Funding: Sloan foundation, NSF

Valuing diverse research products

In Altmetrics: Value all research products (Nature 2013), Heather Piwowar notes that funding agencies are increasingly conscious of the value of diverse research products and claims, “Altmetrics give a fuller picture of how research products have influenced conversation, thought and behaviour.”

Altmetrics.com, also under the Digital Science umbrella, compiles web analytics for papers. Plum Analytics rates individual researchers and whole departments. Microsoft's Academic Search ranks the top authors, publications, conferences, and journals in a given field. Google Scholar keeps profiles for individual researchers.

Given plenty of performance data at the individual level in a highly competitive field, the result will be a Moneyball of science, in which scientists rack up statistics like baseball players.

But, if alternative metrics are to have real value, it will be in re-aligning incentives in scholarly work towards the desired outputs: not just solid reproducible results, but all sorts of specialized supporting contributions that are often undervalued today.

Social Science

There are two contenders for the Facebook or Linked-In of science, although whether such a thing is needed remains to be seen. ResearchGate seems to have made the pleasing insight that the co-authorship graph is a ready-made social network.

Mendeley is a reference manager and academic social network which surpassed 2 million users before being bought by Elsevier in April of 2013 for something like $69M-$100M. With the buyout, some believe an opportunity to create an iTunes for scientific papers was lost.

The link to openness is this: Whatever value comes from these sites (also, to some extent, from open lab notebooks) would be from the ability to build collaboration across organizational boundaries.

Collaborative tools

The journal Push whose topic is described mysteriously as Research & Applied Theory In Writing With Source is an experiment in doing something really different. The entire journal, the source code for its web site, every issue and all articles live in a single GitHub repository and adopt the workflow that implies. Sadly, Push is empty, so far.

SciGit, a collaborative tool for writing scientific papers, is another experiment in layering functionality on top of version control. Authorea is a collaborative scientific typesetting system, also with git on the back end.

Integrating code and prose is great way to communicate and replicate computational methods. In the Python community IPython notebooks fills this niche, while in R there's Knitr and Sweave. These executable documents can be version controlled, collaboratively edited and automated.

If papers are transforming into web-native formats, peer review should follow suit.

  • Hypothes.is will be an open platform for the collaborative evaluation of knowledge. Emphasizes reputation and evaluation using an annotation tool for web pages.

    Team: heavy on start-up experience and light on scientists.

    Technology: Python, javascript, coffeescript. Source on GitHub.

    Funding: Sloan, Shuttleworth and Mellon Foundations.

  • Publons is a platform for open quantifiable peer review, either pre- or post-publication, inspired by Stack Overflow. The founders describe their ambitions for the project in The Future of Academic Research.

    Technology: Postgres, Python/Django, Boostrap, and D3.js. No open repos.

    Team: Based in Wellington, New Zealand. Emerged from startup accelerator Lightning Lab.

Organizations

A number of organizations advocate and develop tools for open science.

Parting thoughts

That's my attempt to get a handle on open science. Of course, it's incomplete and lopsided. In wrapping up, here are a few stray thoughts:

An open system interfaces and interacts with other systems. Network effects kick in as more capabilities plug in. A virtual lab, made up of people distributed across the globe could crowdsource funding (Microryza, Consano), farm out experiments (Science Exchange, Assay Depot), and develop an open challenge (DREAM, nature open innovation pavilion) to analyze the data.

In the future, the humans at the center of that virtual lab may be replaced by an algorithms. “...robotic scientists crawling the web of literature extracting, combining and computing on existing research, deducing new knowledge, finding gaps and inconsistencies, and proposing experiments to resolve them.

Many open science pioneers cite the open source software movement as an inspiration and as a foundation on which to build upon. Open science needs open source tools. Reproducibility practically requires it. The success of open source demonstrates that the open model can work.

In order to succeed, innovation needs to be compatible with and complimentary to existing infrastructure. As tempting as it is to start tossing Molotov cocktails, the better strategy is to quietly pry off functionality and perform it competently, offering improvements that are hard for closed incumbents to match.

Open science is fundamentally about increasing trust, within science but also the public and policy-makers' trust in science. It may offer a solution for reproducibility problems, and maybe a better realization of a science that aspires to truth in a way that is dependable and verifiable.

Thanks to Brian Bot, my colleague at Sage Bionetworks, who introduced me to many of the open science projects linked here and for reading a draft of this post and giving insightful feedback.

Wednesday, December 11, 2013

DREAM conference

Apparently, live-blogging isn't my thing. This post has been sitting on my desktop getting stale for entirely too long...

I'm just back (a month ago) from the conference for the DREAM challenges in Toronto. It was great to see the Dream 8 challenges wrap up and another round of challenges get under way. Plus, I think my IQ went up a few points just by standing in a room full of smart people applying machine learning to better understand biology.

The DREAM organization, led by Gustavo Stolovitzky, partnered with Sage Bionetworks to host the challenges on the Synapse platform. Running the challenges involves a surprising amount of logistics and the behind-the-scenes labor of many to pose a good question, data collection and preparation, provide clear instructions, and evaluate submissions. Puzzling out how to score submissions can be a data analysis challenge in it's own right.

HPN

The winners of the Breast Cancer Network Inference Challenge applied a concept from economics called Granger Causality. I have a thing for ideas that cross over from one domain to another, especially between economics to biology. The team, from Josh Stuart's lab at UCSC calls their algorithm Prophetic Granger Causality, which they combined with data from Pathway commons to infer signaling pathways from time-series proteomics data taken from breast cancer cell lines.

The visualization portion of the same challenge was swept by Amina Qutub's lab at Rice University with their stylish BioWheel visualization based on Circos. Qutub showed a prototype of a D3 re-implementation. I'm told that availability of source code is coming soon. There's a nice write-up Qutub bioengineering lab builds winning tool to visualize protein networks and a video.

BioWheel from Team ABCD writeup

Whole-cell modeling

In the whole-cell modeling challenge, led by Markus Covert of Stanford University, participants were tasked with estimating the model parameters for specific biological processes from a simulated microbial metabolic model.

Since synthetic data is cheap, a little artificial scarcity was introduced by giving participants a budget with which they could purchase data of several different types. Using this data budget wisely then becomes an interesting part of the challenge. One thing that's not cheap is running these models. That part was handled by BitMill a distributed computing platform developed by Numerate and presented by Brandon Allgood.

Making hard problems accessible to outsiders with new perspectives is a big goal of organizing a challenge. One of the winners of this challenge, a student in a neuroscience lab at Brandice, is a great example. They're condidering a second round for this challenge next year.

Toxicogenetics

Yang Xie's lab at UT SouthWestern's QBRC added another win to their collection in the Toxicogenetics challenge. Tao Wang revealed some of their secrets in his presentation.

From Tao Wang's slides

Wang explained that their approach relies on extensive exploratory analysis, careful feature selection, dimensionality reduction, and rigorous cross-validation.

Keynotes

Trey Ideker presented a biological ontology comparable in some ways to GO called NeXO for Network Extracted Ontology. The important difference is this: whereas GO is human curated from literature, NeXO is derived from data. There's a paper, A gene ontology inferred from molecular networks, and a slick looking NeXO web application.

Tim Hughes gave the other keynote on RNA-binding motifs complementing prior work constructing RBPDB, a database of RNA-binding specificities.

DREAM 8.5

Since the next round of challenges is coming sooner than DREAM's traditional yearly cycle, it's numbered 8.5. There are three challenges that are getting underway now:

Discussion

During the discussion session, questions were raised about on whether to optimize the design of the challenges to produce generalizable methods or for answering a specific biological question. Domain knowledge can give competitors advantages in terms of understanding noise characteristics and artifacts in addition to informing realistic assumptions. If the goal is strictly to compare algorithms, this might be a bug. But in the context of answering biological questions, it's a feature.

Support was voiced for data homogeneity (No NAs, data plus quality/confidence metric) and more realtime feedback.

It takes impressive community management skills to maintain a balance that appeals to a diverse array of talent as well as providers of data and support.

Monday, November 25, 2013

Vaclav Smil

I am an incorrigible interdisciplinarian. I was trained in a broad range of basic natural sciences (biology, chemistry, geography, geology) and then branched into energy engineering, population and economic studies and history. For the past 30 years my main effort has gone into writing books that offer new, interdisciplinary perspectives on inherently complex, messy realities.

On Writing

Hemingway knew the secret. I mean, he was a lush and a bad man in many ways, but he knew the secret. You get up and, first thing in the morning, you do your 500 words. Do it every day and you’ve got a book in eight or nine months.
Bill Gates reads you a lot. Who are you writing for?
I have no idea. I just write.

On Apple

Apple! Boy, what a story... When people start playing with color, you know they’re played out.