Sunday, June 10, 2012

Scientific imaging with Hanchuan Peng

Hanchuan Peng of Janelia Farm spoke on bioimaging at ISB a couple weeks back. He's doing some very cool work mining microscopy images doing registration - aligning individual cells across images. They've created a 3D atlas of C. elegans which tracks every cell. The still pictures don't don't do it justice. Check out the movies.

By localizing and registering neural fibers in 2,954 fly brains, Peng's group constructing this wiring diagram of the fly's 100,000 neurons.

More

Tuesday, June 05, 2012

Scaling higher education

In the fall of last year, over 100,000 students signed up for Andrew Ng's Machine Learning Class and more than 12,000 of them completed the course. Sebastian Thrun and Peter Norvig taught Artificial Intelligence with similarly impressive numbers.

I was one of the thousands in the Machine Learning class. I had so much fun with that, I also took Daphne Koller's Probabilistic Graphical Models. That one was a quite a bit harder, covering some fairly advanced stuff at least for my few remaining brain cells. But, I finished! For the PGM class, 6702 took the first quiz and 1441 took the final - pretty good retention for such a ball-buster of a class.

This spring, at least two new companies offering online courses were founded. Andrew Ng and Daphne Koller founded Coursera. Partnering with professors from Princeton, Penn, University of Michigan, and Berkeley, they've broadened their course catalog from a base in computer science to include classes in history, mathematics and even poetry. Sebastian Thrun, founder of Udacity, calls his approach University 2.0 and speaks of "democratizing higher eduction" and "empowering students" especially in the developing world where access to higher education is more limited. Almost three quarters of Coursera's students are outside the US, in countries like Brazil, Britain, India and Russia.

The classes are surprisingly fun. The formula boils down to two key elements:

  • short segments
  • interaction

In the mold of Khan academy, the lectures are broken into short segments of 10 to 15 minutes, which fit nicely into busy schedules. Short quizzes test the student's understanding. The courses have social aspect, as well. Online forums provide a place for questions and a sense of camaraderie while struggling through difficult concepts. Meetups and study groups have sprung up in several cities across the world.

The programming exercises are where the real fun begins. Students write code that implements the crux of an operation, filling in the blanks in provided boilerplate code. Grading works a bit like unit testing. Progressing through the assignment by getting tests to pass gives gratifyingly immediate feedback. Completing an assignment results in working code for handwriting recognition, spam classification, image processing or recognizing an action from kinect position sensing data.

Thomas Friedman says, Let the revolution come:

Welcome to the college education revolution. Big breakthroughs happen when what is suddenly possible meets what is desperately necessary.

With the cost of tuition rising, and public funding falling, the timing might be right for some disruptive innovation in higher education. And, the skills on offer are in high demand. One proposed business model is to offer classes for free and charge employers for access to the data.

Refactoring the university classroom to function at internet scale meshes with the building momentum behind open access journals that some are calling an Academic Spring as well as with citizen science projects like Galaxy Zoo.

Increasing openness in academics, in both teaching and research, can only reduce friction in the process of transferring technology from the lab to production and may help engage the public with science. Look for lots of interesting developments in the next few years, as technology knocks a new door into the ivory tower.

More

Updates

Coursera continues to generate lots of news, signing up 12 new universities, including the University of Washington (yay, UW!), attracting the attention of Bill Gate, and being described as The Single Most Important Experiment in Higher Education by the Atlantic and The Beginning of the End for Traditional Higher Education by Fortune and Reshaping Education on the Web by the NYT.

Sunday, June 03, 2012

Working in academics

The benefits of working in academics are:

  • Important and interesting work
  • Opportunities for development and growth
  • Freedom and fun

Those paying attention will recognize Daniel Pink's elements of motivation - autonomy, mastery, and purpose.

The downside? Having to explain to your spouse why you're not making as much money as so-and-so, who has basically the same skills as you. Or worse yet, why you make less than some other so-and-so who can barely tie his own shoes.

That is, unless your spouse is in academics, too. In which case, God help you.

Explaining this is not fun, but if the three elements above are in abundant supply, a fairly convincing case can be made. If those factors start to run low... well, your spouse might be right.

More

Tuesday, May 08, 2012

Design Philosophies of Developer Tools

Clojure hacker Stuart Sierra wrote an insightful piece on the design philosophies of developer tools. His conclusions are paraphrased here:

Git is an elegantly designed system of many small programs, linked by the shared repository data structure.

Maven plugins, by contrast, are not designed to be composed and the APIs are underdocumented. With Maven, plugins fit into the common framework of the engine, which strikes me as maybe a more difficult proposition than many small programs working on a common data structure.

Of the anarchic situation with Ruby package management (Gems, Bundler, and RVM) Sierra says, “layers of indirection make debugging harder,” and “The speed of development comes with its own cost.”

Principles

  • Plan for integration
  • Rigorously specify the boundaries and extension points of your system

...and I really like this idea:

  • The filesystem is the universal integration point
  • Fork/exec is the universal plugin architecture

Wednesday, May 02, 2012

Phenotypic constraints drive the architecture of biological networks

Biological networks are often compared to random networks in terms of properties like degree distribution, clustering, robustness and over-representation of network motifs. In a talk this morning, Areejit Samal, a Postdoctoral Fellow in the Price lab at ISB, proposed a new null model based on Markov Chain Monte Carlo (MCMC) sampling to generate realistic benchmark ensembles for metabolic networks and gene regulatory networks. Based on this improved background model, the conclusion is that “phenotypic constraints drive the architecture of biological networks”, which is also the title of the talk.

Applied to the metabolic model of E. coli, the sampling technique involves two steps. In the swap step, a reaction is removed from the network and a new reaction is drawn randomly from the KEGG database to replace it. The proposed new network is then tested by flux-balance analysis. If the new network is viable, it is accepted. If not, it is rejected. So, only networks that are biochemically plausible are sampled.

The technique has implications for systems and synthetic biology as well as evolutionary theory. By approximating the space of all possible networks that might lead to a viable functioning organism, we can better understand what properties these networks have, and maybe better design new networks. By changing the acceptance criteria to enforce viability in a second environment, the algorithm nicely models the evolutionary emergence of modularity.

Samal also showed the same general algorithm applied to the gene regulatory network controlling flowering in Arabidopsis.

I especially appreciated seeing a really cool application of Markov Chain Monte Carlo sampling, a topic covered only a couple weeks back in the Probabilistic Graphical Models class.

Links

Maximum expected utility decision rules

Week 6 of the Daphne Koller's Probabilistic Graphical Models class looks at Decision Theory, which integrates the concept of utility functions from economics into our models. I'm getting flashbacks of Dr. Welsh's Econ 101 in Schwab Auditorium.

In the Influence network below, we're making a decision about whether to found a company. Our success is determined by market conditions, which we can't observe. But, we can observe the results of a survey, which will give us some information on which to make our decision. So, how would you find the optimal decision rule?

The factor fm represents the probability that market conditions (var 1) are good (3), fair (2) or poor (1).

fm.var =  1
fm.card =  3
fm.val = [0.5 0.3 0.2]
PrintFactor(fm)
1 
1 0.500000
2 0.300000
3 0.200000

The factor fsm represents the results of a survey (var 2) given the underlying market conditions (var 1).

fsm.var = [2 1]
fsm.card = [3 3]
fsm.val = [0.6 0.3 0.1 0.3 0.4 0.3 0.1 0.4 0.5]
PrintFactor(fsm)
2 1 
1 1 0.600000
2 1 0.300000
3 1 0.100000
1 2 0.300000
2 2 0.400000
3 2 0.300000
1 3 0.100000
2 3 0.400000
3 3 0.500000

The factor fufm represents the utility function, given the market conditions (var 1) and the decision to found (var 3) a company, yes (2) or no (1).

fufm.var = [3 1]
fufm.card = [2 3]
fufm.val = [0 -7 0 5 0 20]
PrintFactor(fufm)
3 1 
1 1 0.000000
2 1 -7.000000
1 2 0.000000
2 2 5.000000
1 3 0.000000
2 3 20.000000

Compute mu of F, S by taking the factor product of factors representing the market, survey and utility function, then summing out over all possible market conditions (var 1).

PrintFactor(FactorMarginalization(FactorProduct(FactorProduct(fm, fsm), fufm), [1]))
2 3 
1 1 0.000000
2 1 0.000000
3 1 0.000000
1 2 -1.250000
2 2 1.150000
3 2 2.100000

We can build our decision rule by walking down all possible survey results (var 2) and selecting the decision to found or not which maximizes expected utility. See the red circles at the bottom of the diagram.

We can make a decision factor, which depends on the survey (var 2). We'll fill in dummy values, for now.

fd.var = [3 2]
fd.card = [2 3]
fd.val = ones(1,prod(fd.card))

Now, we can use the code from the programming assignment to compute the optimal decision rule and the expected utility it yields.

marketI.RandomFactors = [fm fsm]
marketI.UtilityFactors = [fufm]
marketI.DecisionFactors = [fd]
[meu optdr] = OptimizeMEU(marketI)
meu =  3.2500
optdr =
  scalar structure containing the fields:
    var  =       3   2
    card =       2   3
    val  =       1   0   0   1   0   1

Sunday, April 29, 2012

The Scalable Adapter Design Pattern for Interoperability

When wrestling with a gnarly problem, it's interesting to compare notes with others who've faced the same dilemma. Having worked on an interoperability framework, a system called Gaggle, I had a feeling of familiarity when I came across this well-thought-out paper:

The Scalable Adapter Design Pattern: Enabling Interoperability between Educational Software Tools, Harrer, Pinkwart, McLaren, Scheuer, IEEE TRANSACTIONS ON LEARNING TECHNOLOGIES, 2008

The paper describes a design pattern for getting heterogeneous software tools to interoperate with each other by exchanging data. Each application is augmented with an adapter that can interact with a shared representation. This hub-and-spokes model makes sense because it reduces the effort from writing n(n-1) adapters to connect all pairs of applications to writing one adapter per application.

Scalability refers to the composite design pattern, implementing (what I would call) a more general concept, that of hierarchically structured data. If you've ever worked with large XML documents, calling them scalable might seem like an overstatement, but I see their point. XML nicely represents small objects, like events, as well as moderately sized data documents. The same can be said of JSON.

Applications remain decoupled from the data structure, with an adapter mediating between the two. The adapter also provides events when the shared data structure is updated. A nice effect of the hierarchical data structure is that client applications can register their interest in receiving events at different levels of the tree structure.

The Scalable Adapter pattern combines of well established patterns - Adapter, Composite and Publish-subscribe yielding a flexible way for arbitrary application data to be exchanged at different levels of granularity.

The main difference between Scalable Adapter and Gaggle is that Gaggle focused on message passing rather than centrally stored data. The paper says, "it is critical that both the syntax and semantics of the data represented in the composite data structure...", but they don't really address what belongs in the "Basic Element" - the data to be shared. Gaggle solves this problem by explicitly defining a handful of universal data structures. Applications are free to implement their own data model, achieving interoperability by mapping (also in an adapter) their internal data structures onto the shared Gaggle data types.

The Scalable Adapter paper breaks the problem down systematically in terms of design patterns, while Gaggle was motivated by the software engineering strategies of separation of concerns and parsimony, plus the linguistic concept of semantic flexibility. It's remarkable that the two systems worked out quite similarly, given the different domains they were built for.