Showing posts with label rant. Show all posts
Showing posts with label rant. Show all posts

Thursday, May 08, 2014

How to grep git commits

From the There's-got-to-be-a-better-way Department, comes this tale of subterranian Git spelunking...

It all started because I lost some code. I had done my duty like a good little coder. I found a bug, wrote a test that reproduced it and filed an issue noting the name of the failing test. Then, some travel and various fire drills intervened. Memory faded. Finally, getting back to my routine, I had this bug on my docket. Already had the repro; should be a simple fix, right?

But where was my test? I asked Sublime to search for it, carefully turning off the annoying find-in-selection feature. Nothing. Maybe I stashed it. Do git stash list. Nothing that was obviously my test. Note to self: Use git stash save <message> instead of just git stash. I did git show to a several old stashes, to no avail. How am I going to find this thing?

A sensible person would have stopped here and rewritten the test. But, how much work are we willing to do to avoid doing any work? Lots, apparently.

So, I wrote this ridiculous bit of hackery instead. I call it grep_git_objects:

import os
import sys
import subprocess

if len(sys.argv) < 2:
    print "Usage: python grep_git_objects.py <target>"
    sys.exit()

target = sys.argv[1]

hashes = []
for i in range(0,256):
    dirname = ".git/objects/%02x" % i
    #print dirname
    for filename in os.listdir(dirname):
        h = "%02x%s" % (i, filename)
        hashes.append(h)

        cmd = "git cat-file -p %s" % h
        output = subprocess.check_output(cmd.split(' '))
        if target in output:
            print "found in object: ", h
            print output

Gist: grep_git_objects.py

Run it in the root of a repo. It searches through the .git/objects hierarchy to find your misplaced tests or anything else you've lost.

You'd think git grep would do exactly that. Maybe it does and I'm just clueless. I hope so, 'cause otherwise, why would you have a git grep command? In any case, git grep didn't find my missing test. The above bit of Python code did.

Once I found it, I had the hash code for a Git object. Objects are the leaves of Git trees. So, I grepped again for the hash code of the object and found the tree node it belonged to, working my way up a couple more levels to the commit.

With commit in hand, I could see by the "WIP on develop" message that is was indeed a stash, which I must have dropped by mistake. Anyway, that was fun... sorta.

Tuesday, April 22, 2014

The damaging effects of hypercompetition

In Rescuing US biomedical research from its systemic flaws (PNAS 2014), authors Bruce Alberts, Marc Kirschner, Shirley Tilghman and Harold Varmus recount the ways in which research funding is broken, in particular for life sciences.

PI's pour vast amounts of time and energy into competing for a shrinking pool of funds. Pressure for results causes a shift to short-term thinking and engenders a conservatism poorly suited to producing breakthroughs. In much the same way as the corporate sector abandoned fundamental research in favor of product development, the current funding climate overvalues "translational" research over basic science. The situation is especially dire for young scientists facing the prospect of years as an underpaid and overworked post-doc with faint hopes of ever landing a faculty position. The internet is littered with goodbye-academia letters. [1 2 3 4 5 6]

The Alberts paper proposes a few solutions, some sensible and one in particular that doesn't make much sense to me.

Broadening career paths for scientists is a sound idea. In this respect, life science can borrow from information technology. Computer science departments and the tech industry have a long history of exchange. The barriers in biotech are higher, but the flow of people and ideas between academics and industry can only be a good thing. Probably the biggest factor in diversifying away from the shrinking tenure track, is matching expectation with reality.

Managing expectations seems more sensible than trying to match supply and demand by restricting entry into PhD programs based on the reasoning that the system is training too many PhDs. Isn't too many PhD's a good problem to have? Is educating people a bad thing? Scientifically trained people are a tremendous asset. If we're not using a valuable resource effectively, having less of that resource doesn't seem like an ideal solution. With some creative thinking, productive uses for that talent can and are being found.

I'm not convinced that the pyramid shaped structure of science is the problem. Harnessing the idealism, curiosity and naive overconfidence of youth works. What doesn't work in absence of perpetual growth is the expectation that everyone at the bottom of the pyramid will eventually be at the top. But, what's wrong with a system in which people are supported for a time on research grants then continue onto many diverse career paths?

Getting the mix of competition and camaraderie right can be the difference between a situation that's nurturing rather than toxic. Economics teaches that incentives matter. But, re-engineering science will require a nuanced understanding of what those incentives are, both in crass terms of money and prestige but also the more subtle psychology around freedom, inner drive and higher purpose.

Since nobody asked me

If I was in charge of funding science, I'd place the majority of my bets on longer term grants for small focused labs where the PI does hands-on science and training of new scientists. Longer term grants of 5 years or so would give researchers time to do actual science in teams of 2-8 people. The article recognizes the success of the Howard Hughes Institutes model of explicitly encouraging scientists to pursue risky ideas - "the free play of free intellects".

I'd be reluctant to fund 40 person labs because neither mentoring nor creative thinking scales up to that size particularly well. I'd avoid large multi-center grants. Funding agencies seem to favor this sort of thing perhaps as a safer bet, they're a recipe for infighting and politicking. They risk getting all the disadvantages of both collaboration and competition with little of the benefits.

Research has a great track record of paying off in the long run. But, Bill Janeway, author of Doing Capitalism in the Innovation Economy says that there has to be a surplus in the funding system to absorb the unavoidable costs of uncertainty, both in terms of when the payoff comes and who captures the gains. In the past, this surplus has come from monopolistic corporations like Bell or Xerox, from the taxpayer, or from flush venture capital. In The Great Stagnation, Tyler Cowen posits that we've picked the low-hanging fruit of the current wave of technology - that we're in a period of stagnation while the last wave is fully digested and the stage can be set for the next wave. To that end, maybe some hedge-fund wizard can cook up some financial innovation for reliably financing risky, long-term projects with unevenly distributed payouts.

[1] Goodbye academia, I get a life
[2] Goodbye Academia
[3] Goodbye academia? Hello, academia
[4] The Big Data Brain Drain
[5] Why So Many Academics Quit and Tell
[6] On Leaving Academe
[7] Science: The Endless Frontier, Vannevar Bush, 1945
[8] The Decline of Unfettered Research, Andrew Odlyzko, 1995

Wednesday, January 30, 2013

Javascript style objects in Python

As I've gotten older and crankier, one of the things I've gotten cranky about is Object Oriented Programming. Sometimes, it seems like rigid class hierarchies just get in the way. Sure, it's cool that the compiler can prove type-safety properties about your program, but whether the boilerplate is worth the benefit is situation dependent. When it just has to work once, coding for maintenance is a poor investment - quick-n-dirty is the way to go.

I've found class hierarchies in dynamic scripting languages to be especially unnecessary. Why use a language where you're not supposed to worry about typing to build elaborate user-defined type hierarchies?

In those quick-n-dirty scenarios, what I want is a big stinking bag of properties. Ideally, I don't care whether the values are data or functions. In other words, just like it's done in Javascript.

Here's one way to get something like that, in Python:

Sunday, June 03, 2012

Working in academics

The benefits of working in academics are:

  • Important and interesting work
  • Opportunities for development and growth
  • Freedom and fun

Those paying attention will recognize Daniel Pink's elements of motivation - autonomy, mastery, and purpose.

The downside? Having to explain to your spouse why you're not making as much money as so-and-so, who has basically the same skills as you. Or worse yet, why you make less than some other so-and-so who can barely tie his own shoes.

That is, unless your spouse is in academics, too. In which case, God help you.

Explaining this is not fun, but if the three elements above are in abundant supply, a fairly convincing case can be made. If those factors start to run low... well, your spouse might be right.

More

Tuesday, March 06, 2012

Ingenuity

When I was a little whelp, I had a brief and unsuccessful engagement at a bioinformatics startup in the bay area called Ingenuity. The company had a cool idea, a knowledge base for molecular biology, and smart and creative people to implement it. I was tasked, by a long-haired, Tufte-toting, Stanford grad-student, with developing new and rich UI elements with the budding new technology at the time, dynamic HTML. In particular, I was to implement a search bar that could automatically suggest terms from an ontology - the autocomplete feature.

There was even a prototype that sort-of worked on the right version of Netscape when the server 30 feet away. I pulled my hair out trying to get it working consistently across browsers. A better engineer, say John Resig, might have pulled it off, but I had to admit defeat. Like many DHTML toys of the time, it just wasn't ready for production. So, I recommended scrapping the idea in favor of a simple "google-box". This was not well received and my tenure was not long in coming to a close.

For years after, I'd argue on every project for simple, stripped-down web UIs. "Look at Google," I'd say, pointing to the minimalist search box. Trying to do anything advanced in a browser, I'd warn, was an invitation to cross-browser compatibility issues and nightmarish debugging sessions.

Meanwhile, as if to make me look like a bozo (as if I need any help), Google introduced autocomplete, AKA Google Suggest, first as a Google Labs project, in 2004 and finally rolled out autocomplete on the main Google page in 2008. So much for my "Look at Google" argument. Not that it matters now, but I feel slightly vindicated by the fact that it took even the mighty Google this long to deploy a feature that I was basically shit-canned for failing to implement in 2000. ...not that I'm bitter. But, anyway, back in the present day...

Doug Basset, Chief Scientific Officer at Ingenuity gave a demo last week at the ISB, where I now reside, showing off Ingenuity Variant Analysis, a new tool built on top of their knowledge base that helps find disease-causing genetic variations in resequencing data.

The basic trick is to filter down the millions of genetic variants found in any individual genome to those consistent with a given condition, starting with it's frequency in the population and genetic properties like homo- or heterozygosity. The neat part comes next.

For each variant surviving this far, the program traverses the graph of facts in the knowledge base. Of course, it will find known associations between variants or their host genes and disease. What's better, it can also find relationships a few degrees removed from direct implication. Say, A turns off B with regulates C. The biological process to which C belongs runs amok in disease X. Suddenly, a variant in a functional domain of gene A looks like an interesting candidate. If it works, you end up with a handful of genes with enough evidence to warrant further investigation.

The products's flash-based UI is very slick and modern, with drop-shadows, ghosting and barber-pole progress bars. Tables have spiffy little sparkline graphics. Right in the middle of the demo, a search dialog popped up and there was the autocomplete feature, mocking me. It's certainly no big deal these days. My current project has an autocompleting search box, too, thanks to jQuery and Solr. But, I guess the memory of flubbing that gig still has a little sting left in it.

Thursday, February 09, 2012

Twitter feature request

For a long time, I just didn't get Twitter, thinking that nothing of substance could fit in 140 characters. The way I've come to use it is like scanning the headlines on a newspaper, following links to read more about interesting items. I follow folks that are interested in the same nerdy topics I am, in effect like a personal slashdot or hacker news with the riff-raff filtered out, or rather the interesting people filtered in.

Still, my cluster of interests is eclectic and my interests tend to only partially overlap with the people I follow. So, what I want is a feature that will help automatically weed out the stuff I don't care about.

I mostly read tweets on my phone using the Android version of TweetDeck. What if I could swipe right on a tweet for something I like, and swipe left to shit-can a lame tweet? The app could keep track of what I've liked or not and use that as an ever-growing training set to classify new tweets as interesting or dreck. The same data could also be used to find people I should follow or interesting tweets from people I don't necessarily want to follow, etc.

In my case, the app would soon learn to terminate sports-related tweets with deadly efficiency. Also, for you foursquare users, any tweet that starts with "I'm at..." would be ruthlessly eliminated. It would work great for facebook feeds as well, deleting all those "Your sister-in-law just unlocked the Cattywampus badge on whatever-game".

This must already exist. Is this what prismatic is? Maybe, I'll find out if I get my invite!

Update

Looks like I'm not the only one wanted to waste time more efficiently. Seattle hacker Joel Grus built a classifier for Hacker News stories using naive Bayes. Joel gave a nice lightning talk at the latest Seattle Data/Analytics/Machine-Learning MeetUp.

Saturday, May 14, 2011

HTC Incredible internal memory

My phone, an HTC Droid Incredible, may be hopelessly antiquated by the standards of true mobile hipsters. Still, it came with a generous 8GB internal storage. Too bad the SD card is a chintzy 2GB. These days, you get more than 2 gigs on an abacus. It seems like Android wants to use internal storage for the OS and apps, reserving the SD card for media, which makes that 8GB/2GB split a little awkward. I filled that 2GB right up with choice sides of Miles and 'Trane in no time. And, what do I need with 8 gigs worth of apps? What am I running, Bloatpad 2.0? So, anyway, I decided I wanted to use the empty 6 plus gigs on the internal storage for some Thelonious. So, can I do that?

Well, the Help/How to thing at HTC says, "...Music only plays audio files saved on the storage card...". Well, I use another media player anyway - Meridian. Then there's an article called Programmitically accessing internal storage (not SD card) on Verizon HTC Droid Incredible (Android), which says, "To be quite honest, the internal storage is a joke. Just think of it as a flash drive..."

But, it turns out, you can access music and other media on the internal storage. You just have to know that the internal storage is mounted as /emmc. Maybe they shoulda called it /WTF.

Sunday, February 13, 2011

The Tiger Mom and A Clockwork Orange

True disciple is doing what you want.

A wise friend once told me that. Amy Chua, better known as the Tiger Mother, wrote about discipline (from a different point of view) in Why Chinese Mothers Are Superior.

What Chinese parents understand is that nothing is fun until you're good at it. To get good at anything you have to work, and children on their own never want to work, which is why it is crucial to override their preferences. [...] Tenacious practice, practice, practice is crucial for excellence; rote repetition is underrated in America. Once a child starts to excel at something -- whether it's math, piano, pitching or ballet -- he or she gets praise, admiration and satisfaction. This builds confidence and makes the once not-fun activity fun. This in turn makes it easier for the parent to get the child to work even more.

For what it's worth, Chua's book is apparently less strident and more nuanced than the WSJ article. Anyway, like her methods or not, I have a lot of sympathy for a parent trying to teach her kids about delayed gratification, that you can do difficult things if you try, and that hard work pays off.

If it's true that mastering a complex skill takes 10,000 hours of practice, then the persistence to push through those hours is a fairly important lesson to learn early. Recent research has caused a reappraisal in how much talent arrises from innate genius versus how much is the product of effort, practice and persistence.

What brought to mind my old friend's remark about disciple was Paul Buchheit's take: motivation can be either intrinsic or extrinsic. Amy Chua is teaching her kids to be extrinsically motivated, to respond to the praise and admiration of others. You do it because you are told to. You put your energy into chasing the approval of external authorities. In contrast, he describes intrinsic motivation like this:

To the greatest extent possible, do whatever is most fun, interesting, and personally rewarding (and not evil).

Follow your heart, as hippy moms tell their children. Buccheit says, "I'm kind of lazy, or maybe I lack will power or discipline or something. Either way, it's very difficult for me to do anything that I don't feel like doing." Sounds familiar. "The intrinsic path to success is to focus on being the person that you are, and put all of your energy and drive into being the best possible version of yourself."

The difference between intrinsic and extrinsic motivation is easily recognized in the moral dimension. In A Clockwork Orange, Anthony Burgess imagines the transfer of aesthetic sense from creation to violence. Deprived of outlet, creativity turns destructive. The main thrust of the story is an examination of attempts to impose an external morality by force versus growing an internal morality.

Paul Buccheit was the software developer that originated Google's gmail. For myself, and I'm sure lots of others, a key attraction to programming was the ability to create in a powerful medium without asking anyone's permission. The creative freedom, the feeling that the authorities hadn't (yet) figured out how to lock things down was incredibly inspiring.

It's easy to see how that attraction to technology meshes with Daniel Pink's elements of motivation — autonomy, mastery, and purpose. If you've got root and a compiler, you've got autonomy. And it's all about a pissing contest of mastery. (This might partially explain the gender ratio in the field.) And technology is rife with appeals to higher purpose, from the open-source movement to the digital media that helped fuel the revolutions in Tunisia and Egypt.

The Tiger Mom demands mastery before autonomy, leaving purpose firmly in the hands of the parent. Buccheit puts autonomy first, trusting in a natural sense of direction to lead to mastery and purpose. In terms of Maslow's hierarchy of needs, Amy Chua has the "esteem" level covered, but stops short of the top level - intrinsic self-directed creativity.

Some suggest that American society erects a border fence at the entry to the highest level. Well, nobody ever mentions why it's always drawn as a pyramid, but there's probably a reason. Not everyone gets to be at the top. Usually, that's the realm only of the elite.

David Cain of Raptitude writes under the title How to Make Trillions of Dollars, that the fundamentals of being a self-directed person are these:

Creativity. Curiosity. Resilience to distraction. Patience with others.

Cain defines self-reliance as “an unswerving willingness to take responsibility for your life, regardless of who had a hand in making it the way it is”.

Again, we're basically talking about the top levels of the pyramid. But, I like the addition of resilience to distraction. Discipline is a hard sell in a culture the promotes immediate gratification and nonstop indulgence. It's hard to hear your own voice over the clamor of consumer culture and expectations from family, boss and everyone else. Hearing it is nearly impossible while drowning in distractions like twitter and facebook. And here's something else to remember about online amusements:

If you're not paying, you're not the customer; you're the product.

At a recent data mining conference I saw rooms full of marketers ready to slice, dice and mash up your personal data to more precisely target advertising. Resisting this attack is an essential skill of modern life. A healthy cynicism is a necessary defense mechanism. Hearing yourself think is only going to get harder.

The flawed idea implicit in consumer culture is what you consume is what you are. But, valuing consuming over creating or doing is inevitably a dead end. The reason I'm not so hot on the iPad is that it's a device for consuming. The old macs were (marketed as) tools for programmers, graphic artists, musicians and film-makers -- in other works doers, builders and creators.

Purpose has to come from values. The recent travails of the financial sector show what happens when motivations or at least incentives become disconnected from morals and values.

Of course, lots of technology has a purpose no higher than selling golf clubs on the internet. And technology, itself, can be a distraction. It's easy to get caught up in a rat race of the latest whizzy buzzword laden language, tool or application framework dujour. For years, I've had a half-joking theory that the true purpose of the internet is to absorb the excess productivity of mankind.

Creative, conceptual work driven by autonomy, mastery and purpose pursued with uninterrupted concentration. That begins to answer the question, how do we get some motivation, apply it to something good and inspire the same in those around us, especially our ungrateful screaming offspring.

You can argue one way or another about whether a 13 year old has the foresight to be intrinsically motivated. I certainly didn't have the wherewithal at that age to set a long term goal. And pushing intrinsic motivation on your kids sounds to me something like imposing democracy by force. I doesn't make that much sense.

It takes determination to undertake exhausting frustrating efforts whose payoff is distant and uncertain. Most things that are worth doing are hard. The drive and courage to try anyway is what makes real progress possible. That doesn't come easily, and neither does the judgement necessary to gauge what is worth while against the scale of your own values.

From my current position, I don't particularly want to lecture anyone on how to succeed in life. I respond negatively to coercion and can be a champion slacker, both of which were much to the detriment of my academic career. Still, here it is:

  • Value creation over consumption.
  • Surround yourself with creators.
  • Do what you love, do it a lot, and do it hard.

Monday, November 15, 2010

Tech Industry Gossip

Welcome to the new decade: Java is a restricted platform, Google is evil, Apple is a monopoly and Microsoft are the underdogs

I mostly try to ignore tech industry gossip. But, there's a lot of it, lately. And underlying the shenanigans are big changes in the computing landscape.

First, there's the flurry of lawsuits. This is illustrated nicely by the Economist in The great patent battle. There are several similar graphs of the patent thicket floating around. IP law is increasingly being used as a tool to lock customers in and competitors out. We can expect the (sarcastically named) "Citizens United" ruling on campaign finance to result in more of this particular kind of antisocial behavior.

Google and Facebook are in a pitched battle over your personal data and engineering talent. Google engineers, apparently, are trying to jump over to Facebook prior to what promises to be a huge IPO.

Apple caused quite a kerfuffle by deprecating Java on Mac OS X. After remaining ominously silent for weeks, Oracle seems to have lined up both Apple and IBM behind OpenJDK. Apache Harmony looks to be a casualty of this maneuvering. Harmony, probably not coincidentally, is the basis for parts of Google's Android and Oracle is suing Google over Android's use of Java technology.

Microsoft seems to be waning in importance along with the desktop in general. Ray Ozzie, Chief Architect since 2005, announced that he was leaving, following Robbie Bach of the XBox division and Stephen Elop, now running Nokia. I spoke with one MS marketing guy who said of Ozzie, "Lost him? I'd say we got rid of him!" A lesser noticed departure, that of Jython and Iron Python creator Jim Hugunin may also be telling. Profitable stagnation seems to be the game plan there.

Adobe's been struggling to the point where the NYTimes asked where does Adobe go from here? They took a beating over flash performance and rumors circulated briefly of a buyout by Microsoft.

The cloud is where a lot of the action in software development is moving. Mobile has been growing in importance by leaps and bounds ever since the launch of the iPhone. Cloud computing and consumer devices like smart phones, tablets, and even Kindles are complementary to a certain extent. The economics of cloud computing are hard to argue with. (See James Hamilton's slides and video on data centers.)

Another part of what's changing is a swing of the pendulum away from openness and back towards the walled gardens that most of us thought were left behind in the ashes of Compuserve and AOL. Ironically enough, Apple has become the poster child of walled gardens, with iTunes and the app store. ...the mobile carriers even more so. And the cloud infrastructures of both Microsoft's Azure and (to a lesser degree) Google's App Engine are proprietary. Out of the big 3, Amazon's EC2 is, by far, the most open. Mark Zuckerberg says, "I’m trying to make the world a more open place." But, to Tim Berners-Lee, Facebook and Apple threaten the internet.

There's plenty of money in serving the bulk population. That's why Walmart is so huge. My fear is that in a rush to provide "sugar water" to consumers, the computing industry will neglect the creative people that made the industry so vibrant. But, not to worry. Ray Ozzie's essay Dawn of a new Day does a nice job of putting into perspective the embarrassment of riches that technology has yielded. We're just at the beginning of figuring out what to do with it all.

Monday, September 27, 2010

How to send an HTTP PUT request from R

I wanted to get R talking to CouchDB. CouchDB is a NoSQL database that stores JSON documents and exposes a ReSTful API over HTTP. So, I needed to issue the basic HTTP requests: GET, POST, PUT, and DELETE from within R. Specifically, to get started, I wanted to add documents to the database using PUT.

There's CRAN package called httpRequest, which I thought would do the trick. This wound up being a dead end. There's a better way. Skip to the RCurl section unless you want to snicker at my hapless flailing.

Stuff that's totally beside the point

As Edison once said, "Failures? Not at all. We've learned several thousand things that won't work."

The httpRequest package is very incomplete, which is fair enough for a package at version 0.0.8. They implement only basic get and post and multipart post. Both post methods seem to expect name/value pairs in the body of the POST, whereas accessing web services typically requires XML or JSON in the request body. And, if I'm interpreting the HTTP spec right, these methods mishandle termination of response bodies.

Given this shaky foundation to start with, I implemented my own PUT function. While I eventually got it working for my specific purpose, I don't recommend going that route. HTTP, especially 1.1, is a complex protocol and implementing it is tricky. As I said, I believe the httpRequest methods, which send HTTP/1.1 in their request headers, get it wrong.

Specifically, they read the HTTP response with a loop like one of the following:

repeat{
  ss <- read.socket(fp,loop=FALSE)
  output <- paste(output,ss,sep="")
  if(regexpr("\r\n0\r\n\r\n",ss)>-1) break()
  if (ss == "") break()
}
repeat{
 ss <- rawToChar(readBin(scon, "raw", 2048))
 output <- paste(output,ss,sep="")
 if(regexpr("\r\n0\r\n\r\n",ss)>-1) break()
 if(ss == "") break()
 #if(proc.time()[3] > start+timeout) break()
}

Notice that they're counting on a blank line, a zero followed by a blank line or the server closing the connection to signal the end of the response body. I dunno where the zero thing comes from or why we should count on it not being broken up during reading. Looking through RFC2616 we find this description of an HTTP message:

generic-message = start-line
                  *(message-header CRLF)
                  CRLF
                  [ message-body ]

While the headers section ends with a blank line, the message body is not required to end in anything in particular. The part of the spec that refers to message length lists 5 ways that a message may be terminated, 4 of which are not "server closes connection". None of them are "a blank line". HTTP 1.1 was specifically designed this way so web browsers could download a page and all its images using the same open connection.

For my PUT implementation, I fell back to HTTP 1.0, where I could at least count on the connection closing at the end of the response. Even then, socket operations in R are confusing, at least for the clueless newbie such as myself.

One set of socket operations consists of: make.socket, read.socket/write.socket and close.socket. Of these functions, the R Data Import/Export guide states, "For new projects it is suggested that socket connections are used instead."

OK, socket connections, then. Now we're looking at: socketConnection, readLines, and writeLines. Actually, tons of IO methods in R can accept connections: readBin/writeBin, readChar/writeChar, cat, scan and the read.table methods among others.

At one point, I was trying to use the Content-Length header to properly determine the length of the response body. I would read the header lines using readLines, parse those to find Content-Length, then I tried reading the response body with readChar. By the name, I got the impression that readChar was like readLines but one character at a time. According to some helpful tips I got on the r-help mailing list this is not the case. Apparently, readChars is for binary mode connections, which seems odd to me. I didn't chase this down any further, so I still don't know how you would properly use Content-Length with the R socket functions.

Falling back to HTTP 1.0, we can just call readLines 'til the server closes the connection. In an amazing, but not recommended, feat of beating a dead horse until you actually get somewhere, I finally came up with the following code, with a couple variations commented out:

http.put <- function(host, path, data.to.send, content.type="application/json", port=80, verbose=FALSE) {

  if(missing(path))
    path <- "/"
  if(missing(host))
    stop("No host URL provided")
  if(missing(data.to.send))
    stop("No data to send provided")

  content.length <- nchar(data.to.send)

  header <- NULL
  header <- c(header,paste("PUT ", path, " HTTP/1.0\r\n", sep=""))
  header <- c(header,"Accept: */*\r\n")
  header <- c(header,paste("Content-Length: ", content.length, "\r\n", sep=""))
  header <- c(header,paste("Content-Type: ", content.type, "\r\n", sep=""))
  request <- paste(c(header, "\r\n", data.to.send), sep="", collapse="")

  if (verbose) {
    cat("Sending HTTP PUT request to ", host, ":", port, "\n")
    cat(request, "\n")
  }

  con <- socketConnection(host=host, port=port, open="w+", blocking=TRUE, encoding="UTF-8")
  on.exit(close(con))

  writeLines(request, con)

  response <- list()

  # read whole HTTP response and parse afterwords
  # lines <- readLines(con)
  # write(lines, stderr())
  # flush(stderr())
  # 
  # # parse response and construct a response 'object'
  # response$status = lines[1]
  # first.blank.line = which(lines=="")[1]
  # if (!is.na(first.blank.line)) {
  #   header.kvs = strsplit(lines[2:(first.blank.line-1)], ":\\s*")
  #   response$headers <- sapply(header.kvs, function(x) x[2])
  #   names(response$headers) <- sapply(header.kvs, function(x) x[1])
  # }
  # response$body = paste(lines[first.blank.line+1:length(lines)])

  response$status <- readLines(con, n=1)
  if (verbose) {
    write(response$status, stderr())
    flush(stderr())
  }
  response$headers <- character(0)
  repeat{
    ss <- readLines(con, n=1)
    if (verbose) {
      write(ss, stderr())
      flush(stderr())
    }
    if (ss == "") break
    key.value <- strsplit(ss, ":\\s*")
    response$headers[key.value[[1]][1]] <- key.value[[1]][2]
  }
  response$body = readLines(con)
  if (verbose) {
    write(response$body, stderr())
    flush(stderr())
  }

  # doesn't work. something to do with encoding?
  # readChar is for binary connections??
  # if (any(names(response$headers)=='Content-Length')) {
  #   content.length <- as.integer(response$headers['Content-Length'])
  #   response$body <- readChar(con, nchars=content.length)
  # }

  return(response)
}

After all that suffering, which was undoubtedly good for my character, I found an easier way.

RCurl

Duncan Temple Lang's RCurl is an R wrapper for libcurl, which provides robust support for HTTP 1.1. The paper R as a Web Client - the RCurl package lays out a strong case that wrapping an existing C library is a better way to get good HTTP support into R. RCurl works well and seems capable of everything needed to communicate with web services of all kinds. The API, mostly inherited from libcurl, is dense and a little confusing. Even given the docs and paper for RCurl and the docs for libcurl, I don't think I would have figured out PUT.

Luckily, at that point I found R4CouchDB, an R package built on RCurl and RJSONIO. R4CouchDB is part of a Google Summer of Code effort, NoSQL interface for R, through which high-level APIs were developed for several NoSQL DBs. Finally, I had stumbled across the answer to my problem.

I'm mainly documenting my misadventures here. In the next installment, CouchDB and R we'll see what actually worked. In the meantime, is there a conclusion from all this fumbling?

My point if I have one

HTTP is so universal that a high quality implementation should be a given for any language. HTTP-based APIs are being used by databases, message queues, and cloud computing services. And let's not forget plain old-fashioned web services. Mining and analyzing these data sources is something lots of people are going to want to do in R.

Others have stumbled over similar issues. There are threads on r-help about hanging socket reads, R with CouchDB, and getting R to talk over Stomp.

RCurl gets us pretty close. It could use high-level methods for PUT and DELETE and a high-level POST amenable to web-service use cases. More importantly, this stuff needs to be easier to find without sending the clueless noob running down blind alleys. RCurl is greatly superior to httpRequest, but that's not obvious without trying it or looking at the source. At minimum, it would be great to add a section on HTTP and web-services with RCurl to the R Data Import/Output guide. And finally, take it from the fool: trying to role your own HTTP (1.1 especially) is a fool's errand.

Thursday, July 01, 2010

Science funding and productivity

These are interesting times for the practice and funding of science. The traditional model of fee-for-subscription peer-reviewed academic journals is looking more and more outdated. Scientific funding is increasingly competitive and dependent on salesmanship and networking rather than scientific merit.

We Must Stop the Avalanche of Low-Quality Research argues that scientists are drowning in a sea of mediocre papers that nobody reads.

In economic terms, attention is the scarce resource. Electronic publishing is dirt cheap, so it makes sense to publish even weak or negative results. But human attention is expensive and the peer review process is time consuming and unfunded. There needs to be a better mechanism for ranking the quality and importance of papers, so that scarce attention can be allocated efficiently.

Certainly, counting papers is as poor a metric of scientific output as counting lines of code is of programmer productivity.

Scientists should be scientists, not fund raisers. Real Lives and White Lies in the Funding of Scientific Research details the tyranny of grant applications.

One proposed improvement is a track system, in which a researcher would be placed into a funding category and reviewed for productivity every five years and moved up or down to higher or lower tracks accordingly. Emphasis would shift from plans to outcomes.

Stanford bioengineering professor Steven Quake makes a similar point in the New York Times:

As we consider the monumental challenges facing our generation — climate change, energy needs and health care — and look to science for solutions, it would behoove us to remember that it is almost impossible to predict where the next great discoveries will be made — and thus we should invest broadly and let scientists off their leashes.

One has to wonder how well science funding will hold up in the face of the gaping government deficits in most western countries.

Meanwhile, China is becoming scientific superpower.

Luo Minmin, 37, a neurobiologist, returned to China six years ago after getting his PhD from the University of Pennsylvania and completing a postdoctoral research stint at Duke. Luo said he has a big budget at NIBS and greater research freedom than he would have in the United States. "If I had stayed in America, the chances of making a discovery would have been lower," he said. "Here, people are willing to take risks. They give you money, and essentially you can do whatever you want."

Related links

Wednesday, May 26, 2010

Attention, Intelligence, Creativity and Flow

Most coders are aware of the importance of pure uninterrupted concentration. Creative work of any kind requires focuses attention. The state of flow happens when the spotlight of attention is completely focused on an activity. A piece by Jonah Lehrer, Attention and Intelligence, reminded me of that happy bubble so easily popped by meetings, spouses and pointy-haired bosses. He writes, "Our mind has strict cognitive limitations - selective attention helps us compensate."

Herbert Simon said, "A wealth of information creates a poverty of attention."

William James famously wrote, "Everyone knows what attention is... It implies withdrawal from some things in order to deal effectively with others."

Mihaly Csikszentmihalyi tells us, "Flow describes a state of experience that is engrossing, intrinsically rewarding and outside the parameters of worry and boredom." The components of Flow are:

  • Clear goals.
  • Attention is focused on a limited stimulus field. There is full concentration, complete involvement. Focus of awareness is narrowed down to the activity itself.
  • A loss of self-consciousness, action and awareness merge.
  • Immediate feedback; behavior can be adjusted as needed.
  • Balance between ability and challenge.
  • A sense of control and serenity; freedom from worry about failure.
  • Timelessness; thoroughly focused on the present.
  • Intrinsic motivation; the experience becomes its own reward, resulting in effortlessness of action.

Not entirely unrelated is Daniel Pink's analysis of motivation as:

  • Autonomy: The urge to direct our own lives
  • Mastery: The desire to get better and better at something that matters
  • Purpose: The yearning to do what we do in the service of something larger than ourselves

From Palm Sunday by Kurt Vonnuget:

Most of my adult life has been spent in bringing to some kind of order sheets of paper eight and a half inches wide and eleven inches long. This severely limited activity has allowed me to ignore many a storm. It has also caused many of the worst storms I ignored. My mates have often been angered by how much attention I pay to paper and how little attention I pay to them.
I can only reply that the secret to success in every human endeavor is total concentration. Ask any great athlete.
To put it another way: Sometimes I don't consider myself very good at life, so I hide in my profession.
I know what Delilah really did to Samson to make him as weak as a baby. She didn't have to cut his hair off. All she had to do was break his concentration.

All this is a long way of saying multitasking sucks, or as someone with a fistful of yen might say, "What was that? This is not a charade. We need total concentration."

Links

Wednesday, February 03, 2010

Upgrade MacBook Pro Memory?

I was thinking of upgrading my 2007 MacBook Pro with more RAM. It came with 2GB, and the specs say it can take up to 3GB, although some online sources say they can successfully install 4GB. Apparently, these older MacBooks map "system functions", I guess meaning IO mapping and ROM into the region between 3GB and 4GB.

... at least 3 GB of RAM should be fully accessible, while when 4 GB of RAM installed, ~700 MB of of the RAM is overlapping critical system functions, making it non-addressable by the system.

OK, so no 4GB for me, but what about replacing one of the 1GB sticks with a 2GB stick for a total of 3GB? It turns out that if I did that, I'd take a small performance hit.

All Intel Core Macs support dual channel memory access if matching modules are installed. The customary estimate is that this gives a 6% - 8% real world performance benefit. The modules do not have to be the same brand. That means it is quite possible but not 100% guaranteed, that adding a 3rd party SODIMM to an Apple supplied SODIMM of the same size will make a matched pair.

Verdict: don't bother...

Thursday, December 10, 2009

Microformats

Jeff Atwood, who writes the well-known Coding Horror blog, took on the topic of Microformats recently. His misguided comments about the presumed hackiness of overloading CSS classes with semantic meaning (actually their intended purpose) had people quoting the HTML spec:

The class attribute, on the other hand, assigns one or more class names to an element; the element may be said to belong to these classes. A class name may be shared by several element instances. The class attribute has several roles in HTML:
  • As a style sheet selector (when an author wishes to assign style information to a set of elements).
  • For general purpose processing by user agents.

Browsers work great for navigation and presentation, but we can only really compute with structured data. Microformats combine the virtues of both.

There are at least a couple of ways in which the ability to script interaction with web applications comes in handy. For starters, microformats are a huge advance compared to screen-scraping. The fact that so many people suffered through the hideous ugliness of screen-scraping proves that there must be some utility to be had there.

Also, web-based data sources have a browser-based front-end and also often expose a web service. Microformats link these together. A user can find records of interest by searching in the browser, embedded microformats allow the automated construction of a web service call to retrieve the data in structured form.

Microformats aren't anywhere near the whole answer. But, the real question is how to do data integration at web scale using the web as a channel for structured data.

See also

Monday, November 09, 2009

Computational representation of biological systems

Computational representation of biological systems by Zach Frazier, Jason McDermott, Michal Guerquin, Ram Samudrala is a book chapter in Springer's Computational Systems Biology. It introduces basic data warehousing concepts along with a data warehousing effort targeted at biology called Bioverse.

They contrast data warehousing with online transaction processing. OLTP entails frequent concurrent updates. Updates traditionally look like bank machine operations or travel reservations. Data warehousing, in contrast, typically updates only occasionally in an additive way as new data arrives or annotations are added. The star schema, which supports efficient subsetting and computing of aggregates (min, max, sum, count, average), centers on a table of atomic data elements called facts surrounded by related tables holding different types of search criteria called dimensions.

Bioverse nicely illustrates several of the main problems challenges with data warehousing. First, it's data (54 organisms) appears to have been last updated in 2005. Also, we must choose at what granularity to create the fact table based on the questions we expect to ask, but questions come at many scales.

Hierarchical data occurs throughout the Bioverse. Representation of these structures is particularly difficult in relational databases.

They go on to cite a method for supporting efficient hierarchical queries from Kimball, R., Ross, M., (2002). The Data Warehouse Toolkit: The Complete Guide to Dimensional Modeling.

In addition to hierarchies, graphs and networks are common structures in biological systems including protein­protein interaction networks, biochemical pathways, and others. However, the techniques outlined for trees and directed acyclic graphs are no longer appropriate for graphs. Answering any more than very basic graph queries is hard in relational databases.

So, with all these drawbacks, you have to wonder whether the relational database is a good basis for data mining applications. Especially with the prevalence of networks in biological data. I'm a lot more intrigued by ideas along the lines of NoSQL schema-less databases, Dynamic fusion of web data and the web as a channel for structured data.

Wednesday, August 26, 2009

Using R and Bioconductor for sequence analysis

Here's another quick R vignette, in case I pick this up later and need to remind myself where I got stuck. I was trying to use R for a bit of basic sequence analysis, with mixed results.

First, install the BSgenome package, which is part of Bioconductor. Get GeneR while you're at it.

> source("http://bioconductor.org/biocLite.R")
> biocLite("BSgenome")
> biocLite("GeneR")

Follow the instructions in the document How to forge a BSgenome data package. You'll need to get fasta files from somewhere such as NCBI's Entrez Genome. Another nice data source is Regulatory Sequence Analysis Tools.

I created a BSgenome package for our favorite model organism Halobacterium salinarum NRC-1, which I named halo for short. Now, I can ask what sequences make up the halo genome and find out how long they are.

> library(BSgenome.halo.NCBI.1)
> seqnames(halo)
[1] "chr"     "pNRC200" "pNRC100"
> seqlengths(halo)
    chr pNRC200 pNRC100 
2014239  365425  191346
> length(halo$chr)
[1] 2014239

There are a few things I wanted to do next. First, I wanted to load a list of genes with their coordinates. That should allow me to quickly get the sequence for each gene, or get sequence of upstream regions for regulatory motif finding. Second, if I'm going to find any new protein coding regions, I'd like to have a function that could take a stretch of DNA and find ORFs (open reading frames). As far as I can tell, all there is to ORF finding is searching each reading frame for long stretches that start with a methionine (AUG) and end with a stop codon (UAG, UGA, and UAA ). Maybe there's more to it than that.

This is where I left off. GeneR seems to use an entirely different way of encoding sequence based on buffers. I have to admit to being a little disappointed. I hope it's just my cluelessness and there's really a reasonable way to do this kind of thing in R and Bioconductor.

Related stuff from Blue Collar Bioinformatics

Monday, August 10, 2009

Autocompletion and Swing

I can remember being asked to implement cross-browser autocompletion in 2000 and telling my employers that it couldn't be done. We got a prototype working on Netscape (remember when that was the most advanced browser?) but it was buggy on Internet Explorer (remember when IE was the bane of every web developer's existence? Wait, some things never change...) and latency was way too high for most users. Anyway, I didn't last long in that gig, but I still think that for practical purposes at the time I was right. Of course, things are different now.

Type something into Google or Amazon's search box and you'll get a nice drop-down list of possible completions. For a biological example, check out NCBI's BLAST. Make sure the database chooser reads "Nucleotide collection" and start typing "Pyrococcus". Nice, huh?

Several ways to give poor abandoned Swing an autocompleting upgrade are documented in a java.net article. Sadly, they all seem to suffer from one deficiency or another.

The JIDE common layer an open source library that spun out of Jide's commercial offerings seems to be the most stable, but isn't nearly as convenient as the javascript versions. GlazedLists does a nice job, but it's currently (still) broken on OS X. Out of the solutions I found, GlazedLists seems most promising, especially if that bug gets fixed.

I also checkout out the Substance Java look & feel. It looks really sharp. The developer has done some really slick transitions -- highlights that fade in and out or components that expand like icons on OS X's dock. It's probably great on Windows, but unfortunately, it didn't seem very stable on the Mac.

There should be some good lessons to be learned from the failure of Swing. It seems apparent that several developers who are a lot smarter than me have tried to get Swing to cough up a decent UI. The results seem to be consistently limited. Not that some aren't impressive; they are. But the limits to the success of some very good developers speaks very loudly.

More attempts...

  • AutoCompleteCombo by Exterminator13 (haha) -
    Exception in thread "AWT-EventQueue-0" java.lang.ArrayIndexOutOfBoundsException: -1
     at java.util.ArrayList.get(ArrayList.java:323)
     at dzone.AutoCompleteCombo$Model.getElementAt(AutoCompleteCombo.java:476)
  • An autocomplete popup by Pierre Le Lannic, which seems to work as long as you don't care about upper case.
  • Java2sAutoTextField from Sun

Tuesday, March 24, 2009

Great example of what not to do

In case anyone was wondering, one of these web sites got it right. The other got it wrong. First, have a look at the query page from the bug database on java.net: Now, consider the front page of alltheweb.com: Some poor shlep put a lot of work into that query page for the bug database. I'm tempted to shake my head and say, "WTF?", but I know I've been guilty of this kind of thing. To be fair, java.net has a much easier search box, hidden off to the side where bozos like me won't find it. So, here's a reminder not to do shit like that.

Friday, March 20, 2009

More Hacking NCBI

Writing scripts to interface with NCBI's web site has it's challenges. Getting data from the UCSC genome browser is simpler.

If you need a list of complete genomes, that can be had from the NCBI Genome database. One form of list is the genlist.cgi script. The type parameter seems to be a flag that limits the list to chromosomes, plasmids, or organelle specific sequences. The name parameter seems to be there only for looks. So far, I haven't figured out how to make genlist spit out either XML or text.

Two other scripts can produce text output, lproks and leuks.

These two can be scripted like this using parameters like these: view=1 dump=selected p3=11:|12:Green Algae. This information is available by ftp from ftp://ftp.ncbi.nih.gov/genomes/genomeprj/. There are 3 lproks.txt files, which look to correspond to the three tabs Organism info, Complete genomes, Genomes in progress. lproks_1.txt is the one we want. There's a lot of good information in the ftp directories to plunder.

There seems to be yet a third script: GenomesGroup.cgi. This one is linked from the Virus genomes page.

If I really wanted to suffer, I'd look into NCBI's source. Does anyone know where the source of lproks.cgi or genlist.cgi are? Is that part of the NCBI C++ Toolkit? (which is on macports here.) Maybe it's buried in NCBI's ftp site? Maybe I should ask the NCBI Information Engineering Branch? Maybe I need to start doing something more productive!

Thursday, March 12, 2009

Split split

Apparently, there's some disagreement about what it means to split a string into substrings. Biological data frequently comes in good old fashioned tab-delimited text files. That's OK 'cause they're easily parsed in the language and platform of your choice. Most languages with any pretention of string processing offer a split function. So, you read the files line-by-line and split each line on the tab character to get an array of fields.

The disagreement comes about when there are empty fields. Since we're talking text files, there's no saying, "NOT NULL", so it's my presumption that empty fields are possible. Consider the following JUnit test.

import org.apache.log4j.Logger;
import static org.junit.Assert.*;
import org.junit.Test;

public class TestSplit {
  private static final Logger log = Logger.getLogger("unit-test");

  @Test
  public void test1() {
    String[] fields = "foo\t\t\t\t\t\t\tbar".split("\t");
    log.info("fields.length = " + fields.length);
    assertEquals(fields.length, 8);
  }

  @Test
  public void test2() {
    // 7 tabs
    String[] fields = "\t\t\t\t\t\t\t".split("\t");
    log.info("fields.length = " + fields.length);
    assertEquals(fields.length, 8);
  }
}

The first test works. You end up with 8 fields, of which the middle 6 are empty. The second test fails. You get an empty array. I expected this to return an array of 8 empty strings. Java's mutant cousin, Javascript get's this right, as does Python.

Rhino 1.6 release 5 2006 11 18
js> a = "\t\t\t\t\t\t\t";
js> fields = a.split("\t")
,,,,,,,
js> fields.length
8
Python 2.5.1 (r251:54863, Jan 17 2008, 19:35:17)
>>> a = "\t\t\t\t\t\t\t"
>>> fields = a.split("\t")
['', '', '', '', '', '', '', '']
>>> len(fields)
8

Oddly enough, Ruby agrees with Java as does Perl.

>> str = "\t\t\t\t\t\t\t"
=> "\t\t\t\t\t\t\t"
>> fields = str.split("\t")
=> []
>> fields.length
=> 0

My perl is way rusty, so sue me. but, I think this is more or less it:

$str = "\t\t\t\t\t\t\t";
@fields = split(/\t/, $str);
print("fields = [@fields]\n");
$len = @fields;
print("length = $len\n");

Which yields:

fields = []
length = 0

How totally annoying!