Showing posts with label data. Show all posts
Showing posts with label data. Show all posts

Sunday, October 26, 2008

Transparency of data

What's the difference between national polls and scientific data?

As this article at Pollster.com points out, the difference is transparency. The article takes the example of climate change modeling as one instance where a set of people with a big heap of quantitative data and statistical models share the data and the models' assumptions.

Interestingly, it's only since one month ago that clinical trials were required by the FDA to make some basic data accessible--September 27 of this year. But actually, it seems like this does not give the opportunity to re-examine the raw data--only a kind of summary of demographics and outcomes. The intent is to stop people from concealing negative trials.

But, to take the pollster.com point in another direction, shouldn't drug trials be more transparent than political polls, which are run for profit by people who have a financial interest in concealing their raw data and their weighting methods (e.g., for "likely voter" screens)?

Oh, right. So are drug trials. Sorry.

But the ultimate transparency, and one that seems like it's long overdue, is for raw clinical trial data to be open-source, so that people with interests other than profit can examine that data and re-analyze it after the original academicians have published the initial report.

Sunday, May 4, 2008

Favorability



[Note: click on the graphs to get a full view including more-or-less readable data labels].

Here's a shout out to the graphic designers at the New York Times, who often produce great pieces of informational design, and who illustrated this op-ed about the black vote and the white vote in the Democratic primary. (I hope one of the data labels is in error: it cites the end of the graph as April 2, when the graph would really only be truly relevant if the end was May 2.) The article and the graphic make an important point: while the media has been fussing about whether Obama can win over white working-class men (many of whom will not vote for Clinton in the general election either), fewer observers of this political spectacle have been paying attention to the black votes that Clinton has been more or less deliberately throwing away and probably permanently losing.

The NYT article and graph are about a very specific question.

Contrast this to the more general and more common political discussion of whether a candidate is viewed favorably or unfavorably overall. Here's my quickly Excel-graphed illustration of Obama's approval ratings from November until now.


This graph uses overall national "approval" polling data from Rasmussen (the raw data are here) to show Barack Obama's approval ratings over time. The graph shows that the primary season has probably not had that much impact on how the overall electorate views Obama: mainly, people's views have on average become more certain (more "very" and less "somewhat"), but have not changed whether they like or dislike Obama.

The biggest shift came in late February where his "favorable" ratings got as high as 56% and his "unfavorable" ratings as low as 42%. In other words, the monumental flux of this campaign has been about 8% of voters who moved the center line between "kinda like" and "kinda don't like" back and forth.

My graph shows the effects of the political circus, the "who's up/who's down" tallies of cable news--and reveals a much more stable and enduring divide among voters, the one that persists election after election and actually does transcend personality. The smaller fluxes in "favorability" of any given candidate may or may not be important overall.

The NYT graph shows something more important: the actual effects of Hillary Clinton's behavior on a specific part of her base, and what could be one effect of her tactics if she were to win the primary. There is a significant inference here--the assumption that favorability ratings drive turnout. Maybe, maybe not. It may be that black voters would dislike her but vote for her anyway, which would probably be a rational choice. And there are a lot of things left unexamined in this graphic: it compares one group's view of one candidate with another group's view of another rather than comparing both groups' views of both candidates, which would likely be a more nuanced and less dramatic picture. Nonetheless, this single comparison and the clear presentation of the difference is much more interesting and reveals more significant shifts than where her favorability/unfavorability ratings have been going overall. (Not much changed.)

Sometimes smaller questions yield bigger answers.

Tuesday, April 22, 2008

Meta-analyses and pollster.com





The race for Minnesota's US Senate seat, US data on hypothetical McCain-Obama matchup for the general presidential election, and Pennsylvania Democratic primary polling, as shown by Pollster.com

I have a new addiction.

Pollster.com
is the website political junkies have been jonesing for even before we knew what it was. It clusters the results of polls that ask the same question--like, Who are you going to vote for? or, Do you think the country is on the right track? Then it puts them together into a single graph with a unifying trend line. It's imperfect--I can't satisfy myself that the trend line weights for sample size--but it's a lot better than reading the polls one by one.

The medicine parallel is in what we call meta-analyses--when we try to figure out a medical question by combining a number of studies that try to answer that question. Even if a bunch of smaller studies contradict each other, the idea is that by combining a number of studies you get the effect of having one huge study, and in this kind of data (with simple results like "worked better" vs "worked the same"), sample size is all. Thus, if you can create a meta-analysis that has the effect of creating one very large set of data, the answer those data give may be more reliable.

Like any kind of statistics, the problems get more complex as you try to get around the simplest problems. For instance, even when you weight for sample size, you can't throw less-reliable studies and more-reliable studies together and act like they're equivalent--so an ideal meta-analysis gives the data from more reliable studies more weight in the final result. But judging quality, and the appropriate weight given to difference between studies, begins to become a bit more subjective the more you try to fine-tune this problem. (What is quality? And how much weight does which measure of quality get?)

Pollster.com doesn't seem to do any weighting for sample size or other aspects of reliability, but even their relatively straightforward trend line is better than a lot of nonsense political handicapping you hear on TV politics talk shows. By giving more raw data and by combining large sets of data, these graphs and datasets allow you to begin cutting through some of the worst excesses of data-mining by stupid or biased pundits. In other words, you can be your own pundit.

I am aware, of course, that one of the ways that I manage to avoid coherent political action is by being a political junkie--an observer rather than a participant. Another thing that I need to change a little bit in the coming years.

PS:
Wikipedia on meta-analysis
Pollster.com on their trend-line method