One thing I'm struck by, the more I work on medical projects, is how long everything takes. Now, partly this is a a reflection on tempo of academic research, and partly it's that these projects are challenging. But, frustratingly, in many cases there isn't a good reason - could often happen much faster, if we put our minds to it.
That begs the question of how we can more rapidly translate medical research into the clinic, where it can actually help someone.
One major limiting factor is the time it takes to get the resources to get a new project off the ground. If you need grant funding, the proposal itself can take months to put together, especially if you need to generate pilot data. Then the process of your submitted proposal being assessed will take months more, with the associated chance of rejection in most cases. And by the time a funded project gets started, hires the relevant people etc, over a year will have passed since you were first in principle ready to get started. There are now funding streams which try to speed things up, but more are needed.
This couples to another problem, that medical research is often very expensive. It's hard for funders to dole out a research grant quickly if it's for millions of pounds. They have a responsibility to the public/their donors/etc to ensure the money is used wisely. It is much easier for a funder to fund compact projects costing £100k than it is to dole out multi-million pound grants. Cheaper ways to do the research would help accelerate this process.
Then there's the research project itself. This can take a lot of time for all sorts of reasons. And I'm afraid that academics are often very bad at focusing on getting stuff done promptly. This is a mistake if one is concerned with maximising the value-per-unit-time one creates. At every stage, delays creep in because people are busy etc. If it takes a month to arrange a telecon to discuss something, that's 8% of a year wasted. Such delays add up. And I think academia can lack urgency, which in medical research is another mistake. If one aspires to develop something that can save 100 lives a year, a delay of 12 months has killed 100 people...
And even once all of the above is dealt with, we may well not even have begun the process of translation. There needs to be evidence that the research outcomes will work in the clinic (as opposed to a research lab), there needs to be a way of turning the research into a product that can then be deployed (e.g. commercialisation). Perhaps this even needs to be built into the research project from the word go, so the definition of a successful project is that something new has been deployed to the clinic. If all you've done is draw some conclusions and write a paper, maybe who cares?
Every single one of these steps is slow, ponderous, and flaky. What exactly a better system looks like isn't (yet) clear, but I'm pretty sure we want something a lot closer to an exponential organisation.
So, what should we do about all this?
Well, even a simple consideration of process optimisation would help. For example:
1. Remove all unnecessary steps. If there is some funding already in place (e.g. for pilot studies), the whole grant-application step can sometimes be avoided. If infrastructure is already in place (e.g. for data acquisition, sample collection), there is no need to have to build it.
2. Parallelise as many steps as possible. If we have a hundred candidate biomarkers for a disease, why not set up a project that can systematically test them all at the same time. And for clinical trials, multi-arm, multi-stage trials end up being hugely efficient in terms of treatments tested per unit time.
3. Make each remaining step as fast as possible.
How hard can it be...?
Showing posts with label the future of research. Show all posts
Showing posts with label the future of research. Show all posts
Thursday, 28 January 2016
Monday, 31 March 2014
Big Data - in the process of growing up
A very good article was published in the FT online a few days ago. Its title is 'Big Data: are we making a big mistake', and it's a commendably well thought out discussion of some of the challenges of Big Data. And perhaps a bit of a warning to not get too carried away :-)
You should go and read the full article, because it's got loads of great stuff in it. A few of the concepts that leapt out at me in particular were the following.
One of the interesting things about modeling data in order to make predictions (as opposed to explain the data), is that features that are correlated to (but not causal for) the outcome of interest are still useful. But, the FT article makes the really good point that even in this case, correlations can be more fragile than genuinely causal features. This is because while a cause is likely to remain a cause, correlations can more easily change over time (covariate drift). This doesn't mean we can't use correlation as a way to inform predictions, but it does mean that we need to be much more careful and be aware that the correlations may change.
The article also discusses sample variance and sample bias. This is in many ways the crux of the matter for Big Data. In principle, very large data sets offer us the chance to drive sample variance towards zero. But it really has much less to offer in terms of sample bias, and indeed (as the article points out) many of largest data sets, because of the way they're generated, are actually very vulnerable to high levels of sample bias. This is not to say that one can't have both (the large particle physics and astrophysics data sets are great examples where both sample variance and sample bias are addressed very seriously), but it is a warning that just because your data set is huge, it doesn't mean that is is free from sample bias. Far from it.
I've felt for a while that 'Big Data' (and data science) are currently going through a rapid phase of growing up, which I suppose is pretty inevitable because they're in many ways new disciplines. They've gotten up to speed on the algorithmic/computer science side of things very rapidly, but are still very much in the process of learning many of the lessons that are well-known to the statistics community. I think this is just a transition phase (these things take a bit of time), but it seems clear that a big part of the maturation of Big Data/data science lies in getting fully up to speed on the knowledge and skills of statistics.
You should go and read the full article, because it's got loads of great stuff in it. A few of the concepts that leapt out at me in particular were the following.
One of the interesting things about modeling data in order to make predictions (as opposed to explain the data), is that features that are correlated to (but not causal for) the outcome of interest are still useful. But, the FT article makes the really good point that even in this case, correlations can be more fragile than genuinely causal features. This is because while a cause is likely to remain a cause, correlations can more easily change over time (covariate drift). This doesn't mean we can't use correlation as a way to inform predictions, but it does mean that we need to be much more careful and be aware that the correlations may change.
The article also discusses sample variance and sample bias. This is in many ways the crux of the matter for Big Data. In principle, very large data sets offer us the chance to drive sample variance towards zero. But it really has much less to offer in terms of sample bias, and indeed (as the article points out) many of largest data sets, because of the way they're generated, are actually very vulnerable to high levels of sample bias. This is not to say that one can't have both (the large particle physics and astrophysics data sets are great examples where both sample variance and sample bias are addressed very seriously), but it is a warning that just because your data set is huge, it doesn't mean that is is free from sample bias. Far from it.
I've felt for a while that 'Big Data' (and data science) are currently going through a rapid phase of growing up, which I suppose is pretty inevitable because they're in many ways new disciplines. They've gotten up to speed on the algorithmic/computer science side of things very rapidly, but are still very much in the process of learning many of the lessons that are well-known to the statistics community. I think this is just a transition phase (these things take a bit of time), but it seems clear that a big part of the maturation of Big Data/data science lies in getting fully up to speed on the knowledge and skills of statistics.
Tuesday, 16 October 2012
The Sage Bionetworks - DREAM Breast Cancer Prognosis Challenge
I've been competing in the Sage Bionetworks/DREAM Breast Cancer Prognosis Challenge. The submission deadline was a few hours ago (early start for those of us in the UK!), so I thought now was a good time to share some of my thoughts on what has been a very interesting experience. I enjoyed it a lot and I think the folks at Sage Bionetworks are really onto something with this as a concept.
The goal of the challenge is to develop machine learning models that can predict survival in breast cancer. We've been given access to a remotely-hosted R system on which to develop our models, and (on said system) use of molecular and clinical data from the Metabric study of breast cancer. We run our models on this remote system and they're scored using concordance index, a nonparametric statistic for survival analysis that is sensitive to the ranking of predictions.
What makes this challenge a bit different is that it's both competitive and also collaborative. Not only are we competing against one another to get the best-performing model, but once someone has submitted a model, I can download it and inspect their code to see how it works! This is very ambitious (and certainly not without its issues), but aims to create a hybrid competition/crowdsourcing approach that can produce very strong solutions to the scientific goal of interest.
Having put a lot of hours into working on this challenge over the last few months, I have developed some opinions on it. So, in no particular order, here are my thoughts on the challenge:
First is the phenomenon of 'sniping'. Someone else can spend a month developing an awesome model, but once it's been submitted to the leaderboard I can download it straight away, spend 30 minutes applying my favourite tweak and then resubmit the (possibly improved) new model, jumping above the hardworking other competitor on the leaderboard. Of course, overall this leads to better models, which is the collective aim of the challenge. But I think care needs to be taken to ensure that credit (and reward in general) is given where it's due. It can be a bit dissatisfying when this happens to you!
The other consideration is that after a while of sharing models, we end up with a monoculture. Examinations of the high-ranking models over the last couple of weeks show that almost all the models are based on those of the Attractor Team (with some chunks of my own code scattered around, I was gratified to see!). This is probably not surprising, as the Attractor Team won both the monthly incremental prizes, but it's probably an indication that we've got about as far as we can with the challenge when this happens. Now is probably a good time to stop :-)
So, what would I change? I might suggest something a bit more like the following:
This structure uses initial competition to generate a lot of good ideas, then uses a second stage of competition to combine/evolve those ideas. It then has a final, collaborative phase where everyone who wants to pulls everything together to produce, publish and release the code for the challenge's solution to the problem. The challenge doesn't take too long to complete and the contributors get rewarded for their efforts in various ways.
--------
This post has turned into a long one and I hope I've communicated the intended positive tone. I enjoyed the Sage/DREAM BCC a great deal and I think this is a hugely powerful way of getting answers to scientific problems. I'm certainly going to take a look at whatever the next challenge is (I know there are some in the pipeline) and I would certainly recommend you doing the same.
The goal of the challenge is to develop machine learning models that can predict survival in breast cancer. We've been given access to a remotely-hosted R system on which to develop our models, and (on said system) use of molecular and clinical data from the Metabric study of breast cancer. We run our models on this remote system and they're scored using concordance index, a nonparametric statistic for survival analysis that is sensitive to the ranking of predictions.
What makes this challenge a bit different is that it's both competitive and also collaborative. Not only are we competing against one another to get the best-performing model, but once someone has submitted a model, I can download it and inspect their code to see how it works! This is very ambitious (and certainly not without its issues), but aims to create a hybrid competition/crowdsourcing approach that can produce very strong solutions to the scientific goal of interest.
Having put a lot of hours into working on this challenge over the last few months, I have developed some opinions on it. So, in no particular order, here are my thoughts on the challenge:
- Incentives for academics. In addition to some small financial prizes along the way, the big prize on offer is co-authorship on a journal paper. This is a very big incentive for academics (such as myself) who want to compete in a challenge like this. I'm very enthusiastic about the whole concept and would love to join in with future challenges. However, in order to justify spending my work time on it, there needs to be some kind of academic return. Co-authorship fits the bill nicely! Currently, I think the top couple of teams (?) get this prize, but I think extending this would be a good plan. Certainly, my experience in this challenge is that there are many more than 2 academic teams who have contributed significantly to the success of the challenge.
- Incentives for non-academics. Of course, it's also important to have rewards on offer for non-academics. The real strength of such a challenge comes from having a diverse community of competitors. I presume the small financial prizes are nice in this regard; it'd be really interesting to hear from some of the non-academic competitors what their views are on this.
- Sharing of code. This has been a very innovative (and brave!) aspect of the challenge. I don't think the organisers quite nailed every aspect of it, but I think the general approach is very powerful and certainly worth persisting with. I wonder if the sharing should be constrained in some way - perhaps code can only be accessed 48 hours after it is submitted?
- Blitzing the leaderboard? In this challenge we could make as many submissions as we liked to the leaderboard (of which I'm as guilty as anyone :-) ). This worries me as it could lead to a lot of over-fitting. Maybe in future challenges there should be a limit - say 5 submissions per day?
- Challenge length. In total it was approx 3 months long. 2 - 3 months feels about right to me.
- Competitive vs. collaborative. Another research model that's relevant here is the Polymath Project. Essentially, one can imagine a sliding scale between competition and collaboration. Polymath lives at one end, with things like the Netflix Prize and Kaggle competitions at the other. This challenge lives somewhere in the middle. I like the idea of blending the two concepts.
- Populations of ideas vs. monoculture. A competition is great for generating a wide range of ideas. Once people start sharing, I expect (as happened in this challenge) the pool of ideas tends towards more of a monoculture.
- Building an ongoing community. This challenge has been a great way of starting up a research community (a smart mob :-) ). It would be great to harness this community in an ongoing basis.
Sharing code
Sharing code means sharing ideas, and this has allowed us to benefit from each others ideas during the challenge. I'm sure this has led to better overall results. However, it has also has some quirks that might need tweaking.First is the phenomenon of 'sniping'. Someone else can spend a month developing an awesome model, but once it's been submitted to the leaderboard I can download it straight away, spend 30 minutes applying my favourite tweak and then resubmit the (possibly improved) new model, jumping above the hardworking other competitor on the leaderboard. Of course, overall this leads to better models, which is the collective aim of the challenge. But I think care needs to be taken to ensure that credit (and reward in general) is given where it's due. It can be a bit dissatisfying when this happens to you!
The other consideration is that after a while of sharing models, we end up with a monoculture. Examinations of the high-ranking models over the last couple of weeks show that almost all the models are based on those of the Attractor Team (with some chunks of my own code scattered around, I was gratified to see!). This is probably not surprising, as the Attractor Team won both the monthly incremental prizes, but it's probably an indication that we've got about as far as we can with the challenge when this happens. Now is probably a good time to stop :-)
So, what would I change? I might suggest something a bit more like the following:
A possible model for future challenges
The 21st Century Scientist Speculative Future Challenge (21SFC) would look like this:
Stage 1 (initial competition) - a month long competition to top the leaderboard. No-one can access other people's code and at the end of the month, a prize is awarded on the basis of a held-out validation set. After the deadline, all code for Stage 1 is made available.
Stage 2 (competition/code sharing) - another month long competition to top a new leaderboard. Everyone has access to the Stage 1 models, but Stage 2 code is either unavailable or only accessible 48 hours after is has been submitted. At the end of the month, a prize is awarded on the basis of a held-out validation set.
(it might be worth re-randomising the training, test, validation sets for stage 2)
Stage 3 (collaboration) - A non-competitive stage. The aim here is to work as a team to pull together everything that has been learned, produce 1 (or a small number) of good, well-written models and to publish a paper of the results.
The author for the paper is "21SFC collaboration", with an alphabetised list of people given. There can be different ways to qualify for authorship:
- Placing in the top-n in either stage 1 or 2
- Making significant contributions in stage 3 (the criteria for this would need to be established)
--------
This post has turned into a long one and I hope I've communicated the intended positive tone. I enjoyed the Sage/DREAM BCC a great deal and I think this is a hugely powerful way of getting answers to scientific problems. I'm certainly going to take a look at whatever the next challenge is (I know there are some in the pipeline) and I would certainly recommend you doing the same.
Tuesday, 17 July 2012
Open access science
The UK's publicly-funded scientific research is going open access.
This is a Very Good Thing. And I'm quite impressed that the UK government has been pretty decisive about this; I was emailed by the MRC (my main funder) a while ago, telling me that publishing in an open access way was now a condition of funding.
Leaving aside the potential practical complications in making this happen (which other people have considered more than I have), why is this such a good thing? There are a number of reasons.
Firstly, there's the basic consideration of who's paying. In this case, it's the tax-paying UK public. So it seems entirely reasonable that they should have access to the research that they've paid for. It actually seems faintly ridiculous that this was ever not the case (although this was because pre-Web, the cost of distribution wasn't insignificant).
Secondly, and very importantly, it accelerates the pace of scientific research (which relates to this blog post). When I write a scientific paper, I want it to be accessible as rapidly as possible to as many people as possible, and as easy to access as possible. This means that my research can be read, assessed and acted upon as rapidly as possible, which means that people can benefit from my work and/or find ways to refine the ideas therein. Faster is better.
Thirdly, there's a very important consideration of public engagement with science. Science and scientific research are getting more and more complex with time, both because we're already discovered a lot of the easy stuff and also because we come up with progressively more clever ways in which to advance. This is great, but it does mean that it's increasingly difficult for the non-specialist to understand a lot of the good science that goes on. There are all sorts of efforts that are going on to try to address this, but one really good one is to try to make sure that anyone can access any piece of research.
My anecdotal view is that governments are in general pretty rubbish at having a clue about internet-related developments. Politicians are very busy people, making it hard to keep up with developments. I also suspect that the demographic profile of the current generation of politicians means there is a high proportion of Hapless Techno Weenies. But in this instance, they seem to have correctly identified an important internet trend and acted on it. Kudos for that.
This is a Very Good Thing. And I'm quite impressed that the UK government has been pretty decisive about this; I was emailed by the MRC (my main funder) a while ago, telling me that publishing in an open access way was now a condition of funding.
Leaving aside the potential practical complications in making this happen (which other people have considered more than I have), why is this such a good thing? There are a number of reasons.
Firstly, there's the basic consideration of who's paying. In this case, it's the tax-paying UK public. So it seems entirely reasonable that they should have access to the research that they've paid for. It actually seems faintly ridiculous that this was ever not the case (although this was because pre-Web, the cost of distribution wasn't insignificant).
Secondly, and very importantly, it accelerates the pace of scientific research (which relates to this blog post). When I write a scientific paper, I want it to be accessible as rapidly as possible to as many people as possible, and as easy to access as possible. This means that my research can be read, assessed and acted upon as rapidly as possible, which means that people can benefit from my work and/or find ways to refine the ideas therein. Faster is better.
Thirdly, there's a very important consideration of public engagement with science. Science and scientific research are getting more and more complex with time, both because we're already discovered a lot of the easy stuff and also because we come up with progressively more clever ways in which to advance. This is great, but it does mean that it's increasingly difficult for the non-specialist to understand a lot of the good science that goes on. There are all sorts of efforts that are going on to try to address this, but one really good one is to try to make sure that anyone can access any piece of research.
My anecdotal view is that governments are in general pretty rubbish at having a clue about internet-related developments. Politicians are very busy people, making it hard to keep up with developments. I also suspect that the demographic profile of the current generation of politicians means there is a high proportion of Hapless Techno Weenies. But in this instance, they seem to have correctly identified an important internet trend and acted on it. Kudos for that.
Tuesday, 10 July 2012
Rapid Research Prototyping
Does research have to be this slow?
I find the pace of a lot of research projects frustrating. Sure, some things just take time, but I have become increasingly suspicious that a lot of bottlenecks in research are problematic simply because we haven't taken the time to find a faster way of doing something.
This can be very challenging. Experiments and data, for example, can be a lengthy and painstaking process. Data analysis, simulations/modeling, discussions with collaborators and even writing the paper can all take many months to complete. And of course we try to be efficient, but should we really care so much if things take a while to get done.
Yes.
And here's why.
Consider what science is. Science is the generation of new and improved memes for describing the natural world, whose fitness is judged using empirical evidence. It's a memetic process. This means that we should be thinking in terms of evolution of ideas. And one of the key ways in which we can accelerate any evolutionary process is the shorten the generation time. Because the sooner you get new research into the public domain, the sooner other people can benefit from it and the sooner you can get feedback.
There is also a second key point here. Scientific ideas gain most of their value from being tested by other people. And this can't happen until it's been released into the public domain. We should be thinking of our newly-minted science meme as no more a than prototype, that needs to be poked and prodded by as many other people as possible before it can even start to be thought of as being robust.
The idea of spending years or decades on a scientific magnum opus is the wrong plan; getting your work into the public domain is everything.
This absolutely does not mean lowering our standards with regards to quality; there is so much research being produced nowadays that we need to avoid drowning each other in mediocre research. But a single researcher or group can only make a science meme so good. Beyond a certain point, your idea needs to be tested by other people. At that point, faster is better. Much, much better.
What form, then, should science take in the 21st century. It should be about rapid research prototyping - the production and dissemination of new high-quality prototype science memes, as rapidly as possible.
Publish early, publish often. And optimise your bottlenecks.
I find the pace of a lot of research projects frustrating. Sure, some things just take time, but I have become increasingly suspicious that a lot of bottlenecks in research are problematic simply because we haven't taken the time to find a faster way of doing something.
This can be very challenging. Experiments and data, for example, can be a lengthy and painstaking process. Data analysis, simulations/modeling, discussions with collaborators and even writing the paper can all take many months to complete. And of course we try to be efficient, but should we really care so much if things take a while to get done.
Yes.
And here's why.
Consider what science is. Science is the generation of new and improved memes for describing the natural world, whose fitness is judged using empirical evidence. It's a memetic process. This means that we should be thinking in terms of evolution of ideas. And one of the key ways in which we can accelerate any evolutionary process is the shorten the generation time. Because the sooner you get new research into the public domain, the sooner other people can benefit from it and the sooner you can get feedback.
There is also a second key point here. Scientific ideas gain most of their value from being tested by other people. And this can't happen until it's been released into the public domain. We should be thinking of our newly-minted science meme as no more a than prototype, that needs to be poked and prodded by as many other people as possible before it can even start to be thought of as being robust.
The idea of spending years or decades on a scientific magnum opus is the wrong plan; getting your work into the public domain is everything.
This absolutely does not mean lowering our standards with regards to quality; there is so much research being produced nowadays that we need to avoid drowning each other in mediocre research. But a single researcher or group can only make a science meme so good. Beyond a certain point, your idea needs to be tested by other people. At that point, faster is better. Much, much better.
What form, then, should science take in the 21st century. It should be about rapid research prototyping - the production and dissemination of new high-quality prototype science memes, as rapidly as possible.
Publish early, publish often. And optimise your bottlenecks.
Wednesday, 9 May 2012
What's the point of a scientific paper?
Academics write research papers. It's a major way we disseminate our ideas and, increasingly, our continued career progression and funding depends upon it. We live in an era of metrics and impact factors.
Because of this, I've been thinking recently about what a scientific paper is, what is its purpose and how we might improve upon this.
I have also been thinking recently about science as a memetic process, that is to say in terms of evolving populations of ideas. I think this is a useful train of thought, which can give us some interesting insights into how science works and how to make it work better.
Thinking about things in this way, I have come to the conclusion that a scientific paper should contain a small number of high-quality, useful, interesting ideas. Plus whatever evidence is required to back those ideas up. And that's it. They should be as compact as possible and as simple as possible. I'd even go so far as to say that one should be able to communicate the point of a paper in the abstract.
Nowadays, the scientific literature is very large and keeping up with the new papers even in a small area is challenging. I skim read at least a couple of hundred abstracts a day (RSS feeds are awesome for this), and it isn't going to get any better. But if the paper contains a small number of well-supported ideas that are well-communicated in the abstract and title, I can grasp them more easily and pick out the papers I want to read in greater detail.
I think that the point of a scientific paper is to be a communication channel for well-supported, clearly-stated scientific ideas. And the more succinct and high signal-to-noise the better.
Because of this, I've been thinking recently about what a scientific paper is, what is its purpose and how we might improve upon this.
I have also been thinking recently about science as a memetic process, that is to say in terms of evolving populations of ideas. I think this is a useful train of thought, which can give us some interesting insights into how science works and how to make it work better.
Thinking about things in this way, I have come to the conclusion that a scientific paper should contain a small number of high-quality, useful, interesting ideas. Plus whatever evidence is required to back those ideas up. And that's it. They should be as compact as possible and as simple as possible. I'd even go so far as to say that one should be able to communicate the point of a paper in the abstract.
Nowadays, the scientific literature is very large and keeping up with the new papers even in a small area is challenging. I skim read at least a couple of hundred abstracts a day (RSS feeds are awesome for this), and it isn't going to get any better. But if the paper contains a small number of well-supported ideas that are well-communicated in the abstract and title, I can grasp them more easily and pick out the papers I want to read in greater detail.
I think that the point of a scientific paper is to be a communication channel for well-supported, clearly-stated scientific ideas. And the more succinct and high signal-to-noise the better.
Tuesday, 9 February 2010
The Information Hierarchy
Rands In Repose posted an interesting article which included a concept called the Information Hierarchy (also known as Wisdom or Knowledge Hierarchy) which I'd not previously encountered.
The idea is this: information can be classified in a 4-level hierarchy.
In the original version of this process, every stage was carried out by people. This no longer has to be the case, however. Much data gathering is now automated to at least some degree. Even if scientists are ultimately responsible for building and running the experiments/instruments, a lot of the heavy lifting is now carried out by automated or semi-automated systems, with data reduction carried out by software pipelines.
I would argue that we are also able to automate aspects of the second level of the hierarchy, the production of information. Specifically, I think one can regard statistical modeling and machine learning as doing just that. We live in an era of phenomenal scientific data production, so we now routinely use (and create) statistical methods for extracting the useful information from these giant data-sets.
So I think this begs an interesting question: I wonder how much of this process we might ultimately be able to automate, and in what ways? (and what would the implications be of automated systems capable of the Knowledge and Wisdom levels?)
The idea is this: information can be classified in a 4-level hierarchy.
- Data - the raw material of knowledge
- Information - data that have been organised/presented
- Knowledge - information that has been acquired and understood
- Wisdom - distilled and integrated knowledge and understanding
In the original version of this process, every stage was carried out by people. This no longer has to be the case, however. Much data gathering is now automated to at least some degree. Even if scientists are ultimately responsible for building and running the experiments/instruments, a lot of the heavy lifting is now carried out by automated or semi-automated systems, with data reduction carried out by software pipelines.
I would argue that we are also able to automate aspects of the second level of the hierarchy, the production of information. Specifically, I think one can regard statistical modeling and machine learning as doing just that. We live in an era of phenomenal scientific data production, so we now routinely use (and create) statistical methods for extracting the useful information from these giant data-sets.
So I think this begs an interesting question: I wonder how much of this process we might ultimately be able to automate, and in what ways? (and what would the implications be of automated systems capable of the Knowledge and Wisdom levels?)
Tuesday, 12 January 2010
Science and the Internet
There are some interesting points in a recent online article by Martin Rees (president of the Royal Society and a very well-regarded astrophysicist). The article as a whole is very interesting and well worth a read, but a few ideas particularly grabbed me.
- the Internet enables wider participation in front-line science
- it allows new styles of research (for example, mining large publicly available data-sets)
- scientific discoveries can now be made by 'brute force' number crunching (e.g. exhaustive computational searches), as well as the more traditional methods of experiment, insight (and I would add theoretical calculation to the list)
- the Internet gives us faster access to resources, so we can get science done more quickly. For example, literature searches are much easier and faster to do when the papers are online and can be found via Google Scholar or similar.
Subscribe to:
Posts (Atom)