Showing posts with label advocacy. Show all posts
Showing posts with label advocacy. Show all posts

Wednesday, August 16, 2006

decaying filesystems

Wikipedia is a big collection of interconnected articles, each with its own edit history.
If we view each article+history as one object, then it will come as no surprise that links between articles link to the most recent version of the article.
But if instead we view each version of an article as a separate article, a different view comes out. Instead of wikipedia being a big collection of interconnected articles, we have a big collection of interconnected stacks of articles, where each stack represents the article and its previous versions.
Now why do links between articles automatically point to the 'top of the stack'? The article may be changed out of all recognition from what it was further down the stack (earlier in its history).
In fact, we arrive right at the paradox of the heap: when does an article undergoing many small incremental changes actually cross the boundary between being a modification of its earlier self to being a totally new article?
The answer is of course, that it doesn't. The boundary is entirely artificial, and almost entirely not useful.
So this is the fix: wikipedia article links should link through to the version of the article that is relevant to the text linking to it. Intuitive no?
This leads to some nice behaviours: because links can be redirected to different versions of an article you know you're pointing at what you intended to point at. Also, there is more emphasis on the age and stability of the information you're viewing.
But perhaps it would be tedious to have to keep updating links as articles improved over time. Perhaps there could be mechanisms that tracked people's browsing paths and updated links automatically. Or perhaps it would form the basis of a new recommendation system: major edits would gain approval from the community by getting linked to.
This system also suggests another improvement: branching articles. Each modification to an article is actually a branch. Vandalized branches would soon die, while community-approved branches would blossom.
You could even have a system whereby short branches with low link counts (or a high proportion of links from ancient but successful branches, representing abandoned links) could be migrated to lower quality media, or even disposed of: a sort of garbage collection for high-level human-readable information.
Branches are also excellent in that they solve the article renaming, moving, merging and splitting problems at a stroke: because the mechanisms to redirect links quickly are already in place, it is comparatively cheap to perform these operations.
The possibilities are endless. And it would be a job to get right, but it's something we could benefit from a lot I think.

Monday, September 19, 2005

letting things come to me

Since I've become more aware of the Web2.0 revolution, I've noticed more and more web technologies that I'd been thinking about years ago rising to the surface. A key example was an idea I had about online libraries where you list your books, and then other people can browse your library and if appropriate make requests to exchange or borrow books. It would be like an enormous communal, virtual bookshelf, except you'd get to read real books instead of reading off a computer screen (which despite what anyone says is still less comfortable than reading a book! at least to a majority of users...).
Well, now we have Listal, LibraryThing and AllConsuming. I'm not sure who owns them and whether I really want to commit my data to them, but the services are definitely there and doing what I thought such services should do years ago.
Which brings me to my main point. With services like these, I shouldn't have to think about committing my data. My data should reside where I want it to, and I should allow these services access to my data on my terms and conditions.
We need some sort of a platform for maintaining information, and then transforming it and submitting it. Preferably in a relatively extensible and/or standardized way. XML and XSLT style technologies seem to be screaming out to be used in this sort of position.
In addition, we need some sort of voluntary code of conduct whereby we can be reasonably assured that we can reliably dictate the terms under which such services can use our data. Maybe some open source datakeeper software modelled on recent digital rights management advances, so that as well as records companies being able to control our rights on the music we license from them, we can also revoke other organisation's rights on the data they license from us.
Digital rights management isn't necessarily a bad thing, but biased towards the goals of the powerful it is clearly not a good thing.
So the data landscape of the future? Data residing in multiple incarnations on various storage devices across the world, controlled by open source datakeeper software allowing only authorised people to access it, and transform it, using flexible tools. The owners of the data - you and me in addition to the organisations and corporations - empowered by our data's newfound mobility and flexibility.

Friday, September 02, 2005

should we be suspicious of google?

"Google's mission is to organize the world's information and make it universally accessible and useful." - Google's mission statement

I would be a lot more comfortable if it was Google's mission to assist in creating the technologies that will free the world's information. As it is Google has simply asserted that it wishes to become the monopolistic broker for what is fast becoming planet earth's most valuable resource. Attached to the information Google wants to organise and serve to us is, potentially, almost all the value (of any kind) tied up on the planet. This isn't so hard to believe when you see the proliferation of gadgets, and monitoring gadgets, everywhere, and their increasing connectivity to open networks.

I'm an optimist, and I don't think Google can survive much longer behaving in the way it does. Just as I believe Microsoft is a dying monster, serving poorly crafted computing products to the illiterate computing masses out there, I believe Google will eventually go the same way, except I think Google will probably have left a much more valuable legacy in terms of experience, lessons learnt and contribution to the internet and activities undertaken thereon.

What we need is decentralized, open-source search. The days of concentrating this kind of resource in the hands of a single company really has to be over. If we stop kidding ourselves this is clearly the most sensible, and the only possible, way forward. One Google is a single point of failure for one of the most important resources on the internet.

Distributed search composed of hundreds of flavours of search engine scattered all over the planet, even on people's home computers and househouse appliances etc. would be a much stronger system, and would in my opinion have the potential to be a lot more trustworthy than our current centralized solutions.

Check out lucene/nutch. I will soon, it seems to have developed a little in the last few months.

Maybe we'll all be saved.

Google Maps and GMail and all the various Google-Wows are surely amazing, but not as scary as what they all inevitably point to.

So where does that leave all the big internet portals?

In my mind it has to be something along the lines of: decentralize-and-open-your-source or bust. There's no reason why one company can't produce a bulk of the technology and even benefit in a big way financially from it, but I think in this day and age it mustn't do it behind closed doors.

Security and trust demand it, and morality is begging for it.

Roll on the next few years!