Sunday, 13 January 2008

Collaborative Knowledge Generation

Here's an interesting idea, Open Source Learning:





Creating textbooks using a 'Knowledge Eco-System'. Open source tools and content. But how do we get the right people on the content i.e. how have domain knowledge and can create a useful and accurate textbook? This is a significant issue. Here is an example of a problem. I recently went to see Blade Runner: The Final cut, and this make me do some searching on the film to find out the actors in the film. In Daryl Hannah's Wikipedia page, someone edited her page with the following inaccurate information "She has recently had several notable roles, but was not in the movie Kill Bill, as is commonly believed.". This is wrong! The information was corrected - "She has recently had a number of notable roles including the Kill Bill series.". This is confirmed by another source, in case you haven't seen either film. She played Ellie Driver, a one eyed assassin. Richard argues for peer review at the end - I agree!




Which brings me on to the issue of Wikipedia - here's a talk by Jimmy Wales:






"Giving people access to the sum of all knowledge". Providing its accurate that is. I do like the idea of Wikipedia, but it always will have limits to its usefulness, for all the tools for tracking changes provided. For an example, see above. Oliver Kamm makes a number of valid criticisms of Wikipedia here and here.




Harry Barnes gives a Reasonable response to Kamm, but Jimmy Wales doesn't. No Jimmy, you must respond clearly and coherently to criticisms such as these.






Mena Trott on Blogs

Get your hankies out:





ok, its an entertaining talk and she makes some important points on the postive effects blogs can have.

Copyfarleft?

Yep, Copyfarleft. No really. No honestly I'm not kidding, I think they are serious. See:




Its as if the soviet union never existed, and all the failures which led to the downfall of that empire are not failures at all. Kleiner talks about 'art' in the context of Copyleft GPL, but this licence was designed for software only - not for artistic expressions such as music, plays etc. A better way for these is the Creative Commons. Artists should look to licenses based on this, not the GPL! The concept of 'copyjustright' seems to make sense for artistic works. The DRM issue is not directly related to SCO vs IBM case re: Linux and trying to make a direct link makes no sense at all. A piece of music is not the same as a piece of software (the former does not need to be maintained, the latter does). This misunderstanding makes the article pretty incoherent. Companies get involved in open source (or free) software development, largely because of the technical and economic benefits which accrue - they don't 'own' the software, property in terms of free and open source software is something of a fuzzy issue.



The article puts forward the notion that Copyfarleft could be applied to artistic works. I'm interested in the idea as applied to software. Copyfarleft discriminates against a) groups and b) individuals. Any license based on this concept (assuming it could be enforced legally) would reduce the number of people using and developing the software, and therefore reduce the effectiveness of the open source development model (less eyeballs on code). It isn't going to work for software.

Tuesday, 11 December 2007

Privacy Vs. Logging

Bit of a dilemma this. In a rather interesting article, Bin Tan and Zhai talk about the problems in providing both privacy and recording information to assit personalised search. Bit like having your cake and eating it!



There was a famous (infamous) case of a New York times reporter who identified somebody using their searches from a log AOL released without thinking about the privacy issues completely t e.g. properly anonymising data. Really, really silly!



Ask have taken a different view. They allow users to choose an option not to save their searches. This is fine, but how do you use information from the user to improve searches and personalise them? I'm not sure that Ask have fully throught through the implications.



References


Bin Tan, X.S. & Zhai, C. (2007). Privacy Protection in Personalised Search, SIGIR Forum, 41(1), November, 4-17.

Monday, 10 December 2007

You know when you've been tango'd

Been thinking about video search on youtube recently, particularly as more and more searches on Google seem to pick up videos from the source. I wondered if some Classic Ad's from the 1990's where available, and found them very quickly on youtube with the search string 'tango ads'. Here is the first classic 'orange man' one:





Much further down the list was the clown's Tango ad, you know the one with the narrator who sounds like John Major:





All accessed via a simple search string. Does depend however on good indexing on the part of the user - wonder how long it will take before real video indexing takes place on very large collections such as these. Finally heres another video from that search:





Weird!

Tuesday, 4 December 2007

Making music with veg

Now heres a novel application of Vegetables - music!. An example (any vegetarians of a nervous disposition are advised not to watch the video):





Wonder what they can do with a bit of broccoli? Brings in a whole new problem for Music IR - on organic material!

Tuesday, 13 November 2007

Temporal Information Visualisation

Hans Rosling, a professor in Public Health from Sweden, shows how you can take temporal information to make sense of your data - in the cases below the relationship between developed and non-developed countries. He comes up with answers which might surprise you. This is his 2006 TED talk:






Now for his 2007 TED talk (watch out for his sword trick at the end - had me cringing!)





I blogged about the relationship between information visualisation and information retrieval. What amazes me about these presentations is that Hans is able to present very complex data in a way which is very easy to comprehend for a non-expert. Case of a 'visualisation intermediary' at work here?